Less than one-half of one percent of all aspiring authors reach the best-seller list. Publishers receive thousands of manuscripts a year and even the literary titan like T.S. Eliot erred when he rejected the now-famous novel “Animal Farm” by George Orwell.
Now, Two scientists, Jodie Archer and Matthew Jockers claim to have created a new artificial intelligence (AI) algorithm, called the “bestseller-ometer”, that uses Big Data to predict the next best-seller with more than 80% accuracy. For a struggling publishing industry where the economic stakes couldn’t be higher for selecting the next winning novel, such a quantitative tool may be a windfall.
In their new book, “The Bestseller Code: Anatomy of the Blockbuster Novel,” from St. Martin’s Press, these researchers from the Stanford Literary Lab, describe how they analyzed over 25,000 novels and determined some 2799 relevant features for training their algorithms to detect bestsellers. Despite excitement by the predictive success of their algorithm, it is not without its critics. Skeptics argue that such a code may not detect truly new ideas and would create a stale “literary painting by numbers”.
Literary DNA: Theme and Plotline
Despite criticism, Archer and Jockers work is one of many that are attempting to quantify the humanities. Their findings are based upon identifying common patterns from a large number of novels, revealing a type of linguistic DNA that is a blueprint for what makes a novel a bestseller or not. By comparing bestsellers to other novels they identified two fundamental factors that distinguish the two: topics and emotional plotline.
With fine-grain word analysis, their topic-modeling algorithm creates word clouds – frequent words together with adjacently connected words. This collection of connected words is used to infer context and relevant topics (e.g., bar could be a place to drink or related to an exam). By applying this procedure throughout the entire novel, the algorithm quantifies the relative topic proportion. Archer and Jockers found that the blockbuster novels by authors such as Danielle Steel or John Grisham always contain the same proportions of topics: 1/3 of sentences are dedicated to what the author knows (e.g., either domestic family issues in the case of Steel, and legal issues in the case of Grisham). On average, non-bestselling novels have more topics, thereby lacking the same focus and developing unneeded subplots, that tend to frustrate the expectations of the readers. They also show that not only the number of themes matter, but also the theme itself: modern technology, jobs and the workplace, and human closeness.
AI and The Shape of Winning Plots
The other factor is plotline. Every aspiring writer knows the importance of capturing and maintaining the attention of the reader with the plot rhythm. It is the ebbs and flows of the novel’s plotline and language between characters throughout the book that compels the reader to turn the page. By quantifying the emotional impact, they could compare the plotline curves of different novels. They found that two recent blockbuster books, Fifty Shades of Grey and The Da Vinci Code, two books with completely different themes and plots, have nearly identical plotline trajectories while differing from nearly all other novels.
The work of Archer and Jockers is not unique. Sentiment and emotional analysis is an active field, and Archer and Jockers have fierce competition. A recent arXiv preprint by Reagan et al. at the University of Vermont’s Complex Systems Center Computational Story Lab, describe their Hedonometer and provide a comprehensive and exhaustive study of the emotional trajectories for the entire Project Gutenberg’s fiction collection. They demonstrate that in the existing fiction corpus, the essential building blocks consist of a set of six core emotional trajectories. They applied their work to a large set of classic work, making it possible to summarize work with respect to such trajectories: for George Orwell, it is possible to extract the animal farm summaries. Works such as this are putting literature on a quantitative footing, challenging traditional discourse in humanities.
Beyond the Bestseller
The technology behind the bestseller-ometer and Hedonometer is not only revolutionizing the book publishing industry. Deep Learning algorithms can use multiple data types (e.g., audio, video, images, text, etc.) to train large artificial neural networks for detecting sentiments and emotions. As an example, Facebook has recently developed Deep Text to understand the topic areas in social media data from users posts and associated photos. This system would allow Facebook to create more detailed user profiles, targeting their interests. Also, their tool will be used to power more realistic chat-bots, that will offer far more convincing conversations.
As the success of AI applications overcomes different endeavors of our lives, we may be forgiven to recall the stark warnings of Aldous Huxley, one of the literary giants of the twentieth century; this is indeed a “Brave New World”!