Monday, March 17, 2014

Murder and NLP: The Taman Shud Case, Part 3


In previous posts, I've shown that the Tamam Shud Cipher (TSC) is almost certainly an initialism, a sequence of initials corresponding to English text(s), and that a large corpus of prominent literature does not contain the text from which the TSC was derived. This leads to the next question: Given the strong possibility that the text was written by someone (presumably non-famous) who handled the book in which it was found, is it possible for us to decode it? Are initialisms based on English text in general decodable? In my discussion here, I interpret the potentially ambiguous handwriting in a certain common way. Other interpretations may be valid, and we can consider that statistically, but I focus for now on this interpretation of the handwriting:

WRGOABABD
WTBIMPANETP
MLIABOAIAQC
ITTMTSAMSTGAB

Example TSC Readings

Consider the following:

(1)
We rarely go onto Australia’s beaches and bed down.
Wade the beaches in majestic peace and nervously enter the Pacific.
My love, I am blessedly opened, and I am quite certain
In the truth. Mercifully, the sleeper awakens me, stirs the girl, and blossoms.

(2)
Western radar groups operate Australian bases and base defenses.
Weapons testing base is militarily prepared. American nationals electrified the perimeter. Military liaisons in Adelaide boarded on an international aircraft. Queensland considering increases to territorial military trainees. South American mercenaries sent to Guam and Borneo.

Both of these texts are lucid (if not sparkling) English text that correspond to the TSC (give or take alternate interpretations of the handwriting). They took me about 15 minutes each to write. One is a love poem (somewhat like the text of the Rubaiyat in which the TSC was found) and one looks like something a Soviet spy might send to his superiors. A group of writers could surely come up with endless TSC solutions on these or practically any other themes and probably never, duplicate each other's work. The simple fact demonstrated here is: The TSC, like most initialisms, is undecodable into the original text because it has virtually limitless solutions. And therefore, if the TSC is not found in some previously existing text, the source text that generated it will never be known.

Grammatical Analysis

It is easy to lay out, analytically, why most initialisms from English text have endless numbers of solutions. Most of what I say here will apply to a great number of other languages, but I discuss English specifically.

English vocabulary falls into two rough subsets. There are function words, generally drawn from closed classes of words, and content words, generally drawn from open classes of words. Nouns, verbs, adjectives, and adverbs are content words. They comprise the vast majority of all of the words in English. Pronouns, conjunctions, prepositions, and articles are function words. Those parts of speech have only about 3 to 90 words each.

For all but the rarest letters in English, you can come up with a very long list of any of the content word classes that begin with that letter. For the function words, this is not possible. For example, there are only three coordinating conjunctions: and, or, and but. There are no coordinating conjunctions in English that begin with 'z', ‘t’, or ‘e’.

Take any typical passage in English text, and write down the initialism. You may then freely generate almost endless alternative readings of the initialism by changing the content words to other examples of the same part of speech.

"The platypus is one of the strangest animals in Australian fauna."
"The pioneer is one of the strongest archetypes in American fiction."
"The pastegh is one of the sweetest aliments in Armenian food."

It's easy to do this with the content words in virtually any sentence. So even if we kept the function words the same (as in the previous examples), we can still generate many sentences with the same initialism and warp the meaning entirely.

It is relatively difficult to play the same trick with function words, because there are so few options to swap in. However, we could choose different function words in different locations, in effect moving the pivot of function words to another location and then manipulate the content words in their new positions.

"Tommy, play in our old toyroom since Annie is acting funny."

Most function words begin with relatively common letters in English, which is true almost by definition. These letters can be used in other words, content or function, and so the location of function words in an initialism cannot be pinned down.

There are rare letters that could greatly restrict the freedom of recombination shown above. A sequence like XXQXX in an initialism might have actually no valid readings in English. However, the TSC has only one rare letter, Q. Does the Q, or any other pattern in the TSC significantly constrain the range of possibilities for TSC readings? If we can't determine for absolute certain the reading of TSC, can we meaningfully narrow it down?

Learning from Examples

Using the same Project Gutenberg corpus which was searched for matches of TSC substrings, we can search for shorter substrings to see which phrases might match them. Substrings of length 6 are useful for providing multiple matches (at least 7 unique readings) for each position in the TSC. If the particular letters in the TSC constrain the possible readings significantly, then we should see repeated patterns in the Gutenberg matches.

It should be noted first that the Gutenberg corpus has multiple copies of some texts within it, which inflates the counts unnaturally. This observation notwithstanding, there is no substring of length 6 in the TSC that has any single reading which comprises the majority of its Gutenberg hits. In other words, whichever reading we guess to be correct, it is wrong in the majority of cases – over 60%, in fact. Therefore, the would-be sleuth who writes a reading of the TSC and feels that their match is sure to be right is being seduced by the fallacy that the solution they have in mind is rare in matching the text. The sequence that achieves the best match is "do well to bear in mind", which still only covers 38% of the matches for DWTBIM, and is one of 59 different readings found in the Gutenberg corpus. For any 6gram in TSC, whichever guess you offer for the correct reading, you will probably guess wrong.

Can we do better trying to pin down exact words? If we use the 6gram readings from the Gutenberg corpus and tally (counting each reading just once, even if it appeared multiple times) how often particular words are used to fill the specific positions in TSC, and call the share that each word has for that position the derived probability. In these values, we see the same inherent ambiguity as indicated above. Every position allows at least 7 different words to stand for that letter, and in very few cases is the most common case more than 20% frequent, meaning that whatever word we guess in that position, we will probably guess differently than the original text. There is just one case where the most common case rises slightly above 50%: the A in position 24 is filled 53% of the time by the article “a”, a poor, and in any case ambiguous, starting point for interpreting the text. Almost all of the 33 derived probabilities that exceed 20% are exceedingly function words: “a”, “the”, “and”, “in”, “of”, etc. These give not even the slightest indication of topic, genre, or even tense or person.

Four words have a derived probability between 20% and 34% and provide a slight indication of tense and person: There are two such occurrences of “my” (positions 14 and 21), and one each of “is” (position 23) and “am” (position 29). These effectively indicate three votes for the correct TSC reading being written in the first person and two votes for the present tense. These are difficult to interpret as probabilities, however, since each of these votes is, in any case, less than 50% probable, and there’s no clear prior probability of tense and person for a random unknown text. It should be noted that a first person text still contains many third person references, and a text that is primarily in past or present tense still may contain many instances of the other. Therefore, we have a glimmering of an indication that TSC may be an initialism of a first person text, but this is far from conclusive, and in no other way helpful regarding the content. One related observation: There are no occurrences of Y in the TSC, so there are no second person pronouns (“you”, “your”, “yours”) although the second person can be spoken of through circumlocution without those words.

Finally, the derived probability of “quite” for the Q in position 30 is 43%, high among the derived probabilities, but still short of 50%, and utterly ambiguous regarding genre, topic, or content.

Summary

Cumulatively, we have conclusive evidence that the TSC is an initialism, no source text has been identified in literature, and if the source text is not found in an older source, it cannot be decoded into the original text.

This is, most importantly, a strongly negative result for those who have hoped that the TSC could be deciphered, helping to solve the mystery of the Somerton Man, which is still quite a bizarre story even if the TSC is left aside. It still leaves an intriguing situation in which the TSC, which came to light because of the Somerton Man, is effectively a second mystery, which one might have found and glossed over if it were not associated with a dead body, but has been elevated in importance because of the body. There are doubtlessly countless books in the world’s libraries that have mysterious scribbles inside, and no one pays them any particular attention. (I have found some in my older books, written in my own hand, mysterious to me years after they were written.) However, since the case has gotten so much attention, I’ll devote one more post to examining what the TSC might represent – why an initialism might have been written, and what remaining, however slight, possibilities exist for obtaining a definite reading of its content.

Saturday, March 8, 2014

Murder and NLP: The Taman Shud Case, Part 2

In my last post, I showed that the Tamam Shud Cipher (TSC) is almost certainly an initialism, a sequence of initials taken from English text. So what does it say, and why was it written?

It should be noted first that two hypotheses, not exclusive of one another, regarding the other facts of the case are:

H1) That Somerton Man was a Cold War spy who was operating for the Soviets or perhaps the Americans.
H2) That the nurse whose phone number was found in the book was a potential lover of the Somerton Man, and the Rubaiyat was given to him by her to share the romantic aspects of the poetry. She admitted having given another copy of the Rubaiyat to another man in 1945. That man and book both turned up during the investigation.

In that light, some possibilities for the source text behind the Tamam Shud Cipher:

P1) Passage(s) from existing literature.
P1a) Written out to help the writer memorize that literature.

P2) Something written by a person who once had the book.
P2a) A message sent by a Cold War spy.
P2b) A message written out for a friend or lover to read.
P2c) A rough draft for original poetry or prose using initials as shorthand.
P2d) Written out to help the writer memorize their own composition.

P3) Scratchwork for solving a puzzle.

We can investigate each of these possibilities, although definitive answers may prove elusive. I'll devote the rest of this post to beginning an examination of possibility (P1).

If the TSC can be found, in its entirety, in the initials of any previously-existing work of literature, then we can be sure we've found the answer. Why? The probability of two texts sharing the same initials is a function of their length, L, in words, and is approximately:

pMatch = 1 / 10L

The number 10 occurs here because the probability of two randomly-chosen initials from English text being equal is about 1/10, not the 1/26 you would see if all letters occurred equally likely. The exact parameter may be experimentally derived and certainly isn't exactly 10.0, but it's so close that I use 10 to make the math easier to comprehend. The shortest line in the TSC has 9 letters and the longest has 13. Finding a random passge of 9 words to match the shortest line has a probability of about one in a billion, while a passage matching the longest line has a probability of about one in ten trillion. Meanwhile, the probability of an accidental match for the entire 44 TSC is a dicey 10-44, which is essentially zero. If we find any existing text that matches the entire TSC, then it is definitely the text used to produce the TSC.

Note: This does not mean that if a person tries to write out their own original match for the TSC and succeeds that they have found the answer. In fact, it's not that hard to write text to match a given set of initials, which produces the illusion that someone can make quick and easy progress towards figuring it out. I will show in a later post how easy it is to concoct bogus matches after the fact.

Of course, finding matches for shorter substrings is easy, for sufficiently short length. Obviously, matches for substrings of length 1 and 2 are trivial to find, and by searching larger volumes of text, one finds matches of length 3, 4 and so on.

Work has been done in the past on trying to find exact matches for the TSC in a handful of major literary works, including the King James Bible. After an initial, failing search through digital copies of 40 books and 18 collections of poetry, I conducted a rather massive search which is as follows:

Project Gutenberg creates digital transcriptions of literary works. For my search of TSC matches, I downloaded by torrent the April 2010 collection of 29,500 books. Of these, 22,353 books are in text format. I preprocessed these to create an index of their initialisms, which amounts to 1.3 billion words of original text. I searched this for the longest matches that exist for TSC substrings, to arrive at these results:

R1) There is no exact match for the entire TSC, or any of its individual lines.
R2) The longest matches are of length 8, and there were twelve of these. It is apparent that these are entirely coincidental for the following reasons:

a) Many of them contradict one another, by matching the exact same subsequence of TSC.
b) They each begin in the middle of a sentence, and continue on into the middle of the next sentence.
c) We would expect, from the aforementioned formula, to find about ten matches of length L=8 and one match of length L=9 by sheer coincidence.
d) One of the matches is from 1963, many years after TSC was written down.

The matches were from nine works of literature, and one work each from science, economics, and reference. The matches are posted separately here.

In a nutshell, the Project Gutenberg corpus does not contain a match, and this should set the tone for any future searches. The works in this corpus are not only large in number but particularly central in literary importance. It is hard to characterize "literary importance" formally, but the reach of this corpus is impressive. When one thinks of authors predating 1950 who are likely to be taught in university literature courses, the number of works this corpus has is vast, although not comprehensive. For many famous works, it includes more than one edition. There are also translations of English works into other languages (although, recall, TSC is almost certainly an initialism of English text) and translations of literature in other languages into English.

Until a match is found, we can never prove that there is no match. Whatever corpus of n books we search, it's always possible that the exact match will be found in book n+1. It remains possible that TSC matches a book, poem, newspaper article, or other text not in this Gutenberg Corpus, but a lack of matches in 22 thousand books certainly suggests that searching any additional one – or one hundred – books chosen arbitrarily will yield a very low probability of finding a full, exact TSC match.


In an upcoming post, I'll discuss possibility (P2), and how to use a large corpus (now that we have one) as a resource for trying to decode the TSC. We can perhaps find information about the words or genre that are encoded within it. And that will also open up paths to searching for exact matches.

Murder and NLP: The Taman Shud Case, Gutenberg Matches

Appendix:

Longest matches of substrings of the Tamam Shud Cipher among the initials in a corpus of 22,353 Project Gutenberg books. These are all of length 8. There were many shorter matches and no longer matches.


Title, Author, Substring, Passage

"The Iceberg Express", David Magie Cory
cittmtsa: cake i think the mermaid took somewhat after

"On Laboratory Arts", Richard Threlfall
cittmtsa: care is taken to make the strokes as

"Showell's Dictionary of Birmingham", Thomas T. Harman and Walter Showell
cittmtsa: Church in this town, Mr. Thomas Smallwood, an


"Beatrix", Honore de Balzac
cittmtsa: contemplating in turn the marshes the sea and

"The Nail", Pedro de Alarçon
iaboaiaq: is a beautiful one, and I am quite

"The Evolution of Modern Capitalism", John Atkinson Hobos
ittmtsam: in the textile metal transport shipping and machine

"Northanger Abbey", Jane Austen
ittmtsam: it together that miss thorpe should accompany miss

"The Trouble with Telstar", John Berryman
ittmtsam: itching to take me to see a man

"Pensées", Blaise Pascal
mtsamstg: make them saint augustine montaigne s'bond the genealogy

"A Woman for Mayor", Helen M. Winslow
mtsamstg: motioned the stenographer and miss snow to go

"Ernest Maltravers", Edward Bulwer-Lytton
tpmliabo: that point my life is a bad one

"Carette of Sark", John Oxenham
ttmtsams: than twenty miles there soon after midnight steal

Sunday, February 23, 2014

Murder and NLP: The Taman Shud Case, Part 1

Computational Linguistics Murder Mystery Theatre

Cracking the code found near the body of a dead man could help reveal the identity of his killer.

On December 1, 1948, the body of a dead man was found on Somerton Beach near Adelaide, South Australia. The identity of the man (henceforth, Somerton Man) was nowhere to be found among his belonging, and no one ever came forth to identify him. He seemed otherwise healthy, and the cause of death was possibly a deliberate poisoning. The identity of his killer, if there was one, has not been established. It is possible that he was a Soviet spy, and possible that he was killed for reasons pertaining to espionage, but that is speculative. The case remains one of the strangest unsolved crimes in Australia's history.

A slip of paper was found in the dead man's pocket, containing only the words "Tamam Shud." This was identified as the final line of Omar Khayyam's Rubaiyat, and a man came forward saying that he found a copy of the book in his car near the beach the same day the body was found. It turned out to be the same copy the slip of paper was torn from. Written in the book was the phone number of a nurse who lived nearby, and denied knowing the man, but various claims have been made that she did know him and was lying.

What does this strange murder story have to do with Natural Language Processing?! Also written in the book was a sequence of letters, in five lines. There is some ambiguity in the handwriting, but one reasonable reading of four of them is as follows:

WRGOABABD
WTBIMPANETP
MLIABOAIAQC
ITTMTSAMSTGAB

In addition, a line which begins like the third line above was written and crossed out between the first and second.

These letters are not in any obvious way comprehensible, and it has been supposed that the letters might represent a cipher that, if broken, would shed some light on the case. From here forward, I will call this the Tamam Shud Cipher, or TSC, although it is not clear that it is actually is a cipher, an intentionally coded message, in the literal sense of the term. The TSC has not been decoded in a highly convincing way in whole or in part.

Following many hypotheses regarding the nature of the cipher, previous work, including that of the students of Derek Abbott at the University of Adelaide, has suggested that the letters may be initials from an English text.

In previous work, the frequency of letters in the TSC was compared to letter frequencies in other collections of text. First, with samples of several languages, and next with the initial letters of words in those languages. The second is not the same as the first, because letters occur with different frequencies in different positions in words; for example, 'e' is the most common letter in English, but 's' and 't' begin more words than 'e'. Abbott’s students found that the TSC letter frequencies match those of English initials significantly better than those of English text overall, and also better than initials or text in any of several languages. I have performed similar tests using different source texts and reach the same conclusion.

This provides evidence that the TSC is, in fact, an initialism, a sequence of initials from some specific text – in this case, a short English text. However, this evidence falls short of proof. It shows that an English initialism is the best of the possibilities that were tested, and a quite a few were tested, but it leaves open that an untested possibility would fit the TSC letter frequency just as well or better. It also leaves unexamined the possibilty that the letters in TSC are initials taken from English text but are possibly in another sequence.

And so, more definitive evidence that TSC is an initialism from a specific English text is to examine short sequences of letters in TSC and measure how they rank compared to the sequences of initials from English text in general. Grammatical patterns in English make some sequences of initials more common than the same initials in another sequence, so by performing this study with the TSC and variations of TSC with its letters scrambled in random order, we should see how likely it is that TSC preserves the expected sequences.

We can use an arbitrarily large corpus of English to generate the initial-letter ngrams for English, but the TSC itself is short, and therefore it samples the space of ngrams very sparsely. This means that many kinds of statistical metrics will show a mismatch between TSC ngrams and corpus ngrams even if they are initialisms from the same language. For example, Pearson correlations of bigrams from a known, but short, English initialism and the corpus will come up negative due to the sparseness of the short string’s ngram matrix.

A useful metric that is more sensitive when the string we are testing against a corpus is short is to generate the ngrams within the string, and calculate the mean of how high those ngrams rank among the corpus ngrams. This, in effect, gives the string credit for containing common initial ngrams, but doesn’t punish it for lacking other initial ngrams because it is simply too short to “get around to” them.

I generated 1,000 random shuffles of the TSC, and for the TSC and each shuffle, I calculated the mean rank of the string’s initial ngrams in terms of those generated from a corpus of 5 million words of English literature, for n from 2 to 5. If the sequence of letters in the TSC are initials selected at random from English texts, or if they were generated by some other means altogether, then the mean ranks for the TSC should be about 50th percentile in terms of the mean ngram ranks for the 1,000 random shuffles of TSC. If, however, the TSC was generated as an initialism from a specific English text, then its mean ngram ranks should be significantly higher than 50th percentile. The results follow:

N         TSC Percentile
2          85.6
3          92.3
4          96.4
5          93.6

We see that the results are convincing, particular when n=4 (the matrices begin to become sparse for n=5, weakening the result). For many scientific purposes, 95th percentile is offered as a standard of proof, and these results are fairly convincing that the TSC is an initialism, in correct order, of English text.

However, notes also that the TSC is written as a series of lines which may be linguistically unrelated to one another. If the four lines of TSC are lines of poetry, separate sentences, or in any way excerpts of a longer text, then the ngrams that are generated across the boundaries of lines potentially introduce noise. The idea that the lines are separate entities is further validated by the fact that the crossed-out line occurs in a different order than the similar line which is not crossed out.

So we can repeat the analysis, comparing TSC’s ngrams to those generated from random shuffles of TSC, but excluding the ngrams of TSC that start on one line and end on another. If TSC is not derived from sequences of English initials, we would, again, expect to see it rank about 50th percentile in mean ngram rank among the random shuffles. What we see, instead, greatly strengthens the previous result.

N         TSC Percentile
2          96.9
3          99.2
4          99.2
5          99.2

It is exceedingly unlikely that any other method of generating lines of letters would show this regularity if these lines were not initialisms corresponding to one or more short English texts.


And so, this turns the focus deeper, given that the Tamam Shud cipher very likely is an initialism of some short English text(s), what does it say, and what does that say about the case? The story goes on in my next post.

Saturday, February 15, 2014

APIs for Multilingual NLP

Coming soon, I will be able to announce the availability of APIs that provide text analytics in the world's major business languages.

A placeholder to that site is at gistology.com, showing with a map the scope of the countries whose languages will be covered. There'll be more to say about this soon.

Is Google Racist?

Like many service available online, Google's search bar offers completions that suggest the rest of a user's query before they have to type it all out. Potentially, you could enter a 25-letter query in just a few keystrokes, saving a little time – and who doesn't love saving time?

Sometimes, though, you see something like this:

In this case, I deliberately "baited" Google by starting out with a query that I knew would elicit a result like this, but it certainly took the bait. Maybe someone is wondering why black people are more likely than other races to have to sickle-cell anemia, but Google guesses instead that they have a racist question that want to research.

So Google was wrong, and fails to save the person a couple of seconds. But it does something more than that. It exposes the user to a couple of stereotypes that range from unflattering to extremely offensive. And by moderating the query a little, you can find a treasure trove of other stereotypes according to Google completions:

Californians are fake, stupid, and weird. Texans are stupid idiots. New Yorkers are rude and arrogant. (Actually, most groups of people are rude, if you go by the Google completions.) Americans are obese and ignorant. Chinese people are smart. Jews are rich. Asians are bad drivers. At least, this is what Google completions suggest, in response to a partially-typed query.

Why do these completions exist? Who is actually putting these stereotypes forth as true? You'd have to know Google's backend architecture in detail to answer that exactly, but some experimentation indicates the following:

1) People who type queries into Google. Common queries are more likely to appear as completions.

2) People who create web pages. Some completions are oddly-worded and unlikely to have originated as queries, but appear verbatim in various web pages.

3) People who see these queries and then select them. This is where the algorithm becomes insidious. The intention of a completion is to help someone save the effort of typing. But an interesting completion might sidetrack someone from their original purpose and click on it just to see what it's about. For example:

Benjamin Franklin had syphilis? Maybe some student had to write a term paper on the great thinker and now they've been given the lurid suggestion that the man had a sexually-contracted disease! If true, it's certainly not why Franklin became famous, but there it is as one (in fact, two) of the top four relevant facts about the man. Never mind his efforts in founding the United States, publishing newspapers, discovering the electrical nature of lightning and so on: He had V.D.! That's bad enough if it were true, but in fact, it appears not to be! If you follow the links these completions lead to, none of them have any evidence that Franklin had syphilis – just people asking if he did. But I'll admit, the completion made me want to click on it. And that's exactly the problem. I just "voted" for the Franklin - syphilis link, elevating it a little bit higher than the other possibilities. If syphilis had started off as the #4 completion, people like me vote it up to a higher position on the list, so more sensational queries tend to "win".

As a consultant for a news aggregator startup, I once had access to the data of which headlines people clicked on. An unmistakable trend was that headlines with exciting, sensational words in them were clicked on more often. This was true even if the word was simply being used as a metaphor ("Interest Rates Explode", "President Attacks His Critics").

It doesn't require that a lot of people believe that Franklin had syphilis or that any of those stereotypes are listed are true. It only requires that enough people type that query (or click on webpages that assert or even ask the question), it gets somewhere on the list of completions (maybe #10) and then other people see that completion, get intrigued, and vote it up. In fact, a lot of the people who vote up the racial stereotypes could even be people who are incredibly offended by them.

Benjamin Franklin is dead, but these sorts of completions exist for living celebrities, too. I did a search for 3 former NFL quarterbacks, and one of them had a completion for "gay". The man in question has denied being gay. How many people, every day, see this completion? Google is in effect spreading rumors about the man's personal life, just as they flash before our eyes a number of racial stereotypes.

So what is Google's role in this? Surely not that someone at Google decided that these racial stereotypes are useful suggestions. Google only built the system. The data voted for these completions to rank so high. If Google's algorithms work so well in general, then they can claim neutrality on these questions and say, Sorry, but these are the completions lots of people type and choose.

Except Google isn't neutral. Google does sanitize their completions in many cases.

If you type "scarlett johansson photos..." into Google's search bar, and then add any letter of the alphabet, it will show you completions. Any letter, that is, besides "n". "scarlett johansson photos n" produces no completions at all. Why? And so what? The reason why is that Google has specifically censored the completion that would allow the word "nude" to appear. I've had access to the logs of queries from search engines, and I guarantee you that the completions Google shows for "scarlett johansson photos" are not more common than "scarlett johansson nude" or "scarlett johansson photos nude". In fact, Google shows no completions for the word "nude" all by itself, even though it shows completions for "Mohorovičić discontinuity"... which one do you think people are searching for more?

So Google isn't completely neutral. They let the data vote for itself sometimes, even usually, but they censor some completions on, apparently, the suspicion that the results would be offensive or non-family-friendly.

Given that, there's not much excuse for letting these racial slurs show up. If it's offensive to suggest that an actress has taken her clothes off, it's certainly more offensive to allow the data to promote the racial stereotypes listed above. "It's just data" is a valid excuse for a person or company who uses big data as a tool. But once human hands go to work in the system, selecting what does and doesn't show up, those hands start to take some of the blame for the whole system. One imagines that these stereotypes have simply been below Google's radar, and that the "Don't Be Evil" company would want to censor those completions once they're aware of them.

Sunday, September 25, 2011

Laughing Out Loud


This may make you laugh.

And laughs are something people like to share. When people communicate via social media, they type "laughs." In a sample of a million words of Twitter messages in ten different languages, I found that about 0.5% of all "words" are laughs – "haha", "LOL", or other ways of typing out a chuckle.

Do people everywhere laugh equally?

Not on your life.

In a study of ten Western languages (English, German, Dutch, Norwegian, Swedish, Danish, French, Spanish, Italian, and Portuguese), I found enormous differences in the frequency of Twitter-laughs.

The Germans laugh least, with Twitter laughs making up under 0.1% of all words.

Other languages of Northern Europe were somewhat more prone to laughs than German. In increasing order of laugh frequency, Norwegian, French, Swedish, English, and Danish all came in below 0.4%.

And then there are the happy Latins. Laughing just more than the Danes, Portuguese has 0.5% laughs, and that's nothing compared to the Italians who Twitter-laugh in 0.9% of words. But the runaway laugh champions are Spanish speakers who type Twitter laughs for 1.4% of words.

The North-South pattern is noteworthy, but is broken by the Dutch, who out-laugh their neighbors like they're misplaced Latins, finishing way up at 0.8%.

The Dutch withstanding, the North-South trend is sharp and undeniable, as this color-coded map makes clear.

What's even funnier, the languages where people laugh more often, they also type longer laughs.

While three "ha"s are the preferred laugh in Spanish and Italian, they are not afraid to laugh longer. The five-ha laugh ("jajajajaja") is more common in Spanish than the two-ha laugh is in German. This graph shows how laugh length occurs in the five most-spoken languages. While Spanish runs away with the championship here at every length, notice that English is actually the runner-up for the two-ha laugh ("haha") with Italian strongly preferring a three-ha approach ("ahahah").

When you take into account the length as well as frequency of laughs, Spanish Twitter has 24 times more laughing than German, as measured in character count. This is not a subtle difference!

So, why is all of this happening? It's clear that more Twitter laughs come from the warmer and sunnier countries.This is true not only in Europe but also in the Americas, where the most speakers of English, Spanish, and Portuguese live. Statistically speaking, the laugh statistic is highly correlated with the latitude of the corresponding European capital (farther south: r=0.66), how sunny that city is (more sun: r=0.74), and inversely with the suicide rate (r=-0.74; this is the same if you choose the U.S., Mexico, and Brazil instead of the U.K., Spain, and Portugal).

So is it as simple as this: Warm, sunny weather makes people laugh a lot and immune to depression?

That may be part of it. But another idea to consider is that in Germany and Scandinavia Twitter is used comparatively more often for business and relatively less often for chatting. When one subtracts the social chat, then naturally less laughter remains.

Overall, it's not clear how much Twitter reflects life as a whole. Until we plant microphones everywhere and monitor all human communication, studies like this will just be suggestive of larger truths. But insofar as it goes, this study of Twitter laughs serves to support a lot of existing cultural stereotypes.