Skip to content

9.5 — Scripts, Alphabets and the Languages of the World

There are only a handful of ways to write a language down, every script in current use descends from a small number of ancestors, and a few have never been read at all.

The kinds of writing system

What are the four types of script?

Sorted by what one written symbol represents.

Logographic: one symbol per word or meaningful unit. Chinese characters. It requires thousands of symbols — literacy in Chinese needs around 3,000 to 4,000 characters — but the symbols are independent of pronunciation, which is why speakers of mutually unintelligible Chinese varieties can read the same text.

Syllabic: one symbol per syllable. Japanese hiragana and katakana, Cherokee. Works well for languages with simple syllable structures and badly for languages like English, which would need a symbol for strengths.

Alphabetic: one symbol per consonant and vowel, written in sequence. Latin, Greek, Cyrillic.

Abjad: consonants written, vowels mostly not. Arabic and Hebrew. It works because in Semitic languages the consonants carry the root meaning and the vowels carry grammar, which a reader can supply from context — the same way an English reader handles rd as read or road from the sentence around it.

And a fifth, which is the one most Indian scripts use.

What is an abugida, and why does it matter for Indian scripts?

A system where each consonant symbol carries an inherent vowel, and other vowels are shown by marks attached to it.

In Devanagari, is not k — it is ka, with the a built in. To write ki you add a mark: कि. To write ku: कु. To write a bare consonant with no vowel, you add a special mark called the virama or halant: क्.

This is a genuinely different design from an alphabet, and it fits the languages well. The unit on the page corresponds to a spoken syllable, so the script maps onto how the language sounds rather than onto an abstract sequence of segments.

The other feature is that the consonants are arranged phonetically, not arbitrarily. The traditional order groups them by where in the mouth they are produced — velars, palatals, retroflexes, dentals, labials — and within each group by whether they are voiced and aspirated. It is a phonetic chart, worked out over two thousand years ago, and it is why Sanskrit grammarians were able to describe articulation with precision that European linguistics did not reach until the nineteenth century.

Almost every Indian script — Devanagari, Bengali, Gujarati, Gurmukhi, Odia, Tamil, Telugu, Kannada, Malayalam, Sinhala — plus Thai, Lao, Burmese, Khmer, Tibetan and Javanese, descends from Brahmi and shares this structure.

Where did all the alphabets come from?

Almost all of them from one place, once.

A rock surface carved with rows of Brahmi script characters from an Ashokan edict
Brahmi script on an Ashokan edict, third century BCE. Every Indian script in use today descends from this, as do the scripts of Thailand, Cambodia, Myanmar and Tibet. Image: Wikimedia Commons.

The Proto-Sinaitic script of around 1800 BCE, described in 2.1, is the ancestor. Its letters began as pictures used for their initial sound — a picture of a house, bayt, standing for b.

From it: Phoenician, spread by traders around the Mediterranean. From Phoenician: Greek, which added vowels by reusing consonant letters the Greeks did not need; Aramaic, which spread east across the Persian empire; and Hebrew and later Arabic.

From Greek: Latin, via the Etruscans, which is what you are reading; and Cyrillic, devised in the ninth century for Slavic languages.

From Aramaic, most probably: Brahmi, though this is debated and an indigenous Indian origin has been argued.

So the letter A is a fallen-over ox's head. Phoenician aleph was a pictogram of an ox; the Greeks rotated it and used it for a vowel; it has been the first letter ever since because it was first in the Phoenician order, which is also why the word alphabet is alpha plus beta, which are aleph and bayt, ox and house.

The independent inventions of writing are very few: Mesopotamian cuneiform, Egyptian hieroglyphs, Chinese, and Mesoamerican. Everything else is descended or inspired.

The scripts

Why does Chinese have so many characters, and how do people learn them?

Because it is logographic, and because the system was never reformed towards phonetics.

A dictionary lists over 50,000 characters, most of them historical. Functional literacy needs around 3,000; a well-read adult knows perhaps 8,000.

They are not memorised as arbitrary pictures. Around 80 to 90 per cent are compound characters made of two parts: a radical giving a rough semantic category, and a phonetic component suggesting the sound. The character for ocean combines the water radical with a component pronounced yáng. Knowing the roughly 200 radicals and a few hundred phonetic elements makes most characters partially guessable.

Simplified characters were introduced in mainland China from the 1950s, reducing stroke counts to raise literacy. Taiwan, Hong Kong and Macau kept the traditional forms, which is why the same text can look different depending on where it was printed.

Typing was the great problem and it was solved elegantly. Modern input types the pronunciation in Latin letters — pinyin — and the software offers the matching characters. It works so well that a common complaint is character amnesia: people who recognise characters fluently but can no longer write them by hand.

Why does Japanese use three scripts at once?

Because it borrowed a writing system designed for a completely different kind of language and had to patch it.

Chinese characters — kanji in Japanese — were adopted from around the fifth century. But Chinese has little grammatical inflection and Japanese has a great deal: verbs conjugate, particles mark grammatical role. Characters could carry the content words and had no way to write the endings.

So two syllabaries were developed from simplified characters. Hiragana writes grammatical endings, particles and native words without kanji; katakana writes foreign loanwords, onomatopoeia and emphasis.

A typical sentence uses all three, plus Latin letters for abbreviations and Arabic numerals. It looks chaotic and it does real work: because Japanese is written without spaces between words, the switch between scripts signals where words begin and end.

Hangul, the Korean script, is the opposite story and is frequently called the most scientifically designed writing system in existence. It was created deliberately in 1443 under King Sejong, with a stated purpose of making literacy accessible to ordinary people. The consonant shapes are schematic diagrams of the mouth position that produces them, and related sounds have related shapes. Letters are grouped into syllable blocks. It can be learned in a few hours.

Which scripts have never been deciphered?

Several, and each represents knowledge that is simply gone.

The Indus script, from the Harappan civilisation of about 2600–1900 BCE. Around 4,000 inscribed objects survive, mostly seals, with very short texts — the average is about five signs, and the longest known is 26. That brevity is the core problem: there is not enough material to find patterns, and no bilingual text exists. It is not even settled whether it encodes a language at all rather than a system of marks. The language it might encode is also disputed, with Dravidian and Indo-Aryan both proposed and the question entangled with modern politics.

Linear A, from Minoan Crete. Its successor Linear B was deciphered in 1952 by Michael Ventris, an architect working in his spare time, who established that it recorded an early form of Greek. Linear A uses many of the same signs and does not yield Greek, so the sound values are partly known and the language is not.

Rongorongo, from Easter Island, of which about two dozen inscribed objects survive.

The Voynich manuscript, a 15th-century illustrated book in an unknown script, which has resisted a century of professional cryptanalysis. It may be an elaborate hoax, though the parchment dates and the statistical properties of the text are more language-like than random.

How was Egyptian deciphered?

By a bilingual inscription and a twenty-year competition.

The Rosetta Stone, found by French soldiers in 1799, carries the same decree in three scripts: hieroglyphic, demotic Egyptian, and Greek. Greek was readable, so the content was known immediately. What was not known was how hieroglyphs worked.

The assumption for centuries had been that they were purely symbolic — each sign an idea. Thomas Young established that the cartouches, the oval rings around royal names, spelled sounds phonetically.

Jean-François Champollion completed it in 1822. He was fluent in Coptic, the late form of Egyptian written in Greek letters and still used liturgically, which gave him a living form of the language to test against. He established that the system was mixed — some signs stand for sounds, some for meanings, some clarify which reading is intended.

He is reported to have run to his brother's office, said "I've got it", and collapsed.

The languages

Which languages have the most speakers?

By native speakers, roughly: Mandarin Chinese, Spanish, English, Hindi, Bengali, Portuguese, Russian, Japanese.

By total speakers including second-language, English is first by a wide margin, followed by Mandarin, Hindi and Spanish. The gap between the two lists is the whole story of English: it has roughly three times as many second-language speakers as native ones, which is a ratio no other major language approaches.

Hindi's figures depend on a definitional argument, since the census counts a number of related varieties — Bhojpuri, Awadhi, Rajasthani and others — under Hindi, and whether those are dialects or languages is exactly the question from 9.4.

Which languages are hardest to learn, and for whom?

The question only makes sense relative to what you already speak, which is the useful part of the answer.

The US Foreign Service Institute publishes estimated hours for English speakers to reach professional working proficiency, and the categories are informative. Around 600–750 hours for Spanish, French, Italian, Dutch, Portuguese — close relatives. Around 900 for German, which has the same family but more grammar. Around 1,100 for Hindi, Russian, Greek, Hebrew, Thai, Vietnamese. Around 2,200 for Arabic, Chinese, Japanese and Korean.

The difficulty is distance, not complexity. Japanese is not intrinsically harder than Spanish; it is further from English in vocabulary, grammar and script. For a Korean speaker, Japanese is one of the easier languages available.

Features that reliably cost time: a different script, tones where the pitch changes the word, grammatical gender or noun classes with no semantic logic, case systems, and sounds that do not exist in your first language — which get harder to hear, not just to produce, after early childhood.

What are the strangest features found in languages?

Each of these is ordinary to its speakers and startling to everyone else.

Evidentiality: some languages, including several in the Amazon and the Caucasus, require every statement to mark how you know it — witnessed, inferred, or reported. You cannot say "it rained" without specifying whether you saw it.

Absolute spatial reference: Guugu Yimithirr in Australia has no words for left and right, only compass directions. Speakers say "the cup is to your north", and they maintain accurate orientation continuously, including indoors and in unfamiliar places.

Whistled languages: Silbo Gomero in the Canary Islands encodes Spanish as whistles that carry several kilometres across valleys. Similar systems exist in Türkiye, Mexico and West Africa.

Click consonants, in southern African languages such as Xhosa and !Xóõ, the latter of which has one of the largest consonant inventories of any known language.

Pirahã, in the Amazon, has been claimed to lack number words, colour terms and recursive embedding — the last of which would contradict a central claim about universal grammar. The claims are strongly disputed and the debate has been unusually heated, which is itself worth knowing: it is a live argument about whether any feature of language is truly universal.

What comes next

The next Part turns to the stories people told before any of this was written down — myth, and the surprisingly consistent shapes it takes across the world.