Communication in human language can be signed, spoken, or written. Let’s take a look at the spoken language.
Sounds in Language
Spoken language is made up of sounds, which are organized into patterns that convey meaning.
Phonemes are the smallest units of sound that can distinguish meaning. Phonology is the study of how phonemes are organized and used in language.
The human anatomy is such that we can distinguish several hundred phonemes. No other animal has anywhere near this capability.
Exercise: But what about parrots?
Exercise: Research how the evolution of the human vocal tract and brain enabled the production and perception of such as massive number of phonemes. (A good place to start is Mithen’s The Language Puzzle, Chapter 5.) What were the selective pressures that led to this capability? How does this compare to other animals, such as birds, that can produce a wide variety of sounds?
Some spoken languages have only a few dozen phonemes. Some have over a hundred. They are impossible to count exactly, given that some counters may not distinguish short and long vowels, tones, morae, broadness, or other distinctions. Additionally, some dialects introduce new phonemes. Therefore most reports of phoneme inventory have a range rather than a precise number:
Language
Approx. Phoneme Count
Notes
Rotokas
~11–12
Hawaiian
~13–18
Depends on whether ā, ē, ī, ō, ū are counted as separate phonemes from a, e, i, o, u
Spanish
~25
20 consonants, 5 vowels
Italian
~30
23 consonants, 7 vowels
Japanese
~20–26
Depending on mora/long vowel count and whether palatalized consonants are counted separately
German
~36–48
English
~36–46
Depending on dialect and vowel analysis
Russian
~40–45
Large consonant count comes from palatalized/non-palatalized pairs, e.g. /tʲ/ vs. /t/
Arabic (Standard)
~34
Irish
~49–72
High count due to broad/slender consonant distinctions
Here’s some vocabulary for talking about sounds, roughly ordered from the smallest units up to larger structures:
Phone: any distinct speech sound, regardless of whether it distinguishes meaning in a given language.
Phoneme: the smallest unit of sound that can distinguish meaning in a language.
Allophone: one of the different phones that can realize the same phoneme without changing meaning, e.g., the aspirated [pʰ] in pin and the unaspirated [p] in spin are both allophones of the English phoneme /p/.
Minimal pair: a pair of words that differ by only one phoneme, e.g., pat vs. bat. Minimal pairs are how linguists confirm that two phones are separate phonemes rather than allophones of the same phoneme.
Voicing: whether the vocal cords vibrate during the production of a sound (voiced, e.g., /b/, /z/, /d/, /g/, /v/, /ð/) or not (voiceless, e.g., /p/, /s/, /t/, /k/, /f/, /θ/).
Consonant: a speech sound that is articulated with complete or partial closure of the vocal tract.
Vowel: a speech sound that is produced without significant constriction of the vocal tract.
Diphthong: a single vowel sound formed by gliding from one vowel quality to another within the same syllable, e.g., the vowel in face, /eɪ/.
Syllable: a unit of organization for a sequence of speech sounds, typically consisting of a vowel sound with optional surrounding consonants.
Mora: a unit of syllable weight, or length, used to measure the timing of syllables in some languages.
Phonotactics: the rules governing which phoneme sequences are allowed in a language, e.g., English allows words to start with str- but not with rtk-.
Suprasegmentals: features such as stress, tone, duration, juncture, and intonation that apply to syllables or larger units rather than to individual phonemes.
One more....
Telephone: tele- (far) + -phone (sound) or “far sound,” a device for transmitting speech sounds over a distance.
Classifying Sounds
There are quite a few dimensions on which sounds can be classified. For consonants, we have place, manner, and voicing. The place of articulation refers to where in the vocal tract the sound is produced, while the manner of articulation refers to how the sound is produced. For example, the sound /p/ is a voiceless bilabial plosive, meaning it is produced by closing both lips and releasing a burst of air without vibrating the vocal cords. Vowels are classified according to height, backness, and roundness. Details forthcoming. You’ll find that there’s more to sounds than just vowels and consonants!
The IPA
The International Phonetic Association (IPA) is the major organization for phoneticians. Their aim is “to promote the scientific study of phonetics and the various practical applications of that science.” They are best known for creating and maintaining the International Phonetic Alphabet (IPA), a standardized system for representing the sounds of spoken language.
Here is a screenshot of the IPA Chart, taken in 2026 (full attribution below). The alphabet is revised from time to time. Visit the official IPA Interactive Chart for an interactive version where you can click on each sound for further information and to hear each sound.
Here’s a little diagram showing where sounds can be made (taken from Eric Dunn):
This allows us to classify consonants:
Bilabial: Made with both lips together, e.g., /p/, /b/, /m/.
Labiodental: Made by touching the lower lip to the upper teeth, e.g., /f/, /v/.
Dental: Made with the tongue against the upper teeth, e.g., /t̪/, /d̪/.
Interdental: Made with the tongue against or between the teeth, e.g., /θ/, /ð/.
Alveolar: Made with the tongue at the alveolar ridge just behind the upper teeth, e.g., /t/,/d/, /s/, /n/, /l/, and the trilled r of Spanish, /r/).
Postalveolar: Made just behind the alveolar ridge, e.g., /ʃ/, /ʒ/, /tʃ/.
Palatal: Made by raising the tongue body to the hard palate, e.g., /j/.
Retroflex: Made with the tongue curled back toward the palate, e.g., /ʈ/, /ɖ/.
Velar: Made by raising the back of the tongue to the soft palate/velum, e.g., /k/, /g/, /ŋ/.
Uvular: Made by raising the back of the tongue to the uvula, e.g., /q/, /ʁ/.
Pharyngeal: Made by constricting the pharynx, e.g., /ħ/, /ʕ/.
Glottal: Made with constriction at the glottis or vocal cords, e.g., /h/.
Manner of Articulation
While place tells us where a consonant is made, manner of articulation tells us how, specifically, how much the airflow through the vocal tract is obstructed. Manners are usually listed from most obstructed to least:
Plosive (Stop): complete closure of the vocal tract, then a released burst of air, e.g., /p/, /b/, /t/, /d/, /k/, /g/.
Affricate: a stop immediately released into a fricative at the same place, e.g., /tʃ/ (church), /dʒ/ (judge).
Fricative: a narrow constriction forcing air through with audible friction, e.g., /f/, /v/, /s/, /z/, /θ/, /ʃ/.
Nasal: complete mouth, but air escapes through the nose instead, e.g., /m/, /n/, /ŋ/.
Trill: an articulator vibrates rapidly against the place of articulation, e.g., Spanish rr /r/.
Tap/Flap: a single, quick contact between articulator and place, e.g., American English tt in butter, /ɾ/, or the single r of Spanish pero.
Approximant: articulators come close together but not close enough to create turbulence, e.g., /w/, /j/, English /ɹ/. In a Lateral approximant, air flows around the side(s) of the tongue rather than over the center, e.g., /l/.
You may also hear the traditional cover term liquid, which lumps together the laterals and the rhotics (/l/- and /r/-type sounds)—it’s a useful phonological grouping, but not itself a manner category on the IPA chart the way, say, "fricative" is.
Aspiration (an extra puff of breath after a stop is released, as in the English p in pin [pʰɪn], versus the unaspirated p in spin [spɪn]) is not a manner in its own right—it’s a finer-grained detail within the stop manner, and in English it’s purely allophonic (see Vocabulary Time, above): it never distinguishes one word’s meaning from another’s.
Exercise: (Classic) Hold your hand in front of your mouth and say "pin," then "spin." Can you feel the difference in airflow on the /p/?
Exercise: Give examples of pairs of words in Hindi showing that aspiration variants are true phonemes rather than allophones.
Voiced Consonants
Voicing is the third dimension in addition to manner and place. Combining all three dimensions is what gives us fully specific labels like "voiceless bilabial plosive" for /p/, or "voiced alveolar fricative" for /z/.
Exercise: For each of the following, name the voicing, place, and manner: /m/, /ʃ/, /d/, /ŋ/, /f/.
Vowels
Vowels are classified along three independent dimensions, based on the position and shape of the tongue and lips:
Height
Close (High): tongue raised close to the roof of the mouth, e.g., /i/ (beet), /u/ (boot)
Mid: tongue in a middle position, e.g., /e/ (French été), /o/ (Spanish solo)
Rounded: lips pushed forward and rounded, e.g., /u/, /o/, /ɔ/
Unrounded: lips relaxed or spread, e.g., /i/, /ɛ/, /æ/, /ɑ/
Note that this classifies individual vowel qualities, or monophthongs. A diphthong is a quick glide between two of these vowel qualities within a single syllable. Here are some common diphthongs in English:
/aɪ/ as in my, high, fly, light, pie
/aʊ/ as in now, loud, town, cow
/eɪ/ as in say, face, day, play, break
/oʊ/ as in go, no, show, home, loan, though
/ɔɪ/ as in boy, coin, toy
/ɪə/ as in here, beer, or deer, fear
/eə/ as in air, hair, or bear, chair
/ʊə/ as in fur or sure
Exercise: How many dipthongs are officially recognized in different dialects of English?
Suprasegmentals
Not every meaningful sound feature belongs to an individual segment. Suprasegmentalsare (prosodic) features that extend over units larger than a single phoneme, such as syllables, words, or phrases, rather than acting as isolated consonant or vowel segments. Common suprasegmental features include stress, tone, intonation, length, and juncture:
Stress: extra emphasis (loudness, pitch, and length) on one syllable of a word. In English, stress placement can distinguish a noun from a verb: RE-cord (noun) vs. re-CORD (verb); PRO-duce (noun) vs. pro-DUCE (verb).
Tone: pitch used to distinguish word meaning, not just to convey emphasis or emotion. In Mandarin, for example, ma can mean mother, hemp, horse, or scold depending entirely on the pitch contour used.
Intonation: pitch patterns that apply across a whole phrase or sentence rather than a single word to convey attitudes, emotions, or grammatical structures, e.g., the rising pitch that often signals a yes/no question in English.
Length (duration): the relative duration of a sound or syllable, which can differentiate words in certain languages.
Juncture: the pausing or transition features that distinguish between word boundaries (such as ice cream vs. I scream).
Exercise: Say the English sentence You’re going to the store once as a statement and once as a question (by changing only your intonation, not the word order). What changes, physically, about how you say it? Make questions in which stress is placed on a single word, but a different word for each sentence question.
The UCLA inventory has 451 languages covered so far, and has identified over 500 consonants and over 200 vowels.
One fun fact: no vowel is common across all languages.
Exercise: Which vowels appear most frequently across different languages?
Exercise: List some vowels or consonants that seem to appear in only one or in a very small number of languages.
Exercise: Which languages have the fewest phonemes? Which have the most?
Exercise: Find some consonant pairs easy to distinguish by speakers of one language but difficult for speakers of another language.
You’ll notice IPA symbols written between either / / or [ ], depending on the language context:
Slashes, / /: a phonemic transcription. This represents sounds at the level of the phoneme.For example, English /p/ is written the same way whether it’s aspirated or not, because that difference never changes meaning in English.
Square brackets, [ ]: a phonetic. This captures the actual, physical sound in as much detail as is useful, including allophones. So in English, we need to capture aspired [pʰ] and unaspirated [p] in square brackets.
So the same English word can be transcribed two ways depending on what you want to show: pin is phonemically /pɪn/, but phonetically [pʰɪn].
Exercise: For the words top and stop, give both a phonemic and a phonetic transcription of the initial consonant, showing whether aspiration is present.
Evolution of Spoken Language
Let’s look at the origins and evolution of spoken language.
Origin Theories
Roughly speaking, we have origin theories that are (1) Gesture-First, (2) Vocal-First, and (3) Multi-Modal.
Gesture-First Theory
Maybe human language did not start off spoken. The Gesture-First Theory (also called the gestural origin hypothesis) argues that language first emerged as a system of manual and bodily gestures, with vocal speech added on later. Evidence and arguments offered in favor include:
Ape gesture is flexible; ape vocalization isn’t. Chimpanzees and bonobos use manual gestures intentionally, flexibly, and toward specific goals (e.g., reaching, pointing at food they want another individual to fetch), and their gestures show far more voluntary control than their vocal calls, which tend to be fixed, largely involuntary, and tied to emotional/arousal states (like an alarm call). If our closest living relatives already show voluntary, flexible communication mainly through gesture, this suggests that a shared ancestor with humans may have relied on gesture, too.
Mirror neurons. Neurons in the Macaque’s premotor cortex (area F5) fire both when a monkey performs a hand action and when it merely watches another individual perform the same action. Area F5 is considered a homolog of Broca’s area in humans, a region central to human language production and planning and observations of hand and arm movements. This overlap between an action/gesture-observation system and the brain region most associated with language suggests these systems may share a deep evolutionary history.
Sign languages are full languages. Natural sign languages (like ASL) have the full expressive and grammatical complexity of any spoken language, which proves that the gestural/visual channel is entirely capable of carrying complete linguistic structure, not just a simplified stand-in for speech.
Gesture comes early developmentally. Human infants typically point, wave, and use other communicative gestures for months before they produce their first words, and even congenitally blind people (who have never seen anyone gesture) spontaneously gesture while speaking to other blind people. This suggests a tight, perhaps ancient, link between gesture and the language faculty, independent of vision.
The theory posits that the shift from gesture to speech may have happened as hominins increasingly needed their hands free (for tool use, carrying food or infants, etc.) and that the advantages of vocal communication (works in the dark, around corners, over longer distances, and while the hands and eyes are busy with something else) offered selective advantages.
Vocal-First Theory
The Vocal-First Theory (or continuity-based theory) posits that human language evolved primarily from primate vocalizations, rather than gestures. Proponents argue that the vocal-auditory channel was the primary medium for early communication, with gestures playing a secondary role.
Evidence and arguments in favor of this theory include:
Primate vocalizations are flexible. While less flexible than gestures, many primate calls show some degree of voluntary control and can convey specific meanings, suggesting a potential foundation for vocal communication.
Infant vocal development. Human infants produce a wide range of vocalizations (cooing, babbling) before they develop meaningful gestures, indicating a predisposition for vocal communication.
Neuroanatomical support. Brain regions involved in vocal production and perception are highly developed in humans, supporting the idea that vocal communication has deep evolutionary roots.
Sound Changes
Vowel shifts, consonant changes, stress shifts occur as a natural drive for efficiency and for reasons of identity marking. The clearest examples for English speakers come from comparing the Romance branch of Indo-European (descended from Latin) with the Germanic branch (which includes English), since both diverged from the shared parent language, Proto-Indo-European (PIE).
Grimm’s Law, part 1: PIE voiceless stops → Germanic voiceless fricatives
Seen in the Germanic branch (English, German, etc.), not the Romance branch. Latin, being Italic rather than Germanic, keeps the original PIE consonant, so Latin/English cognate pairs show this nicely.
/p/ → /f/: Latin pater → English father
/t/ → /θ/: Latin tres → English three
/k/ → /h/: Latin centum → English hundred
Dyēus Ph₂tḗr, Zeus, and Jupiter
Speaking of pater: reconstructed PIE had a chief sky-god, *Dyḗus Ph₂tḗr, literally "Sky Father." That name is still recognizable in not one but two familiar names for gods. In Greek, it survived almost unchanged as Zeus Patḗr (Ζεῦ πάτερ), the vocative form used in Homer, which contracted over time into simply Zeus. In Latin, the same compound (Dyēu-pater) fused and shifted sound-by-sound into Iuppiter—the source of our Jupiter. (An older, less-fused Latin form, Diespiter, is also attested, and preserves the original compound a bit more transparently.) The same PIE root, *dyew- ("sky, to shine"), shows up all over the Indo-European family without the "father" part attached: Latin deus ("god") and dies ("day"), Sanskrit deva ("god," as in Dyaus Pitṛ, the Vedic sky-father), and even English Tuesday, named for the Germanic god Tiw (Old Norse Týr)—himself a descendant of the very same root as Zeus and Jupiter, several millennia and several sound changes removed.
Grimm’s Law, part 2: PIE voiced stops → Germanic voiceless stops
/b/ → /p/: Latin labium (lip) → English lip
/d/ → /t/: Latin decem → English ten (also duo → two)
/g/ → /k/: Latin ager (field) → English acre
Grimm’s Law, part 3: PIE voiced aspirated stops → Germanic voiced stops (but → Latin voiceless fricatives)
These are especially fun because they show both branches changing away from the PIE original, just in different directions. Note that these are written with slashes, /bʰ/, /dʰ/, /gʰ/, not brackets—unlike the allophonic aspiration on English stops discussed under the IPA section, above. In PIE (and in daughter languages that still preserve this system, most famously Sanskrit), aspiration on a voiced stop is not predictable from context the way it is in English; a plain voiced stop and its aspirated counterpart are separate phonemes, capable of distinguishing one word from another all by themselves.
/bʰ/ → /b/ in Germanic, /bʰ/ → /f/ in Latin: PIE *bʰrā́tēr → Latin frāter, English brother
/dʰ/ → /d/ in Germanic, /dʰ/ → /f/ in Latin: PIE *dʰwer- → Latin foris (doorway), English door
/gʰ/ → /g/ in Germanic, /gʰ/ → /h/ in Latin: PIE *ghóstis → Latin hostis (stranger, enemy), English guest
Example: Sanskrit bhrātr becomes English brother (bh → b). Similarly, dh → d as in Sanskrit dʰātr → English father, and gh → g as in Sanskrit agha → English egg.
Consonant cluster reduction
A later, purely English development (16th–17th century), unrelated to Grimm’s Law. English spelling still reflects the old pronunciation even though the sound itself dropped out of speech.
Loss of /k/ before /n/: knight, knee, know, knife
Loss of /w/ before /r/: write, wrist, wrong, wrap
Loss of /l/ before certain consonants: walk, talk, half, calm
Loss of final /b/ after /m/: comb, climb, limb, thumb, dumb, numb, bomb
Lenition (intervocalic voicing)
Common in the development of Spanish, French, and other Western Romance languages from Latin; a voiceless stop between vowels becomes voiced.
/p/ → /b/: Latin sapere → Spanish saber
/t/ → /d/: Latin vita → Spanish vida; Latin pater → Spanish padre
/k/ → /g/: Latin amica → Spanish amiga
Diphthongization
Common in Spanish: certain short, stressed Latin vowels split into two-vowel sequences.
Latin short e → ie: terra → tierra, tempus → tiempo
Latin short o → ue: porta → puerta, bonus → bueno
Vowel change at word endings
Common in Italian, where Latin’s word-final case endings collapsed into a small, regular set of vowels.
Latin -us/-um → Italian -o: lupus → lupo, amicum → amico
Latin -a → Italian -a (unchanged): rosa → rosa
End-of-word erosion
Common in French, which lost most of Latin’s word-final sounds over time.
Loss of the final unstressed -e in pronunciation, though it’s often kept in spelling: table is pronounced [tablə] in Old French but [tabl] in Modern French.
Loss of most final consonants in pronunciation: Latin digitus → French doigt, pronounced [dwa], with the final -t silent.
Vowel fronting (the "Anglo-Frisian Brightening")
An early, shared change in the ancestors of English and Frisian: Germanic /a/ fronted to /æ/.
Proto-Germanic *fadēr → Old English fæder → Modern English father
Proto-Germanic *that → Old English þæt → Modern English that
Vowel diphthongization (Great Vowel Shift and after)
English long vowels raised and eventually broke into diphthongs; still ongoing in most varieties of English today.
/eː/ → /eɪ/: Middle English name [naːmə] → Modern English name [neɪm]
/oː/ → /oʊ/: Middle English go [ɡoː] → Modern English go [ɡoʊ]
Umlaut (i-mutation)
A vowel is fronted or raised because of an /i/ or /j/ sound in the following syllable. The plural connection: in Proto-Germanic, many plural nouns were formed by adding a suffix containing /i/—singular *fōts "foot," plural *fōtiz "feet." That /i/ in the suffix dragged the stem vowel ō forward in the mouth to ē, giving *fōtiz a fronted stem even before the ending eventually disappeared. Once the -iz ending itself eroded away (a case of "end-of-word erosion," above) and stopped being pronounced, the fronted vowel was all that was left to signal "plural"—which is why foot/feet don’t just add -s like a regular English plural. The same process, on adjectives instead of nouns, produced some irregular comparatives.
foot / feet (Old English fōt / fēt, from Proto-Germanic *fōts / *fōtiz)
mouse / mice
goose / geese
man / men
old / elder (comparative, from Proto-Germanic *althiz)
full / fill (the verb fill comes from Proto-Germanic *fulljan, "to make full," with the -j- triggering umlaut)
Metathesis
Two adjacent sounds swap places.
Old English brid → Modern English bird
Old English wæps → Modern English wasp
Old English þridda → Modern English third (compare three)
Assimilation
A sound becomes more like a neighboring sound.
Latin in- + possibilis → English impossible (/n/ → /m/ before /p/)
Latin in- + legalis → English illegal (/n/ → /l/ before /l/)
Latin in- + regularis → English irregular (/n/ → /r/ before /r/)
Exercise: Grimm’s Law has a famous exception, discovered by Karl Verner and known as Verner’s Law. Look it up: what environment does it depend on, and why did it initially look like an exception to Grimm’s Law rather than a rule of its own?
Here’s shrot showing the evolution of English:
Mapping Sounds to Writing
English is an interesting case study. It is often said to have about 44 phonemes: roughly 24 consonants and 20 vowels (including diphthongs), though as we saw above, the exact count varies by dialect and analysis. English has been written with more than one alphabet. The earliest surviving English writing used a runic alphabet, a Futhark, brought over from the Germanic mainland. Later, missionaries brought the Latin alphabet, which had only five vowel letters and no letters at all for a couple of English consonant sounds—so scribes patched the gap by borrowing two letters, þ (thorn) and ƿ (wynn), straight out of the runic alphabet. Latin letters did their best mapping onto Old and Middle English sounds, and increasing French influence after the Norman Conquest introduced still more spelling quirks (like qu for /kw/ and ch for /tʃ/).
From Futhark to Latin Letters
The runic script used across the early Germanic-speaking world is called the Elder Futhark—the name comes from the sounds of its first six letters, F-U-Þ-A-R-K, the same way "alphabet" comes from the Greek Alpha-Beta. It had 24 letters, each standing for a single sound:
Rune
Name
Sound
ᚠ
fehu
/f/
ᚢ
uruz
/u/
ᚦ
thurisaz
/θ/
ᚨ
ansuz
/a/
ᚱ
raidho
/r/
ᚲ
kaunan
/k/
ᚷ
gebo
/g/
ᚹ
wunjo
/w/
ᚺ
hagalaz
/h/
ᚾ
naudiz
/n/
ᛁ
isaz
/i/
ᛃ
jera
/j/
ᛇ
eihwaz
/eː/ or /ç/
ᛈ
perthro
/p/
ᛉ
algiz
/z/
ᛊ
sowilo
/s/
ᛏ
tiwaz
/t/
ᛒ
berkano
/b/
ᛖ
ehwaz
/e/
ᛗ
mannaz
/m/
ᛚ
laguz
/l/
ᛜ
ingwaz
/ŋ/
ᛟ
othala
/o/
ᛞ
dagaz
/d/
When Germanic-speaking peoples (Angles, Saxons, Jutes) settled in Britain, they brought a version of this script with them, which grew into the Anglo-Saxon Futhorc (named the same way, from its first six sounds F-U-Þ-O-R-C). Old English had more vowel sounds than the original 24 runes could represent, so scribes kept inventing new runes, eventually reaching 29 (33 in Northumbria). Here is the standard 29-rune Futhorc:
Rune
Name
Sound
ᚠ
feoh
/f/
ᚢ
ur
/u/
ᚦ
þorn
/θ/
ᚩ
os
/o/
ᚱ
rad
/r/
ᚳ
cen
/k/
ᚷ
gyfu
/g/
ᚹ
wynn
/w/
ᚻ
hægl
/h/
ᚾ
nyd
/n/
ᛁ
is
/i/
ᛄ
ger
/j/
ᛇ
eoh
/iː̯o/
ᛈ
peorð
/p/
ᛉ
eolh(x)
/ks/
ᛋ
sigel
/s/
ᛏ
tir
/t/
ᛒ
beorc
/b/
ᛖ
eh
/e/
ᛗ
mann
/m/
ᛚ
lagu
/l/
ᛝ
ing
/ŋ/
ᛟ
eðel
/œː/
ᛞ
dæg
/d/
ᚪ
ac
/ɑ/
ᚫ
æsc
/æ/
ᚣ
yr
/y/
ᛡ
ior
/iː̯o/
ᛠ
ear
/æːa/
Exercise: Compare the two tables above. Which runes kept the same shape and sound? Which sounds needed brand-new runes in the Futhorc, and why do you think Old English needed them but Proto-Germanic didn’t?
Once Christian missionaries arrived and started copying manuscripts in the Latin alphabet, the Futhorc mostly fell out of everyday use, surviving mainly for inscriptions, charms, and the runic letters (þ and ƿ) that got folded into the new Latin-based spelling system.
The Great Vowel Shift
Then, roughly between 1400 and 1700 (with most of the action happening in the 15th and 16th centuries), English underwent the Great Vowel Shift: nearly every long vowel in the language shifted upward or, for the two vowels already at the top, broke into a diphthong. Spelling had already mostly settled down by the time this happened (largely thanks to the printing press, introduced to England in 1476), so English ended up with spellings that reflect Middle English pronunciation while the words themselves are now said very differently. Here’s the whole shift, vowel by vowel:
Middle English /iː/ (as in time) became Modern English /aɪ/
Middle English /eː/ (as in meet) rose to become Modern English /iː/
Middle English /ɛː/ (as in meat) rose to /eː/ and then, for most words, on up to merge with the /iː/ above (a few words, like great, break, and steak, stopped one step early and ended up as /eɪ/ instead)
Middle English /aː/ (as in name) rose to become Modern English /eɪ/
Middle English /ɔː/ (as in boat) rose to become Modern English /oʊ/
Middle English /oː/ (as in moon) rose to become Modern English /uː/
Middle English /uː/ (as in mouse) became Modern English /aʊ/
Here are some fun examples:
Word
Middle English IPA
Modern English IPA
time
tiːm(ə)
taɪm
mice
miːs
maɪs
meet
meːt
miːt
feet
feːt
fiːt
meat
mɛːt
miːt
sea
sɛː
siː
great
ɡrɛːt
ɡreɪt
name
naːm(ə)
neɪm
make
maːk(ə)
meɪk
boat
bɔːt
boʊt
road
rɔːd
roʊd
moon
moːn
muːn
food
foːd
fuːd
mouse
muːs
maʊs
house
huːs
haʊs
Notice how meet and meat end up pronounced identically today, even though they started from two different Middle English vowels and are still spelled differently.
Exercise: Keep practicing your IPA skills. Then read aloud the IPA transcription for Middle English.
Like English, a lot of languages have adopted the Latin alphabet, even when the sounds don’t match the letters perfectly. Watch the following video to see if the Latin alphabet is a good fit for Tlingit.
Exercise: How badly does the Latin alphabet fit English? How well does it fit Spanish? Why the difference? What other alphabets are used for English?
Exercise: Research the click sounds of Zulu. How are they represented in the Latin alphabet? How are they represented in the IPA?
There are some very rare sounds that humans can produce, but are difficult for people whose native languages do not include them:
Recall Practice
Here are some questions useful for your spaced repetition learning. Many of the answers are not found on this page. Some will have popped up in lecture. Others will require you to do your own research.
What is a phoneme?
The smallest unit of sound that can distinguish meaning.
What is phonology?
The study of how phonemes are organized and used in language.