What is Phonemes?
Phonemes are the smallest units of sound in a language that can change the meaning of a word. For example, changing the /b/ in bat to /p/ makes it pat. Each language has its own set of phonemes, and they are key to how we hear, process, and produce speech.
The Basic Idea
Unless you’re one of the lucky few, you’ve probably found yourself belting out the lyrics to a song with complete confidence, only to later realize that you got it wrong. Maybe you sang “Hold me closer, Tony Danza,” instead of “Hold me closer, tiny dancer,” in Elton John’s Tiny Dancer, or rapped “You broccoli that you are better now” instead of “You probably think that you are better now,” in Post Malone’s Better Now.
While it might be embarrassing, the reason that we make these mistakes and mishear words and phrases—known as mondegreens—is because of phonemes. Phonemes are the smallest units of sound that make up every word, usually used to distinguish one word from similar-sounding ones. For example, the phoneme /f/ in fit distinguishes it from words like hit, sit, and lit. Unfortunately, these tiny units of sound often sound similar to others. As the building blocks of language, our brains are decoding phonemes unconsciously and very quickly, this can lead us to decode a phrase incorrectly and sing the wrong words.1
Phonemes are not the same as letters. The same letter can represent more than one phoneme depending on the word in which it is used. For example, take the letter “c”: in country, it makes a /k/ sound, but in ceiling, it makes a /s/ sound. This also means that the same sound may be represented by different letters. The phoneme /f/ might be represented by an “f” like in the word fax or by “ph” in the word phase. Now you may better understand why it’s so difficult to learn new languages!2
It would then be found that the words, vowels, and phonemes are so many ways of ‘singing’ the world. The initial form of language, therefore, would have been a kind of song.
— Maurice Merleau-Ponty, French philosopher who spent a portion of his career exploring language and perception3
Key Terms
Linguistics: The scientific study of language that explores the various properties that make up a language. Linguists are interested in the comparison of languages and how languages evolve over time.4
Grapheme: A written symbol that we use to represent phonemes. Graphemes take the form of letters like “f” or “p,” which represent phonemes such as /f/ or /p/.2
Digraph: A combination of two letters that together represent a single phoneme. Digraphs are a type of grapheme, such as “ph” or “sh.”2
Allophone: The different ways that a phoneme can be pronounced depending on its phonetic context. Sometimes, the allophone is determined by where the letters are positioned in a word. For example, although the words hit and little both use the phoneme /t/, its pronunciation differs in emphasis and voicing. Allophones are also caused by different accents and dialects. The phoneme /w/ in water sounds different if a British person is speaking versus an American.5
Minimal Pair: Two words that are very similar and differ only by a single sound/phoneme. The differing sound must occur in the same position in the word. For example, hat and cat vary by their first sound, /h/ and /k/ respectively. Minimal pairs show us why pronunciation and enunciation are important.6 The minimal pair test is used in linguistics to see if the two sounds are distinct phonemes.7 For example, while you may think the words night and knight have a different starting phoneme because of the different spelling, they actually sound the exact same and have the same phonemes.
Mondegreen: A misheard word or phrase, which can cause people to repeat it incorrectly. This is common for song lyrics, or for children who are still learning language.8
History
From very early on, before the field of linguistics was formally established, early grammarians and philosophers were interested in the sounds that make up language. In the 4th century BCE, Sanskrit scholars were already exploring phonemes, with grammarian Panini developing a treatise on the ancient Indo-European language Sanskrit. Panini’s treatise included grammatical rules that governed the language, as well as notation of basic sound units.9
However, it wasn’t until 1873 that the term phoneme was used. French linguist Antoni Dufrice-Desgenettes introduced the term at the Société de linguistique de Paris, to describe a sound of speech.10 Dufrice-Desgenettes wanted a concise term to describe sounds in language instead of having to say “a letter of the spoken alphabet.” A few years later, Polish linguist Jan Niecisław Baudouin and his student Mikołaj Kruzewski developed the idea of phonemes further. Rather than simply representing a physical sound, they proposed that phonemes are mental categories used to distinguish meaning. They showed how changing just one phoneme could alter the meaning of a word, but that sometimes, slightly different sounds were still categorized under the same phoneme. They defined a phoneme as a mental representation of a group of sounds.
In 1888, the International Phonetic Alphabet (IPA) was created in order to provide a unique symbol for each phoneme. Danish linguist Otto Jespersen was the first person to suggest the idea of a phonetic alphabet in a letter to French linguist Paul Passy, who founded the International Phonetic Association. Jespersen wanted a way to standardize the representation of spoken language, thus leading to the birth of the IPA.11
In the late 1940s, Dr. Franklin S. Cooper of Haskins Laboratories, a research organization dedicated to the study of spoken and written language, created the Pattern Playback machine. The machine allowed a systematic study of the interaction between individual phonemes. It made it possible for researchers to discover key features and cues that help people recognize speech sounds. For example, say the words bat and pat out loud. These words are a minimal pair, but to help us distinguish between the two words, there is a longer delay in vocal cord vibration after the /p/ sound in pat than after /b/ sound in bat. The delay gives us a cue of which word is being spoken. Researchers at Haskins Laboratories defined this cue as the voice-onset time. This tool led to the development of speech perception research.12
The study of phonemes has continued to evolve, especially with the developments in technology that have led to the creation of digital and artificial intelligence (AI)-driven tools to analyze speech sounds. Today, phoneme research lives at the intersection of linguistics, neuroscience, and AI. What began as a way to describe the sounds of ancient languages has become a gateway to understanding how we think and communicate.
People
Antoni Dufrice-Desgenettes
A self-taught French linguist and teacher, whose passion for studying language was sparked by his world travels. Dufrice-Desgenettes is thought to have coined the term phoneme to describe a speech sound, making it easier to discuss in academic circles. Although his exploration into phonemes didn’t go any further, his terminological innovation paved the way for other linguists to conduct further research into units of sound across languages.13
Jan Niecisław Baudouin
A Polish linguist who developed the idea of phonemes by suggesting that instead of being mere physical sounds, they were structural entities that helped people distinguish meaning between words. Baudouin developed a theory with his student, Mikołaj Kruzewski, proposing that phonemes are mental representations of speech sounds. Baudouin spent most of his career working in comparative linguistics, researching how language and linguistic structure affect perception in different cultures.14
Otto Jespersen
A Danish linguist who was a leading expert on analyzing the rules of grammar for English. His understanding of grammar revolutionized the way that English was taught in Europe. His most significant contribution to phonemes was his suggestion that a phonetic alphabet be established to provide a universal framework to compare languages.15
Dr. Franklin S. Cooper
An American physicist who began his career researching radiation therapy. Together with American scientist Caryl Haskins, he co-founded Haskins Laboratories in the mid-1930s. During World War II, Cooper worked at the Office of Scientific Research and Development, where he was asked to create a program for the development of prosthetic devices for blinded veterans, including a reading machine. This sparked his interest in speech and he later developed the Pattern Playback machine, which allowed for the systematic study of phonemes.16
behavior change 101
Start your behavior change journey at the right place
Impacts
Phonemes may be small units of sound, but their influence is enormous. From early childhood literacy to speech recognition and assistive technologies, phonemes form the foundation of how we teach, support, and use language to interact today.
Helping children to read
Long before children develop a true understanding of language and its rules, they begin by sounding out letters and short words. If you’ve ever tried to teach an infant to read, or watched someone else do it, you’ll notice that we emphasize phonemes to help them distinguish the sounds in the word. If you are trying to get them to say the word cat, you would break it into /k/, /æ/, and /t/, and then repeat the full word.
In fact, phonemic awareness is the first step in teaching children to read. Phonemic awareness is the ability to notice and work with individual sounds in spoken words. Children need to identify other words with similar phonemes. A teacher might ask children, “Ball starts with /b/. What other words start with /b/?” As children get better at it, the teacher can move on to more complex words or more difficult questions, like figuring out what words rhyme with each other.17
Improving communication for people with hearing or speech impairment
Did you know that hearing and speech assistive technologies, such as cochlear implant systems, rely on the processing of phonemes to work? Cochlear implants convert sounds into electrical signals that are then sent to the brain. Cochlear implant systems classify individual phonemes in noisy environments and amplify those sounds to help the wearer distinguish what word has been said.18
Additionally, recent research suggests that technologies focused on phonemes may help individuals with speech impairment as a result of motor impairment, known as dysarthric speech. There is a wide range in severity of people with dysarthria, and the sounds produced by individuals with motor impairment can differ significantly from typical speech. This makes it difficult for traditional systems, often trained on “normal” speech, to function effectively. New research suggests speech impairment technology should be trained through phoneme-to-phoneme learning to better support people. AI tools would be trained by breaking words into individual phonemes and learning each one individually.19
Talk-to-text AI technology
Have you ever wondered how talk-to-text technology works? Talk-to-text is a form of AI tool trained on human speech. These systems are built by analyzing vast amounts of data, allowing them to associate different pitches and frequencies with certain words and phrases. When you speak into your phone’s microphone, it analyzes these features and predicts which phonemes you used. Then, based on the sequences of phonemes, it predicts which words you said—using language models and context to resolve ambiguity (e.g., whether you said “their” or “they’re”). Finally, it converts what it thinks you said into text. All of this happens in a matter of seconds, but it wouldn’t be possible without an understanding of phonemes!20
Controversies
While phonemes are widely used in linguistics, education, and speech technology, there is still considerable debate about how they work—and who they serve. From questions about their psychological reality to concerns about equity and inclusion, phonemes sit at the center of several important controversies.
Are phonemes fixed entities?
In the late 1870s, Polish linguist Jan Niecisław Baudouin developed the theory that phonemes aren’t just physical sounds, but mental categories and universal entities that are imbued with meaning. However, some linguists question whether phonemes exist as fixed categories in the mind or are just analytical constructs imposed by researchers. Treating phonemes as fixed categories creates a clear boundary between phonemes such as /b/ and /p/ and does not account for ambiguous sounds that fall between them.
Some linguists, who fall on the skeptical side of the phoneme debate, point to the fact that sounds change over time and vary from one another, suggesting that phonemes are not true psychological categories, but constructs imposed on language. If phonemes are not fixed categories, then speech recognition tools are not actually decoding distinct phonemes but making probable guesses. For example, if it detects a sound that is approximately 80% like a /b/ and only 20% like a /p/, it would predict that a /b/ has been spoken, unless the rest of the sequence of sounds suggests otherwise.21,22
How important is phonemic recognition for communication?
Although it is widely accepted that phonemic awareness is important in teaching children how to read, there is some debate about how important phoneme recognition is for effective communication.
Even when we don’t hear someone properly, we often can still deduce what they said. We use other clues, like the rest of the words we heard, context, and intonation, to make a likely guess for what we heard and respond accordingly. For example, if you’re working in a restaurant and a customer comes up to you and says, “Could you tell me where the bathroom is? I need to wash my hands,” but you heard the phoneme /p/ at the start of the word bathroom, instead of /b/, you’d still be able to understand and respond correctly. Because we can usually comprehend speech even when phonemes are missing or unclear, phoneme recognition may not be as important as some linguists and educators believe.23
Do phoneme tools perpetuate bias?
Phoneme charts, linguistic transcriptions, and speech technologies often reflect the phonemic inventories of widely studied, colonial, or high-resource languages—especially English. The IPA, for example, is largely based on the sounds used in English, but was developed to make it easier to compare phonemes across languages, hence the term “international.” However, there are sounds used in some languages that do not neatly correspond to the categories represented in the IPA.
Attempts to create universal frameworks may actually perpetuate bias by excluding marginalized languages. Speech recognition tools like Siri or Alexa have received a lot of criticism for performing worse when analyzing speech from individuals with accents or underrepresented groups. In fact, one study found that five different speech recognition programs from global companies like Apple and Microsoft were twice as likely to incorrectly transcribe words from Black speakers compared to white speakers. To get these tools to understand what they are saying, people from marginalized groups often have to adjust the way they speak, conforming to mainstream American speech patterns. This raises critical questions about whose speech is considered “standard” and whether our linguistic tools and phoneme categories are truly inclusive—or simply reinforcing existing power dynamics in communication.24
Case Studies
The multilingual phoneme recognition tool
Many languages are at risk of extinction, with only a few speakers remaining. Languages hold significant cultural value, and preserving them contributes to a diverse and rich culture. However, to do so, we may need to rely on digital tools.
In 2020, researchers at Carnegie Mellon University created the tool Allosaurus, a multilingual phoneme recognizer. This means that it is able to recognize sounds from multiple languages. In 2021, researchers used this tool to transcribe under-resourced languages at risk of extinction, such as Bukusu, spoken in Kenya, and Saamia, which is spoken in Uganda. Unlike conventional systems that rely on vast training data in one language, Allosaurus was designed to generalize across many languages by learning at the phoneme level.
The researchers found that even with the analysis of only around 1,000 utterances in these languages, Allosaurus was able to label phonemes with far fewer mistakes than other traditional speech recognition tools. They found that the tool was able to recognize approximately 80% of the phonemes across more than 2,000 languages. Phoneme recognition is an important first step in developing speech recognition tools, which allow for the documentation of language. By documenting near-extinct languages like Bukusu and Saamia, multilingual phoneme recognition tools can help preserve linguistic diversity with minimal data and infrastructure.25
Mind-reading through phonemes
AI has revolutionized what we thought was possible, and soon, AI tools may in fact be able to read our minds. Preliminary experiments have shown that brain-computer interfaces—technology that allows communication from our brains to an external device, like a speech generator—may be able to analyze brain waves to decode which phonemes someone is thinking.
In 2023, researchers conducted an experiment using an electroencephalogram (EEG) to train a neural network to identify phonemes from participants’ brain activity. Participants were asked to imagine speaking different words and sounds while the EEG recorded their brain activity. They found that the neural network was able to decode the phonemes with up to 97% accuracy in some participants.
By focusing on phonemes instead of whole words, the system gained the granularity needed to detect subtle shifts in speech planning. For individuals with severe speech impairment, such as those who have experienced strokes or paralysis, technology that can decode brain waves into sounds may allow them to communicate more easily. The potential for the technology to analyze brain waves and generate speech means that the individuals could communicate without using their voice or even any physical movement.26
Related TDL Content
Soundtracking a Better Customer Experience
Sounds not only allow us to communicate with one another, but they can also have a strong impact on our mood. Have you ever been bored waiting on hold with an institution when a pleasant tune comes on, instantly making you feel better about having to wait? That’s because sound has a profound impact on our perception of the world. In this article, we explore the psychological and behavioral effects of sound. We conducted a study to test the impact of seven different tunes on individuals’ stress levels.
Speech Recognition
If you want to learn more about how technologies like talk-to-text tools work beyond the analysis of phonemes, read this article by Mariana Ontañón. We explore the wide use of speech recognition tools across industries such as healthcare, the automotive industry, and customer service, which are leveraged to make communication more efficient.
Sources
- Encyclopaedia Britannica. (2025, May 31). Phoneme. https://www.britannica.com/topic/phoneme
- Craiker, K. N. (2022, June 28). Phoneme: Definition and meaning. ProWritingAid. https://prowritingaid.com/phoneme
- Merleau-Ponty, M. (1962). Phenomenology of perception (C. Smith, Trans.). Routledge & Kegan Paul. (Original work published 1945)
- Encyclopaedia Britannica. (2025, May 19). Linguistics. https://www.britannica.com/science/linguistics
- Encyclopaedia Britannica. (2025, June 3). Allophone. https://www.britannica.com/topic/allophone
- EnglishClub. (n.d.). Minimal pairs. https://www.englishclub.com/pronunciation/minimal-pairs.php
- Pretto, A. (2025, January 30). The Minimal Pairs Test – A tool to assess speech discrimination & optimize cochlear implant maps. MED-EL Professionals Blog. https://blog.medel.pro/rehabilitation/the-minimal-pairs-test/
- Nordquist, R. (2020, June 17). What is a mondegreen? ThoughtCo. https://www.thoughtco.com/what-is-a-mondegreen-1691401
- Encyclopaedia Britannica. (2025, June 22). Ashtadhyayi. https://www.britannica.com/topic/Ashtadhyayi
- TranslationDirectory.com. (2008, December). Phoneme. https://www.translationdirectory.com/articles/article1882.php
- Encyclopaedia Britannica. (2025, June 20). International Phonetic Alphabet. https://www.britannica.com/topic/International-Phonetic-Alphabet
- Haskins Laboratories. (n.d.). Pattern Playback. https://www.haskinslaboratories.org/pattern-playback
- Mugdan, J. (2011). On the origins of the term “phoneme.” HAL Archives Ouvertes. https://hal.science/hal-00941718/document
- Encyclopaedia Britannica. (2025, June 22). Jan Niecisław Baudouin de Courtenay. https://www.britannica.com/biography/Jan-Niecislaw-Baudouin-de-Courtenay
- Encyclopaedia Britannica. (2025, June 3). Otto Jespersen. https://www.britannica.com/biography/Otto-Jespersen
- National Academy of Engineering. (1999). Franklin S. Cooper (1908–1999). https://www.nae.edu/187832/FRANKLIN-S-COOPER-19081999
- Reading Rockets. (2025). Phonological and phonemic awareness. WETA Public Broadcasting. https://www.readingrockets.org/reading-101/reading-and-writing-basics/phonological-and-phonemic-awareness
- Chu, K., Collins, L., & Mainsah, B. (2021). A causal deep learning framework for classifying phonemes in cochlear implants. Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 6498–6502. https://doi.org/10.1109/icassp39728.2021.9413986
- Ee, W., Im, S., Do, H., Kim, Y., Ok, J., & Lee, G. G. (2025). DyPCL: Dynamic phoneme-level contrastive learning for dysarthric speech recognition. Proceedings of the 2025 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT 2025), 4701–4712. https://doi.org/10.18653/v1/2025.naacl-long.240
- Bansal, L. (2023, September 28). Evaluate the best speech to text models. Clarifai. https://www.clarifai.com/blog/evaluate-the-best-speech-to-text-models
- Schiller, L. M. G. B., & Pierrehumbert, J. (2015). The phoneme as a discrete unit. In J. Goldsmith, J. Riggle, & A. C. L. Yu (Eds.), The handbook of phonological theory (2nd ed., pp. 104–134). Wiley-Blackwell.
- Pisoni, D. B. (1997). Some thoughts on “normalization” in speech perception. Speech Communication, 22(2-3), 165–173. https://doi.org/10.1016/S0167-6393(97)00013-3
- Norris, D., McQueen, J. M., & Cutler, A. (2000). Merging information in speech recognition: Feedback is never necessary. Behavioral and Brain Sciences, 23(3), 299–325. https://doi.org/10.1017/S0140525X00003241
- Lloreda, C. L. (2020, July 5). Speech recognition tech is yet another example of bias. Scientific American. https://www.scientificamerican.com/article/speech-recognition-tech-is-yet-another-example-of-bias/
- Siminyu, K., Li, X., Anastasopoulos, A., Mortensen, D., Marlo, M. R., & Neubig, G. (2021). Phoneme recognition through fine-tuning of phonetic representations: A case study on Luhya language varieties (arXiv:2104.01624). arXiv. https://doi.org/10.48550/arXiv.2104.01624
- LaRocco, J., Tahmina, Q., Lecian, S., Moore, J., Helbig, C., & Gupta, S. (2023, December 18). Evaluation of an English language phoneme-based imagined speech brain–computer interface with low‑cost electroencephalography. Frontiers in Neuroinformatics, 17, Article 1306277. https://doi.org/10.3389/fninf.2023.1306277



















