A
ñ
Second Language

Comprehensible Input: Why Understanding 95% Is the Magic Number

You have probably had the experience from both sides. Pick up a novel in a language you know well and a single unfamiliar word per page barely slows you down. You guess its meaning from the sentence, keep reading, and by the third encounter you own the word without ever cracking a dictionary. Now try the reverse: open a text where every other sentence contains three or four words you have never seen. Within minutes you are not reading anymore. You are decoding, grinding through a puzzle with too many missing pieces, and the meaning of the passage has evaporated. The difference between those two experiences is not a matter of willpower or study habits. It is a matter of percentages, and the research has pinpointed where the threshold lies.

The idea that language is acquired primarily through understanding messages, rather than through drills or grammar explanations, belongs to the linguist Stephen Krashen, who formalized it in the early 1980s as the input hypothesis. Krashen argued that learners progress when they are exposed to language that is slightly beyond their current competence, a level he called i+1: the "i" being the learner's existing knowledge, and the "+1" being the next small step. "We acquire language in only one way," Krashen wrote, "by understanding messages, or obtaining 'comprehensible input' in a low-anxiety situation." The elegance of the idea is its simplicity. You do not learn a language by studying its rules. You learn it by understanding things said in that language, provided the gap between what you know and what you are hearing is not too wide.

But how wide is too wide? That question moved from theory to measurement in 2000, when Marcella Hu and Paul Nation published a study that remains one of the most cited in vocabulary research. They gave learners a fiction text and systematically varied the percentage of words the learners knew, then measured how well they understood the story. The results drew a sharp line. At 95 percent coverage, meaning one unknown word in every twenty, some learners achieved adequate comprehension, but most did not. It was a minimum threshold, not a comfortable one. At 98 percent coverage, one unknown word in every fifty, the vast majority of readers could follow the text without assistance. [1] That three-point gap between 95 and 98 percent turned out to be the difference between struggling and reading freely.

Those numbers become vivid when you translate them into vocabulary sizes. Paul Nation calculated that to reach 95 percent coverage of general English text, a learner needs roughly 4,000 word families. To reach the 98 percent threshold for unassisted reading, the number climbs to between 8,000 and 9,000 word families. [2] A word family includes a base word and its common inflections and derivations, so "run," "runs," "running," and "runner" all count as one family. For a beginning learner staring at 9,000 as a target, the number can feel paralyzing. But it also clarifies the task. Language acquisition is not a vague journey. It is a countable one, and the research tells you exactly how many words you need before the language starts to teach itself.

The relationship between vocabulary knowledge and comprehension is not a cliff but a slope. Norbert Schmitt and colleagues demonstrated this in a 2011 study involving 661 participants from eight countries. They found a clear linear relationship: as the percentage of known words in a text increased, comprehension increased in a steady, predictable line. [3] There was no sudden threshold where understanding clicked on like a light switch. Every percentage point of vocabulary coverage purchased a corresponding increment of comprehension. The practical takeaway is that every word a learner acquires makes the next text slightly more penetrable, and slightly more likely to yield incidental learning of the words that remain.

Honest science, however, does not stop at confirmation. In 2023, Benjamin Kremmel and colleagues attempted to replicate the 98 percent threshold finding with Sri Lankan learners of English. Their results were more ambiguous. While vocabulary coverage predicted comprehension, the study could not cleanly reproduce the specific 98 percent threshold as a universal tipping point. [4] The implication is not that the earlier research was wrong, but that the exact number may shift depending on the learner population, the type of text, the learner's first language, and how comprehension is measured. Ninety-eight percent is a useful benchmark, not a law of physics. The underlying principle, that you need to understand the overwhelming majority of what you encounter for acquisition to happen, remains robust even if the precise cutoff has some flex in it.

Why does comprehensible input work at all? The mechanism is what researchers call incidental learning. When you understand most of a message, the few unknown elements are constrained by context. Your brain does not need to be told what a new word means. It infers the meaning from the surrounding language, the situation, the visual environment, the logic of the sentence. Each encounter with the word in a slightly different context sharpens the inference, building a richer and more durable mental representation than any definition could provide. This is how children acquire their first language: not through explicit instruction, but through thousands of hours of contextualized exposure where the world itself provides the meaning. Comprehensible input recreates that process for second-language learners, provided the input stays within the zone where inference is possible.

But Krashen's model, for all its influence, has always had a conspicuous gap. It treats language acquisition as something that happens to the learner through input alone. Merrill Swain, a Canadian applied linguist, noticed that French immersion students in Canada who received years of comprehensible input in French still produced grammatically inaccurate speech. They understood everything but could not produce it cleanly. That observation led Swain to propose the output hypothesis: the idea that producing language, not just receiving it, forces learners to notice gaps in their own knowledge and push toward precision. "Negotiating meaning needs to incorporate the notion of being pushed toward the delivery of a message that is not only conveyed, but that is conveyed precisely, coherently, and appropriately," Swain wrote. Output, in other words, is not just evidence that learning has occurred. It is part of the learning mechanism itself.

Michael Long added a third piece with the interaction hypothesis. Long argued that the back-and-forth of real communication, where speakers negotiate meaning, ask for clarification, and repair misunderstandings, is what connects input to acquisition most efficiently. "Negotiation of meaning facilitates acquisition because it connects input, internal learner capabilities, particularly selective attention, and output in productive ways," Long wrote. Interaction forces the learner to pay attention to exactly the features of the language that are causing the breakdown, which is precisely where growth happens. Pure input gives you the raw material. Interaction tells you what to do with it.

The most pointed critique of Krashen's framework arrived in 2025, when Nguyen and Doan published a review in Frontiers in Psychology arguing that the comprehensible input hypothesis is "conceptually flawed, empirically outdated, and practically insufficient." [5] Their argument drew on neuroscience: language production recruits broader neural circuitry than comprehension alone, engaging motor planning, articulatory processes, and self-monitoring systems that passive listening never activates. If acquisition depends on building neural pathways, and production builds pathways that comprehension does not, then a model that ignores output is missing a significant part of the picture. The critique was sharp, but it was also careful to distinguish between dismissing input and recognizing its limits.

Krashen's defenders have a response, and it is not a weak one. In 2021, Karen Lichtman and Bill VanPatten published an assessment of Krashen's legacy in Foreign Language Annals. Their conclusion was that many of Krashen's core ideas "have evolved and are still driving SLA research today, often unacknowledged and under new terminology." [6] The notion that comprehensible input is central to acquisition has been confirmed repeatedly under different labels: input processing theory, usage-based approaches, emergentist models. Researchers who would never cite Krashen by name are building on foundations he laid. The input hypothesis was not wrong. It was incomplete, and the field has spent four decades filling in the missing pieces.

The synthesis that has emerged from these decades of debate is not complicated. Input is necessary but not sufficient. You cannot acquire a language without massive exposure to comprehensible messages, and the 95-to-98-percent coverage range describes the zone where that exposure becomes maximally productive. But input alone leaves gaps, particularly in production accuracy and in the kind of rapid, flexible language use that real communication demands. Interaction and output fill those gaps. The strongest models of language acquisition treat all three, input, output, and interaction, as components of a single integrated process rather than competing theories.

This is where game-based language learning enters with a genuine structural advantage. A well-designed language game does not just deliver comprehensible input, though it does that naturally through visual context, adaptive difficulty, and narrative scaffolding. It also requires output, in the form of dialogue choices, written responses, and vocabulary challenges. And it creates interaction, through NPC conversations that adjust to the learner's level and provide immediate, contextual feedback. A 2025 meta-analysis of mobile language-learning games found an overall effect size of g = 0.962, and vocabulary-specific gains reached g = 1.251. [7] Those are large effects by any standard in educational research, and they reflect the fact that games address all three components of acquisition simultaneously rather than focusing on input or output in isolation.

There is another dimension worth noting. A 2020 study published in Scientific Reports by Frances, Martin, and Dunabeitia found that contextual diversity, encountering a word across different situations rather than repeating it in the same context, improved retention without requiring more total exposures. [8] A word encountered in a shop transaction, then in a quest dialogue, then in a character description, builds a richer memory trace than the same word drilled twenty times on a flashcard. Games generate contextual diversity as a natural byproduct of their structure. Every new scene, character, and quest is a new context, and the vocabulary travels with the player across all of them.

For learners, the practical implication of the 95 percent finding is both liberating and clarifying. It means that at the earliest stages, when you know few words, you need material specifically designed to keep you in the comprehensible zone: graded texts, simplified dialogues, games that introduce vocabulary incrementally and surround new words with enough context to make them guessable. As the applied linguist Gianfranco Conti has put it, "Masses of research have evidenced that with average-ability learners any L2 input that is less than 98% comprehensible is very unlikely to be conducive to learning." The goal is not to struggle heroically through material you do not understand. The goal is to find material where you understand nearly everything, so that the few unknown elements can be absorbed through context rather than lost in noise.

The magic number, then, is not a ceiling. It is a floor. Understanding 95 percent of what you read or hear is not the end of learning. It is the beginning of the conditions under which learning happens most naturally. Below that threshold, you are fighting the text. Above it, the text is teaching you. Every word you acquire raises your coverage percentage, which makes the next piece of input slightly more comprehensible, which makes the next word slightly easier to pick up. It is a virtuous cycle, and the research across multiple decades and populations converges on the same core truth: give the brain enough context to work with, ask it to produce as well as receive, and it will do what it evolved to do. It will learn the language.

Citations

  1. 1.
    Unknown Vocabulary Density and Reading ComprehensionHu & Nation, Reading in a Foreign Language, 2000
  2. 2.
    How Large a Vocabulary Is Needed For Reading and Listening?Nation, Canadian Modern Language Review, 2006
  3. 3.
    The Percentage of Words Known in a Text and Reading ComprehensionSchmitt, Jiang & Grabe, Modern Language Journal, 2011
  4. 4.
  5. 5.
  6. 6.
    Was Krashen Right? Forty Years LaterLichtman & VanPatten, Foreign Language Annals, 2021
  7. 7.
    Do Mobile Games Improve Language Learning? A Meta-AnalysisChen et al., Computer Assisted Language Learning, 2025
  8. 8.