How Spaced Repetition Makes Vocabulary Stick
In 1885, a German psychologist named Hermann Ebbinghaus sat alone in his study and memorized lists of nonsense syllables -- meaningless three-letter combinations like DAX, BUP, and ZOL -- then tested himself at precise intervals to see how quickly he forgot them. What he discovered was both elegant and brutal: memory decays in a steep logarithmic curve. Within twenty minutes of learning, retention dropped to 58.2 percent. After one hour, 44.2 percent. After a single day, just 33.7 percent remained. By thirty-one days, the number had cratered to 21.1 percent. [1] This was not a rough estimate. It was a mathematical function that described, with startling precision, how human memory leaks information over time. One hundred and thirty years later, Murre and Dros replicated his experiment almost exactly, spending seventy hours on learning and relearning tasks, and confirmed that the curve's shape holds. [1] The forgetting curve is not a metaphor. It is a measurement, and every vocabulary word you learn is subject to it.
The implications for language learners are immediate and uncomfortable. If you study fifty new words today and do nothing else, you will likely remember fewer than eleven of them a month from now. Cramming the night before a test can produce impressive short-term performance, but the curve shows why those words evaporate within days. The largest meta-analysis of distributed practice ever conducted, encompassing 839 assessments from 317 experiments across 184 published articles, confirmed that spaced learning consistently and substantially outperforms massed learning. [2] This is not a contested finding. It is one of the most replicated results in all of cognitive psychology, cutting across age groups, material types, and testing conditions. The question has never been whether spacing works. The question is how to calibrate it.
The answer turns out to be surprisingly specific. A landmark study by Cepeda and colleagues tested over 1,350 participants with spacing gaps of up to three and a half months and retention tests up to one year later. They found that at any given retention interval, increasing the gap between study sessions first improved and then gradually reduced test performance, tracing an inverted U-shape. The optimal gap was approximately 20 percent of the desired retention interval for shorter delays, falling to about 5 to 10 percent for year-long retention. [3] If you want to remember a word for a month, you should review it after about six days. If you want to remember it for a year, the optimal first review comes at roughly two to four weeks. This finding demolished the intuition that more frequent review is always better. There is a sweet spot, and overshooting it wastes time while undershooting it lets the curve win.
Paul Pimsleur recognized this decades before the data fully caught up. In 1967, he published "A Memory Schedule" proposing graduated intervals for vocabulary review: 5 seconds, 25 seconds, 2 minutes, 10 minutes, 1 hour, 5 hours, 1 day, 5 days, 25 days, 4 months, and 2 years. His insight was disarmingly simple: if you remind a student of a word before they have completely forgotten it, their chances of remembering will increase, and after each such recall it will take them longer and longer to forget again. Five years later, Sebastian Leitner, a German science journalist, turned this principle into something anyone could use with a stack of index cards. His system, published in 1972, sorted flashcards into boxes. A correct answer moved the card to the next box, which was reviewed less frequently. An incorrect answer sent it back to box one. No computer required. Research has shown that spaced repetition systems like Leitner's can improve long-term retention by up to 200 percent compared to cramming.
The computer changed everything. In December 1987, Piotr Wozniak, a Polish student of molecular biology, released the SM-2 algorithm, the first software-based spaced repetition scheduler. SM-2 tracks three properties for each flashcard: a repetition number counting how many times the card has been reviewed, an easiness factor starting at 2.5 that adjusts based on how difficult the card feels, and an inter-repetition interval measured in days. The first review comes after one day, the second after six days, and each subsequent interval is the previous interval multiplied by the easiness factor. If you rate a card as difficult, the easiness factor drops, compressing the schedule. If you rate it as easy, the factor rises, stretching the intervals out. If you fail a card entirely, it resets to day one, but its easiness factor is preserved so the system remembers it was always hard. This elegant feedback loop meant that for the first time, a learner's entire vocabulary could be managed by an algorithm that adapted to their individual performance.
SM-2's most important descendant is Anki, the open-source flashcard application created by Damien Elmes in 2006. Anki adopted a modified version of SM-2 and became the most widely used spaced repetition tool in the world. Among American medical students, 86.2 percent report using Anki, with 66.5 percent using it daily. The r/Anki subreddit for medical students has over 109,000 registered users, exceeding the approximately 89,000 active medical students in the entire United States. A 2025 cohort study of 130 first-year medical students found that Anki users scored significantly higher on every preclinical exam: 6.4 percent higher on Course I, 6.2 percent on Course II, 7.0 percent on Course III, and a striking 12.9 percent higher on the comprehensive board-style exam, with a p-value of 0.003. [4] Significant positive correlations were found between the number of matured cards and exam scores. The tool works, and the data is unambiguous.
In 2023, Anki took a significant step forward by integrating FSRS, the Free Spaced Repetition Scheduler, a next-generation algorithm based on the "Three Component Model of Memory." Where SM-2 relies on a single easiness factor and a deterministic interval formula, FSRS uses machine learning to analyze each user's complete review history and find the parameters that best predict their individual recall patterns. Early comparisons show that FSRS requires fewer reviews than SM-2 to achieve the same retention level, which means less time grinding flashcards for the same result. The evolution from Pimsleur's hand-crafted schedule to Leitner's cardboard boxes to SM-2's easiness factors to FSRS's neural networks traces a clear arc: each generation gets closer to the theoretical optimum described by the forgetting curve research.
Duolingo took a different path entirely. In 2016, Duolingo researchers Burr Settles and Ben Meeder published a model called half-life regression, or HLR, which combined psycholinguistic theory with machine learning trained on 13 million student learning traces. [5] Rather than tracking a single easiness factor per card, HLR estimates a "half-life" for each word-learner pair: the point at which the probability of recall drops to 50 percent. By 2020, Duolingo had scaled this into a system called Birdbrain, which processes data from over one billion exercises completed daily to predict when each learner will forget specific words. Birdbrain dynamically generates lessons that target items at a predicted recall probability, aiming for a sweet spot around 93 percent chance of getting the answer right. Compared to Duolingo's original Leitner-based system, HLR reduced prediction error by more than 45 percent and produced a 9.5 percent increase in daily retention for practice sessions and a 12 percent increase for overall activity. [5]
The meta-analytic evidence for spaced repetition in second language learning specifically is now overwhelming. Kim and Webb's 2022 meta-analysis examined 98 effect sizes from 48 experiments involving 3,411 participants and found that spaced practice had a medium-to-large effect on L2 learning, with Hedges' g of 0.76 on immediate posttests and 1.15 on delayed retention tests. [6] That delayed-test number is particularly important. It means the advantage of spacing actually grows over time, precisely because massed practice produces short-term performance that evaporates, while spaced practice produces durable retention that holds up weeks and months later. The study also found that shorter spacing was as effective as longer spacing on immediate tests but significantly less effective on delayed tests, confirming Cepeda's finding that the optimal gap depends on how long you need to remember.
Real classrooms, not just laboratories, confirm the effect. A 2025 meta-analysis by Mawson and Kang examined 31 effect sizes from 22 reports involving more than 3,000 students in authentic classroom settings and found a moderate effect favoring distributed over massed practice, with Cohen's d of 0.54 and a 95 percent confidence interval of 0.31 to 0.77. [7] Larger effect sizes were associated with longer retention intervals and fewer re-exposures, which aligns with the theoretical prediction that spacing becomes more valuable as the stakes shift from short-term performance to long-term retention. The fact that this holds up in messy, real-world classrooms, not just controlled laboratory conditions, is what makes the finding actionable for anyone learning a language on their own.
After ten days of consistent spaced-repetition practice, learners in one study were able to recall approximately 79.77 percent of designated target words. Compare that to Ebbinghaus's 21.1 percent after thirty-one days without review, and the magnitude of the intervention becomes clear. Spaced repetition does not eliminate forgetting. It manages forgetting, intercepting each word just before it slips below the recall threshold and pushing it back above the line. Over multiple cycles, the intervals between reviews grow longer and longer until the word is effectively permanent. The system is not magic. It is engineering applied to a biological process that we now understand with remarkable precision.
Yet spaced repetition has real limitations that its most enthusiastic advocates tend to downplay. The most significant was documented by Nakata and Elgort in 2021. In their study, Japanese learners of English encountered 48 novel vocabulary items in context. Spacing significantly improved explicit knowledge, meaning conscious recall of word meanings and form-meaning matching. But it had no effect whatsoever on tacit semantic knowledge, the kind of automatic, context-sensitive word processing measured by semantic priming tasks. [8] In other words, spaced repetition made learners better at knowing what a word means when asked directly, but it did not make them faster or more fluent at processing that word in natural context. Massed practice was equally effective for developing this implicit dimension of word knowledge.
This distinction between explicit and tacit knowledge points to a deeper structural limitation of most SRS tools. The standard flashcard paradigm presents a word on one side and a translation on the other, what researchers call the pair-associate paradigm. This addresses exactly one dimension of word knowledge: the form-meaning link. But Paul Nation identified at least nine dimensions of word knowledge, spanning form (spoken, written, and morphological), meaning (referential, associative, and conceptual), and use (grammatical functions, collocations, and register constraints). A flashcard that teaches you that "perseverar" means "to persevere" tells you nothing about what preposition it takes, what register it belongs to, what words commonly appear beside it, or how it sounds in a sentence. Most SRS tools test recognition rather than production, meaning learners may recognize a word when they see it but struggle to produce it in speech or writing. Spacing has been shown to be "more obvious for receptive than for productive vocabulary knowledge," a gap that matters enormously for anyone whose goal is actual communication.
This is where the case for combining spaced repetition with contextual learning becomes compelling. Paivio's Dual Coding Theory suggests that combining visual, auditory, and contextual input creates stronger memory traces than any single channel alone. Seeing a word on a flashcard, hearing it spoken by a character, and encountering it embedded in a story or game create multiple retrieval pathways that reinforce one another. A study using gamified spaced repetition in an interactive setting found that gradually increasing the level of processing facilitated the transfer of vocabulary from working memory to long-term memory. A 2023 systematic review of gamified learning tools found that 76 percent of studies measuring academic achievement reported positive outcomes, with particular strength in engagement and motivation, where 80 percent of studies found positive effects. Game-based language learning provides exactly what SRS lacks: meaningful context, communicative purpose, emotional engagement, and productive practice opportunities.
The complementary model is straightforward. Spaced repetition handles what it does best: scheduling optimal review timing, building reliable form-meaning links, ensuring breadth of coverage across a learner's entire vocabulary, and delivering the explicit knowledge that forms the foundation of word recognition. Contextual learning, whether through games, stories, conversations, or immersive media, handles what SRS cannot: providing meaningful encounters between reviews, developing collocational and grammatical knowledge, building depth rather than just breadth, requiring productive use rather than passive recognition, and cultivating the tacit processing speed that separates someone who knows a word from someone who can actually use it. Neither approach alone is sufficient. SRS without context produces learners who can pass vocabulary tests but freeze in conversation. Context without SRS produces learners who feel comfortable but plateau because they never systematically consolidate the words they encounter.
The trajectory of the field points clearly toward integration. Duolingo's Birdbrain already blends SRS scheduling with contextual exercises. Anki's FSRS uses machine learning to optimize the scheduling half of the equation. The next frontier is systems that embed algorithmically timed review into rich, narrative-driven contexts, so that the review interval is respected but the encounter itself carries meaning, emotion, and communicative purpose. The research is unambiguous on both sides: spaced repetition produces large, durable effects on vocabulary retention, and contextual learning develops the deeper dimensions of word knowledge that flashcards cannot reach. The question is no longer which approach works. The question is how to weave them together so that every review is also an experience, and every experience is also a review.
Hermann Ebbinghaus sat alone with his nonsense syllables and discovered a curve that has withstood 140 years of scrutiny. The curve has not changed. Human memory still leaks at the same rate it did in 1885. What has changed is our ability to fight back. From Pimsleur's hand-written schedule to Leitner's cardboard boxes to Wozniak's SM-2 algorithm to Duolingo's half-life regression trained on billions of data points, each generation of tools has gotten more precise at intercepting forgetting at exactly the right moment. The science says space your reviews, and it says so with 839 assessments, 317 experiments, and 3,411 second-language learners backing it up. [2] [6] But the science also says that a word truly learned is a word that lives in context, not just in a flashcard box. The most effective vocabulary strategy is not SRS or contextual learning. It is SRS and contextual learning, working in concert, each covering the other's blind spots. The forgetting curve is steep, but it is not destiny. It is a problem with a known solution, and the solution keeps getting better.
Citations
- 1.Replication and Analysis of Ebbinghaus' Forgetting CurveMurre & Dros, PLOS ONE, 2015
- 2.Distributed Practice in Verbal Recall Tasks: A Review and Quantitative SynthesisCepeda, Pashler, Vul, Wixted & Rohrer, Psychological Bulletin, 2006
- 3.Spacing Effects in Learning: A Temporal Ridgeline of Optimal RetentionCepeda, Vul, Rohrer, Wixted & Pashler, Psychological Science, 2008
- 4.Exploring the Impact of Spaced Repetition Through Anki Usage on Preclinical Exam PerformanceWinter et al., Journal of Medical Education and Curricular Development, 2025
- 5.A Trainable Spaced Repetition Model for Language LearningSettles & Meeder, Proceedings of the 54th Annual Meeting of the ACL, 2016
- 6.The Effects of Spaced Practice on Second Language Learning: A Meta-AnalysisKim & Webb, Language Learning, 2022
- 7.The Distributed Practice Effect on Classroom Learning: A Meta-Analytic Review of Applied ResearchMawson & Kang, Behavioral Sciences, 2025
- 8.Effects of Spacing on Contextual Vocabulary Learning: Spacing Facilitates the Acquisition of Explicit, but Not Tacit, Vocabulary KnowledgeNakata & Elgort, Second Language Research, 2021