Why Most Language Apps Fail (and What the Research Says Works)
The global language learning market was valued at $85.1 billion in 2025, with 1.5 billion people actively studying a second language and two-thirds of them using digital platforms. [1] Duolingo alone reported 103 million monthly active users and $748 million in revenue in 2024. Babbel, Busuu, Rosetta Stone, and dozens of smaller competitors serve tens of millions more. By any measure of access and adoption, language learning has never been more democratized. And yet, by the most commonly cited estimate, only about 1 percent of language learners ever reach fluency. That gap between adoption and outcome is not a marketing problem or a willpower problem. It is a design problem, and the research makes the diagnosis surprisingly clear.
Before sharpening the critique, it is worth stating plainly what language apps get right. They are inexpensive or free, available on every phone on earth, and require no scheduling, commute, or social anxiety to begin using. They have brought spaced repetition, one of the most empirically validated learning techniques in cognitive science, to hundreds of millions of people who would never have encountered an Anki deck. Students using spaced-repetition software score 6.2 to 10.7 percent higher on standardized exams compared to those using traditional study methods. [2] Apps have also radically lowered the barrier to entry. A teenager in rural Brazil can start learning Mandarin at midnight for zero cost. That achievement is real and should not be dismissed. The problem is not that apps are useless. The problem is that they stop working far sooner than their users expect.
The best available research suggests that apps deliver genuine results at the beginner level. The Vesselinov and Grego study, funded by Duolingo and published in 2012, found that beginner Spanish learners needed an average of 34 hours of Duolingo use to cover the equivalent of one university semester, gaining 8.1 points per hour on the WebCAPE placement test. [3] A 2021 study by Jiang and colleagues in Foreign Language Annals found that learners who completed beginning-level Duolingo courses reached reading proficiency comparable to four semesters of university study, in roughly half the time. [4] At Michigan State University, Loewen and colleagues found that 59 percent of participants who used Babbel for 12 weeks improved their oral proficiency by at least one ACTFL sublevel, rising to 75 percent among those who studied 15 or more hours. [5] These are meaningful results. They are also, without exception, results that describe beginner-level gains in predominantly receptive skills.
The Jiang study is particularly revealing for what it did and did not measure. Learners reached Intermediate Low in reading and Novice High in listening, both strong showings. But speaking and writing were not assessed at all. [4] The Loewen study, one of the few to measure speaking outcomes, found only modest gains of one sublevel, exclusively at beginner levels. [5] The Lord study at the University of Florida compared Rosetta Stone learners to classroom learners in conversation tasks and found that the app-only group routinely gave up and reverted to English, while classroom learners used more Spanish and actively negotiated meaning through clarification requests and comprehension checks. [5] As Shawn Loewen of Michigan State put it, "Despite the fact that millions globally are already using language learning apps, there is a lack of published research on their impact on speaking skills." The evidence base for apps is thinner than most users assume, and it almost entirely describes the easiest part of the journey.
The dropout numbers tell a story that the efficacy studies cannot. Despite Duolingo's industry-leading next-day retention rate of 55 percent, roughly half of all users churn by Day 7. Learning apps as a category have one of the lowest long-term retention rates of any mobile app vertical, at just 1.76 percent. [1] Of Duolingo's 103 million monthly active users, only 9.5 million are paid subscribers, meaning roughly 91 percent use the app casually enough that they see no reason to pay. These are not the numbers of a product that is transforming its users' abilities. They are the numbers of a product that is very good at getting people to start and very poor at getting them to continue. The question is why.
One popular answer is that people simply lack discipline, but the research suggests a more structural explanation. Duolingo's gamification improvements in early 2022 reduced its churn rate from 47 percent to 37 percent, a genuine achievement in product design. [1] Streaks, XP, leaderboards, and hearts create powerful short-term engagement loops. But a 2024 meta-analysis in Frontiers in Psychology found that while gamification can promote intrinsic motivation initially, extrinsic reward systems may "create a stage of boredom that could limit learning the target language." [6] Students showed positive attitudes toward gamified apps at basic levels but not at advanced levels. Shallow gamification, defined as "simplistic and superficial application of game elements without transforming the core experience," was identified as a determinant of failure. The implication is uncomfortable for the industry: the engagement mechanics that drive Day 1 retention may actively undermine Year 1 progress.
The theoretical framework for this problem has been established for over fifty years. In 1971, Edward Deci demonstrated that subjects who received monetary rewards for solving puzzles spent significantly less time on the task once rewards were removed. This overjustification effect, formalized within Self-Determination Theory, predicts that expected external incentives decrease intrinsic motivation, and once those incentives stop, the original interest does not return. Applied to language apps, the prediction is that users who maintain long streaks driven by XP and leaderboard mechanics may not be developing genuine intrinsic motivation to learn the language at all. They are playing a meta-game of streak maintenance. When the novelty of that meta-game fades, and novelty always fades, there is nothing underneath to sustain continued effort. Self-Determination Theory identifies three innate psychological needs for sustained motivation: autonomy, competence, and relatedness. Most language apps satisfy competence at the beginner level through easy wins, but they struggle with autonomy, since drill sequences are rigidly linear, and they largely fail on relatedness, since the learner is alone with a screen.
The most consequential failure, though, is not motivational. It is structural. Second language acquisition research has spent decades identifying what drives real proficiency, and drill-based apps systematically omit most of it. Stephen Krashen's comprehensible input hypothesis holds that language is acquired when learners understand messages slightly above their current level. [7] This is perhaps the best-known idea in applied linguistics, and apps nominally implement it through leveled content. But Krashen himself emphasized that the input must be meaningful and delivered in low-anxiety situations "containing messages that students really want to hear." Isolated sentence translation drills, the backbone of most app curricula, rarely meet that standard. The sentences are decontextualized, the topics are arbitrary, and there is no communicative purpose. You are not trying to understand something you care about. You are trying to get the answer right.
Comprehensible input is necessary but not sufficient. Michael Long's Interaction Hypothesis, first proposed in 1981 and refined in 1996, argues that language proficiency is promoted by face-to-face interaction, specifically through the negotiations for meaning that occur when communication breaks down: confirmation checks, clarification requests, comprehension checks, and repair sequences. [7] These breakdowns are not obstacles to learning. They are the mechanism of learning. When a conversation partner says "Wait, do you mean X or Y?" the learner receives precisely targeted feedback on the gap between their intended meaning and their current ability. Most language apps provide zero negotiation of meaning. There is no interlocutor, no breakdown, no repair, and no genuine communicative pressure. The learner taps a correct answer from four choices and moves on, having engaged in recognition rather than the meaning-making that drives acquisition.
Merrill Swain's Output Hypothesis addresses the other side of the equation. Observing French immersion students in Canada who had received years of comprehensible input and developed strong listening and reading skills, Swain found that their productive abilities, speaking and writing, remained weak. Her conclusion was that learners need to produce output, not just receive input, because production serves three cognitive functions: noticing gaps between intended meaning and current ability, testing hypotheses about how the language works, and engaging in metalinguistic reflection. [7] Drill-and-tap exercises in apps require recognition, which is a fundamentally different cognitive operation from generation. Selecting the correct translation from four options is not the same as constructing a sentence from scratch under communicative pressure. The gap between these two tasks is the gap between knowing a word exists and being able to use it when you need it.
The depth of processing framework, established by Craik and Lockhart in 1972, explains why this matters at the level of memory formation. Deeper levels of processing, semantic analysis, contextual integration, and personal relevance, produce stronger and longer-lasting memory traces than shallow processing like phonetic repetition or visual pattern matching. [7] Matching a word to its translation is shallow processing. Encountering that word in a narrative context where you need it to understand what happens next, or using it to solve a problem that matters to you, is deep processing. Research on emotional enhancement of memory adds a further dimension: emotionally arousing experiences are encoded more strongly, with emotional moments during narrative perception increasing cohesion across functional brain modules including the hippocampus and amygdala and predicting better recall fidelity. [8] Repetitive drills are emotionally flat by design. They do not create the stakes, surprise, or personal investment that drive durable memory encoding.
All of these factors converge on what researchers call the intermediate plateau, and it is where the app model breaks down most visibly. The CEFR B1-B2 transition is where most app users stall, and the proportion of users decreases sharply as language levels go up. More than half of language app users consider themselves beginners. Duolingo's own research caps its strongest proficiency claims at Intermediate Mid on the ACTFL scale, roughly CEFR B1. Beyond that, there is no published evidence of app-only learners progressing to advanced proficiency. One theory for the plateau is that at equilibrium, learners forget infrequently-used words as quickly as they learn new ones, and improvement simply stops. The structural explanation is more damning: at intermediate levels, learners need connected discourse competence, repair strategies, spontaneous output under pressure, and deep contextual processing, and none of these are things that drill-based apps are designed to provide.
The research on game-based language learning offers a way through this impasse that is more than theoretical. A 2025 meta-analysis of mobile language-learning games published in Computer Assisted Language Learning found an overall large effect size of Hedges' g = 0.962, substantially outperforming non-game conditions. [8] This is not a marginal improvement. An effect size approaching 1.0 means that the average learner in a game-based condition outperformed roughly 83 percent of learners in the control condition. The key distinction here is between gamification, bolting points and badges onto existing drill formats, and genuine game-based learning, where the game mechanics and the language learning are inseparable. In an immersive game, vocabulary is not a list to be memorized. It is a tool you need to navigate a world, talk to characters, and solve problems. The comprehensible input is embedded in visual context. The output pressure comes from genuine communicative need. The emotional engagement comes from narrative stakes, not streak anxiety.
Consider what an immersive language learning game provides against the checklist of what SLA research says works. Comprehensible input is delivered in rich visual and narrative context, making it meaningful rather than arbitrary. Interaction and meaning negotiation occur through dialog with characters who respond to the learner's choices. Output is pushed by gameplay that requires the learner to produce language, not just recognize it. Depth of processing is guaranteed because the language serves a purpose the learner cares about, advancing a story, completing a quest, building something. Emotional engagement is built into the medium: games are narrative experiences with characters, conflict, and resolution. Spaced repetition is embedded in the progression system, with vocabulary recurring across varied meaningful contexts rather than in isolated flashcard reviews. A 2020 study in Nature found that encountering words across multiple meaningful contexts improves recall more than simply increasing the number of exposures, which maps directly to vocabulary encountered across different game scenarios rather than repeated on the same flashcard. [2] This is not gamification bolted onto flashcards. It is language acquisition through genuine play.
None of this means that apps have failed. They have succeeded at something important: making language learning accessible, affordable, and low-stakes enough that hundreds of millions of people have tried it. Spaced repetition works. Low-anxiety environments work. The ability to practice for five minutes on a bus works. What the research clearly shows is that these strengths are necessary but insufficient, and that the drill-based, gamification-driven model has a ceiling that most learners hit well before they can hold a conversation, read a novel, or function in a professional setting in their target language. The critique is not that apps are useless. It is that they are a strong foundation that was never designed to be the whole building.
The future of language learning is not more streaks. It is not more XP, more badges, or more leaderboard anxiety. The research points consistently toward the same cluster of principles: meaningful context, comprehensible input, pushed output, emotional engagement, social interaction, and spaced repetition embedded in varied contexts. The question for the next generation of language learning tools is whether they can deliver all of these at the scale and accessibility that apps pioneered. Immersive game-based learning is the most promising synthesis because it does not require learners to choose between engagement and effectiveness. The game is the lesson. The story is the input. The quest is the output pressure. The characters are the social context. And the world is the spaced repetition system, surfacing vocabulary not because an algorithm says it is time, but because the narrative demands it. The $85 billion question is not whether people want to learn languages. They clearly do. The question is whether we can build tools that match the depth of what research says they need.
Citations
- 1.Language Learning Market Size & Share Report, 2025-2034GM Insights, 2025; Growth.Design Case Study; Business of Apps, 2026
- 2.
- 3.An Effectiveness Study of Duolingo for Learning SpanishVesselinov & Grego, 2012
- 4.How Well Does Duolingo Teach? Comparing Learner Proficiencies Across Languages and University CoursesJiang, Rollinson, Plonsky, Gustafson & Pajak, Foreign Language Annals, 2021
- 5.Emerging Technologies: Mobile-Assisted Language Learning Apps and Their Impact on L2 SpeakingLoewen et al., Foreign Language Annals, 2020
- 6.Gamification, Motivation, and Language Learning Outcomes: A Systematic ReviewFrontiers in Psychology, 2024
- 7.
- 8.The Effectiveness of Mobile Language-Learning Games: A Meta-AnalysisComputer Assisted Language Learning, Taylor & Francis, 2025