Can AI Replace Cultural Fluency?
In 2017, a Palestinian man in the West Bank posted a photo of himself on Facebook leaning against a bulldozer at a construction site. His caption, written in Arabic, said "Good morning." Facebook's automatic translation algorithm converted this into Hebrew as "Attack them" and into English as "Hurt them." Israeli police, acting on the translated post, arrested the man and held him for hours before the error was discovered. A Facebook spokesperson later acknowledged the failure: "Unfortunately, our translation systems made an error last week that misinterpreted what this individual posted." The incident was treated as an embarrassing glitch, a one-off malfunction in an otherwise impressive system. But the error was not random. It was structural. The algorithm processed characters and predicted word sequences. It did not understand context, intent, or the social meaning of a man smiling next to heavy machinery on his way to work. It translated words. It did not understand what was said.
To be fair, AI translation has made genuine and remarkable progress. Google Translate supports over 240 languages and processes billions of words daily. For common language pairs like English-Spanish, accuracy rates can reach 94 percent on straightforward text. Real-time translation earbuds now cost less than a pair of good headphones. For travelers ordering meals, tourists reading museum placards, or aid workers distributing supplies, these tools are legitimately useful and sometimes lifesaving. Neural machine translation has moved from a novelty to a utility in less than a decade, and the pace of improvement has been staggering. Dismissing this progress would be both inaccurate and ungrateful.
But the accuracy numbers, impressive as they are, conceal a deeper problem. A 2025 study by Anik and colleagues found that up to 47 percent of contextual meaning is lost in machine translation, even when the individual words are rendered correctly. Google Translate's average accuracy across language pairs sits around 82.5 percent, a number that sounds serviceable until you consider what lives in the remaining 17.5 percent. That gap is not filled with random errors. It is filled with tone, register, social hierarchy, implied meaning, humor, and the accumulated cultural knowledge that separates a sentence from a communication. When Appen tested three major large language models across 20 languages and 24 dialects, not a single translation was ready to publish without human editing. [7] Every one required intervention. The machines are getting better at words. They are not getting better at meaning.
Consider Japanese, where the problem is not vocabulary but social architecture. The keigo honorific system encodes three distinct levels of politeness: teineigo for general courtesy, sonkeigo for elevating the person you are addressing, and kenjougo for humbling yourself. The verb "to say" alone has at least four different forms depending on the social relationship between speaker and listener: iu in casual speech, iimasu in polite form, ossharu when honoring the other person's speech, and mousu when deferring about your own. A 2022 study at the University of Colorado found that Google Translate produces inconsistent formality levels for identical inputs, sometimes rendering a sentence in casual register and sometimes in formal register with no discernible pattern. [4] In Japanese business culture, using the wrong level of keigo is not a grammatical mistake. It is a social offense, a signal that you either do not understand or do not respect the relationship. An algorithm that randomly toggles between formality levels is not just inaccurate. It is dangerous.
The problem deepens with Chinese, where entire social operating systems resist translation altogether. Mianzi, often rendered in English as "face," is not a word but a universe. It encompasses social standing, personal dignity, family honor, and the invisible ledger of respect that governs every professional and personal interaction. Face can be given through public praise. It can be lost through a careless remark in the wrong setting. It can be fought for across generations. The related concept of guanxi, usually translated as "relationships" or "connections," describes a web of mutual obligation, reciprocal favors, and social trust that structures Chinese business in ways that have no Western equivalent. When Khoong and colleagues studied Google Translate's performance on Chinese medical instructions in a 2019 study published in JAMA Internal Medicine, they found that while 81 percent of translations were accurate, 8 percent contained errors that were clinically harmful. [1] Eight percent. In a medical context, that means roughly one in twelve translations could hurt a patient. The words were translated. The meaning, the cultural context in which those words function, was not.
Korean presents a variation of the same challenge, with seven distinct speech levels that encode not just politeness but the speaker's precise social relationship to the listener, their relative ages, their professional hierarchy, and the formality of the setting. [5] Choosing the wrong speech level in Korean is not like using "you" instead of "thou" in English, a quaint anachronism. It is a serious social breach that can damage relationships and signal disrespect. AI translation systems, trained on text corpora that flatten these distinctions into a single register, consistently fail to preserve them. Spanish offers a parallel example: the distinction between tuteo, voseo, and ustedeo, the three systems for addressing "you," carries social information that AI compresses into nothing. Argentina's vos and Mexico's tu signal different cultural identities, different class markers, different degrees of intimacy. An algorithm that outputs one where the other is expected has not made a translation error. It has made a cultural one.
Then there are the words that refuse to cross linguistic borders at all. The Portuguese saudade describes a longing for something absent, a bittersweet ache for a person, place, or time that may never return, and it carries within it an entire national temperament. The Danish hygge, often mangled into English as "coziness," actually encodes a philosophy of intimate togetherness, candlelight, and deliberate simplicity that resists reduction to a single English term. The Japanese wabi-sabi names the beauty found in imperfection and impermanence, a cracked teacup more moving than a perfect one. The Spanish duende, as Federico Garcia Lorca described it, is the dark, earthy power of art that moves you beyond reason. The Korean han holds within it a profound blend of sorrow, regret, anger, and longing born from collective historical trauma, a word that carries centuries of national suffering. The Chinese yuan fen describes the fateful coincidence that brings two people together, a concept poised between destiny and luck. Each of these words is a compressed worldview, a piece of cultural philosophy disguised as vocabulary. No algorithm translates them because they are not translation problems. They are understanding problems.
The cognitive linguist Lera Boroditsky, working at UC San Diego, has spent her career demonstrating that language does not merely label the world but actively shapes how speakers perceive it. "The beauty of linguistic diversity is that it reveals to us just how ingenious and how flexible the human mind is," she has said. "Human minds have invented not one cognitive universe, but 7,000." Each of those 7,000 languages, she argues, constitutes its own cognitive toolkit, a distinct way of categorizing time, space, color, causality, and social relationships. Speakers of Kuuk Thaayorre, an Aboriginal language in Australia, use cardinal directions instead of relative ones, saying "move your cup to the north-northeast a little" rather than "move it to the left." The result is that Kuuk Thaayorre speakers maintain an extraordinarily precise sense of orientation that English speakers simply do not develop. Translation between these cognitive systems is not word substitution. It is an act of rebuilding meaning from different raw materials.
The cultural bias embedded in AI systems makes the problem worse, not better. A 2024 study published in PNAS Nexus tested large language models across 107 countries and found that every version of GPT reflected the values of English-speaking and Protestant European countries. [2] The bias was not a bug in one model or one training run. It was consistent across all versions tested, a structural feature of systems trained predominantly on English-language internet text. Cultural prompting, explicitly instructing the model to consider a specific cultural perspective, reduced the bias for 71 to 81 percent of countries tested, but the fact that a model needs to be told to consider non-Western values reveals how deeply the default assumptions run. When an AI translates between languages, it is not operating from a neutral position. It is filtering every culture through the lens of one.
The linguist Jenny Thomas coined the term "pragmatic failure" to describe "the inability to understand what is meant by what is said." The distinction is crucial. What is said is the surface level of language, the words and grammar that AI can increasingly handle. What is meant is everything underneath: the implication, the social context, the shared cultural knowledge that allows a speaker to say one thing and communicate another. When a Japanese colleague responds to your proposal with "That would be very difficult," they are not describing a logistical challenge. They are saying no. When a British friend says "That's quite good," the word "quite" is doing the opposite of what an American ear expects: it is dampening the praise, not amplifying it. A 2025 study published in Frontiers in Education found that large language models systematically struggle with exactly this kind of pragmatic competence, the ability to interpret indirect speech acts, understand implicature, and navigate the gap between literal meaning and intended meaning. [6] AI processes what is said. Cultural fluency is the ability to understand what is meant.
Humor offers a particularly revealing test of this gap. A 2025 study published in the MDPI journal Digital compared humor retention rates across different translation systems and found that even the best-performing large language model, GPT-Ex, preserved humor in only 62.94 percent of cases, while traditional neural machine translation managed just 50.12 percent. [3] A pun that works in one language almost never works in another, because puns depend on the specific sound patterns and double meanings of individual words. But humor is not just wordplay. It is social intelligence made audible. A joke calibrates the relationship between speaker and listener, tests shared assumptions, and navigates sensitive topics with indirection. When AI loses the joke, it is not just dropping a punchline. It is losing the social information the joke was designed to carry. An 8 percent accuracy drop on medical terminology can harm a patient. A 37 percent failure rate on humor can cripple a relationship.
The real-world consequences of these failures extend well beyond awkward dinner conversations. In a Norwegian incident that became legendary in translation circles, the Norwegian Olympic team used Google Translate to order groceries from a Korean supplier and accidentally requested 15,000 eggs instead of the 1,500 they needed, a tenfold error produced by a number-classifier system that AI handled incorrectly. In 2009, when Hillary Clinton presented Russian Foreign Minister Sergei Lavrov with a symbolic "reset button" to signal a new beginning in U.S.-Russia relations, the Russian word on the button did not mean "reset." It meant "overcharged." The gaffe, caused by a translation error, became an international news story and a metaphor for the very diplomatic clumsiness it was meant to repair. In medical contexts, the stakes are highest: when the phrase "your child is fitting," describing a seizure, was translated into Swahili as "your child is dead," the consequences were not symbolic. They were devastating.
What cultural fluency actually requires cannot be scraped from a text corpus, no matter how large. It requires years of immersion, the slow accumulation of pattern recognition that allows you to read a room in another culture the way you read one in your own. It requires the embodied knowledge of knowing when to bow and how deeply, when to pour tea for someone else before filling your own cup, when silence is comfortable and when it is pointed. It requires the social intelligence to know that in some cultures, arriving exactly on time signals eagerness while in others it signals disrespect for the host who is still preparing. A 2024 study published in the ACL Anthology examined how the politeness of prompts affects large language model performance across English, Chinese, and Japanese, and found that the optimal level of politeness differed by language — impolite prompts degraded output quality, but the threshold varied culturally in ways the models could not anticipate. [8] The finding confirms what anyone who has lived abroad already knows: cultural competence is not a database lookup. It is a way of being in the world that develops through sustained, attentive participation in another culture's daily life.
Michele Hutchison, a translator who has won and been shortlisted for the International Booker Prize, captures the distinction precisely: "A translator translates more than just words; we build bridges between cultures, taking into account the target readership every step of the way." That metaphor is the right one to end with. AI can carry the bricks. It can move words from one language to another with increasing speed and decreasing cost. For routine tasks, for getting the gist, for surviving as a tourist, the bricks are enough. But building a bridge requires understanding both sides of the river: knowing what each culture values, how each culture communicates, what each culture leaves unsaid. That understanding is not a technical problem awaiting a technical solution. It is a human capacity that develops through the patient, difficult, irreplaceable work of learning to think in another language and live in another culture. The 27.2 percent of idioms that AI simply omits entirely, the 47 percent of contextual meaning that vanishes in translation, the seven speech levels of Korean and the three registers of Japanese honorifics and the untranslatable saudade of Portuguese longing — these are not flaws to be engineered away. They are evidence that language is more than information transfer. Language is culture made audible, and fluency is the only way to hear it.
Citations
- 1.Assessing the Use of Google Translate for Emergency Department Discharge InstructionsKhoong et al., JAMA Internal Medicine, 2019
- 2.Cultural Bias and Cultural Alignment of Large Language ModelsTao, Kizilcec et al., PNAS Nexus, 2024
- 3.Jokes or Gibberish? Humor Retention in Translation with NMT vs. LLMMDPI Digital, 2025
- 4.Honorific Language in Japanese and Its Effects on Translator SystemsUniversity of Colorado, 2022
- 5.
- 6.How Inclusive Large Language Models Can Be? The Curious Case of PragmaticsFrontiers in Education, 2025
- 7.
- 8.Should We Respect LLMs? A Cross-Lingual Study on Prompt Politeness and LLM PerformanceYin et al., ACL Anthology, 2024