Hello there, subtitle sleuths, menu explorers, and last-language holdouts…
A game can keep receiving updates while one of its language versions becomes harder to fund. The players are still there. So are the new quests, item descriptions, and instructions that someone needs to make understandable. AI game localization offers a possible way to keep serving that audience. The uncomfortable part is deciding what counts as serving them.
During the audience Q&A in Gridly’s Beyond the Buzzwords webinar, localization consultant Tamara Tirják described a practice she has seen in mobile games: when some markets underperform, teams may replace human translation with AI text instead of removing those languages from the game.
That is a more interesting proposition than another promise to generate words faster. An existing audience could keep access to a game it already enjoys. It also raises a question that belongs to players as much as production managers: what experience survives the change?
Tirják did not identify the games or provide before-and-after quality results. Her account supports examining this choice, without proving either a successful rescue or a quality downgrade. The case for AI should survive that uncertainty. Useful, reviewed language support can be worth preserving even when the production method changes.
Keeping the language means keeping an audience
The webinar brought together Tirják, AI and machine-translation implementation consultant Cristina Anselmi, and Gridly’s Michael Souto , with Gridly CEO Anna Albinsson moderating. Their practical discussion repeatedly returned to the work surrounding a model: context, evaluation, defined responsibilities, and realistic budgets.
Souto argued that savings could fund more language coverage and more time for specialists to do creative work. Anselmi described testing a smaller translated sample, having humans review it, gathering evidence, and then deciding whether a larger investment makes sense. Tirják’s observation adds the maintenance problem: a studio may already have an audience whose next update needs translating.
The reasonable objection to demanding identical production methods everywhere is money. A team with finite resources may have to choose among reduced scope, a different workflow, and withdrawing support. A reviewed AI translation could be the best available option. Players gain nothing from an argument about ideal production conditions if its practical outcome is a language they can no longer use.
Even Valve’s localization guidance tells developers to prioritize languages using traffic, sales, and community feedback. It also warns against equating language with geography. A language audience can live across multiple markets; one country’s revenue figure cannot describe all of it.
🦊 Kiki:
Keep the language if you can make it work. I want every version to have brilliant dialogue, lavish voice acting, and someone personally checking the joke I liked. My wish list requires its own finance department. Unfortunately, the studio has an actual budget, and the player waiting for the next quest cannot spend my excellent intentions. So test the cheaper route properly and let the result earn its place. If reviewed AI text keeps that person playing, I will take the useful outcome. I can continue campaigning for the lavish voice acting after I locate the money I was so confidently spending.
🍪 Chip props a fallen language sign upright and wedges a small stone beneath its base.
Good enough for which part of the game
Tirják’s answer to the question of quality was about purpose. Understanding player feedback for internal analysis asks something different of a translation from delivering spoken dialogue with heavy characterization. A rough internal summary can be useful even when nobody would want those sentences in the final game.
Players encounter those distinctions through ordinary actions. Can they understand what an ability does? Does the quest tell them where to go? Does a conversation preserve who is speaking and what the relationship means? Can they read the text before the subtitle disappears?
These are practical questions for evaluating a version, not evidence that AI necessarily fails them. Human translation also needs context and testing. The proposed change should be judged against the experience the game requires, with the same willingness to catch errors from either method.
Steam already separates interface, subtitles, and full audio in its in-game language declarations. It also distinguishes a translated store page from a translated game. That gives a useful starting point for a player-facing explanation of what support actually includes, without promising identical dubbing budgets in every language.
The quality review also has to reach the running interface. Unity’s pseudolocalization guidance describes testing with artificial text to reveal problems such as insufficient space, unsupported characters, and text that never enters the localization system. Its menu example shows a Spanish translation growing beyond the space available.
Pseudolocalization helps find technical risks before real translation. It cannot tell you whether the villain’s threat sounds convincing. Linguistic judgment and interface testing answer different questions, and a playable version needs both.
🦊 Kiki:
I will forgive an awkward sentence before I forgive an instruction that gets me killed. That is my extremely sophisticated evaluation framework, developed through years of possessing a health bar. If the cheaper version preserves the quest, the characters, and the buttons, show it to me. I might like it. If the repair plan is to wait until players explain why the door instruction sends them into lava, somebody has quietly assigned us a job. In my imaginary production meeting, I would ask where to submit my invoice. Apparently the compensation package is an apology post and another attempt at the checkpoint.
🍪 Chip follows a crooked arrow, stops at the edge of a puddle, and rotates the sign toward dry ground.
Context has to survive the next update
Anselmi pushed back on treating the generic model itself as the entire problem. She emphasized how prompts are managed, how context and translation memory are supplied, and where checks happen before and after generation. A translation memory stores earlier translations that can help keep later work consistent.
For a maintained language, that history is especially valuable. Established item names, character relationships, and interface terms should travel into the next batch. Otherwise, a fluent new translation may make an old mechanic look like a different feature.
Consider a hypothetical update that renames an ability. The team needs to know which tutorial, inventory entry, and quest reference depend on that name. Generating a new sentence does not tell the team whether every affected place was updated. That relationship has to be represented in the content system and checked in the game.
Gridly’s discussion of why AI translation quality breaks down at scale describes the operational cost of fragmented context and manual preparation. Its companion explanation of agentic localization focuses on systems that select context and carry project state across tasks. Those are relevant approaches to this maintenance problem, provided teams test how they behave on their own content.
The Gridly Localization Agent is one example of that approach. Gridly says it can assemble project context, create reviewable instructions, and let teams validate a sample before processing a batch. As of September 7, its public page advertises a fall 2026 launch with early demos available. A studio evaluating it still needs evidence from its own game and target languages.
Tirják also urged teams to plan for failures and recheck the underlying model as it changes. A useful approval last month cannot answer whether this month’s configuration still handles the game’s terminology. Someone needs the authority and time to stop a bad batch, correct its cause, and retest it.
A quality score needs a definition
The webinar’s measurement discussion was more useful than a universal pass mark. Anselmi recommended combining automatic measures with standardized human evaluation. Souto pointed to previous issue data as a comparison. Teams need to know which terminology or style requirements they are testing before announcing that a system followed them.
The Multidimensional Quality Metrics framework, or MQM, offers one way to organize that judgment. Its guidance distinguishes minor defects from errors that seriously damage understanding or usability. Acceptance thresholds depend on the requirements and content; the scorecard it describes rejects content containing a critical error regardless of the aggregate score. MQM is a translation-evaluation framework, rather than a certification that a game is ready to ship.
⭐ Byte: MQM’s guidance explains that its scoring values can be scaled to resemble percentages, commonly using a maximum value of 100. The resulting score reflects selected error penalties and normalization. Reading it as the percentage of correctly translated words would misstate what it measures. For a comparison between workflows, keep the evaluation rules and content coverage comparable, report serious errors separately, and record human review time. The webinar provides no measured language-retention result from which to calculate a studio’s savings.
A practical trial could compare an AI-assisted version of a representative update with the existing workflow. Include difficult dialogue and gameplay instructions alongside routine text. Have qualified reviewers assess it using agreed criteria, check it inside the build, and record the effort required to correct it. That is a proposed test for the studio’s decision, with passing conditions chosen before the results arrive.
The bill continues after the first batch
Anselmi was blunt about return on investment: implementation requires investment first, including expertise, experimentation, and training. Tirják pointed to late content, repeated changes, and weak source text as problems AI cannot simply compensate for. An inexpensive model call can still sit inside an expensive production process.
For the language-support decision, compare the cost of maintaining a usable version over actual updates. Include preparing context, reviewing output, checking the build, and correcting new failures. A one-off demonstration cannot show whether the same team can keep delivering that work throughout the game’s update cycle.
Then describe the benefit honestly. Keeping an existing language available is valuable, even if it never becomes the game’s biggest market. Faster delivery may also help. Neither benefit automatically proves that AI increased revenue, prevented players from leaving, or paid for itself. Those outcomes need their own evidence.
🦊 Kiki:
I love a bargain enough to be dangerous around a sale banner. That is why I want the boring invoice, including setup, review, and fixes. A workflow can absolutely pay for itself. It can also deliver a very affordable first draft followed by an imaginary meeting titled Why Is Everyone Still Here At Midnight. Small teams deserve the first result, and they deserve an honest trial that catches the second. Please let the localization team finish counting before somebody orders the celebratory cake. I have bought enough discounted nonsense to recognize when the accessories cost more than the thing.
🍪 Chip places a tiny repair kit beside the budget sheet and watches the paper curl around it.
Keep support alive beyond the announcement
The strongest case for AI here is modest enough to test: maintain a useful language version that would otherwise be difficult to sustain. It deserves a serious trial, without requiring a claim that every language can be served cheaply or that every previous task should become automated.
For Game Cookies, the conditions are concrete. Set the expected experience before changing the workflow. Preserve the context needed to deliver it. Check meaning and usability before release. Keep responsibility for corrections after the update. Those conditions leave room for different tools and budgets while protecting the point of localization: people being able to play.
A team may discover that AI handles routine material well while important scenes still require substantial human work. It may also discover that the proposed process costs too much once reviewed. Both results are useful. They let the studio decide what it can support based on actual work instead of a demonstration’s first impression.
🦊 Kiki:
Give the tool a fair trial, then keep the people who can tell you when it is failing. I want the language menu to get longer, and I want existing entries to mean something after the next update. My loyalty belongs to the person trying to understand the game, whether a human, a model, or a carefully managed combination produced the words. If AI helps a small team keep that promise, fantastic. I will bring the cookies. If management cancels the checking and calls the remaining text “supported,” I will need those cookies back. The repair crew has earned lunch.
🍪 Chip carries the repair kit through the reopened route and leaves the language sign standing behind him.
🍪 Missed part of the gaming week?
The Sunday Cookie Box delivers five handpicked Game Cookies stories in one free Sunday email.
⚙️ Stay curious about the languages your favorite games can keep supporting.
⚙️ Keep asking whether the next update has been checked where players will actually read it.
⚙️ And remember: a language menu earns its place one playable update at a time.
🦊 Kiki · 🍪 Chip · ⭐ Byte · 🦁 Leo
Tips, leaks, and suspicious language menus: contact us here!


