Beyond COMET: Rethinking MT Quality Evaluation in Game Localization | Allcorrect Games

Beyond COMET: Rethinking MT Quality Evaluation in Game Localization

Valentine Pronin

Beyond COMET: Rethinking MT Quality Evaluation in Game Localization

What Actually Counts as a “Good” Translation?

Hi everyone, I’m Valentin Pronin. I started my localization career as an editor, then became Head of Linguists, and now I’m the R&D Lead at Allcorrect. That journey means I’ve seen translation quality from basically every angle—from fixing typos by hand to building the automated pipelines we use today.

Allcorrect has been on the market for 20 years. We were pioneers in multilingual localization, and for the last 18 years we’ve focused on video games. Every single month, we process more than 2000 orders from 200 different clients, working on everything from indie Steam titles and mobile hits all the way up to the massive releases everyone is waiting for.

Today I want to talk about something that gives producers, developers, and QA teams absolute existential dread:

“What exactly is a ‘good’ translation?”

Is it a translation with zero typos? One that perfectly matches the source, word for word? Or one that sounds so natural players think it was written in their language—even if you had to completely change a joke to make it land?

As a producer, you probably just want the headache to go away. You constantly have to ask: “Which MT engine should we use?” or “How do we measure the quality of the post-editing we’re paying for?”

Getting a translation evaluation wrong costs you money three times: once when you pick the wrong vendor or engine, again when you pay for extra LQA testing, and a third, massive time when you fail platform certification or get review-bombed by players.

So today, I’ll show you why standard metrics are broken, and how we finally built a system that gives you the truth.

Level 1: The Atomic Age

Let’s rewind a bit. Back in the day, the industry evaluated translation using standard benchmarks—things like LISA QA or MQM. These are “atomistic” systems: you look at a text, find a specific bug, log the error type, and give it a penalty weight.

It’s incredibly detailed. MQM alone has over 190 types of errors. Too much, to be honest.

I actually started out as an editor doing these exact manual evaluations. We would sit there with a massive rubric, trying to score translators. And let me tell you a secret: human evaluation is basically a giant, subjective mess.

We ran an experiment where several experts evaluated the exact same translation. The reproducibility rate was around 23%. Why? Because even with a strict rulebook, one editor looks at a sentence and says:

“This is a critical style error!”

…while another editor looks at the exact same sentence and says:

“Nah, that’s a minor preference issue.”

But the biggest flaw was the sheer scale. If you have a million-word update dropping in a month, paying linguists to run this detailed atomic evaluation on 100% of the text is physically impossible. It would cost more than the translation itself, and you’d miss your deadline by half a year.

So what did we all do? We checked a 5–10% sample and just hoped the rest wasn’t completely broken.

Level 2: The Holistic Era & the Producer’s Nightmare

That wasn’t working. In 2018, we moved to holistic evaluation, pioneered by our colleague Serge Gladkoff from Logrus. The idea was to stop counting every single comma and just ask: is this accurate, and does it sound good?

We scored text on a simple 0 to 9 scale. It was fast, and much more consistent—we could finally predict how a player would feel reading the text. But we still had that massive scaling problem: even faster, we still couldn’t afford to check 100% of a massive update. We went from checking 5% to maybe 20–25%.

And here is where the pain begins for a producer. A holistic approach may ignore compliance. You can have a brilliant, creative translation that players love—but if that creative translator messed up a strict naming convention for a console button, calling it the “Options Key” instead of the “OPTIONS button,” you fail platform certification. Suddenly that “beautiful” text just cost you thousands of dollars in delays.

And let’s be honest about how translation often gets checked on the client side. If you’re a producer who gets a batch of German text and don’t speak German, how do you know it’s good? Usually, one of two things happens:

  • The text is creative and perfectly adapted for the target culture. They check it in Google Translate, it looks completely different from English, and they panic: “We checked this in Google! It doesn’t match! You guys translated it all wrong!”
  • The text is simple UI where there’s only one way to say it. They check it in Google Translate, it matches perfectly, and they get angry: “We checked this in Google! It’s identical! You guys just used Machine Translation and charged us for humans!”

Level 3: The MT Metric Rabbit Hole

So, humans are subjective, and Google Translate is a terrible lie detector. When AI and Machine Translation exploded, I thought: great, let’s just use automatic metrics! I went down the rabbit hole of things like BLEU, COMET, Edit Distance, and Vector Embeddings. Here’s the reality for game devs:

  • They need a cheat sheet. The best metrics require a perfect human reference translation to compare against—but if you already paid for a perfect human translation, why are you running MT in this case?
  • They hate creativity. If you use a metric that doesn’t need a reference, it usually penalizes creative slang and rewards literal, robotic text.
  • They are post-mortem. Metrics like Edit Distance tell you how much an editor had to fix after the job is done. It tells you how much money you burned, but doesn’t help you before you start.

The hard truth? Right now, there’s no silver bullet. If you want to use metrics, you can’t choose just one—you need the full orchestra, plus a good linguist to draw the right conclusions from the numbers.

Human bias is wild, too. Take Indonesian: we do a ton of volume there, and the market is split into two rival camps of translators. They review each other’s work and score it terribly—not because it’s wrong, but because it isn’t their preferred style. Even fully anonymous QA doesn’t help; they can smell the rival style on the text.

We needed an unbiased judge.

Level 4: Back to the Roots — AI Scoring

I realized the old, strict atomistic systems—like LISA and MQM—were actually the right idea. We just needed to take the human bias and the slow speed out of the equation.

We built an AI-powered scoring tool that acts as that unbiased judge. It doesn’t give you a vague “85% match” or a 6–7 accuracy score. It reads the text, identifies the specific error, flags whether it’s a terminology issue or a compliance failure, and assigns a penalty weight.

We finally got the holy grail: atomic, detailed assessment on 100% of the volume, instantly. Whether it’s pure MT, human translation, or post-editing, we know exactly what’s broken and exactly how to fix it.

Here’s a real English-to-German example: scoring scanned the text, gave it 88.15, and flagged exactly two minor errors. First, an accuracy issue—the English word “worldview” was mistranslated as “moral values” instead of the correct German lore term. Second, a tricky grammar mistake: a missing dative article.

The valid alert rate here was 100%. Not a vague vibe check—the editor gets the exact category, the severity, and the reason, then just fixes it and moves on.

Boss Fight: The Blue Protocol Case Study

Let me show you how this actually cures the localization headache. We worked on a massive update for the MMO Blue Protocol: Star Resonance. We had a batch of 1.4 million source words going into four languages—that’s 5.6 million target words to evaluate.

If linguists had evaluated 100% of that text the old way, it would have taken about 95 business days—almost five months of production time. Instead, we ran it all through our AI Scoring system.

Metric Traditional 100% human review AI Scoring (run.loc)
Review timeline ~95 business days (~5 months) 1 month, start to finish
Reviewer/translator agreement ~53% valid alert rate 85.45% valid alert rate
Localization cost Baseline Cut by 58%
Final LQA bug rate Considered excellent at 1–3 bugs / 1,000 words 1–3 bugs / 1,000 words

The linguists looked at the flagged issues and agreed with them 85.45% of the time—they just went in, fixed the exact highlighted spots, and moved on. The final result: we delivered the entire project in exactly one month instead of five, and cut overall localization costs by 58%. When the text went to final LQA, testers found only 1 to 3 bugs per thousand words—excellent even for a slow, 100% human pipeline.

We saved a cosmic amount of time and budget, and more importantly, the producer didn’t have to worry about platform certification fails or general quality.

The Ultimate Loadout: Enter run.loc

This scoring system is the core of our hybrid pipeline called run.loc.

If you’ve tried AI translation, you know it’s usually a black box—you throw text in and pray it doesn’t hallucinate a space marine into your fantasy RPG. run.loc is different: it’s a co-op mode between man and machine. Think of it like your optimal raid party.

01

The Support — Term Extractor & Style Guide

Before we pull the boss, this buffs the party. It vacuums up your terminology and locks in your Style Guide settings. By hardcoding these core project requirements from the start, it literally cannot fail compliance.

02

The Tank — run.loc Engine

Our LLMs take all the aggro. They do the heavy lifting of translating the text, applying your glossaries, and respecting your lore.

03

The Healer — AI Scoring

The system described above watches the party’s health. It scans every single line to flag glitches and compliance drops.

04

The DPS — Human Pass

Our expert linguists step in. They don’t waste time attacking millions of perfect strings—they look at the alerts, execute precision strikes to fix the text, and tune the engine for the next batch.

It cuts out the routine, keeps your terminology perfectly consistent, and clears the map faster than a speedrunner.

Evaluating translation doesn’t have to be a guessing game, and you don’t have to rely on Google Translate to check your vendor’s work. By combining the speed of AI with strict, human quality standards, run.loc takes the headache out of the pipeline.

FAQ

Atomistic systems like MQM define over 190 error types and require a linguist to review every segment manually. On a million-word update, running that on 100% of the text is physically impossible without blowing the budget and the deadline, so most teams fall back to a 5-10% sample and hope the rest holds up.
Holistic scoring is faster and more consistent than atomistic review, but it can miss strict compliance issues, like a mistranslated console button name, and it still can't cover 100% of a large volume. Teams typically move from checking 5% of the text to 20-25%, not the full batch.
Not reliably. If the text was creatively adapted for the target culture, a machine back-translation will look completely different from the source and can wrongly read as an error. If the text is simple UI with only one natural phrasing, it will match Google Translate exactly even when it was translated by a human, which also creates false alarms.
The most reliable metrics need a perfect human reference translation to compare against, which defeats the purpose if you're evaluating MT to avoid paying for full human translation. Reference-free metrics tend to penalize creative, natural-sounding text and reward literal, robotic phrasing, and metrics like Edit Distance only tell you what was fixed after the fact, not what to expect before you start.
AI Scoring applies the same atomic, detailed logic as MQM or LISA QA, but removes human bias and slowness. It reads the text, identifies specific errors, flags whether each one is a terminology issue or a compliance failure, and assigns a penalty weight, delivering detailed assessment on 100% of the volume instantly.
A 5.6-million-word evaluation that would have taken roughly 95 business days of manual review was completed in one month. The valid alert rate reached 85.45%, overall localization costs dropped by 58%, and final LQA still found only 1-3 bugs per thousand words.
run.loc works like a four-role raid party: a Term Extractor and Style Guide lock in terminology and compliance before translation starts, the run.loc engine (LLMs) handles the heavy lifting of translation, AI Scoring continuously flags glitches and compliance drops, and expert linguists step in only to fix the flagged issues and tune the engine for the next batch.

WANT TO BE A GUEST?

Allcorrect Gamedev Show — a podcast about how to build amazing games, thriving studios, and in-demand careers in the gaming industry. Join our next podcast episode! If you are a game studio founder, producer, designer, artist, or composer and know what it takes to create incredible games and make money, we’d love to chat with you.

Contact us for more details!

Contact us

Allcorrect is a game content studio that helps game developers free their time from routine processes to focus on key tasks. Our expertise includes professional game localizations, localization testing, believable voice-overs, and narrative design.

FOLLOW US
  • Instagram
  • Facebook
  • Twitter
  • Linkedin