← All articles

Save Up to 40% Time: MT Post Editing for Small Teams with Arkian

Save Up to 40% Time: MT Post Editing for Small Teams with Arkian

Hands editing machine translation sheets

MT post editing (MTPE) is the practice of a human editor revising machine-generated translation to meet a defined quality target rather than translating from scratch. The single most important decision is choosing the right level of edit for each job: light post-editing for internal or gisting content, full post-editing for anything customer-facing. ISO 18587:2017 offers a starting framework for both, but the level only means something once you track edit distance and post-editing speed against it.


TL;DR:

  • Light post-editing can typically save around 20% in time compared to full post-editing, but savings vary widely based on language pair and content type.
  • Segments that are fluent but contain inaccuracies are often underestimated during triage and require careful checking to avoid costly errors.
  • Using structured formats and real-time TM updates can significantly streamline workflow efficiency and reduce editing time.
  • Editors should prioritize judgment, terminology consistency, and understanding QE scores rather than mere speed to ensure quality in MTPE.
  • Adhering to ISO 18587:2017 standards and tailoring project-specific thresholds improve post-editing outcomes and reduce rework.

Table of Contents

What Is MT Post Editing, and What Standard Should You Follow?

MT post editing is the process of correcting raw machine translation output so it meets a quality bar agreed upon before the project starts, not one improvised segment by segment. Without that agreement, “post-edit this” means five different things to five different editors, and your delivery dates become guesswork.

ISO 18587:2017 is the reference point most localization teams reach for first. It specifies requirements for the post-editing process itself and for the competence of the people doing it, distinguishing full post-editing (aimed at output indistinguishable from human translation) from light post-editing (aimed at output that is merely comprehensible and accurate). It doesn’t hand you a plug-and-play checklist for every content type, but it gives you the vocabulary to write one.

That’s where the project brief comes in. A working MTPE brief should specify:

  • Purpose and audience: who reads this, and what happens if they misunderstand it (a support macro carries different risk than a product label).
  • Do-not-change list: brand terms, legal disclaimers, placeholders, and formatting tags that editors must leave untouched regardless of MT output quality.
  • Annotated examples: two or three segments showing “good enough” versus “needs a full rewrite,” pulled from the actual content, not a generic style guide.
  • Quantitative thresholds: a maximum acceptable edit distance per segment, a target words-per-hour rate, and the error categories that trigger automatic rejection.

Academic reviews of post-editing guidelines have pushed back on treating light and full as two fixed boxes. A frequently cited paper on rethinking MTPE guidelines argues that with modern MT quality, the old binary is blurring, and that situated, project-specific rules produce better outcomes than generic categories. In practice, that means your brief should name the actual risk tolerance for this content, not just cite “light” or “full” as labels.

Pro Tip: Write your do-not-change list before you run a single segment through the MT engine. Editors who discover placeholder rules mid-batch tend to go back and re-check earlier segments, which quietly doubles their time on the job.

Light vs Full Post-Editing: How to Pick the Right Level

Light post-editing fixes only what breaks comprehension or accuracy. Full post-editing brings the text to a quality level a reader can’t distinguish from a professionally translated document, matching tone, register, and terminology throughout. Postediting distinctions between “gisting” and “publishable” output have anchored the industry for over a decade, and they still hold up as a starting frame even as the line between them gets fuzzier.

Use this checklist to decide which level a job needs:

  1. Check visibility. Will this text appear on a public product page, app store listing, or customer email? If yes, lean full. Internal wikis and support ticket summaries can usually stay light.
  2. Check risk. Legal terms, medical instructions, and safety warnings need full post-editing regardless of visibility, because an ambiguous sentence carries real consequences.
  3. Check shelf life. Content that will be updated again in a month (a changelog, a seasonal promo) rarely justifies full post-editing effort.
  4. Check brand exposure. Marketing copy and app store descriptions represent the brand directly; treat them as full even if the volume is small.
  5. Check volume against deadline. If the queue is thousands of segments and the deadline is measured in hours, light post-editing on lower-risk content frees editor time for the segments that actually need care.

Productivity swings are wide enough that you shouldn’t quote a client a flat discount without testing it on their content first. Industry reports point to time savings approaching 40% in favorable conditions, but academic studies frequently find savings reported to vary widely, often modest, depending on language pair, domain, and MT engine quality. A technical manual translated between closely related languages might see strong gains; a creative marketing piece translated into a structurally distant language might see almost none. Price and deadline commitments should be based on pilot batches rather than generic industry averages.

The MTPE Workflow: From Pre-Editing to Final QA

MT post-editing workflow diagram

A workflow that treats MT output as a finished draft to “just clean up” produces inconsistent results. A workflow that treats it as one stage in a pipeline, with checks before and after, produces predictable ones.

Step 1: Pre-edit the source

Clean, unambiguous source text produces cleaner MT output. Before anything touches an engine, flag source segments with unresolved ambiguity, broken sentence structure, or missing context (a UI string with no screen reference, for instance). Fixing these five minutes before translation saves far more than five minutes of post-editing later.

Step 2: Configure the MT engine and feed it context

Generic MT output improves substantially once you feed the engine your translation memory ™, glossary, and a handful of previously approved segments as examples. Practitioner guidance consistently recommends seeding the pipeline with human direction early rather than waiting to fix drift after the fact. If your engine supports custom terminology injection or few-shot examples, load your approved glossary before the first batch runs, not after editors start flagging the same wrong term for the tenth time.

Step 3: Triage every segment before editing it

Not every segment deserves the same attention. A fast first pass should sort output into four buckets:

  • Accept as-is: fluent, accurate, terminology correct.
  • Minimal edit: one or two word-level fixes (a wrong article, a missed capitalization).
  • Rewrite: the meaning is right but the phrasing is stiff, unnatural, or inconsistent with style guide.
  • Retranslate: the MT output is fluent but wrong, or has dropped meaning entirely.

That fourth bucket is the one editors most often underestimate. Fluent, grammatically correct sentences that say the wrong thing are harder to catch than obviously broken ones, because there’s no surface signal telling the editor to slow down.

Step 4: Run final LQA and package the output

Once editing is done, a language quality assurance (LQA) pass checks for consistency across the whole batch, not just segment by segment. Confirm placeholders and formatting tags survived the edit, terminology stayed consistent across the document, and no segment silently reverted to raw MT output during a rushed pass. Only after LQA clears should the files move into their delivery format, whether that’s a translation memory update, a structured JSON file, or a packaged app string set.

Pro Tip: Run your LQA sample from the middle and end of the batch, not just the beginning. Editor fatigue is real, and error rates climb noticeably in the last quarter of a long session.

Optimizing MTPE for Efficiency: Tools, QE, and Memory Loops

The tools stack around MTPE matters as much as the editing skill itself. A CAT tool that surfaces MT output alongside TM matches and glossary hits in the same pane lets an editor make a triage decision in seconds instead of switching between three applications. Without that integration, editors lose time just assembling context, before they’ve fixed a single word.

Quality estimation (QE) scoring, when available, lets you route segments by predicted quality instead of editing everything at the same pace. A segment the engine scores as high-confidence can go straight to a light pass or even auto-acceptance for low-risk content; a low-confidence segment gets flagged for a full rewrite before an editor even opens it. This is selective post-editing in practice: spending editor attention where the MT output is actually weak, not spreading it evenly across a batch where most segments didn’t need it.

A few workflow patterns consistently reduce total effort:

  • Seed with high-quality examples. A handful of translator-approved snippets fed to the engine as reference material measurably improves output on similar content, cutting repeat corrections on recurring phrasing.
  • Use cascaded models for high-risk content. Run a first pass through a general engine, then a domain-tuned or custom-trained model for technical or brand-sensitive segments, rather than forcing one engine to handle everything.
  • Close the feedback loop. Vendor guidance on human-in-the-loop MTPE describes feeding post-editor corrections back into custom engine training, which compounds quality gains over successive batches instead of starting from zero each time.
  • Update TM immediately, not in a batch job at project close. Editors working on segment 500 of a 2,000-segment file benefit from corrections made on segment 50, but only if the TM updates in near real time.

None of this replaces editor judgment. It just means the editor spends judgment on the 20% of segments that actually need it, instead of re-reading fluent, correct output that didn’t need a second look.

Common Post-Editing Pitfalls and How to Avoid Them

The most expensive MTPE mistakes are invisible until they’ve already shipped. Fluent output is the biggest trap: modern MT engines produce grammatically flawless sentences that are confidently, completely wrong, and there’s no stylistic tell warning the editor to double-check. Practitioner analyses list this as one of the four recurring bottlenecks in real MTPE work, alongside triage decisions, priming effects, and pricing mismatches.

Build these checks into your process rather than trusting individual vigilance:

  • Spot-check meaning against source, not just fluency. Pull a random 10% sample and read it against the source text specifically for dropped or inverted meaning, independent of how natural the target reads.
  • Watch for anchoring. Editors who read the MT draft first tend to accept its sentence structure even when a better one exists, simply because the draft primed their thinking. Rotate a portion of segments through blind retranslation as a check.
  • Never assume a flat productivity discount. Effort varies by segment, not just by project, so time-tracking at the segment level (not just the batch level) catches jobs that are quietly running over budget.
  • Audit terminology drift across long batches. A term translated correctly in segment 10 can drift by segment 400 if the glossary isn’t actively enforced, especially across multiple editors on the same file.
  • Check placeholders and tags on every segment, not a sample. A single broken variable tag can break a build; this is the one check that doesn’t scale down to spot-checking.

Pro Tip: If two editors on the same project consistently disagree about whether a segment needs a rewrite, that’s not an editor problem, it’s a brief problem. Go back and tighten the do-not-change list and examples before assigning more volume.

How to Measure MTPE Success: Metrics, ROI, and KPIs

Edit distance and post-editing speed are the two numbers that actually tell you whether MTPE is working. Edit distance measures how much an editor changed relative to the raw MT output, usually as a percentage or character count; post-editing speed measures words or segments processed per hour. Together they show you whether a given content type and language pair are genuinely good candidates for MTPE, or whether editors are quietly retranslating from scratch while the project gets billed as post-editing.

Automatic metrics like BLEU compare MT output to a reference translation and can flag gross mismatches quickly, but BLEU alone doesn’t catch meaning errors that a fluent, structurally similar sentence can still contain. Human LQA remains the check that catches brand voice violations, safety-critical mistranslations, and subtle meaning shifts that a similarity score simply isn’t built to see.

A simple ROI model has three inputs: the per-word cost of MT plus post-editing, versus pure human translation cost, weighted against the time saved and the error rate difference between the two. If light post-editing on internal documentation runs 20% faster than full human translation and the error tolerance for that content is high, the math favors MTPE clearly. If a legal contract shows only a 5% speed gain and any error carries real liability, the case for MTPE weakens fast, and that gain-versus-risk comparison is the entire decision.

For a dashboard, track:

  • Average edit distance per content type, flagged when it exceeds your threshold for that level.
  • Words per hour by editor and by language pair, to catch systemic slowdowns early.
  • Error category counts from LQA samples, especially mistranslation and omission, which BLEU won’t surface.
  • Segments retranslated versus lightly edited, as a proxy for whether your MT engine is actually suited to this content.

Reported time savings for MTPE range from close to zero to roughly 40% depending on domain and language pair, which is precisely why a pilot batch with real KPI tracking beats trusting an industry average for your specific content.

Arkian in Practice: Packaging and Validation That Smooth MTPE

Structured output formats do more for MTPE efficiency than most teams expect, because editors waste real time reconciling inconsistent file structures before they ever touch a translation. Arkian’s platform generates localization strings in JSON, iOS, Android XML, and YAML formats already organized for review, which means an editor opens a file that’s already structured for line-by-line comparison instead of a raw export that needs reformatting first.

A typical workflow through Arkian looks like this:

  • Generate machine translation output for app strings, along with structured metadata, directly from an approved source.
  • Route candidate translations to a post-editor for triage and correction using the do-not-change rules and thresholds set in the project brief.
  • Validate the edited output against the original structure, catching broken placeholders or missing keys before packaging.
  • Package the final iOS .strings files or Android XML files into delivery-ready assets without needing repository access or a full translation management system.

Arkian’s collaboration with the Quiet Harbour app demonstrated this validation-then-packaging approach in practice, producing a coherent localized experience across multiple languages without the team needing a dedicated TMS setup. For teams running voice or narration content alongside text strings, the same validate-and-package pattern applies to TTS audio output, keeping text and voice assets aligned through the same review step.

The core benefit isn’t the MT itself; it’s that editors spend their time editing, not reassembling files or chasing down broken formatting after the fact.

How Does MT Post-Editing Differ From Traditional Translation Revision?

Traditional human translation revision starts from a first draft written by a person who made deliberate stylistic choices, even flawed ones. A reviser’s job is to evaluate those choices and refine them, working from an assumption that the translator understood the source.

MT post-editing starts from a very different kind of draft. The MT engine has no communicative intent; it optimized for statistical or learned patterns, not meaning in context. That changes what the editor is actually checking for. A human reviser mostly asks “is this the best way to say it?” A post-editor has to ask “does this even say the right thing?” first, and only then move to phrasing.

This distinction is why ISO 18587 treats post-editing competence as a distinct skill set from translation or revision competence, requiring editors trained specifically to spot MT-typical error patterns: dropped negation, mistranslated homographs, inconsistent terminology across a batch, and fluent-but-wrong output that a human draft would rarely produce in the same way. A skilled human reviser can sometimes struggle with MTPE at first, precisely because the error patterns don’t match what years of reviewing human drafts trained them to look for.

Pay models often reflect this difference too, with per-word MTPE rates typically set separately from per-word revision rates, since the effort profile (fast on easy segments, slow on deceptively fluent ones) doesn’t map cleanly onto either translation or revision pricing.

Protecting Productivity and Well-Being During MTPE Work

Sustained MTPE work carries a specific kind of fatigue that straight translation doesn’t produce in the same way: constant low-grade vigilance against errors that don’t announce themselves. Reading fluent sentences all day while staying alert for the ones that are fluent and wrong is mentally taxing in a way that’s easy to underestimate when scheduling a day’s workload.

Hands poised for detailed manual editing work

A few practices help editors sustain both speed and accuracy across a full workday:

Build in bucket variety rather than running the same content type for eight straight hours. Switching between triage-heavy batches and edit-light ones gives attention a chance to reset. Cap high-risk, full post-editing sessions to shorter blocks, two to three hours, with breaks between, since error detection rates measurably decline with sustained close reading. Track speed by time of day for individual editors; many people show a clear afternoon dip, and scheduling lower-risk batches during that window protects quality on the segments that matter most.

Pay structures matter here too. A flat per-word rate that doesn’t account for edit-distance variability quietly punishes editors for landing on a batch with more retranslation-level segments, which erodes morale over time even when the editor did nothing wrong. Rate models that account for actual effort, not just word count, keep skilled editors doing MTPE work instead of burning out on it.

What Skills Actually Matter for MT Post-Editors

Triage judgment matters more than raw editing speed in modern MTPE work. The editors who add the most value aren’t necessarily the fastest typists; they’re the ones who can look at a segment for two seconds and correctly sort it into accept, edit, or retranslate without agonizing over borderline cases.

Terminology discipline is the second underrated skill. MT engines drift on terminology across long documents in ways human translators rarely do, because a person remembers what they called something on page one. An editor who checks terminology consistency as a discrete step, not just as a byproduct of reading, catches drift that a casual read-through misses entirely.

QE literacy is becoming a genuine hiring differentiator. Editors who understand what a quality estimation score is actually measuring, and where it tends to be wrong, make faster and better triage calls than editors treating QE output as a black box. Training programs that still teach post-editing as “fix what’s broken” without teaching editors to read engine confidence signals are teaching half the skill set the job now requires.

Payment models should follow this shift. Editors paid purely per word have no incentive to flag systemic MT weaknesses back to the project manager; editors with any stake in overall project quality do. That feedback loop, more than any single tool, is what separates MTPE programs that improve over time from ones that repeat the same errors batch after batch.

— Arkian

Arkian: Structured Localization Without the Repository Headache

Small product teams running MTPE don’t need a full translation management system to get consistent, validated output. Arkian replaces the manual assembly work, chasing file formats, checking placeholders by hand, reconciling TM exports, with a platform that generates multilingual strings, structured metadata, and TTS audio, then validates and packages them into delivery-ready files automatically.

Arkian

That matters most for teams without a dedicated localization engineer on staff. If your current process involves exporting strings, running them through an MT engine, post-editing in a spreadsheet, and manually rebuilding JSON or XML files afterward, Arkian collapses that into one pipeline: generate, post-edit, validate, package, all without needing repository access. It’s built specifically for teams that need localization strings in developer-ready formats without standing up enterprise TMS infrastructure just to ship a few languages.

If your team fits that profile, small, resource-constrained, shipping multilingual content without a dedicated localization stack, check whether Arkian fits your workflow and see what a packaged output actually looks like before your next release.

Where to Verify These MTPE Standards and Practices

ISO 18587:2017 remains the primary standard for post-editing requirements and post-editor competence, and it’s worth reading directly if you’re drafting a formal MTPE policy for your team. The Wikipedia entry on postediting offers a concise overview of the light/full distinction and its history for anyone new to the terminology.

For the argument that rigid light/full categories are outdated, the academic paper on rethinking MTPE guidelines lays out the case for situational, project-specific rules in more depth than any industry blog post will. Practitioner-level guidance on seeding models and human-in-the-loop workflows covers the operational side well, and the Crowdin overview of MTPE best practices is a useful gut check on realistic productivity expectations before you commit to a client-facing time-savings number.

Sources

Made with BabyLoveGrowth’s AI