Arkian logo
Start free
Menu
← All articles

CI Ready Batch Translation for Developers: Keep Plurals, Placeholders

CI Ready Batch Translation for Developers: Keep Plurals, Placeholders

Developer reviewing structured localization resource files

The best way to batch translate strings is to export your source text as structured key-value files, run them through a batch translation API or platform-native tool that preserves placeholders and plural rules, then reimport and validate before shipping. Safe formats include XLIFF, TSV or JSON with stable IDs, and platform-specific packages like Android’s strings.xml. The main failure points are broken plurals, stripped placeholders, and mangled markup.


TL;DR:

  • Batch translation is most effective for high-volume, repetitive content sent through automated APIs, especially when integrated into CI pipelines, rather than one-off quick translations.
  • Export formats like XLIFF, TSV, JSON, and platform-native files must include stable IDs and contextual data to ensure translations survive the round trip without errors in plurals, placeholders, or markup.
  • Proper handling of plurals requires testing according to CLDR rules, as languages like Arabic and Russian need multiple plural forms, unlike English’s simple singular and plural.
  • Automated QA should verify that placeholders, plural forms, and string length constraints remain intact before final import and deployment.
  • Platforms like Arkian streamline the whole process by automating export, translation, validation, and packaging, but batch translation remains unsuitable for culturally nuanced or layout-dependent content.

Arkian
arkian.ai
Simplify Your Localization Pipeline
Arkian automates multilingual scripts, voice outputs, validation, and structured language packages for small teams without complex TMS workflows.
Visit Arkian

Table of Contents

Which tools and APIs support batch translation

The right tool depends on volume, automation needs, and how tightly you want translation tied to your build pipeline.

For large or recurring jobs, cloud batch APIs are built for scale. Google Cloud’s batch translation accepts file-based input from storage buckets and runs asynchronously, which suits nightly or release-triggered jobs rather than quick one-off requests. For tightly integrated iOS workflows, Apple’s platform tools keep translation closer to the source: preparing text for translation in Xcode works directly against your String Catalogs, so you avoid a separate export step entirely.

Desktop apps and quick bulk translators fill the gap for ad-hoc work, a single screen’s worth of copy, a quick check before a demo, but they rarely integrate with version control or preserve IDs the way a pipeline needs.

Pick your approach by answering a few questions:

  • How much text are you translating, and how often does it change?
  • Do you need the job to run unattended inside CI, or is a person triggering it each time?
  • Does your content include sensitive or proprietary text that limits which service can process it?
  • Do you need a shared glossary or translation memory across releases?

A one-time landing page translates fine with a desktop tool. A product shipping weekly in ten languages needs an automated batch pipeline.

What file formats to export and why they matter

Format choice determines whether your translations survive the round trip intact. Each ecosystem has a format that keeps IDs, context, and plural structure attached to the string instead of floating loose in a spreadsheet.

  1. XLIFF works across most translation management systems and is the format Xcode exports natively, making it the default choice for iOS projects.
  2. TSV or CSV with an ID column suits cloud APIs that accept file-based input, as long as the ID column stays untouched through the translation step.
  3. JSON or YAML fits custom pipelines where you control both the export and import code and want full flexibility over structure.
  4. Platform-native formats (.strings, .stringsdict, strings.xml) are the end destination regardless of what you translate through, so your pipeline should always round-trip back to these.

Apple’s guidance on preparing app text for translation recommends exporting String Catalogs as XLIFF, since catalogs already unify source strings, plural variants, and translator context in one file. Enabling “Localization Prefers String Catalogs” early in a project avoids a painful migration later, because catalogs compile down to .strings and .stringsdict automatically at build time.

On Android, string resource documentation is explicit that plural resources must stay as separate nodes rather than collapsing into one flat value, since each language needs its own set of plural categories.

Pro Tip: Keep a permanent, immutable ID column in every export format you use, even formats that do not require one, so you can always reconcile translated output against the original source.

Why plurals, placeholders, and markup break in batch jobs

Three things cause most of the visible bugs after a batch translation job: plural mismatches, lost placeholders, and stripped markup.

Why plurals, placeholders, and markup break in batch jobs — overview diagram

Plural handling trips people up because English has only two categories (one, other), but many languages need more. The CLDR plural rules define these categories per language and recommend testing with minimal-pair examples so a translator or automated system picks the right form for each count. Skipping this step is why apps built for English-only pluralization often show grammatically wrong counts once translated into Arabic, Polish, or Russian.

Placeholders fail when translation tools treat {0} or %1$s as ordinary text and reorder, translate, or drop them. Safe handling looks like this:

  • Use stable, numbered tokens instead of positional-only markers wherever the platform allows it.
  • Never concatenate translated fragments at runtime. Build full sentences with placeholders instead.
  • Confirm every placeholder present in the source string is also present in the translated string before import.
  • Treat HTML tags or styling annotations as protected spans, escaping or marking them so a translation engine does not alter or discard them, as Android’s string resource guide warns for styled and annotated strings.

One of the most common automated checks catches a real structural rule: CLDR’s plural category system means a language can require anywhere from one to six plural forms, so a validation pass that only checks for “singular and plural” will silently miss required forms in many target languages.

A batch workflow you can automate end to end

A reliable pipeline has four stages, and each one can run unattended once it is set up correctly.

  1. Export your strings with IDs and context intact. Pull XLIFF from Xcode String Catalogs, or generate TSV or JSON from a custom script, making sure translator comments travel with each entry.
  2. Submit the batch job. Attach any glossary your team maintains, choose an asynchronous job for large volumes, and point the output to a storage location you can poll or trigger a webhook from. Google Cloud’s batch translation reference notes that async jobs write results and an index file mapping inputs to outputs, which is what your next stage reads.
  3. Run automated QA against the output: check every placeholder survived, every required plural form is present, and no string exceeds a length threshold that would truncate on a small screen. Flag anything suspicious for a human reviewer instead of importing it blind.
  4. Import and verify. Write translated strings back into their platform format, mark any strings that failed validation as stale rather than shipping them, and run a smoke test on a device or emulator in at least one non-English locale before merging.

Pro Tip: Run the QA step as a required check in your pull request pipeline, not as a manual task someone remembers to do before release.

How to integrate batch translation into your release cycle

Where you run this in your lifecycle matters as much as how you run it. Teams that translate only at release time tend to batch large files under time pressure, which is exactly when placeholder and plural bugs slip through unnoticed.

A steadier pattern runs batch jobs nightly against whatever strings changed that day, so by release time only a small delta needs review. Glossaries and translation memory keep terminology consistent release over release, which matters most for product names, legal terms, and recurring UI labels that should never be translated two different ways.

  • Gate high-risk strings (onboarding flows, payment screens, error messages) behind a sampled human review even when the rest of the batch ships automatically.
  • Maintain one glossary source of truth rather than letting each translation job apply its own term list.
  • Use synthetic rendering or device-level checks to catch truncation and layout breaks before a human ever opens the build.
  • Treat a TestFlight-style internal distribution as your last check for render issues a desktop QA pass would miss.

A partner resource on AI tools for translators covers how human translators themselves are adopting these tools, which is a useful read if your pipeline includes a human post-edit stage rather than pure automation.

How Arkian handles batch string translation and packaging

We built our platform around the same export-translate-validate-package sequence described above, minus the integration overhead. We automate the creation of multilingual scripts, voice outputs, and structured language packages without requiring repository access or a full translation management system.

  • We accept iOS .strings, Android XML, JSON, TS, and YAML, so whichever export format your project already produces is usable as-is.
  • We run placeholder and plural validation automatically as part of the job, not as a separate manual step you have to wire up.
  • We package validated output into delivery-ready folders organized by language and file type, ready for a developer to drop into a build.

We demonstrated this approach through our work with the Quiet Harbour app, producing a coherent localized experience across multiple languages from a single source set.

When batch translation is the wrong tool

Batch translation is built for volume, not nuance, and that trade-off shows up fastest in marketing copy, legal text, and anything culturally sensitive. A pricing page or a consent dialog translated word-for-word can be technically correct and still land wrong for the audience reading it. These need a human translator who understands context, not just a string.

Layout-dependent content has the same problem. If a translated string has to fit inside an image or a fixed-width button, no amount of placeholder validation catches a visual overflow. The most durable setup treats batch translation as the first pass and routes anything customer-facing or visually constrained to a targeted human review before release.

— Arkian

Get your strings translated and packaged without the TMS overhead

We handle the batch translation, placeholder and plural validation, and packaging into ready-to-review language files, so you skip the setup most pipelines require before a single string ships. Our localization strings service takes your exported files and returns structured, validated output organized for your build.

Arkian

If you want to see what a batch job looks like end to end, check plans and pricing starting with the Starter tier at 29.00 CAD one-off, or request a sample run on your own string files.

FAQ

How do you translate images in bulk?

Bulk image translation typically means extracting embedded text (through OCR or source layers), translating that text as structured strings, then re-rendering it back into the image or overlaying it in your app’s UI layer. It follows the same export-translate-reimport pattern as string translation, but with an extra extraction step since image text usually is not already in a key-value format.

How long does it take to translate 20,000 words?

Turnaround depends entirely on the method: a batch machine translation job can process that volume in minutes to hours depending on API queue times and file size limits, while human translation of the same volume typically takes days, since professional translators work at a steady per-day word pace. A hybrid approach, machine batch plus human post-edit on flagged sections, lands somewhere in between.

What file formats work best for batch string translation?

XLIFF, TSV or CSV with a stable ID column, and JSON or YAML are the most reliable formats because they keep each string’s identifier, context, and plural structure attached through the translation process. Platform-native formats like iOS .strings and Android’s strings.xml remain the final destination regardless of which format you batch through.

Does Arkian support Android and iOS string files for batch translation?

Yes, we accept iOS .strings, Android XML, JSON, TS, and YAML files directly, and we validate placeholders and plural forms as part of the job before packaging the output for review.

Sources