Build Cross Platform Translation Files in an Afternoon for Small Teams
Build Cross Platform Translation Files in an Afternoon for Small Teams

For cross-platform localization, pick a single canonical source and automate conversion to platform-native files using interchange formats and automated validation. Store your strings in JSON or YAML, or in XLIFF/TMX when a vendor needs a standard handoff format, then run scripted conversion into iOS, Android, and web targets. Validate placeholders, plurals, and encoding before every merge. You can set up the skeleton of this pipeline in an afternoon.
TL;DR:
- Use XLIFF or TMX for vendor and CAT tool handoffs; keep JSON or YAML as the internal source, with platform files generated automatically.
- English uses two plural forms, while Arabic needs six and Polish needs four; validate locale specific categories rather than assuming singular and plural suffice.
- Choose build time generation for compile time safety or runtime loading for hot updates, then compare generated key sets across platforms nightly or per build.
- Google Cloud caps document translation batches at 1GB or 100 million Unicode codepoints; scanned PDFs need OCR first and should not be compressed.
Table of Contents
- Which translation file formats matter and when to use them
- Developer tools for converting and managing translation files
- Building a CI/CD pipeline for multi-platform translation output
- Automated validation checks that catch translation bugs before release
- Handling PDFs, scanned documents, and large translation batches
- A practical ingest-to-package workflow for small teams
- Best practices for keeping translation files in sync across platforms
- Using context and metadata to improve cross-platform translation accuracy
- Version control and collaboration strategies for translation files
- Managing pluralization and gender variations across platforms
- Speed versus fidelity: what small teams should prioritize
- Automate your cross-platform localization with Arkian
- FAQ
- Sources
Which translation file formats matter and when to use them
Not every format does the same job, and picking the wrong one creates friction you will feel for years.
XLIFF and TMX exist for interchange. They are the formats translation agencies and CAT tools expect, and Microsoft’s globalization documentation lists both as standard formats that let tools extract translatable text and reassemble it without losing structure. XLIFF 2.2 defines a trans-unit structure built specifically to carry source text, target text, and metadata together so nothing gets stripped out during a round trip, according to the OASIS XLIFF 2.2 core specification. TMX does the same job for translation memory, and its Level 2 markup preserves inline codes and formatting rather than flattening everything to plain text, per the TMX specification.
Platform-native formats solve a different problem: runtime loading. Android wants XML resource files, iOS wants .strings and .stringsdict, and web apps usually run on JSON or YAML through libraries like i18next. These formats are fast to parse but weaker at carrying translator notes, context, or formatting metadata.
A simple rule covers most projects:
- Choose XLIFF or TMX when handing files to an external vendor, translation agency, or CAT tool.
- Choose JSON or YAML as your canonical source when your whole pipeline is internal and automated.
- Keep platform-native formats as generated output, never as the source of truth.
- Preserve plural categories and placeholder syntax at every conversion step, since this is where fidelity breaks down fastest.
Developer tools for converting and managing translation files
Hand-editing files across five platforms does not scale past a handful of strings, so most teams reach for a mix of APIs, CLI tools, and converters.
- APIs and SDKs handle batch import and export, letting a build step pull the latest strings and push generated files back without a human touching a file manager.
- CLI and TUI tools fit naturally into local development and CI. The open-source LocalizationManager supports JSON,
.resx, Android XML, iOS.strings, PO, and XLIFF in one tool, with commands to scan, validate, and export that drop straight into a pipeline. - Format converters handle the smaller, recurring headaches: ICU MessageFormat to i18next syntax, JSON to YAML, or XLIFF import and export for vendor round trips.
A reliable script follows the same order every time: detect the source format, back up the original file, run validation, then convert. Skipping the backup step is how teams lose translator edits during a bad conversion run.
Pro Tip: Run your converter against one placeholder-heavy file first; placeholder bugs surface faster there than in a full batch.
Building a CI/CD pipeline for multi-platform translation output
A canonical workflow turns a single source file into every platform output automatically, and it holds up whether you have three locales or thirty.
The sequence looks like this: extract strings from code, canonicalize them into your source JSON or YAML, generate platform-specific files, validate the output, package it for review, then publish to each app target. Each step is a discrete, testable unit rather than one brittle script.
- Hook extraction into your version control system so new strings trigger a pipeline run automatically.
- Run generation as a build task, not a manual step, so Android and iOS outputs never drift from the source.
- Store generated artifacts in a dedicated bundle rather than committing them alongside source code.
- Decide early whether you generate files at build time, for compile-time safety, or load translations at runtime, for hot updates without a rebuild.
Translation memory and glossary checks belong in this same pipeline, not as an afterthought. A TM lookup step can flag repeated strings before they go to a translator, and a glossary check can catch inconsistent terminology before it ships. Our translation memory guidance and glossary-building steps cover how these fit into a small team’s automation without requiring a full TMS.
Automated validation checks that catch translation bugs before release
Validation failures are one of the most common sources of post-deployment localization bugs, according to Adobe’s guidance on OCR and translation workflows, which notes how often broken placeholders and malformed tags slip through unchecked files.
Four checks catch most problems before they reach a reviewer:
- Placeholder parity: confirm every
{{variable}},%@, or{count}token in the source appears, unchanged, in every target language. - Plural category coverage: test that each locale’s required plural forms (one, few, many, other) are present, since English’s two-form system undercounts what Polish or Arabic need.
- Schema and well-formedness checks: run XML files through a schema validator and JSON files through a strict parser before they enter a build.
- Encoding checks: verify UTF-8 throughout, since a single mis-encoded file can corrupt an entire app build silently.
Pro Tip: Make these checks a pre-merge gate, not a post-release audit. A failed placeholder check should block a pull request, not a release.
Handling PDFs, scanned documents, and large translation batches
PDFs and scanned documents need different treatment than app strings. Native, text-based PDFs translate cleanly because the text layer is intact. Scanned documents need OCR first, and OCR quality depends heavily on image resolution and layout detection.
- Prefer native PDFs over scanned ones whenever a text-based source exists.
- Use zone-based OCR for scanned documents to preserve column layout and reduce misread text, as recommended by guidance on OCR tools for PDFs.
- Never compress a scanned document before running OCR; lossy compression degrades the image enough to produce hallucinated text.
- Split large batches to respect provider limits. Google Cloud’s document translation service caps batch jobs at 1GB or 100 million Unicode codepoints, and native PDFs are explicitly preferred over scanned ones because scanned files require extra preprocessing and risk losing layout.
Developers converting large documentation sets often extract text into an interchange format like XLIFF before translation, which keeps structure intact through the round trip. A partner guide on converting PDFs to markdown covers a practical preprocessing approach for turning PDF content into clean, machine-readable text before it ever reaches a translation step.
A practical ingest-to-package workflow for small teams
A workflow that mirrors the steps above looks like this in practice: ingest the source file, validate it for placeholder and plural integrity, convert it into each target format, then package the results into a reviewable bundle.

Packaging matters more than it sounds. A small team reviewing a validated, organized bundle can approve translations without digging through a repository or learning a translation management system. That difference in reviewer friction is often the gap between a localization step that ships on schedule and one that stalls for a week.
We built our own localization strings handling and structured packaging around this exact sequence, and our Quiet Harbour case study shows the same ingest, validate, convert, package flow applied to a real app release.
Best practices for keeping translation files in sync across platforms
Drift between platforms is the most common failure mode in multi-platform localization: an iOS string gets a quick fix, and the Android equivalent never catches up.
The fix is structural, not procedural. Keep one canonical source file per locale and treat every platform-native file as generated output, never as something a developer edits by hand. When a string changes, it changes in the source, and the pipeline regenerates every target automatically. Teams that let developers patch a .strings file directly, “just this once,” are the ones that end up with three slightly different English strings across three platforms within a year.
Automated diffing helps catch drift that slips through anyway. A nightly or per-build job that compares key sets across all generated files, flagging any key present in one platform’s output but missing in another, catches most sync failures before a reviewer ever sees them. Pair this with a single source of truth for locale lists, so a new language added to the iOS build cannot silently stay missing from Android.
Our guides on YAML as a canonical source, Android string generation, and iOS .strings output each walk through the platform-specific side of this, but the sync discipline itself is the same regardless of target: one source, generated outputs, automated comparison.

Using context and metadata to improve cross-platform translation accuracy
A string without context is a guess. “Close” could mean shutting a window or standing near something, and a translator working from a bare key-value pair has no way to tell which one you meant.
Metadata fixes this cheaply. XLIFF’s structure was designed to carry exactly this kind of auxiliary information alongside the translatable text, per the OASIS XLIFF core specification, and most platform formats support comments or notes fields that serve the same purpose even outside a formal interchange format. A comment above a JSON key, a note attribute in an Android string resource, or a developer comment in an iOS .strings file all do the same job: they tell a translator or an automated system what the string is for, where it appears, and how long it can safely run before it breaks a layout.
Screenshots and UI context help further, especially for short strings in buttons or labels where ambiguity is highest. A glossary tied into your conversion pipeline catches a different kind of context failure: inconsistent terminology, where “sign in” and “log in” both appear in the same app because two different translators made two different reasonable choices. Our glossary-building guide covers how to wire a termbase into a conversion step so this gets caught automatically rather than in a late-stage review.
Version control and collaboration strategies for translation files
Translation files are code, and they should live under the same version control discipline as everything else in a repository, with one adjustment: the people editing them are often not developers.
The practical answer most small teams land on is to keep canonical source files in version control, but to generate a separate, human-readable review bundle, rather than asking translators or reviewers to work directly in a pull request. A folder of generated files with inline diffs and context notes lets a non-technical reviewer approve changes without repository access or Git familiarity. A suggested structure for organizing these bundles is covered in our multilingual output folder guidance.
Branching strategy matters less than commit discipline here. Treat every string addition or change as its own small commit, tied to a ticket or a clear message, so a sync bug six months later can be traced back to the exact change that caused it. Translation memory exchange, often via TMX, lets a reviewer’s approved edits flow back into future jobs instead of being re-translated from scratch, which is the collaboration pattern our translation memory guide walks through in more detail.
Managing pluralization and gender variations across platforms
Plural rules are not a two-case problem. English has “one” and “other,” but Arabic has six plural categories, Polish has four, and Russian has three, and a conversion pipeline that only checks for singular and plural forms will silently drop categories for half your locales.
ICU MessageFormat and each platform’s native plural syntax, like Android’s <plurals> resource or iOS’s .stringsdict, are built to express these categories, but they need to survive conversion intact. A validation step that checks for the correct plural category count per target locale, rather than assuming two forms everywhere, catches this before release.
Gender variation works similarly and is easy to underestimate in English-source content, where gendered forms rarely need marking. Languages like German, French, and Spanish often need a different adjective or article depending on grammatical gender, and a flat key-value translation file has no field for that distinction unless you build one in deliberately, usually as a separate key per gender variant or a structured placeholder your conversion script expands at build time. Treat both plural and gender handling as schema requirements, checked in the same validation pass as placeholder parity, rather than as edge cases to patch after a bug report arrives from a locale team.
Speed versus fidelity: what small teams should prioritize
Enterprises optimize for fidelity across dozens of locales and a large reviewer bench. Small teams rarely have that luxury, so automation and packaging matter more than process depth. Pick the canonical source, automate the conversion, and ship.
— Arkian
Automate your cross-platform localization with Arkian
The workflow this guide describes, canonical source, automated conversion, validation, packaging, is exactly what we built for small teams that do not want to run a full translation management system.

Our localization strings handling covers JSON, YAML, iOS, and Android output from a single source, our structured packaging delivers review-ready bundles without repository access, and our metadata production service adds the context translators need to work accurately. For teams that also need voice assets, our multilingual voice production service extends the same pipeline to TTS output.
- Review a sample package before committing to a workflow change.
- Check current plans on our pricing page, which includes the Arkian Membership at 19.00 CAD per month alongside the Starter, Pro, and Studio product lines.
- See the approach applied to a real release in our Quiet Harbour case study.
Request a demo or start with a small package to see how the pipeline fits your existing build process.
FAQ
Can I upload a file and have it translated?
Yes, most modern localization tools and platforms accept a file upload and return a translated version, though the format matters: JSON, XLIFF, and .strings files generally convert cleanly, while scanned PDFs need OCR first. Always validate the output for placeholder and plural integrity before using it in production.
What program can I use to translate PDF files?
For native, text-based PDFs, Google Cloud’s document translation service handles batch jobs up to 1GB or 100 million Unicode codepoints directly. Scanned PDFs need OCR preprocessing first, ideally with zone-based detection to preserve layout, since Google Cloud explicitly recommends native files over scanned ones to avoid layout loss.
How do I translate large files for free?
Open-source CLI tools like LocalizationManager let you validate, convert, and export translation files across formats including JSON, XLIFF, and Android XML without licensing cost, though you still need a translation source for the actual text. For large document batches, splitting files to stay under provider limits avoids failed jobs and timeouts.
What file format should I use for cross-platform translation?
Use XLIFF or TMX when handing files to a translation vendor or CAT tool, since both are built for lossless interchange, per Microsoft’s globalization documentation. For internal automation, a single JSON or YAML source that generates platform-native files tends to be simpler to maintain.
How do I keep translation files synchronized across iOS and Android?
Keep one canonical source file and generate both platforms’ outputs from it automatically, rather than editing each platform’s file by hand. An automated diff check that compares key sets across generated files catches drift before it reaches a release.
Sources
- Localization file formats - Globalization | Microsoft Learn
- XLIFF Version 2.2. Part 1: Core
- LocalizationManager (GitHub) — CLI/TUI tool for localization file management