8 Steps to Localize Dialogue Scripts for Developers
8 Steps to Localize Dialogue Scripts for Developers

The fastest reliable way to localize dialogue scripts is to export each line with a stable ID and serve translations from keyed string tables at runtime. That single decision determines whether your localization pipeline scales or collapses under revisions. Start now by exporting your dialogue to a CSV, PO, or JSON file with a permanent ID attached to every line, before you send anything to a translator.
TL;DR:
- Using stable IDs attached to dialogue lines and exporting to CSV, PO, or JSON ensures scalable and accurate localization workflows.
- Unity and Unreal have specific tools for managing string tables and localization data, but the core pattern remains the same across engines.
- Proper workflow discipline includes locking the script before translation, maintaining consistent keys, and handling version control and incremental updates carefully.
- Context notes, idiom adaptation, and font coverage checks are critical to prevent mistranslations and display issues, especially for languages with longer or non-Latin scripts.
- Automating export, translation, and packaging processes with tools like Arkian reduces engineering overhead and improves turnaround for repeated updates.
Table of Contents
- Core pattern: keys, string tables, and the localization pipeline
- Engine-specific implementations (Unity, Unreal, Yarn Spinner): practical steps
- From authored script to localized runtime: a step-by-step workflow
- Automating localization for dialogue scripts
- Practical rules, gotchas, and testing tips specific to dialogue
- Starter checklist and minimal example
- Giving translators the context they need
- Handling cultural nuances and idiomatic expressions in dialogue
- Tools and best practices for quality assurance and testing of localized dialogue
- Version control strategies for dialogue scripts during localization
- Collaborating with translators and localization teams effectively
- Author perspective: trade-offs for small teams vs. large studios
- How Arkian can speed dialogue localization
- Sources
- FAQ
Core pattern: keys, string tables, and the localization pipeline
Every reliable localization system rests on the same architecture regardless of engine: a stable key identifies a line, a string table maps that key to text in each locale, and your runtime code looks up the key instead of hardcoding text. Break this pattern and you get duplicated work, mismatched updates, and lines that silently fail to translate.
Stable line IDs matter because dialogue text changes constantly during development. If your pipeline uses the English text itself as the lookup key, a single word edit orphans every existing translation. A numeric or slug-based ID (quest_003_greeting_01) survives text edits untouched, so translators only see genuinely new or changed content.
Unity’s Localization package implements this with String Tables, where each row maps a Key Id to locale-specific strings, and supports Smart Strings for pluralization and grammar-aware templating. Keeping your collections consistent (same keys across every locale table) prevents the classic bug where French has a line that Japanese never received.
Format choice affects your translator workflow more than most teams expect:
- CSV is the easiest format for translators to open in a spreadsheet and edit in bulk.
- PO carries context comments and plural forms natively, which suits professional translation tools.
- JSON fits custom dialogue systems and web-based tools but needs careful escaping for quotes and line breaks.
- LocRes is Unreal’s compiled runtime format, not something you hand-edit directly.
One overlooked risk: Unity does not preload String Tables by default, so the first access triggers an asynchronous load unless you configure preloading. A dialogue line that fires before that load completes can show a blank string or a loading flicker, which is worth testing explicitly rather than assuming.
Engine-specific implementations (Unity, Unreal, Yarn Spinner): practical steps
The canonical key-to-table pattern looks different in each engine’s tooling, but the underlying logic never changes.
Unity builds around String Table Collections. Each collection holds one table per locale, and each row is addressed by a Key Id. Practical steps:
- Create a String Table Collection per dialogue category (main quests, side quests, UI barks) to keep large games manageable.
- Use Smart Strings for any line with runtime-injected variables, plurals, or gender agreement, so translators work with structured placeholders instead of raw string concatenation.
- Bind LocalizedString entries to TextMeshPro components so text updates automatically when the player switches locale.
- Export a Character Set from your String Tables before recording or finalizing fonts. Unity can generate a Character Set file used to build font atlases, which matters for languages like Korean or Thai that your default font may not cover.
Unreal Engine organizes localization around Localization Targets rather than loose files. The Unreal localization pipeline gathers text into manifests and archives, exports PO files for translation, and compiles the results into LocRes files that the engine loads at runtime. Practical steps:
- Set up a Localization Target per project or DLC pack so text gathering stays scoped.
- Export PO files for your translation vendor, import the completed PO, then run the compile step to produce LocRes.
- For voiced dialogue, use Dialogue Waves to gather per-culture spoken audio, and export Dialogue Sheets as CSV during the pipeline for recording reference.
- Set up continuous localization (CI-triggered incremental exports) once your text volume grows, since incremental diffing reduces repeated translator work on iterative builds, though it requires keys that never change once assigned.
Yarn Spinner takes the lightest-weight approach of the three. It can automatically add stable line IDs to script lines and export the full dialogue as a CSV file for translation. Practical steps:
- Write your base language script first, then run Yarn Spinner’s line-tagging step so every line gets a permanent ID.
- Export the CSV, send it to translators, and reimport the completed file to populate your string tables.
- Decide early whether to use Yarn Spinner’s own localization system or hand the exported strings to Unity’s Localization package instead, since mixing both adds unnecessary complexity.
From authored script to localized runtime: a step-by-step workflow
Dialogue localization goes wrong most often because teams start translating before the script stabilizes. Lock a script draft, tag every line with an ID, and only then hand it to translators. Late script edits after that point should be treated as exceptions, not routine updates, since each one forces a partial retranslation and possibly a re-record.
A working pipeline generally follows this sequence:
- Freeze the script for the current milestone and confirm every line has a stable ID.
- Export the tagged lines to CSV, PO, or JSON depending on your engine and translator tooling.
- Translate the exported file, keeping placeholders and Smart String tokens intact.
- Import the translated file back into your string table collection or Yarn Spinner project.
- Compile or package the result (LocRes for Unreal, table collections for Unity).
- Run QA with pseudo-localization and in-context checks before locking further.
- Produce audio per locale once text is approved, using a consistent file naming convention tied to line IDs.
- Package the final assets for build integration.
For audio, name files after the line ID rather than the English text (quest_003_greeting_01_fr.wav), and place them in per-culture subfolders that mirror your exported dialogue sheet structure. This avoids renaming-induced reimports later, since a text-based filename breaks the moment a line changes.
Pro Tip: Run a pseudo-localization pass (expanding text length and swapping characters) before any real translation lands, so UI truncation bugs surface while they are still cheap to fix.
QA checks worth running on every localized batch include pseudo-localization for layout stress, in-context review inside the actual scene or UI, and subtitle timing checks against the recorded audio length.
Automating localization for dialogue scripts
Manual export-translate-import cycles work fine for a one-time release, but they break down once you ship updates every few weeks. Automation building blocks generally include:
- Extractors that pull new or changed lines automatically instead of re-exporting the entire script.
- Incremental export scripts that diff against the last export and only send what changed.
- TMS connectors for teams already running a translation management system.
- Packaging scripts that convert translated files back into the engine’s native format without manual steps.
Continuous localization triggered by CI reduces translator workload on iterative projects, though it depends entirely on keys that never get reassigned once a line ships. Break that rule once and your incremental diff becomes unreliable.
Small teams without dedicated localization engineers often find that maintaining extractors, TMS connectors, and packaging scripts costs more engineering time than the localization itself. This is where a platform like Arkian fits: it automates the generation of multilingual scripts and voice output, then validates and packages the results into delivery-ready files, without requiring repository access or a full TMS setup.
A tagged, ID-stable dialogue script is the input every localization pipeline depends on, whether you build the automation yourself or use a packaged service; the ID discipline described in Unity’s and Unreal’s own documentation is what makes either path work.
Choose manual workflows when your dialogue volume is small and stable. Choose automation once you are shipping content updates on a recurring schedule and the manual cycle starts eating more time than the localization work itself.
Practical rules, gotchas, and testing tips specific to dialogue
Dialogue localization fails in predictable ways. A few rules prevent most of them.
- Never inline variables into raw strings; use tokenized or Smart String templates so translators can reorder placeholders for languages with different word order.
- Treat pluralization, gender, and conditional phrasing as structured template features, not string concatenation hacks.
- Watch for stale translations: any edit to source text should flag the corresponding translated lines as needing review, not silently keep the old translation attached to a new ID.
- Budget for text expansion. German and Finnish routinely run longer than English, and CJK languages need font coverage your default typeface may lack.
- Export your character set and confirm font coverage before recording localized audio, since a missing glyph is cheap to fix in text and expensive to fix after a voice session is booked.
Pro Tip: Record a short pilot batch of voiced lines in your most complex target language before committing to a full recording schedule. It surfaces font, timing, and phrasing problems while they are still isolated to a handful of lines.
Voiced dialogue costs meaningfully more than text-only translation in both time and money, since it adds casting, direction, and audio QA on top of the translation step itself.
Starter checklist and minimal example
Before sending anything to a translator, confirm the following:
- Every dialogue line has a permanent, human-readable ID.
- Lines are exported to CSV, PO, or JSON with ID, source text, and a context column.
- A String Table Collection (or equivalent) exists per language and stays in sync across locales.
- Fonts are checked against your target character sets before recording.
- Native audio recording is scheduled only after text is locked.
- In-editor locale switching works for quick spot checks.
A minimal CSV layout needs four columns: ID (the stable key), Source (original text), Context (a short note on speaker, scene, or tone), and Translation (left blank for the translator). PO files add plural-form fields natively; JSON needs an explicit structure per key.
Test early using an in-editor locale toggle and a pseudo-localization pass, both of which catch truncation and missing-string bugs long before a real translation arrives.
Giving translators the context they need
A line of dialogue without context is a guess. “Right” could mean a direction, an agreement, or a correction, and a translator working from a bare spreadsheet row has no way to know which. The context column in your export file is not optional polish, it is the difference between a usable translation and a rewrite.
Useful context includes the speaking character, the scene or quest, the emotional tone, and any preceding line the dialogue responds to. Screenshots or short video clips attached to a batch help enormously for humor, sarcasm, or physical comedy that text alone cannot convey.

Smart String and PO comment fields exist specifically for this. Rather than writing a note in a separate document that a translator may never see, attach the context directly to the row or key so it travels with the line through every export and reimport cycle. Teams that skip this step tend to see the same category of error repeatedly: mistranslated pronouns, wrong formality register, and jokes translated literally instead of adapted.
Provide a glossary for recurring terms, character names, and invented vocabulary. A fantasy game with a made-up currency or title needs that term defined once, not reinterpreted differently by every translator who touches the file.
Handling cultural nuances and idiomatic expressions in dialogue
Direct translation breaks idioms constantly. A line built around an English idiom rarely has a literal equivalent in another language, and forcing one produces dialogue that sounds like it was written by a machine rather than a character.
The fix is adaptation, not translation. Give translators permission to replace an idiom with an equivalent one that carries the same tone and meaning in their language, even if the literal words differ completely. This is where context notes matter again: a translator who knows a line is meant to sound sarcastic can choose a sarcastic idiom native to their language instead of translating the English one word for word.
Humor, insults, and slang age quickly and vary sharply by region even within the same language. Spanish in Spain and Spanish in Mexico can require different phrasing for the same joke to land, and a translator working without that regional flag may default to a neutral phrasing that lands flatly everywhere.
Cultural references (holidays, foods, historical figures) sometimes need substitution rather than translation if the reference means nothing to the target audience. This is a judgment call best left to a native-speaking translator working with clear context, not a rule you can encode into a string table.
Tools and best practices for quality assurance and testing of localized dialogue
Localized dialogue needs QA at three separate points: the text itself, the text in context, and the recorded audio if voiced.
Pseudo-localization is the cheapest first check. It artificially expands string length and swaps in accented or wide characters to reveal UI truncation and font coverage problems before a single real translation exists.
In-context review means viewing the translated line inside the actual scene, dialogue box, or subtitle overlay rather than in a spreadsheet. A translation that reads correctly in isolation can still overflow a dialogue box, clash with an on-screen timer, or contradict a visual the translator never saw.
Subtitle timing checks matter specifically for voiced dialogue. Translated text is often longer or shorter than the source, so subtitles need independent timing validation rather than assuming the original timing still fits.
A practical QA pass, in order:
- Run pseudo-localization on all new and changed strings.
- Review flagged strings in context inside the actual build.
- Check subtitle timing against recorded audio for any voiced line.
- Confirm font rendering across every target locale, especially for CJK and right-to-left languages.
Linguistic review by a native speaker remains the step that catches tone and idiom problems no automated check can find. Budget for at least one review pass per locale before a line ships, not just before the final release.
Version control strategies for dialogue scripts during localization
Dialogue scripts change constantly during development, and localization files need a version control strategy that survives that churn without losing translator work.
Keep source dialogue and translated files in the same repository structure when possible, with source text and each locale’s translation as parallel files sharing the same line IDs. This makes diffing straightforward: a change to the English file shows exactly which lines need translator attention, without guesswork.
Tag or branch localization exports at each milestone (alpha, beta, release candidate) so you can trace which translation batch corresponds to which build. This matters most when a bug report references a specific locale and you need to know which script version was live at the time.
Avoid editing translated files by hand outside the import/export cycle. A manual edit to a French string table that never round-trips through your translation file will get silently overwritten the next time you import, and nobody will notice until a tester flags the regression.
For teams without deep repository access or a formal branching strategy, exporting locked snapshots at each milestone and archiving them alongside the build is a lower-overhead alternative that still gives you a traceable history.
Collaborating with translators and localization teams effectively
The teams that get the best translations treat translators as collaborators with questions, not a queue that returns finished files. Build in a channel for translators to flag ambiguous lines and get an answer within the same day, rather than guessing and moving on.
Batch your exports predictably. Sending small, frequent batches with no pattern makes it hard for a translator to build context across a script, while a single enormous batch at the end of development leaves no time for revisions. A milestone-based export schedule, tied to your script freeze points, gives translators a rhythm to work with.
Share the glossary and style guide once, early, and keep it updated rather than repeating context in every batch. Recurring character names, invented terms, and tone guidelines belong in a living reference document translators can check independently.
Give translators visibility into how their work is going to be used when possible. Screenshots, short clips, or even a build they can play through catches far more tone and context errors than a spreadsheet ever will, and it tends to build the kind of working relationship that produces better results on the next batch rather than just the current one.
Author perspective: trade-offs for small teams vs. large studios
Small teams almost always lose time to plumbing rather than translation itself: building extractors, wiring a TMS, writing packaging scripts. Automated packaging that validates and delivers ready-to-use files removes that overhead without giving up editorial control over the source script.
Large studios with dedicated localization staff often justify full TMS integration and staged QA cycles across multiple vendors, since their volume and headcount make that investment pay off.
The practical middle ground I would recommend to most small teams: automate the mechanical parts of the pipeline (export, translate, package) and keep human review concentrated where it actually matters, on context, tone, and the final in-context check.
— Arkian
How Arkian can speed dialogue localization
Small teams rarely need a full translation management system to ship dialogue in multiple languages, they need the export-translate-package cycle to stop eating engineering time. Arkian automates that cycle directly: it generates multilingual scripts, produces voice output, and validates everything into delivery-ready packages without requiring repository access.

Relevant services for dialogue-heavy projects:
- Localization strings for turning exported dialogue into translated, structured string files.
- Multilingual voice production for teams that need voiced dialogue generated and packaged per locale.
- Structured packaging for delivering validated language files in the format your engine expects.
This fits best when your team wants packaged, review-ready output without setting up a TMS or granting repository access to a translation vendor. Check Arkian’s pricing to see current plans, including the Starter, Pro, and Studio product lines and the Membership option, and find the one that matches your release schedule.
Sources
- String Tables | Localization | 1.5.11
- Localization Overview — Unreal Engine 4.27
- Yarn Spinner — Assets and localization
FAQ
What is a script dialogue?
A script dialogue is the written exchange of lines spoken between characters, formatted so a game or application can display or voice each line in sequence. In localization work, each line typically gets a stable identifier so it can be tracked, exported, and translated independently of the surrounding text.
What does “localized language” mean?
A localized language is content adapted for a specific region’s language, cultural conventions, and expectations rather than simply translated word for word. This includes phrasing, idioms, formality, date and number formats, and sometimes cultural references that are swapped for locally relevant equivalents.
What are the four types of dialogue?
Common categorizations describe narrative dialogue, expository dialogue, character-driven dialogue, and functional or system dialogue (menus, prompts, tutorials), though the exact taxonomy varies by source. For localization purposes, the more useful distinction is between scripted narrative lines and dynamic system text, since each needs different handling in your string tables.
What is an example of localization?
Adapting a game’s dialogue, menus, and voice lines from English into French, complete with region-appropriate idioms, reformatted dates, and re-recorded audio, is a typical example of localization. Platforms such as Arkian automate parts of this process for small teams by generating translated scripts and voice output and packaging them for delivery.
Which file formats are commonly used to export dialogue for translation?
CSV, PO, and JSON are the most common formats for exporting dialogue lines to translators, while Unreal Engine compiles translations into its own LocRes format for runtime use. Yarn Spinner exports dialogue as CSV with stable line IDs attached to each row.