AI technical translation QA
Check protected identifiers, numeric values and approval language between source text and a translation before relying on fluent wording.
Synthetic evaluation input
Source EN: Sample for tag V-104. Recorded value 310.5. Draft, not approved. Target DE: Muster fuer Kennzeichen V-140. Erfasster Wert 310,5. Entwurf, nicht freigegeben. Compare identifiers exactly; compare decimal notation by numeric value.
Original Inzonex training example, version 1.0. No customer document or real equipment data. No model has been scored on this page.
Expected fields
| Field | Reference answer |
|---|---|
| source_tag | "V-104" |
| target_tag | "V-140" |
| tag_match | false |
| numeric_match | true |
| approval_status | "not_approved" |
Check a structured response
The comparison runs in your browser against this answer key. It checks exact fields and values, not the accuracy of an explanation or a model's general ability.
Recommended workflow
Separate language review from invariant checks. Compare protected tags and numbers explicitly, then have a qualified reader assess technical meaning.
Technical translation has details that ordinary fluency scoring can miss. An equipment tag should not be translated, decimal separators can change with locale, and an approval qualifier should not disappear. These are useful independent checks even when the prose reads naturally.
The synthetic pair intentionally changes one digit in an equipment identifier. Its decimal notation changes from a point to a comma without changing the numeric value. A correct comparison flags the identifier mismatch but does not call the decimal convention a numerical error.
This packet does not grade translation quality. It evaluates a small set of invariants. Terminology, negation, procedural meaning and regional conventions still require competent language and subject-matter review.
Failure checks
- Flag V-104 versus V-140.
- Recognise 310.5 and 310,5 as the same numeric value in this packet.
- Retain the not-approved qualifier.
What this exercise does not prove
Not a translation certification or a broad language benchmark. The example is intentionally short and uses ASCII transliteration for German umlauts.
For an actual evaluation, keep a separate held-out set, record tool/model version and settings, and log raw outputs, corrections, elapsed time and actual charges. Do not compare tools tested on different inputs as though they ran the same benchmark.
What is checked in AI technical translation QA?
Check protected identifiers, numeric values and approval language between source text and a translation before relying on fluent wording.
What does AI technical translation QA not prove?
Not a translation certification or a broad language benchmark. The example is intentionally short and uses ASCII transliteration for German umlauts.
Related tasks
Match this workflow to your data requirements
Methodology and reuse
These packets and answer keys are original Inzonex educational material, licensed CC BY 4.0. Attribute Inzonex and link to this task page when reusing the packet. The licence does not cover third-party material linked from this site.
NIST AI 600-1: background on generative AI evaluation and risk. This exercise is not NIST-certified.