Translation direction, register, audience, and domain are separate variables. A short MSA email, a Gulf customer-service transcript, a contract, and a classical text should not share one score. Public model specifications can identify candidates, but they do not establish dialect quality or a winner for your material.
Editorial candidates to test
The shortlist below is a set of editorial hypotheses, not a published comparative result. Hosted frontier models are candidates when their documented context and workflow features fit the job. Open-weight Qwen3.6 checkpoints (Apache-2.0) and Llama 4 are additional candidates when open weights, local hosting, licensing, or runtime control matter. A result reported for one translation-tuned variant cannot be transferred to an entire model family.
The September 2026 list names the current release in each family. Claude Opus 5 replaced Opus 4.8 on July 24, 2026 at the same $5 / $25 per 1M tokens, and Anthropic's model guide now tells most users to start with Opus 5. GPT-5.6 Sol is newer than GPT-5.5 and cheaper on both sides: $4 input / $20 output per 1M tokens for GPT-5.6 Sol, against $5 / $30 for GPT-5.5. Gemini 3.1 Pro is still Google's newest Pro model, and Google still labels it Preview. Meta has not released a Llama 5, so Llama 4 remains the current open-weight Llama.
| Candidate | Documented reason to shortlist | What the local test must decide |
|---|---|---|
| Claude Opus 5 | Hosted model with a documented 1M-token context workflow | Meaning, register, terminology, names, consistency, latency, and cost on your files |
| GPT-5.6 Sol | Hosted model and API workflow documented by its provider | Both directions and every required regional register |
| Gemini 3.1 Pro | Hosted model with documented long-context and Google workflow options | Instruction adherence, source coverage, register, and current plan fit |
| Qwen3.8-Max, Qwen3.6 open weights, or an exact translation variant | Hosted API, Apache-2.0 open-weight, or translation-specific deployment may offer more runtime control | The exact checkpoint, tokenizer, serving stack, and license you will deploy |
| Llama 4 / exact checkpoint | Open-weight deployment may fit local-hosting or governance requirements | The exact checkpoint and runtime; no family-wide Arabic rank is assumed |
A reproducible blind evaluation
- Freeze the setup. Record the exact model ID, date, system prompt, temperature, tools, glossary, and output constraints. Use the same instructions and source text for every candidate.
- Stratify the set. Include both translation directions and the registers you will ship: MSA, Gulf, Egyptian, Levantine, Maghrebi, code-switched text, legal terminology, names and transliteration, and classical or religious material when relevant.
- Blind the review. Randomize and relabel outputs so native reviewers do not know the provider. Use more than one reviewer for subjective register judgments.
- Score defined criteria. Measure preservation of meaning, omissions and additions, terminology, register, names, formatting, and reviewer preference. Track disagreement instead of collapsing it into an unsupported winner claim.
- Repeat and report. Rerun a sample to check stability. Publish the test set, prompt, model IDs, scoring rubric, mean scores, reviewer disagreement, latency, and cost separately.
Required stress tests
State the intended audience, country, register, glossary, and transliteration convention explicitly. For legal, medical, religious, or other consequential text, model preference is not a substitute for domain review. If a required dialect has little written standardization, document reviewer disagreement and test real examples from that audience instead of extrapolating from MSA benchmarks.
Context-window size is also not evidence of long-document translation quality. Test full-document terminology, cross-reference preservation, omissions, and recovery from truncation on representative files. If the document exceeds a system's practical limits, use a controlled segmentation and glossary process and verify consistency after recombination.
Calculate your cost →·Compare this model →·Find your model →