What breaks when the user is not speaking English?
Language
Dialects, code-switching, right-to-left layout and text-to-speech, with Arabic as the case we have run in production.
6 posts
Start here
How to evaluate dialect coverage in a speech stack
A vendor language list is a claim about a corpus, not about your users. The method for measuring what a speech model does on the varieties they actually speak.
- When your user changes language mid-sentenceCode-switching is the normal way bilingual people talk, and it breaks pipelines that pick one language per utterance. What fails, and how to test for it.Language9 min
- Shipping speech for a variety with less training dataHow to measure recognition coverage for a language variety the big corpora barely contain, using Gulf Arabic as the case where the gap is documented.Language10 min
- Why some varieties of a language get recognised and others do notEgyptian Arabic is the best-served spoken variety of Arabic, and the reasons are historical rather than linguistic. What that predicts for every other language.Language8 min
- Text-to-speech for agents when the language is hardSynthesis quality is decided upstream of the model that makes the sound. The two stages that break on a hard language, with Arabic as the worked example.Language9 min
- Running ten dialects of one language through one pipelineMost speech stacks assume one user speaks one language. Production breaks that in week one, and the fixes are the same whichever language you start from.Language10 min
