Topic
Evaluation
Testing that an agent does the right thing often enough to ship.
7 posts
- Does your agent work eight times out of eight?Average accuracy is the wrong number for a product. Run the same task eight times and count how often it worked every single time.Safety11 min
- How to evaluate dialect coverage in a speech stackA vendor language list is a claim about a corpus, not about your users. The method for measuring what a speech model does on the varieties they actually speak.Language9 min
- Why some varieties of a language get recognised and others do notEgyptian Arabic is the best-served spoken variety of Arabic, and the reasons are historical rather than linguistic. What that predicts for every other language.Language8 min
- How far you actually get in a day with an in-app agentA working spoken turn takes an afternoon. This is what those hours buy, what week two costs, and why the integration is not on the critical path at all.Integration9 min
- The metrics that tell you an in-app agent worksSeven numbers worth tracking, what each one hides, and why containment is the one that looks best while telling you least about the feature.Business9 min
- Shipping speech for a variety with less training dataHow to measure recognition coverage for a language variety the big corpora barely contain, using Gulf Arabic as the case where the gap is documented.Language10 min
- When your user changes language mid-sentenceCode-switching is the normal way bilingual people talk, and it breaks pipelines that pick one language per utterance. What fails, and how to test for it.Language9 min
