ULAB
A platform for objectively benchmarking Uzbek-language AI models - 17 models, 1,377 test questions
Overview
ULAB (Uzbek Language AI Benchmark) is an internal platform for objectively and consistently measuring how well AI models handle the Uzbek language. The bank planned to roll out an Uzbek-language AI assistant for internal departments and needed to know which model understands Uzbek best. ULAB compared 17 leading AI models across 1,377 questions via an interactive dashboard - moving model selection from opinion to hard numbers.
Scale
Project goals
- Build a ranking of Uzbek-language AI models for the bank.
- Automatic, standardized scoring on 1,377 MCQ (A/B/C/D) questions.
- Independent testing of Uzbek vendors (Kotib LLM, Muxlisa LLM).
- An interactive dashboard and Excel reports for management.
- Reusable infrastructure for future evaluations.
System modules
MCQ Benchmark
Core module: 1,377 questions, 17 models, automatic scoring and a results dashboard.
Dataset generation
A pipeline that auto-generates questions via gpt-4o-mini with quality filtering and deduplication.
Voice benchmark (ULAB-Voice)
Evaluates ASR (speech-to-text), TTS (text-to-speech) and speaker verification.
Vendor testing
Directly tests the Kotib and Muxlisa vendors and produces a comparison report.
Top model results
- Kotib LLM v1 (local vendor)
- 72.8% - best; data stays in-country
- Kimi K2.5 (open source)
- 72.3% - best open-source model
- Mistral Large 2512 (commercial)
- 71.3% - best commercial model
- Cogito 671B / Llama 4 Maverick
- 70.2% / 70.1% - strong open-source alternatives
Key findings
- No model passed 75% - Uzbek remains an unsolved challenge for AI.
- 13 of 14 models sit in 61-72% - so choose on price and latency.
- Reading comprehension (RC) was easiest - nearly all models scored 90%+.
- Fill-in and word meaning were hardest (32-42%) - weak deep semantics.
- Kotib LLM, a local vendor, matched the best open-source models (72.8%).
Voice technologies (ULAB-Voice)
ASR - speech to text
Whisper and Kotib/uzbek_stt_v1 evaluated; CER/WER metrics. Kotib had the lowest error (WER 16.7%).
TTS - text to speech
Assessed via UTMOS, intelligibility WER and human MOS (30+ speakers on Toloka).
Speaker verification
Critical for biometric auth; ECAPA-TDNN baseline, target EER <= 2%.
Recommendations and next steps
- Recommended models: Kotib LLM v1 (72.8%), Kimi K2.5 (72.3%), Mistral Large (71.3%).
- Pilot the chosen model in 3 departments.
- Explore fine-tuning for the bank's domain.
- Re-run the benchmark quarterly (for new models).
Want to know more about this project?
Reach out to the SQB AI team - we'll share the details and show the platform in action.
Contact us