NLP / ChatbotsProduction

ULAB

A platform for objectively benchmarking Uzbek-language AI models - 17 models, 1,377 test questions

17
AI models
1 377
Benchmark questions
72.3%
Top accuracy
Technology stack
PythonOpenAIAnthropicGroqWhisperDashboard
01

Overview

ULAB (Uzbek Language AI Benchmark) is an internal platform for objectively and consistently measuring how well AI models handle the Uzbek language. The bank planned to roll out an Uzbek-language AI assistant for internal departments and needed to know which model understands Uzbek best. ULAB compared 17 leading AI models across 1,377 questions via an interactive dashboard - moving model selection from opinion to hard numbers.

02

Scale

17
AI models
1 377
Benchmark questions
10
Task types
3
Language registers
72.3%
Top accuracy
03

Project goals

  • Build a ranking of Uzbek-language AI models for the bank.
  • Automatic, standardized scoring on 1,377 MCQ (A/B/C/D) questions.
  • Independent testing of Uzbek vendors (Kotib LLM, Muxlisa LLM).
  • An interactive dashboard and Excel reports for management.
  • Reusable infrastructure for future evaluations.
04

System modules

01

MCQ Benchmark

Core module: 1,377 questions, 17 models, automatic scoring and a results dashboard.

02

Dataset generation

A pipeline that auto-generates questions via gpt-4o-mini with quality filtering and deduplication.

03

Voice benchmark (ULAB-Voice)

Evaluates ASR (speech-to-text), TTS (text-to-speech) and speaker verification.

04

Vendor testing

Directly tests the Kotib and Muxlisa vendors and produces a comparison report.

05

Top model results

Kotib LLM v1 (local vendor)
72.8% - best; data stays in-country
Kimi K2.5 (open source)
72.3% - best open-source model
Mistral Large 2512 (commercial)
71.3% - best commercial model
Cogito 671B / Llama 4 Maverick
70.2% / 70.1% - strong open-source alternatives
06

Key findings

  • No model passed 75% - Uzbek remains an unsolved challenge for AI.
  • 13 of 14 models sit in 61-72% - so choose on price and latency.
  • Reading comprehension (RC) was easiest - nearly all models scored 90%+.
  • Fill-in and word meaning were hardest (32-42%) - weak deep semantics.
  • Kotib LLM, a local vendor, matched the best open-source models (72.8%).
07

Voice technologies (ULAB-Voice)

01

ASR - speech to text

Whisper and Kotib/uzbek_stt_v1 evaluated; CER/WER metrics. Kotib had the lowest error (WER 16.7%).

02

TTS - text to speech

Assessed via UTMOS, intelligibility WER and human MOS (30+ speakers on Toloka).

03

Speaker verification

Critical for biometric auth; ECAPA-TDNN baseline, target EER <= 2%.

08

Recommendations and next steps

  • Recommended models: Kotib LLM v1 (72.8%), Kimi K2.5 (72.3%), Mistral Large (71.3%).
  • Pilot the chosen model in 3 departments.
  • Explore fine-tuning for the bank's domain.
  • Re-run the benchmark quarterly (for new models).

Want to know more about this project?

Reach out to the SQB AI team - we'll share the details and show the platform in action.

Contact us