Speech-to-Text
Turkish speech recognition for real call conditions
Mank Speech-to-Text is developed for real call audio — noisy, compressed, accented and naturally paced — not for studio recordings. The model can be adapted with the organization's own audio data.
Why real call audio is harder than laboratory data
- Line compression and narrowband audio remove high-frequency information.
- Background noise, contact-centre acoustics and overlapping speech occur together.
- Accent, dialect and speaking-rate differences create a wide distribution.
- Natural speech is full of hesitation, self-correction and unfinished sentences.
- Institutional vocabulary, product names and abbreviations are absent from general models.
Comparative accuracy evidence
The table below shows results only once the methodology is publishable.
Benchmark evidence
Real Banking Call Center Data Benchmark- Test date
- August 2026
- SECTOR DATA
- Banking
- Dataset duration
- 25h
- Segment count
- 24938
| Model | WER (WORD ERROR RATE) |
|---|---|
| Mank STT | %19,42 |
| Deepgram Nova 3 | %25,12 |
| ElevenLabs Scribe v2 | %29,21 |
| AssemblyAI Universal 3 Pro | %31,11 |
| Speechmatics Enhanced | %39,27 |
| Qwen3 ASR 1.7B | %41,91 |
| Whisper Large v3 | %44,75 |
Results apply only to the stated test set, date and model version. Outcomes change with different audio quality, line conditions and speech types. These values are not a general performance commitment.
Use cases
Call transcription and searchable archives
Quality monitoring and mandatory-disclosure checks
Topic, intent and outcome labelling
Real-time transcription for Voice AI
Operational reporting and trend analysis
Let's measure on your own audio
The most meaningful comparison is a test on your own audio. We can plan a controlled evaluation on a limited data set.
