Skip to content

Speech-to-Text

Turkish speech recognition for real call conditions

Mank Speech-to-Text is developed for real call audio — noisy, compressed, accented and naturally paced — not for studio recordings. The model can be adapted with the organization's own audio data.

Why real call audio is harder than laboratory data

  • Line compression and narrowband audio remove high-frequency information.
  • Background noise, contact-centre acoustics and overlapping speech occur together.
  • Accent, dialect and speaking-rate differences create a wide distribution.
  • Natural speech is full of hesitation, self-correction and unfinished sentences.
  • Institutional vocabulary, product names and abbreviations are absent from general models.

Comparative accuracy evidence

The table below shows results only once the methodology is publishable.

Benchmark evidence

Real Banking Call Center Data Benchmark
Test date
August 2026
SECTOR DATA
Banking
Dataset duration
25h
Segment count
24938
Benchmark evidence
ModelWER (WORD ERROR RATE)
Mank STT%19,42
Deepgram Nova 3%25,12
ElevenLabs Scribe v2%29,21
AssemblyAI Universal 3 Pro%31,11
Speechmatics Enhanced%39,27
Qwen3 ASR 1.7B%41,91
Whisper Large v3%44,75

Results apply only to the stated test set, date and model version. Outcomes change with different audio quality, line conditions and speech types. These values are not a general performance commitment.

Use cases

Call transcription and searchable archives
Quality monitoring and mandatory-disclosure checks
Topic, intent and outcome labelling
Real-time transcription for Voice AI
Operational reporting and trend analysis

Let's measure on your own audio

The most meaningful comparison is a test on your own audio. We can plan a controlled evaluation on a limited data set.

Try Now