Language Mixing
One conversation moving naturally between multiple languages.
Sahara v2.5 understands conversations as they actually happen across African languages, accents and natural switches from one language to another.
One conversation moving naturally between multiple languages.
Speech recognition across languages, accents and regional dialects.
Real-time speech intelligence built for natural African conversations.
People often mix languages in one breath: a Kano trader switches between Hausa and English; a Nairobi patient uses Swahili and English; a Lagos colleague blends Yoruba and Pidgin. Global models see this as an error, but Sahara-v2.5 understands it. It’s the first bilingual language-mixing support across 12 African languages, expanding the Swahili-English feature from v2 to the whole continent.
Sahara v2.5 expands bilingual speech intelligence with 12 models built for conversations that naturally move between languages.
Sahara v2.5 can follow one conversation as it moves naturally across English, Kinyarwanda and French without resetting context.
Sahara v2.5 extends from understanding African speech to generating natural voice responses across 13 supported languages.
Seven additional languages expand Sahara v2.5 from 56 to 63 total languages, bringing African voice AI to more speakers, markets and real-world use cases.
Sahara v2.5 improves recognition quality for real-world African speech, helping voice systems understand conversations more reliably across practical environments.
From collections and healthcare to courtrooms and clinical documentation, Sahara is already delivering measurable outcomes across real African workflows.
BRANCH
Sahara-powered Branch collections agents recovered more than ₦1.2 million in delinquent loans in one week, with record after-hours and weekend repayments, outperforming human agents on delinquent loans over 356 days.
BRANCH
Collaborating with Intron to build a Branch-aligned collections bot was a rewarding experience. Customers engaged naturally even after hours and on weekends. Conversations were human and effective at delivering real payments on delinquent loans, sometimes over 2 years old. Adding local language and language-mixing support in this v2.5 release helps us reach even more customers.
AUDERE
Youths across South Africa can now speak naturally with Audere’s Self-Cav reproductive health WhatsApp chatbot using voice to express themselves freely instead of being confined to text chat.
ARM
Using Intron AI models, we’ve seen significant improvement in transcription and summaries compared to models we previously explored. Their systems capture context and nuance better, leading to more accurate results.
OGUN STATEOver one year working with Intron, deployment has grown from an initial 3 pilot courts to 9 courts, with the goal of bringing all 18 courts across the state on board. Yoruba is frequently spoken during testimony, and Sahara v2.5 helps ensure those segments are properly captured in court records.
NAIROBIA 14-minute Swahili-English doctor-patient consultation in Nairobi is converted into a structured clinical note in less than 30 seconds, reducing documentation burden so physicians can focus more on patient care.
PAMOInternet connectivity had blocked the adoption of many transformative solutions. Intron’s offline medical scribe deployment has transformed clinical operations. At over 2,000 consultations using voice AI and a 4× documentation speed improvement, the team is now benefiting from AI without depending on reliable internet.
GOOEY.AI
On 4 of 6 Nigerian languages, independent Medical Question Answering benchmarks for the Gates Foundation and CLEAR Global showed Intron to be on par with or better than transcription alternatives including Gemini and Meta’s Omnilingual LLM. Sahara v2.5 is one of the first dedicated models for African code-switching they have evaluated.
Five teams. Four sectors. One challenge: build voice-first products that work for how Africa actually speaks.
Explore the challenge ↗Voice bookkeeping for African market traders who run their businesses by talking instead of typing.
SautiLedger is a voice bookkeeping app for African market traders who run their business by talking, not typing. A trader just says her sale out loud, in whatever mix of Pidgin, Yoruba, and English feels natural, and the app logs it, reads it back, and can tell her how much she made that day. It's built with one hard rule: it never guesses when money is involved — if anything is unclear, it asks instead of assuming.
Sahara handles the speech-to-text step that turns a trader's spoken sale into text the app can safely act on.
Sahara roughly halved Whisper's error rate on real market speech and led by an even wider margin on wild code-switched audio. It also produced far more usable, correctly parsed transactions than Whisper did — 47% exact versus 13–27%.
Multilingual hospital triage that understands patients and speaks responses back in their language.
“Nina maumivu ya kifua…”
Sauti Yetu helps hospitals communicate with patients who don't speak the local language, right at triage, when it matters most. A patient speaks freely in their own language, and the system figures out what's wrong, how urgent it is, and which department they need, then fills out the intake form automatically. It even speaks the health worker's reply back to the patient in their own language.
Sahara transcribes what the patient says and pulls out the clinical details in English for the health worker to read, then Sahara TTS speaks the worker's reply back to the patient in their own language — a genuine closed loop on both ends of the conversation.
Across 27 real clinical-speech clips, Sahara's error rate was 12.6%, roughly a third of Whisper's and Meta MMS's 47–52% — the clearest gap of any benchmark in this set.
Students explain their reasoning naturally while the system grades their thinking rather than their accent.
“The answer changes because the rate increased after…”
GradrAI tackles a real problem in African schools: students often think and explain best in a mix of local languages and English, but exams only reward English fluency. GradrAI lets a student answer a question, then explain their reasoning out loud in whatever language mix feels natural. It grades the actual thinking, not the accent or the language used to say it.
Sahara transcribes the student's spoken reasoning, which then gets graded on content only.
On genuinely code-switched speech, Sahara had the lowest error rate of the three models tested — 0.33 versus Gemini's 0.45 — though Gemini was more reliable on single-language audio. No one model won everywhere, which the team was upfront about.
Smallholder farmers control irrigation, monitoring and farm systems using natural voice commands.
SAOT — Sustainable Agriculture Optimization Tool — lets smallholder farmers in Nigeria control real farm hardware, irrigation, intrusion detection, and weather monitoring just by speaking naturally in code-switched Yoruba-English or Pidgin. A farmer says something like “start irrigation for my farm,” and the system acts on it directly — no app, no typing, no English required.
Sahara replaced the team's existing speech-recognition layer, letting them isolate exactly how much better transcription quality improved farm-command accuracy.
Swapping in Sahara cut transcription errors by 73% compared to their previous setup and pushed successful command execution from 94% to 98%, with the biggest gains on Yoruba-English.
Rain expected this evening. Best planting window: tomorrow morning.
Conversational bookkeeping for market traders, street vendors and small shop owners.
Koko Business AI is a voice bookkeeping app for African market traders, street vendors, and small shop owners who don't have time to type out every sale. A trader taps the mic and just talks. Koko figures out what was sold, confirms it with the trader, and updates a live dashboard of sales, expenses, and profit automatically. No forms, no typing — just a normal conversation.
Sahara transcribes the trader's voice, tuned for African accents and multilingual, code-switched speech, before the app extracts the transaction details.
Tested on 20 real code-switched clips, Sahara had the lowest average error rate and outperformed the other two models on more than half the clips outright.
Preliminary benchmark data across code-switched ASR, monolingual African-language ASR, text-to-speech and accented English speech recognition.
Metric: Word Error Rate (WER). Lower is better.
| Language | Sahara 2.5 | Sahara 2.1 | Omni ASR 7B LLM | ElevenLabs | Gemini 3.6 Flash |
|---|---|---|---|---|---|
| Pidgin | 27.28% | 42.19% | 32.52% | 42.71% | 29.32% |
| Swahili | 31.11% | 43.27% | 35.08% | 39.03% | 32.38% |
| Afrikaans | 26.45% | 59.14% | 30.43% | 26.04% | 36.90% |
| Hausa | 28.08% | 45.56% | 33.93% | 41.01% | 41.71% |
| Amharic | 17.86% | 50.35% | 52.60% | 51.80% | 44.44% |
| Kinyarwanda | 25.31% | 56.34% | 39.13% | 46.58% | 60.68% |
| Yoruba | 29.12% | 62.23% | 42.71% | 51.05% | 51.81% |
| Zulu | 41.28% | 81.01% | 43.70% | 62.02% | 48.04% |
| Igbo | 35.17% | 79.62% | 69.92% | 60.98% | 72.58% |
| Luganda | 43.87% | 89.95% | 85.33% | 90.73% | 95.97% |
Metric: Word Error Rate (WER). Lower is better.
| Language | Sahara V2.5 | Omnilingual LLM 7B | Eleven Labs | Gemini 3.6 | Sahara V2 |
|---|---|---|---|---|---|
| ALL LANGS AVERAGE | 29.00% | 34.64% | 42.91% | 54.82% | 26.76% |
| Swahili | 10.53% | 10.11% | 13.38% | 21.75% | 14.16% |
| Pidgin | 14.40% | 17.54% | 17.02% | 11.69% | 18.21% |
| Kinyarwanda | 10.49% | 13.49% | 31.31% | 45.80% | 10.14% |
| Zulu | 14.62% | 14.96% | 35.76% | 35.06% | 16.01% |
| Afrikaans | 11.50% | 16.68% | 29.51% | 39.09% | 22.10% |
| Hausa | 21.72% | 24.74% | 24.80% | 29.51% | 28.18% |
| Luganda | 17.08% | 17.24% | 24.47% | 65.92% | 19.43% |
| Kikuyu | 29.48% | — | — | — | — |
| Amharic | 20.45% | 21.35% | 42.84% | 41.25% | 24.85% |
| Dholuo | 34.38% | 42.61% | — | — | — |
| Tigrinya | 31.24% | 48.67% | — | — | — |
| Yoruba | 31.71% | 33.89% | 32.58% | 69.70% | 34.56% |
| Akan | 28.27% | 38.98% | 60.41% | 56.41% | 27.97% |
| Igbo | 37.28% | 44.32% | 20.13% | 72.23% | 47.25% |
| Wolof | 30.90% | 32.57% | 63.61% | 65.58% | 37.82% |
| Fulfulde | 34.12% | 42.87% | 53.82% | 71.78% | 47.14% |
| Nupe | 69.02% | 83.32% | 94.19% | 101.31% | — |
| Kanuri | 74.72% | 85.60% | 99.87% | 95.17% | — |
Blank cells indicate that no value was supplied for that model-language pair.
Metric: Mean Opinion Score (MOS). Higher is better.
| Language | Human Audio | Sahara-TTS-v2.5 | Sahara-TTS-v2.4 | Gemini | ElevenLabs | OmniVoice TTS |
|---|---|---|---|---|---|---|
| Kinyarwanda | 4.848n=120 | 4.670n=100 | 4.348n=120 | 2.655n=120 | — | 4.138n=120 |
| Hausa | 4.636n=297 | 4.448n=249 | 4.273n=281 | 3.798n=119 | 2.744n=111 | 2.862n=259 |
| Amharic | 3.717n=392 | 3.698n=378 | 3.675n=399 | 3.475n=212 | — | 3.045n=389 |
| Yoruba | 4.138n=376 | 4.033n=369 | 3.901n=384 | 2.806n=185 | — | 2.772n=385 |
| Igbo | 3.882n=217 | 4.105n=197 | 4.055n=217 | 1.680n=20 | — | 3.044n=219 |
| Pidgin | 4.210n=383 | 4.259n=387 | 3.695n=380 | 4.438n=198 | — | 3.444n=378 |
| Oromo | 3.510n=184 | 3.765n=179 | 3.674n=193 | 3.957n=47 | — | 3.654n=216 |
| Shona | 4.037n=27 | 3.388n=16 | 3.600n=24 | 2.191n=23 | — | 2.487n=23 |
| Swahili | 3.565n=299 | 3.631n=282 | 3.715n=305 | 3.704n=112 | 3.826n=96 | 3.601n=309 |
| Accented English | 4.007n=540 | 4.028n=559 | 3.921n=544 | 4.119n=186 | 4.124n=184 | 3.925n=558 |
n = number of audio samples. Blank scores are shown as em dashes.
Metric: Word Error Rate (WER). Lower is better.
| Category | Sahara V2 | Sahara V2.5 | Azure Speech Recognition | deepgram nova v3 medical | Google Gemini 2.5 flash | OpenAI gpt-4o-mini-transcribe | OpenAI Whisper-large-v3 | Meta Omni CTC 7B V2 |
|---|---|---|---|---|---|---|---|---|
| Names | 10.16% | 11.87% | 45.29% | 50.27% | 52.16% | 62.00% | 52.34% | 43.95% |
| In-the-Wild | — | — | — | — | — | — | — | — |
| Finance | 6.09% | 6.67% | 22.93% | 53.31% | 26.66% | 36.39% | 24.84% | 37.34% |
| Medical | 14.44% | 13.33% | 24.30% | 21.77% | 24.20% | 27.21% | 26.21% | 46.95% |
| Call Center | 15.36% | 13.96% | 24.95% | 23.50% | 23.41% | 23.90% | 24.69% | 57.63% |
| Legal | 14.02% | 14.34% | 25.04% | 26.52% | 21.91% | 41.98% | 25.34% | 46.78% |
| General | 8.80% | 9.73% | 15.17% | 18.09% | 15.74% | 17.23% | 14.21% | 20.82% |
| Robustness | 7.85% | 7.15% | 32.61% | 1.96% | 32.04% | 37.44% | 115.91% | 70.41% |
| AVERAGE | 10.96% | 11.01% | 27.18% | 27.92% | 28.02% | 35.16% | 40.51% | 46.27% |
The “In-the-Wild” row is blank in the supplied workbook and is kept blank here.
Deploy voice bots that speak the language of your customers, reducing call centre volume and improving response times.
Explore Fintech ↗Medical dictation designed to understand clinical terminology across regional accents.
Enable customers to check balances, confirm transactions and navigate banking services naturally by voice.
Explore Voice Banking ↗Convert hearings, interviews, depositions and legal proceedings into searchable, structured transcripts.
Please state your name for the record.
My name is...
And where were you at the time?
Bridging the information gap for smallholder farmers through voice-first advisory services in their native languages.
Read Case Study ↗Sahara v2.5 introduces new streaming endpoints for speech recognition and speech generation,
making it easier to build responsive voice agents, live transcription and real-time customer experiences.
Receive transcripts while the speaker is still talking.
Begin playing natural speech before the complete response has been generated.
Support more simultaneous conversations and production workloads.
Improved infrastructure for always-on, business-critical applications.
Build across languages, voices and markets without managing separate speech stacks.
bring voice technology to your organization in the languages your customers speak.
One API to power every voice interaction on the continent.
Across 57 languages, 500+ accents, and bridging languages in real-time.
Bilingual switching across 12 languages with deep intent recognition.
Voice bots with fluent native speech and low-latency endpoints.
Seamless integration with your existing product stack and voice channels.
Bring industry-leading voice technology to your organization in the languages your customers actually speak.
Choose the type of enquiry that best fits what you need. We’ll route your submission to the right team.
Gain deep insights into the evolving technological landscape with our upcoming 2026 Africa Voice AI Report. The report explores the trends, challenges, and opportunities shaping the future of speech interfaces across the continent.