agent · 04voice

Voice AI agents

A voice agent is the same loop as any other agent — model, tools, guardrails — running under a latency budget of about a second and a half per turn, with speech recognition in front and speech synthesis behind. In Zimbabwe two things dominate the design: the cost of the phone line, and whether the speech stack actually understands your callers.

The phone line costs more than the intelligence

Twilio's published Programmable Voice rates for Zimbabwe, accessed 14 September 2026, are US$0.8641 per minute to mobiles and US$0.4090 per minute to landlines, inbound and outbound alike. Speech-to-text (Deepgram's Nova-3 streaming tier is reported at US$0.0048 per minute), synthesis and agent runtime (ElevenLabs' agent overage is reported at US$0.08 per conversation minute) and the model (about US$0.015 per minute of talk at Sonnet 5 prices) together come to roughly a tenth of that. For 1,000 minutes a month through an international carrier the telephony leg alone is about US$864; the rest is under US$100.

ComponentBasisPer 1,000 min
Telephony, international carrier to ZW mobileTwilio ZW rate US$0.8641/min≈ US$864
Speech-to-text≈ US$0.0048/min (reported)≈ US$5
TTS + agent runtime≈ US$0.08/min (reported)≈ US$80
Model≈ 1,500 tokens/min at Sonnet 5≈ US$15
Totalillustrative≈ US$964

Three ways to change that arithmetic, in order of leverage: (1) terminate calls on a local SIP trunk or operator integration instead of an international carrier — the per-minute cost becomes a Zimbabwean interconnect rate, which you must obtain from the operator; (2) move the conversation to WhatsApp voice notes, which are asynchronous, cost nothing per message inside the service window, and still let the customer speak rather than type; (3) use voice only for the first 60 seconds — identify the need, then send a WhatsApp message and end the call.

The latency budget

A single horizontal timeline of about 1.5 seconds for one voice turn: network and carrier 200 milliseconds, end-of-speech detection 300 milliseconds, speech-to-text 150 milliseconds, model first token 500 milliseconds, text-to-speech first audio 200 milliseconds, playback start 100 milliseconds. One voice turn · target ≤ 1.5 s from the caller stopping to the agent starting carrier / network200 ms end-of-speech detect300 ms STT150 ms model → first token500 ms (streamed) TTS first audio200 ms play100 ms ≈ 1,450 ms · any tool call inside the turn adds its own latency
Illustrative budget. Every tool call inside a voice turn must be fast or deferred ("let me check that while we talk").

A turn that takes longer than about 1.5 seconds feels like the line dropped. That budget is spent before the model does anything clever, so voice agents keep tool calls out of the turn where possible: the agent says "let me check that" while the availability lookup runs, and streams its first words before the model has finished. Barge-in — the caller interrupting — must stop synthesis immediately; without it every voice agent sounds like an IVR.

Languages: this is a data problem

Shona is the mother tongue of about 80.9% of Zimbabweans and Ndebele of 11.5% (2022 estimates), and most calls mix English with one of them. Text models handle this reasonably; speech models vary widely. Econet's Ndebele-speaking assistant, introduced in phases in September 2026, was trained on thousands of hours of recorded Ndebele speech — a signal that off-the-shelf recognition was not enough. Before you promise a Shona voice agent, run 200 real recordings through the candidate STT and count the errors on names, numbers and dates, which are the words that matter for a booking.

Where voice fits — and where WhatsApp wins

SituationVoice agentWhatsApp agent
After-hours reception for a clinic or workshopFits: callers expect a phoneOnly if the number is advertised for WhatsApp
Appointment remindersExpensive; consent for automated callsWins: utility template at US$0.004
Customers without smartphones or dataFits: the only channelNot possible
Anything with a reference number, address or spellingWeak: errors on names and digitsWins: text is exact and stays in the thread
Poor network areasWeak: audio drops break the turnWins: asynchronous, retries silently

Guardrails for voice

  • Read back every booking, amount and reference before committing, and require a spoken "yes".
  • Send a WhatsApp or SMS confirmation after the call — the caller has no transcript otherwise.
  • Announce that the caller is speaking to an automated assistant at the start, and offer a person on request; recordings and transcripts are personal data under the Cyber and Data Protection Act.
  • Hard cap on call length and on retries to the same number.

Sources (accessed 2026-09-14): Twilio voice pricing for Zimbabwe — twilio.com; Deepgram and ElevenLabs rates as reported by HappyRobot — deepgram, elevenlabs; Claude pricing — platform.claude.com; language shares — CIA World Factbook archive; Econet Ndebele assistant — CITE, 6 Sep 2026. Related: voice costs, barge-in.