What an AI agent costs
Four cost lines: the model (per token), the channel (WhatsApp per template, voice per minute), hosting, and the platform fee if you use a BSP or an agent platform. Everything below is priced from public rate cards accessed 14 September 2026 and turned into one worked example — the booking conversation from the demo. Business-case questions (what a booking is worth to you) belong on businessai.co.zw's ROI calculator; this page is the cost side only.
1. Model: priced per million tokens
A token is roughly four characters. An agent pays for everything it reads (input) and writes (output) on every call — and it makes several calls per conversation. Two features change the bill more than model choice: prompt caching (a repeated prefix — system prompt, tool definitions — is billed at 0.1× the input price) and the Batch API (50% off for jobs that can wait).
| Model | Input / MTok | Cached read / MTok | Output / MTok | Use in agents |
|---|---|---|---|---|
| Claude Haiku 4.5 | US$1.00 | US$0.10 | US$5.00 | High-volume, narrow-scope agents; classification |
| Claude Sonnet 5 | US$2.00 | US$0.20 | US$10.00 | Most production agents (the basis for this page) |
| Claude Opus 5 | US$5.00 | US$0.50 | US$25.00 | Hard synthesis; messy documents; escalation model |
| OpenAI gpt-5.6-luna / terra / sol (reported) | US$0.20 / 2.00 / 5.00 | — | US$1.20 / 12.00 / 30.00 | Comparable tiers; verify on openai.com before budgeting |
Anthropic list prices from platform.claude.com/docs/en/about-claude/pricing (5-minute cache write 1.25×; batch 50% off). OpenAI figures as reported by Morph (morphllm.com/openai-api-pricing) as of 21 Aug 2026. Accessed 2026-09-14.
2. The worked example: one booking conversation
The demo trace has 8 model calls and 5 tool calls. Its stable prefix (system prompt + five tool schemas) is about 3,700 tokens, cached on the first call. Totals across the conversation: 925 uncached input tokens, 35,390 cached input tokens, 635 output tokens. These are illustrative estimates, not measured usage.
| Line | Tokens | Rate (Sonnet 5) | Cost |
|---|---|---|---|
| Cache write (first call) | 3,700 | US$2.5/MTok | US$0.0092 |
| Uncached input | 925 | US$2/MTok | US$0.0019 |
| Cached input reads | 35,390 | US$0.2/MTok | US$0.0071 |
| Output | 635 | US$10/MTok | US$0.0063 |
| Model total, Sonnet 5 | US$0.0245 | ||
| Same conversation, Haiku 4.5 | US$0.0123 | ||
| Same conversation, Opus 5 | US$0.0613 | ||
| Sonnet 5 without caching | US$0.0864 | ||
| WhatsApp: 4 replies inside the 24 h window | US$0 | US$0.0000 | |
| WhatsApp: next-day reminder, utility template to +263 | US$0.0040 | US$0.0040 | |
| Variable cost per booked appointment | US$0.0285 |
Two things to notice. Caching is worth more than model choice: turning it off costs more than moving from Sonnet to Opus. And the channel is free — the only WhatsApp charge is the reminder, because it is sent outside the service window.
3. WhatsApp: per template, by category, by country code
Since 1 July 2025 Meta bills per delivered template message. Free-form replies inside the 24-hour customer service window (or the 72-hour free entry point after a Click-to-WhatsApp ad) cost nothing. +263 is in Meta's "Rest of Africa" market.
| Category | Rest of Africa (+263) | South Africa (+27) | Nigeria (+234) | Typical agent use |
|---|---|---|---|---|
| Marketing | US$0.0225 | US$0.0379 | US$0.0516 | Re-engagement, promotions — avoid in agents |
| Utility | US$0.0040 | US$0.0076 | US$0.0067 | Reminders, order/booking updates, payment requests |
| Authentication | US$0.0040 | US$0.0076 | US$0.0067 | One-time codes |
| Service (free-form, in window) | US$0 | US$0 | US$0 | Every agent reply to an inbound message |
Rate card effective 1 July 2026 as published by SleekFlow (sleekflow.io/blog/whatsapp-business-price); pricing model and window rules from Meta's developer documentation. Utility and authentication rates fall with monthly volume tiers. Verify against Meta's downloadable rate card.
Add the BSP: providers such as Twilio, 360dialog or Infobip charge platform fees on top of Meta's rates — flat monthly, per message, or both. Going direct to Meta's Cloud API removes that layer but adds engineering. Budget the BSP line separately; it can exceed the Meta line at low volumes.
4. Voice: the phone line dominates
Twilio's Zimbabwe voice rates are US$0.8641 per minute to mobiles and US$0.4090 to landlines, inbound and outbound. The speech stack and model add roughly US$0.10 per minute. A local SIP trunk or operator integration replaces the international rate with a domestic interconnect rate — obtain it from the operator; it is the single most important number in any Zimbabwean voice-agent budget. The full breakdown and the WhatsApp-voice-note alternative are on the voice agents page.
5. Hosting and orchestration
The orchestrator is a small web service: it receives webhooks, calls APIs and stores sessions. It does not need a GPU. A small cloud VM or a serverless tier — on the order of US$5–40 a month depending on provider and region (illustrative) — carries thousands of conversations. Where a data-residency requirement means the service and its logs must sit in Zimbabwe, local hosting costs vary and should be quoted directly. Self-hosting the model is a different order of cost and only justified by a hard residency mandate.
6. Monthly bands at scale
| Booking agent volume / month | Model (Sonnet 5) | Model (Haiku 4.5) | Reminder templates | Variable total (Sonnet) |
|---|---|---|---|---|
| 1,000 bookings | US$25 | US$12 | US$4 | US$29 |
| 5,000 bookings | US$123 | US$61 | US$20 | US$143 |
| 20,000 bookings | US$491 | US$245 | US$80 | US$571 |
Excludes hosting, BSP platform fee, build and support. Illustrative.
7. What moves the number
- Cache the prefix. Keep the system prompt and tool list byte-stable; put anything volatile after them.
- Fewer model calls per conversation. Menus in front of the agent; one clarifying question instead of three.
- Right-size the model per step. A small model for routing and extraction, a mid model for the conversation, a large one only for escalations.
- Utility templates, not marketing. US$0.004 vs US$0.0225 to +263; write the template so Meta classifies it as utility.
- Batch anything that can wait — reconciliation, digests — at half price.
- Local telephony for voice, or WhatsApp voice notes instead of calls.
Sources
- Anthropic model pricing, caching and batch multipliers — platform.claude.com/docs/en/about-claude/pricing (accessed 2026-09-14)
- OpenAI API pricing as reported — morphllm.com/openai-api-pricing (accessed 2026-09-14)
- WhatsApp Business Platform pricing model, windows, Rest-of-Africa country list — developers.facebook.com (accessed 2026-09-14)
- WhatsApp per-message rates by market, effective 1 Jul 2026 — sleekflow.io (accessed 2026-09-14)
- Twilio Programmable Voice pricing, Zimbabwe — twilio.com/en-us/voice/pricing/zw (accessed 2026-09-14)
- Deepgram and ElevenLabs per-minute rates as reported — happyrobot.ai (Deepgram), happyrobot.ai (ElevenLabs) (accessed 2026-09-14)