04economics

What an AI agent costs

Four cost lines: the model (per token), the channel (WhatsApp per template, voice per minute), hosting, and the platform fee if you use a BSP or an agent platform. Everything below is priced from public rate cards accessed 14 September 2026 and turned into one worked example — the booking conversation from the demo. Business-case questions (what a booking is worth to you) belong on businessai.co.zw's ROI calculator; this page is the cost side only.

1. Model: priced per million tokens

A token is roughly four characters. An agent pays for everything it reads (input) and writes (output) on every call — and it makes several calls per conversation. Two features change the bill more than model choice: prompt caching (a repeated prefix — system prompt, tool definitions — is billed at 0.1× the input price) and the Batch API (50% off for jobs that can wait).

ModelInput / MTokCached read / MTokOutput / MTokUse in agents
Claude Haiku 4.5US$1.00US$0.10US$5.00High-volume, narrow-scope agents; classification
Claude Sonnet 5US$2.00US$0.20US$10.00Most production agents (the basis for this page)
Claude Opus 5US$5.00US$0.50US$25.00Hard synthesis; messy documents; escalation model
OpenAI gpt-5.6-luna / terra / sol (reported)US$0.20 / 2.00 / 5.00—US$1.20 / 12.00 / 30.00Comparable tiers; verify on openai.com before budgeting

Anthropic list prices from platform.claude.com/docs/en/about-claude/pricing (5-minute cache write 1.25×; batch 50% off). OpenAI figures as reported by Morph (morphllm.com/openai-api-pricing) as of 21 Aug 2026. Accessed 2026-09-14.

2. The worked example: one booking conversation

The demo trace has 8 model calls and 5 tool calls. Its stable prefix (system prompt + five tool schemas) is about 3,700 tokens, cached on the first call. Totals across the conversation: 925 uncached input tokens, 35,390 cached input tokens, 635 output tokens. These are illustrative estimates, not measured usage.

LineTokensRate (Sonnet 5)Cost
Cache write (first call)3,700US$2.5/MTokUS$0.0092
Uncached input925US$2/MTokUS$0.0019
Cached input reads35,390US$0.2/MTokUS$0.0071
Output635US$10/MTokUS$0.0063
Model total, Sonnet 5US$0.0245
Same conversation, Haiku 4.5US$0.0123
Same conversation, Opus 5US$0.0613
Sonnet 5 without cachingUS$0.0864
WhatsApp: 4 replies inside the 24 h windowUS$0US$0.0000
WhatsApp: next-day reminder, utility template to +263US$0.0040US$0.0040
Variable cost per booked appointmentUS$0.0285

Two things to notice. Caching is worth more than model choice: turning it off costs more than moving from Sonnet to Opus. And the channel is free — the only WhatsApp charge is the reminder, because it is sent outside the service window.

3. WhatsApp: per template, by category, by country code

Since 1 July 2025 Meta bills per delivered template message. Free-form replies inside the 24-hour customer service window (or the 72-hour free entry point after a Click-to-WhatsApp ad) cost nothing. +263 is in Meta's "Rest of Africa" market.

CategoryRest of Africa (+263)South Africa (+27)Nigeria (+234)Typical agent use
MarketingUS$0.0225US$0.0379US$0.0516Re-engagement, promotions — avoid in agents
UtilityUS$0.0040US$0.0076US$0.0067Reminders, order/booking updates, payment requests
AuthenticationUS$0.0040US$0.0076US$0.0067One-time codes
Service (free-form, in window)US$0US$0US$0Every agent reply to an inbound message

Rate card effective 1 July 2026 as published by SleekFlow (sleekflow.io/blog/whatsapp-business-price); pricing model and window rules from Meta's developer documentation. Utility and authentication rates fall with monthly volume tiers. Verify against Meta's downloadable rate card.

Add the BSP: providers such as Twilio, 360dialog or Infobip charge platform fees on top of Meta's rates — flat monthly, per message, or both. Going direct to Meta's Cloud API removes that layer but adds engineering. Budget the BSP line separately; it can exceed the Meta line at low volumes.

4. Voice: the phone line dominates

Twilio's Zimbabwe voice rates are US$0.8641 per minute to mobiles and US$0.4090 to landlines, inbound and outbound. The speech stack and model add roughly US$0.10 per minute. A local SIP trunk or operator integration replaces the international rate with a domestic interconnect rate — obtain it from the operator; it is the single most important number in any Zimbabwean voice-agent budget. The full breakdown and the WhatsApp-voice-note alternative are on the voice agents page.

Two horizontal bar charts. Text agent on WhatsApp per 1,000 booking conversations: model about 25 dollars, WhatsApp templates about 4 dollars, hosting about 10 dollars; BSP platform fee varies. Voice agent per 1,000 call minutes via an international carrier: telephony about 864 dollars, speech-to-text about 5 dollars, text-to-speech and agent runtime about 80 dollars, model about 15 dollars. Text agent on WhatsApp · per 1,000 booking conversations model (Sonnet 5, cached)≈ US$25 reminder templates (utility)≈ US$4 hosting / orchestration≈ US$10 BSP platform feevaries by provider (US$0–50+) Voice agent via international carrier · per 1,000 call minutes to Zimbabwe mobiles telephony (Twilio ZW mobile)≈ US$864 speech-to-text≈ US$5 TTS + agent runtime≈ US$80 model≈ US$15
Illustrative, from the sourced price basis on this page. On WhatsApp the model is the cost; on voice the phone line is.

5. Hosting and orchestration

The orchestrator is a small web service: it receives webhooks, calls APIs and stores sessions. It does not need a GPU. A small cloud VM or a serverless tier — on the order of US$5–40 a month depending on provider and region (illustrative) — carries thousands of conversations. Where a data-residency requirement means the service and its logs must sit in Zimbabwe, local hosting costs vary and should be quoted directly. Self-hosting the model is a different order of cost and only justified by a hard residency mandate.

6. Monthly bands at scale

Booking agent volume / monthModel (Sonnet 5)Model (Haiku 4.5)Reminder templatesVariable total (Sonnet)
1,000 bookingsUS$25US$12US$4US$29
5,000 bookingsUS$123US$61US$20US$143
20,000 bookingsUS$491US$245US$80US$571

Excludes hosting, BSP platform fee, build and support. Illustrative.

7. What moves the number

  1. Cache the prefix. Keep the system prompt and tool list byte-stable; put anything volatile after them.
  2. Fewer model calls per conversation. Menus in front of the agent; one clarifying question instead of three.
  3. Right-size the model per step. A small model for routing and extraction, a mid model for the conversation, a large one only for escalations.
  4. Utility templates, not marketing. US$0.004 vs US$0.0225 to +263; write the template so Meta classifies it as utility.
  5. Batch anything that can wait — reconciliation, digests — at half price.
  6. Local telephony for voice, or WhatsApp voice notes instead of calls.

Sources

  1. Anthropic model pricing, caching and batch multipliers — platform.claude.com/docs/en/about-claude/pricing (accessed 2026-09-14)
  2. OpenAI API pricing as reported — morphllm.com/openai-api-pricing (accessed 2026-09-14)
  3. WhatsApp Business Platform pricing model, windows, Rest-of-Africa country list — developers.facebook.com (accessed 2026-09-14)
  4. WhatsApp per-message rates by market, effective 1 Jul 2026 — sleekflow.io (accessed 2026-09-14)
  5. Twilio Programmable Voice pricing, Zimbabwe — twilio.com/en-us/voice/pricing/zw (accessed 2026-09-14)
  6. Deepgram and ElevenLabs per-minute rates as reported — happyrobot.ai (Deepgram), happyrobot.ai (ElevenLabs) (accessed 2026-09-14)