02architecture

How agents work

Six parts, one loop. The model is the only probabilistic component; everything around it is deterministic software you control. Understanding which part does what tells you where cost comes from, where risk lives, and what you can buy off the shelf.

Agent loop diagram: a customer message enters an orchestrator; the model reasons, emits a tool call, the orchestrator checks guardrails, runs the tool, returns the result to the model, and the loop repeats until a reply is sent. Memory and evaluation attach to the orchestrator. customerWhatsApp · voice · web orchestratorruns the loop · executes toolsenforces guardrails · logs traceordinary code model (LLM)reads context → next action toolscalendar · CRM · Paynow · handoff memorycustomer record · notes · history guardrailsscope · permissions · confirm evaluationgolden set · scores context tool call / reply execute result checked before every write
The loop: the model only ever sees text and proposes actions; the orchestrator decides what actually runs.

1. The model

A hosted large language model reads a context — system prompt, tool definitions, conversation history, the latest message and any tool results — and returns either text or a structured tool call. Every call is stateless: the model remembers nothing between calls except what you send it again. That is why agents are priced per token and why the same prefix (system prompt, tools) is sent on every turn — and why prompt caching, which bills a repeated prefix at a tenth of the input price, is the single largest cost lever.

Model choice is a cost–capability trade. On the Claude price list accessed 14 September 2026, a small model (Haiku 4.5) is US$1 per million input tokens and US$5 per million output; the mid model (Sonnet 5) US$2 / US$10; the large model (Opus 5) US$5 / US$25. A booking agent that must follow a scope policy and format three options reliably sits comfortably on the mid tier; a reconciliation agent reading messy PDFs may need the large tier for the hard cases only.

2. The orchestrator

A small service — often a few hundred lines — that owns the loop: receive a webhook from WhatsApp or a telephony provider, load the session, assemble the context, call the model, parse the response, check guardrails, execute tool calls, append results, loop, send the reply, write the trace. Frameworks (LangGraph, the vendor SDKs' tool runners, Vercel's AI SDK) remove boilerplate; they do not remove the need to understand this loop, because every failure you debug will be in it.

orchestrator.py — the whole idea
def run(session, message):
    ctx = build_context(session, message)          # cached prefix + history
    for step in range(MAX_STEPS):                  # stop condition #1
        out = model(ctx)
        trace.log(out)
        if out.tool_calls:
            for call in out.tool_calls:
                policy.check(call, session)        # scope, permission, confirm?
                res = tools[call.name](**call.args)
                ctx.append(tool_result(call, res)) # untrusted data
            continue
        return send(session.channel, out.text)     # stop condition #2
    return handoff(session, "step budget exhausted")

3. Tools

Each tool is a function with a name, a plain-language description and a JSON schema for its arguments. The description is read by the model, so it is part of the prompt and should be written like documentation: what the tool does, when to use it, what it returns, what it must not be used for. Good agents have few tools with narrow contracts — get_availability(service, from, to, window) rather than run_sql(query). Tools are also the permission boundary: an agent cannot do what it has no tool for. The integrations page lists the tools that matter in Zimbabwe: WhatsApp Cloud API, Paynow for EcoCash and OneMoney, calendars, CRMs, and ZIMRA's FDMS for fiscal invoices.

4. Memory

Two kinds. Working memory is the conversation history in the context window; it grows with every turn and is the part that must be trimmed or summarised on long sessions. Long-term memory is a database: the customer record, previous bookings, notes the agent wrote last week. The model reaches it through tools (lookup_patient), which means memory is governed by the same permission rules as everything else. Retrieval — fetching relevant documents into context at request time — is memory for knowledge rather than for people.

5. Guardrails

Layers of control enforced outside the model's discretion. The prompt is the innermost, weakest layer. Around it: a confirmation step before any write; per-turn tool permissions (payments tool disabled unless the conversation reached a stage that needs it); scope and input filters applied before text reaches the model; and an audit log around everything. Each layer catches a different failure, from a hallucinated discount to a prompt injection hidden in a supplier's invoice.

Five nested boxes from outside in: audit log and monitoring; scope and input filtering; tool permissions with least privilege; confirmation before writes; the model at the centre. 5 · audit log + monitoring — every action recorded, traces kept, costs attributed 4 · scope + input filtering — off-topic, injection patterns, PII rules applied before the model sees text 3 · tool permissions — least privilege, read before write, per-turn enabling 2 · confirmation before writes — human or explicit customer "yes" 1 · model + system prompt
The prompt is the innermost, weakest layer. Every layer outside it is code you control.

6. Evaluation

A fixed set of scenarios (a golden set) with expected outcomes, run every time the prompt, tools or model change. Scores: task success, policy compliance, correct handoffs, cost per conversation, latency. Without this you are changing a probabilistic system by feel. The article evaluating agents before you trust them with customers gives a scorecard you can copy.

A token walk-through

Take the demo booking conversation. The stable prefix — system prompt plus five tool definitions — is about 3,700 tokens and is cached after the first call. Across eight model calls the agent reads roughly 925 uncached tokens (new messages and small history increments), about 35,000 cached tokens, and writes about 635 tokens. At Sonnet 5 prices that is around US$0.025 for the whole conversation; on Haiku 4.5, about US$0.012. Without caching the same conversation would cost roughly three times as much. This is why "how many model calls per conversation" and "is the prefix cached" are the first two questions to ask any vendor.

Build or buy, part by part

PartBuyBuildZimbabwe note
ModelAlways — hosted API (Anthropic, OpenAI, Google) or a cloud regionSelf-host only for data-residency mandatesData may need to be described as stored outside Zimbabwe on your POTRAZ licence application
OrchestratorPlatform (Zvino Agent, vendor agent builders) for standard patternsCustom for anything with bespoke tools or approval flowsKeep it small; it must run cheaply and survive intermittent connectivity
ChannelWhatsApp via Meta Cloud API or a BSP; voice via Twilio/Vonage or a local SIP trunk—+263 is priced under Meta's "Rest of Africa"; international voice to ZW mobiles is ≈US$0.86/min
ToolsCalendar, CRM, Paynow SDKsThin wrappers with narrow contractsPaynow supports EcoCash and OneMoney express checkout with a poll URL
GuardrailsContent filters, PII detectorsPolicy, permissions, confirmation gates — always customAutomated decisions with significant effect need consent or legal authorisation
EvaluationTracing/eval toolingThe golden set — nobody can buy your scenariosInclude Shona/Ndebele and mixed-language cases

Next: what it costs, what can go wrong, or watch the loop run.