
What AI actually costs in South Africa
Rand-denominated budgets, dollar-denominated APIs, and the costs South African teams keep missing.
Head of Platform
Amara leads the gateway and routing team at WebParam. She spent a decade building inference infrastructure at scale before turning her attention to the multi-provider problem: how to make switching models a configuration value rather than an integration project.
Areas of focus

Rand-denominated budgets, dollar-denominated APIs, and the costs South African teams keep missing.
A local region doesn't mean local inference, a rand invoice doesn't mean rand exposure, and POPIA doesn't require in-country processing. The three questions South African teams should answer before committing to an AI provider.
Token prices converted to rands are the start of an AI budget, not the end of it. Where the gap opens between the pricing page and the invoice — output ratios, context re-billing, agent loops, forex, card fees and VAT.
Why WebParam replaced its CMS screens with a conversation.
Why WebParam replaced its CMS screens with a conversation.
Why WebParam replaced its CMS screens with a conversation.
Why WebParam replaced its CMS screens with a conversation.
For two years, betting on a single provider was defensible because one model was usually ahead. The frontier has flattened, prices move weekly, and the teams shipping fastest now treat the model as a configuration value. Here is what that changes about how you build.
First-token latency is what users perceive, so streaming isn't optional. But every provider streams differently, and a mid-stream failure surfaces to your app as a truncated response. Here is what normalising the stream at the gateway actually buys you.
See it in your stack
Browse the catalogue, ask Echo what fits your workload, or get started and we'll help you onboard the right models and gateway configuration.