Phony.ai

That "five cents a minute" AI phone agent quote almost never survives contact with a real invoice. This episode tears apart every billable layer — carrier, speech-to-text, LLM, and voice synthesis — so you know exactly what you're paying for in 2026.

Show Notes

That five-cents-a-minute headline price for an AI voice agent looks tidy on a vendor slide deck. The actual invoice almost never matches it. This episode of Phony.ai dissects the real cost structure of an AI phone call in 2026, pulling apart the four separate billable layers that sit beneath any platform's sticker price — and showing precisely where vendor margins hide and self-built stacks silently overrun. The full breakdown is laid out in the 2026 AI phone call per-minute cost breakdown that this episode is based on.

The episode walks through each cost layer in turn and explains why the gap between advertised and actual rates is almost never accidental:

  • Telephony is cheap — until it isn't. The raw carrier leg costs roughly a penny per minute, but platforms routing calls through Twilio's Conversation Relay add up to seven cents per minute on top of that, a charge most bundled bills never surface.
  • Speech-to-text is close to a rounding error. Streaming transcription from providers like Deepgram runs well under a cent per call minute and is rarely where budgets blow out.
  • Text-to-speech is where costs climb fast. Premium voice synthesis can run around eight cents per minute at standard concurrency — and double to sixteen cents when simultaneous-call limits are breached during busy periods. Silence billing (metering hold time and pauses) can quietly add another 15–30% on top.
  • The model layer carries the biggest swing. Depending on whether a call uses a flagship real-time LLM or a lighter text-mediated pipeline, the model cost alone can range from two cents to well over fifteen cents per minute — a variation that bundled pricing makes completely invisible to buyers.
  • Stacked together, an honest mid-range minute costs more than most ads suggest. Combining all four layers realistically lands a delivered minute somewhere between fourteen and twenty-two cents, with model choice explaining most of the spread. Rates below that range are typically a promotional subsidy, a concurrency trap, or a thinner pipeline; rates above it usually reflect compliance add-ons such as HIPAA requirements.
  • Five questions to ask before signing anything. The episode closes with the specific written questions buyers should demand answers to — including which carrier ran the leg, which underlying model providers are in use, and whether you can swap to a lighter model for simpler call intents.

Whether you're evaluating a bundled platform or pricing out your own stack, understanding what each layer actually costs is the only way to hold a vendor accountable — or to know when a low quote is a gift and when it's a clause waiting to activate. For more on a related dimension of AI call design, check out the earlier episode Barge-In Is Not a Feature — It's a Promise You Have to Keep.

Phony.ai

What is Phony.ai?

AI phone and voice agents, explained through the constraints that actually decide whether one works: end-to-end latency and where it comes from, interruption and barge-in handling, telephony plumbing and call control, transfer design, and the disclosure and recording rules around automated calls.

Each episode takes one design decision and works through it concretely — provider-neutral, comparing approaches rather than selling one. Written for teams evaluating, buying or building voice AI who need to know what breaks before it breaks in production. Five or six minutes an episode.

Topics include end-to-end latency and where it comes from, barge-in and interruption handling, telephony and call control, transfer and escalation design, prompt and turn design, evaluation and call review, and disclosure, consent and recording rules.

Produced by Phony.ai, provider-neutral AI phone and voice agents. Full details, services and further reading at https://phony.ai