<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Keencraft]]></title><description><![CDATA[Practical guides on building production AI voice agents with LiveKit, Twilio and LLMs. Latency, CRM integrations and automation from real projects.]]></description><link>https://keencraft.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6ac5fbb32efaa8fb90aaa899/6e7991f1-25b2-4722-b0e3-53bf3af7fb7c.png</url><title>Keencraft</title><link>https://keencraft.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Fri, 09 Oct 2026 10:31:26 GMT</lastBuildDate><atom:link href="https://keencraft.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Why Your App Works in the Demo and Falls Over Under Load]]></title><description><![CDATA[The demo works. One user, a warm laptop, and a presenter who knows which button to press. Then the first week of real traffic arrives, and the same product waits, times out, or returns a stale answer.]]></description><link>https://keencraft.hashnode.dev/app-works-in-demo-falls-over-under-load</link><guid isPermaLink="true">https://keencraft.hashnode.dev/app-works-in-demo-falls-over-under-load</guid><category><![CDATA[backend]]></category><category><![CDATA[performance]]></category><category><![CDATA[PostgreSQL]]></category><category><![CDATA[Redis]]></category><category><![CDATA[System Design]]></category><dc:creator><![CDATA[Muhammad Jareer]]></dc:creator><pubDate>Thu, 08 Oct 2026 08:23:12 GMT</pubDate><content:encoded><![CDATA[<p>The demo works. One user, a warm laptop, and a presenter who knows which button to press. Then the first week of real traffic arrives, and the same product waits, times out, or returns a stale answer.</p>
<p>People call that a scaling problem. Usually it's a path problem: the request was designed for an audience of one.</p>
<p>I lead development on Rawk.ai, a live voice agent builder, where a latency pass cut end-to-end response time from about 900ms to about 320ms. None of that came from a better prompt. It came from the pipeline, caching, and where services ran. Here's the approach, with the patterns in code. The snippets are simplified illustrations, not production source.</p>
<h2>Load is a tail, not an average</h2>
<p>A product can look healthy on average and still fail the people who hit the slow path. Voice callers hang up. Checkout users retry and double-charge themselves. Staff refresh a dashboard and act on stale data.</p>
<p>Three things matter, in this order:</p>
<ul>
<li><p><strong>Speed:</strong> time to the first useful response, not time until a spinner disappears.</p>
</li>
<li><p><strong>Correctness:</strong> a fast stale cache is just a bug with better graphs.</p>
</li>
<li><p><strong>Cost:</strong> a path that calls a model or vendor API on every keystroke gets expensive before it gets slow.</p>
</li>
</ul>
<p>Before optimizing anything, name the exact request you're timing. "The app feels slow" is a complaint. "Time to first audio on a live call" is a job.</p>
<h2>Measure the tail first</h2>
<p>Averages hide the problem, so measure percentiles on production-like data:</p>
<pre><code class="language-python">import statistics
import time

async def time_request(fn, runs: int = 200) -&gt; dict:
    samples = []
    for _ in range(runs):
        start = time.perf_counter()
        await fn()
        samples.append((time.perf_counter() - start) * 1000)  # ms
    q = statistics.quantiles(samples, n=100)
    return {"p50": q[49], "p95": q[94], "p99": q[98]}
</code></pre>
<p>If p50 looks fine and p95 doesn't, you've found where users are actually suffering. Keep this number, because you'll rerun the exact same measurement after every change.</p>
<h2>Where the time actually goes</h2>
<p>Most slow products are a chain of reasonable steps that were never allowed to overlap.</p>
<h3>1. Sequential I/O</h3>
<p>The page waits for the user, then permissions, then the list, then the counts. Each call is fine on its own. Together, they add up:</p>
<pre><code class="language-python">import asyncio

# Before: each await waits for the previous one to finish
async def dashboard_slow(user_id: str):
    user = await get_user(user_id)
    perms = await get_permissions(user_id)
    items = await list_items(user_id)
    counts = await get_counts(user_id)
    return user, perms, items, counts

# After: independent calls run concurrently
async def dashboard(user_id: str):
    return await asyncio.gather(
        get_user(user_id),
        get_permissions(user_id),
        list_items(user_id),
        get_counts(user_id),
    )
</code></pre>
<p>The total time drops from the sum of all four calls to roughly the slowest one. The catch: only run calls together when they're truly independent. If <code>list_items</code> depends on the permissions result, it has to wait.</p>
<h3>2. A cache that's missing, or lying</h3>
<p>Postgres stays the system of record. Redis sits in front of reads that are hot, safe to serve more than once, and expensive to recompute. The real design work is the boundary: which keys exist, how long they live, and which writes delete them.</p>
<pre><code class="language-python">import json
import redis.asyncio as redis

r = redis.Redis()
TTL_SECONDS = 300

async def get_plan_prices(account_id: str) -&gt; dict:
    key = f"plan_prices:{account_id}"
    cached = await r.get(key)
    if cached is not None:
        return json.loads(cached)
    prices = await db.fetch_plan_prices(account_id)  # Postgres is the source of truth
    await r.set(key, json.dumps(prices), ex=TTL_SECONDS)
    return prices

async def update_plan_prices(account_id: str, prices: dict) -&gt; None:
    await db.save_plan_prices(account_id, prices)
    await r.delete(f"plan_prices:{account_id}")  # invalidate on write
</code></pre>
<p>The TTL is a safety net. The <code>delete</code> on write is what keeps the cache honest. A cache without an invalidation rule is how a paid invoice keeps showing as "new lead" in the dashboard your team trusts.</p>
<h3>3. Queries that scan</h3>
<p>If the access pattern was never the index, Postgres will still answer. It'll just answer late. Run <code>EXPLAIN ANALYZE</code> on the slow request's queries and look for sequential scans on large tables.</p>
<h3>4. Retrieval on every request</h3>
<p>Embeddings, vector search, and a model call for a question the last hundred users already asked. Pinecone or pgvector belongs in the path only when the answer actually depends on search. If the fact is already a column, read the column.</p>
<h3>5. Cross-region hops</h3>
<p>When each vendor defaults to a different cloud region, every request picks up tens of milliseconds per hop. No amount of code optimization fixes that. Colocate the services that talk on every request.</p>
<h2>What changed on Rawk AI</h2>
<p>For a voice agent, a 900ms gap after every sentence breaks the conversation. Callers talk over the agent, the agent talks over them, and the call falls apart. Sub-400ms is roughly where a call starts to feel natural.</p>
<p>Getting there was pipeline work:</p>
<ul>
<li><p><strong>Streaming speech-to-text</strong> instead of waiting for a complete transcript</p>
</li>
<li><p><strong>Streaming the model's output</strong> instead of buffering a full paragraph</p>
</li>
<li><p><strong>Starting audio on the first clause</strong>, so the caller hears a response while the rest is still generating</p>
</li>
<li><p><strong>Prefetching tool results</strong> the agent was likely to need</p>
</li>
<li><p><strong>Colocating services</strong> so no extra regions were added to every turn</p>
</li>
</ul>
<p>Each change shaved time off the same measured request. None of them involved the prompt.</p>
<h2>A checklist for your own slow path</h2>
<ol>
<li><p>Name the one request users feel, and time it on production-like data, including p95 and p99.</p>
</li>
<li><p>List every downstream call on that request, in order, with what each returns.</p>
</li>
<li><p>Run independent calls concurrently.</p>
</li>
<li><p>Mark which results are identical across requests and which must never be shared.</p>
</li>
<li><p>Cache only the safe, identical reads, with an explicit delete on write.</p>
</li>
<li><p>Colocate the services that talk on every request.</p>
</li>
<li><p>Re-measure the same request. If the tail didn't move, that wasn't the bottleneck.</p>
</li>
</ol>
<p>And write down what you changed. Undocumented speed is a demo you won't be able to repeat.</p>
<hr />
<p><em>The original version, including how we scope latency audits, is on</em> <a href="https://www.keencraft.tech/blog/when-a-product-falls-over-under-load"><em>keencraft.tech</em></a><em>.</em></p>
]]></content:encoded></item><item><title><![CDATA[Breaking Down the Per-Minute Cost of a Production Voice Agent]]></title><description><![CDATA[Every AI voice call is five services billing at the same time. Platforms advertise one of them as the headline price, which is why the number on the pricing page rarely matches the invoice.
In 2026, a]]></description><link>https://keencraft.hashnode.dev/breaking-down-the-per-minute-cost-of-a-production-voice-agent</link><guid isPermaLink="true">https://keencraft.hashnode.dev/breaking-down-the-per-minute-cost-of-a-production-voice-agent</guid><category><![CDATA[AI]]></category><category><![CDATA[llm]]></category><category><![CDATA[twilio]]></category><category><![CDATA[voice ai]]></category><category><![CDATA[livekit]]></category><dc:creator><![CDATA[Muhammad Jareer]]></dc:creator><pubDate>Wed, 07 Oct 2026 08:19:21 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ac5fbb32efaa8fb90aaa899/93c83fbc-9303-4341-82ab-cd3bef09bea2.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Every AI voice call is five services billing at the same time. Platforms advertise one of them as the headline price, which is why the number on the pricing page rarely matches the invoice.</p>
<p>In 2026, a production voice agent on a standard stack costs roughly $0.07 to $0.15 per call minute. With a premium voice and a frontier model, it can pass $0.30. This post breaks down where each cent goes, works through the math for a managed platform versus your own stack on LiveKit, and covers the costs that only show up once real callers start dialing in.</p>
<p>I've shipped production agents on Vapi, Retell, and LiveKit, so the numbers below are the same ones I use when scoping real builds.</p>
<h2>The five layers billing on every call</h2>
<p>Rates below are US list prices checked in September 2026. Providers change them often, so check the live pricing page before you put a number in front of a client.</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>What it does</th>
<th>Typical provider</th>
<th>Cost per minute</th>
</tr>
</thead>
<tbody><tr>
<td>Telephony</td>
<td>Connects the agent to a real phone number</td>
<td>Twilio</td>
<td>~$0.0085 inbound, ~$0.014 outbound</td>
</tr>
<tr>
<td>Speech-to-text</td>
<td>Transcribes the caller in real time</td>
<td>Deepgram Nova-3</td>
<td>~$0.0048 streaming (promotional), $0.0077 regular</td>
</tr>
<tr>
<td>LLM</td>
<td>Decides what to say and which tool to call</td>
<td>OpenAI, Anthropic, Google</td>
<td>~$0.003 to $0.16, by model</td>
</tr>
<tr>
<td>Text-to-speech</td>
<td>Speaks the reply</td>
<td>Cartesia, ElevenLabs</td>
<td>~$0.015 standard, ~$0.04 ElevenLabs</td>
</tr>
<tr>
<td>Orchestration</td>
<td>Runs the loop: turn-taking, interruptions, tool calls</td>
<td>Vapi, Retell, LiveKit Cloud</td>
<td>$0.05 Vapi, $0.055 Retell, $0.01 LiveKit</td>
</tr>
</tbody></table>
<p>One note on Deepgram: the \(0.0048 Nova-3 streaming rate is promotional. Budget the regular \)0.0077 so your estimate survives the promotion ending.</p>
<h2>A worked example</h2>
<p>Here's an inbound appointment-booking agent with a fast mid-size model, a standard voice, and basic monitoring. I've assumed about $0.02 per minute for the LLM, which is typical for a fast mid-size model on booking calls with short turns.</p>
<table>
<thead>
<tr>
<th>Layer</th>
<th>Vapi</th>
<th>Retell</th>
<th>LiveKit Cloud</th>
</tr>
</thead>
<tbody><tr>
<td>Twilio inbound</td>
<td>$0.0085</td>
<td>$0.0085</td>
<td>$0.0085</td>
</tr>
<tr>
<td>Deepgram Nova-3 (regular rate)</td>
<td>$0.0077</td>
<td>$0.0077</td>
<td>$0.0077</td>
</tr>
<tr>
<td>LLM (fast mid-size, assumed)</td>
<td>$0.0200</td>
<td>$0.0200</td>
<td>$0.0200</td>
</tr>
<tr>
<td>TTS (standard voice)</td>
<td>$0.0150</td>
<td>$0.0150</td>
<td>$0.0150</td>
</tr>
<tr>
<td>Orchestration</td>
<td>$0.0500</td>
<td>$0.0550</td>
<td>$0.0100</td>
</tr>
<tr>
<td>Monitoring and transcripts</td>
<td>$0.0100</td>
<td>$0.0100</td>
<td>$0.0100</td>
</tr>
<tr>
<td><strong>Total per minute</strong></td>
<td><strong>~$0.11</strong></td>
<td><strong>~$0.12</strong></td>
<td><strong>~$0.07</strong></td>
</tr>
</tbody></table>
<p>Swap in ElevenLabs and a bigger model and the managed stacks move toward $0.15 and beyond. The LiveKit figure applies once you're past the plan's included agent minutes, where LiveKit bills $0.01 per agent-session minute. Below that allotment, orchestration is cheaper still.</p>
<p>At 2,000 minutes a month (about 700 three-minute calls), that's roughly $140 to $300 to run the agent, depending on the stack.</p>
<h2>Two layers decide most of the bill</h2>
<p>The voice and the model. A premium voice can cost nearly three times a standard one, and a frontier model can cost many times a small, fast one.</p>
<p>For appointment booking, a fast mid-size model is usually the right call. Callers notice a slow reply long before they notice a slightly less clever sentence. The latency budget for that reply, including the calendar write mid-call, is covered in <a href="https://www.keencraft.tech/blog/production-ai-voice-agents">how to build an AI voice agent that actually books appointments</a>.</p>
<h2>Managed platform vs your own stack: the break-even</h2>
<p>Vapi and Retell are the right choice when you're under a few thousand minutes a month, still testing whether callers will stay on the line with an AI agent, or have nobody to run voice infrastructure. Both can be live in days, and at low volume the platform fee is a fair price for skipping the engineering.</p>
<p>As volume grows, orchestration is the line item that changes. Going from Retell at $0.055 to LiveKit Cloud at $0.01 saves about $0.045 on every minute past the included allotment:</p>
<table>
<thead>
<tr>
<th>Monthly call minutes</th>
<th>Retell at $0.055</th>
<th>LiveKit Cloud at $0.01</th>
<th>Monthly saving</th>
</tr>
</thead>
<tbody><tr>
<td>2,000</td>
<td>$110</td>
<td>$20</td>
<td>$90</td>
</tr>
<tr>
<td>10,000</td>
<td>$550</td>
<td>$100</td>
<td>$450</td>
</tr>
<tr>
<td>50,000</td>
<td>$2,750</td>
<td>$500</td>
<td>$2,250</td>
</tr>
</tbody></table>
<p>To get a payback period, divide the extra engineering cost of running your own stack by that monthly saving. At 2,000 minutes, the saving rarely justifies the work. At 10,000 and above, it starts to, and that's before counting the other reasons to own the pipeline.</p>
<h2>Latency is part of what you're buying</h2>
<p>The bigger reason to go custom is often control, not cost. On Rawk.ai, a voice agent builder I lead development on, we brought end-to-end response time down from about 900ms to about 320ms by owning the pipeline end to end.</p>
<p>Owning the stack also means you keep the call data, pick each provider independently, and don't have to rebuild the agent when a platform changes its pricing.</p>
<h2>Costs that show up on the first invoice</h2>
<p>None of these is large on its own, but together they're why the raw per-minute math is never the real number.</p>
<ul>
<li><p><strong>Testing minutes.</strong> Every test call bills like a real one. Tuning an agent before launch can burn hundreds of minutes.</p>
</li>
<li><p><strong>Calls that go nowhere.</strong> Voicemails, hang-ups, and spam still use telephony, speech-to-text, and orchestration.</p>
</li>
<li><p><strong>Transfers.</strong> A warm transfer keeps two call legs open, so telephony roughly doubles for those minutes.</p>
</li>
<li><p><strong>Monitoring.</strong> Recording, transcripts, and observability are often billed separately, commonly around $0.01 per minute.</p>
</li>
<li><p><strong>Numbers and concurrency.</strong> Each phone number has a monthly fee, and a busy hour can hit a concurrency cap. Vapi includes 10 concurrent lines, then about $10 per extra line per month. Retell includes 20.</p>
</li>
<li><p><strong>Compliance.</strong> On Vapi, HIPAA is a published add-on at $2,000 a month on top of usage, which matters a lot if the calls are clinical.</p>
</li>
<li><p><strong>Ongoing tuning.</strong> Prompts need adjusting once real callers say things nobody scripted.</p>
</li>
</ul>
<p>A realistic budget adds 15 to 20 percent on top of the raw per-minute math for the first two months. The buffer usually shrinks after launch, once testing stops and the prompts settle.</p>
<h2>Where the engineering effort actually goes</h2>
<p>Per-minute cost is mostly set by vendors. Engineering effort is where projects differ, and it scales with how much of the business the agent has to touch.</p>
<table>
<thead>
<tr>
<th>Factor</th>
<th>Simpler</th>
<th>More complex</th>
</tr>
</thead>
<tbody><tr>
<td>Call direction</td>
<td>Inbound only: answer, qualify, book</td>
<td>Inbound plus outbound follow-up with retry rules and calling-hour limits</td>
</tr>
<tr>
<td>Calendar and CRM</td>
<td>Books into one calendar</td>
<td>Reads availability and writes contacts, stages, and notes to HubSpot, Salesforce, or Zoho</td>
</tr>
<tr>
<td>Locations</td>
<td>One business, one number</td>
<td>Multiple locations with separate hours, staff, and isolated data</td>
</tr>
<tr>
<td>Languages</td>
<td>English only</td>
<td>Several languages, or one with thin speech-model support</td>
</tr>
<tr>
<td>Compliance</td>
<td>General business calls</td>
<td>Healthcare calls needing HIPAA-conscious data handling</td>
</tr>
<tr>
<td>Handoff</td>
<td>Takes a message</td>
<td>Warm transfer with context attached</td>
</tr>
</tbody></table>
<p>The biggest hidden driver is usually the CRM, not the voice. Getting a call to sound good is the smaller piece of engineering. Creating the right contact, skipping duplicates, and moving the right pipeline stage is where the hours go. I wrote up that write-path in detail in <a href="https://www.keencraft.tech/blog/crm-automation-hubspot-salesforce-zoho">CRM automation that writes back to HubSpot, Salesforce, and Zoho</a>.</p>
<h2>Sources</h2>
<p>Prices checked September 2026. Vendor pages change, so verify before you quote a rate.</p>
<ul>
<li><p><strong>Retell AI pricing:</strong> voice infrastructure \(0.055/min, most voices \)0.015/min, ElevenLabs \(0.040/min, advertised all-in range about \)0.07 to $0.31/min.</p>
</li>
<li><p><strong>Vapi pricing:</strong> hosting $0.05/min. STT, model, voice, and telephony billed by those providers, or at no Vapi charge with your own keys.</p>
</li>
<li><p><strong>LiveKit pricing:</strong> agent session minutes at \(0.01/min after the plan allotment, Deepgram Nova-3 monolingual at \)0.0048/min on Build and Ship plans.</p>
</li>
<li><p><strong>Deepgram pricing:</strong> Nova-3 monolingual streaming, promotional \(0.0048/min, regular \)0.0077/min.</p>
</li>
<li><p><strong>Softcery voice agent cost calculator:</strong> Twilio US local about \(0.0085/min inbound and \)0.014/min outbound.</p>
</li>
</ul>
<hr />
<p><em>I build production voice agents and backend systems at</em> <a href="https://www.keencraft.tech/services/ai-voice-agents"><em>KeenCraft</em></a><em>. This article was originally published on the</em> <a href="https://www.keencraft.tech/blog/ai-voice-agent-cost"><em>KeenCraft blog</em></a><em>.</em></p>
]]></content:encoded></item></channel></rss>