GamiWays
STT+RAG+LLM+TTS

Cost Simulator — Voice Pipeline

Estimate the real monthly cost of your voice pipeline (STT + RAG + LLM + TTS) accounting for subscription plans, included minutes, top-ups, and parallel stream limits.

Start with a preset

Usage Volume

Set the number of users and session duration — total minutes are computed automatically. Parallel streams define simultaneous capacity and may trigger plan alerts.

Users / month250
1050010k
Avg session duration10 min
2 min30 min120 min
Simultaneous parallel streams10 streams
1 stream50 streams1000 streams
Users/month
250
Avg session
10 min
Total minutes/month
2.5k
250 × 10 min
Parallel streams
10
simultaneous capacity
Formula: Monthly variable cost = (cost/min STT + RAG + LLM + TTS) × 250 users × 10 min/session = 2,500 total min/month. Parallel streams (10) define simultaneous capacity and may require a specific subscription plan.

1 — Speech Recognition (STT)

2 — Knowledge Base (RAG)

Monthly fixed cost: $25/mo (vector DB) — independent of volume

3 — Language Model (LLM)

Assumption: 2 turns/min, ~2,000 input tokens + ~150 output tokens per turn.

4 — Voice Synthesis (TTS)

Simulation Results

Total cost / month (variable)
$455
250 users × 10 min + $25 RAG
Optimal cost / month (with plans)
$455
Best STT + TTS + LLM plan
Cost / user
$1.72
10 min × $0.172/min
Estimated pipeline latency
625ms
STT + RAG + LLM + TTS
Recommended optimal plan for your volume(2,500 min/mois)
STT — Deepgram Nova-3
Pay As You Go
overage $0.00430/min
$11/mo
LLM — Gemini 2.0 Flash
Pay As You Go
overage $0.00052/min
$1/mo

The “Free (Google AI Studio)” plan includes 0 min but allows no overage. Your volume (2,500 min/mo) exceeds this by 2,500 min — “Pay As You Go” plan selected automatically.

TTS — Eleven v4 / v4 Turbo
Scale
$299/mo + overage $0.170/min
$418/mo

The “Free” plan would cost $896.40/mo with overage. The “Scale” plan ($299/mo + 1,800 min included) costs $418.00/mo — saving $478.40/mo.

Variable cost breakdown
TTS 97%
ComponentProvider$/min (optimal plan)Selected plan$/user (plan)Total/mo
STTDeepgram Nova-3$0.00430Pay As You Go$0.043$11
RAGVoyage AI + Supabase pgvector$0.00000Free + Supabase Free$0.00000$0.00+$25
LLMGemini 2.0 Flash$0.00052Pay As You Go$0.00520$1
TTSEleven v4 / v4 Turbo$0.167$0.200Scale$299/mo$1.67$418
Total (optimal plans)$1.72$455
OPTIMAL TOTAL (best plans)—$455
Pipeline latency breakdown (estimated)
75ms
STT
50ms
RAG
300ms
LLM TTFT
200ms
TTS TTFA
Estimated total:625ms⚠ Acceptable but optimizable

Preset comparison at your volume

Estimated monthly cost for 250 users × 10 min/session = 2,500 min/month. 10 parallel streams.

PresetSTTRAGLLMTTS$/userTotal/moLatencyEUStreams
Storygami — Optimal
$11$0.00+$25$1$500$2.05$537625ms∞
Storygami — Budget
$4$0.00+$25$1$31$0.147$62542ms∞
Edugami — Optimal
$5$0.00+$25$1$50$0.225$81769ms∞
Sovereign EU Stack
$26$0.00+$25$1$108$0.542$160686ms60
Self-Hosted (Open-Source)
$0.00$0.00$0.00$2$0.00700$2650ms∞
Full STT comparison
15 engines, WER benchmarks
Full TTS comparison
25 engines, ELO, expressiveness
Interactive V2V Pipeline
Latency & component diagram