Smallest.ai Lightning V3.1
Full voice pipeline (TTS + STT + LLM + S2S) — sub-100ms, 15+ languages
Comparative Scores
Architecture
Highly relevant for GamiWays as a complete voice pipeline alternative to assembling separate components. Hydra S2S (native full-duplex) directly addresses the GamiWays Phase B requirement for interruption-aware conversation. Electron LLM enables a fully integrated pipeline within a single vendor. Recommended for: Phase B evaluation (Hydra S2S), cost-effective production pipeline, HIPAA-compliant deployments.
Analysis
Smallest.ai is a full-stack real-time voice AI platform offering four integrated models: Lightning TTS (sub-100ms, 217 voices, 12 languages, 44.1 kHz, ~$25/1M chars), Pulse STT (38 languages, speaker diarization, emotion detection, ~$0.005/min), Electron LLM (sub-3B, sub-300ms TTFT, 70 languages, OpenAI-compatible), and Hydra Speech-to-Speech (native full-duplex, sub-300ms, early access). SOC 2, GDPR, HIPAA, ISO 27001 compliant. On-premises available on Enterprise.
Strengths
- Full pipeline: TTS + STT + LLM + S2S in one platform
- Sub-100ms Lightning TTS latency
- Hydra: native full-duplex S2S (early access)
- Pulse STT: 38 languages + emotion detection
- Electron LLM: sub-3B, OpenAI-compatible
- SOC 2, GDPR, HIPAA, ISO 27001 compliant
- On-premises enterprise option
Weaknesses
- No ELO score on Artificial Analysis
- Hydra S2S still in early access (waitlist)
- Electron LLM not accessible on pay-as-you-go
- Voice cloning: Enterprise plan only for production
- No lip-sync timestamps
- US-based (GDPR compliance via contractual)
Voice Capabilities
Instant voice cloning in under 10 seconds. No professional equipment required. Enterprise plan only for production use.
Speed and vocal identity are controllable; no explicit TTS emotion control is documented.
How: `speed`, `voice_id`, `model`, `language`, word timestamps.
Validate: Relevant for responsiveness and synchronisation, pair with another layer if performance intent is central.
WebSocket streaming for Lightning TTS and Pulse STT Realtime. Hydra S2S: full-duplex WebSocket (early access, sub-300ms).
No lip-sync timestamps documented for Lightning TTS.
Pricing
Lightning TTS: ~$0.025/1000 chars (~$25/1M chars). Pulse STT: ~$0.005/min (standard), ~$0.008/min (realtime). Electron LLM: Enterprise only. Hydra S2S: Early access.
Sovereignty & Compliance
On-premises deployment available on Enterprise plan. HIPAA Zero Data Retention add-on ($1000/mo on PAYG, included on Enterprise).
Data residency: SOC 2 Type 2, GDPR, HIPAA, ISO 27001 compliant. Headquarters: San Francisco, CA.
Full Voice Pipeline
Smallest.ai offers a 4-model integrated voice stack, replacing an assembled pipeline (external STT + external LLM + external TTS) with a single vendor.
Certifications & Compliance
This sheet's verification
Smallest.ai docs v4.0.0 (June 2026) — Lightning: ~200ms TTFB, 217 voices, 12 languages. Hydra: sub-300ms full-duplex.
Update note: AA Speech Arena sync 2026-08-18: ELO 1024 (rank N/A — Free tier). Name: Lightning v3.1.
This date covers provider-specific prices, capabilities and notes. Comparative benchmarks follow the synchronization and methodology shown above.
ELO benchmarks / indices: Artificial Analysis · Last API sync : 18 August 2026