Find the portal’s entities, people, products, services, topics, questions and issues.
Definitions of the 51 technical terms used across this portal — Voice Pipeline, Video Avatars, metrics, and infrastructure.
51 term(s) displayed
Automatic Speech Recognition — the technology that converts spoken audio into text. Also called STT (Speech-to-Text). Mo…
Short verbal or non-verbal feedback signals produced by a listener to indicate attention and understanding without takin…
The strategic decision between building a capability in-house (open-source models, custom development) versus purchasing…
Character Error Rate — an accuracy metric for STT systems, measuring errors at the character level rather than the word …
The process of splitting a large document into smaller, overlapping text segments (chunks) for RAG ingestion. Chunk size…
A hosted service accessed over the internet via HTTP/WebSocket APIs. Cloud APIs for STT/TTS offer easy integration and n…
A photorealistic AI-generated video avatar that replicates a specific person's appearance, voice, and behavioral pattern…
Speaker diarization is the process of partitioning an audio stream into segments according to the speaker identity — ans…
A comparative ranking system borrowed from chess, used in TTS benchmarks to measure perceived audio quality. Models are …
A dense numerical vector representation of text, audio, or other data in a high-dimensional space, where semantically si…
A single neural model that processes audio input and produces audio output directly, without separate STT → LLM → TTS st…
Full-Duplex Benchmark — an evaluation protocol measuring the conversational naturalness of voice AI systems in real-time…
A communication mode where both parties can speak and listen simultaneously, like a phone call. In voice AI, full-duplex…
General Data Protection Regulation — the EU regulation governing personal data processing. For voice AI, GDPR requires e…
Graphics Processing Unit — specialized hardware originally designed for graphics rendering, now widely used for deep lea…
The process of running a trained neural model on new input data to produce predictions or outputs. In voice AI, inferenc…
The delay between a request and its response. In voice pipelines, latency is the sum of STT processing time + LLM genera…
The synchronization between the movement of a speaker's lips and the audio being played. In video avatar systems, lip-sy…
Large Language Model — a neural network trained on large text corpora to understand and generate natural language. In a …
Vendor lock-in occurs when switching costs (technical, contractual, or operational) make it difficult to change provider…
Mergers and Acquisitions — corporate transactions where companies are combined or purchased. In the voice AI market, M&A…
Mean Opinion Score — a subjective audio quality metric where human listeners rate speech on a scale from 1 (bad) to 5 (e…
A neural network-based audio compression system that encodes audio into compact discrete token sequences and decodes the…
Non-Intrusive Speech Quality Assessment — a deep learning model that predicts speech quality without needing a reference…
New Federal Act on Data Protection (Neue Datenschutzgesetz / Nouvelle Loi sur la Protection des Données) — the Swiss dat…
Deployment model where software runs on servers physically located within an organization's own data center or controlle…
Software whose source code is publicly available and can be freely used, modified, and distributed. In voice AI, open-so…
Latency percentiles used in performance benchmarking. P50 (median) means 50% of requests complete within that time — it …
Perceptual Evaluation of Speech Quality — an ITU-T standard (P.862) for objective speech quality measurement that requir…
Personally Identifiable Information redaction — the automatic detection and removal or masking of sensitive personal dat…
The patterns of rhythm, stress, and intonation in speech. In TTS, prosody control determines how natural and expressive …
Retrieval-Augmented Generation — an architecture that enhances LLM responses by retrieving relevant context from an exte…
Real-Time Factor — the ratio of processing time to audio duration. An RTF of 1.0 means the model processes audio exactly…
Residual Vector Quantization — a neural audio compression technique used in modern TTS and audio codec models. RVQ encod…
Running an open-source model or software on your own infrastructure (cloud VM, on-premise server, or edge device) rather…
Data sovereignty refers to the principle that data is subject to the laws and governance structures of the country or ju…
State Space Model — a class of neural architectures (e.g., Mamba, S4) that process sequences more efficiently than Trans…
Speech Synthesis Markup Language — an XML-based markup language that provides a standard way to annotate text for speech…
In voice pipelines, streaming refers to the ability to process and transmit audio incrementally, without waiting for the…
In video avatar systems, streaming latency is the delay between receiving audio input and displaying the corresponding a…
Speech-to-Text — the conversion of spoken audio into written text. The first stage of a voice pipeline. Key metrics: WER…
A neural network architecture based on self-attention mechanisms, introduced in 2017. Transformers are the foundation of…
Time To First Audio — the latency between sending a text request to a TTS engine and receiving the first audio chunk. TT…
Time To First Token — the latency between sending a prompt to an LLM and receiving the first generated token. TTFT is di…
Text-to-Speech — the conversion of written text into spoken audio. The final stage of a voice pipeline. Key metrics: TTF…
The conversational mechanism that determines when each participant should speak and when to listen. In human conversatio…
Voice Activity Detection — an algorithm that detects the presence or absence of human speech in an audio stream. VAD is …
An AI-generated animated video representation of a person, capable of speaking with synchronized lip movements. Streamin…
The ability to synthesize speech that mimics a specific person's voice characteristics (timbre, accent, prosody) from a …
The end-to-end chain of components that enables a conversational voice AI: STT (transcription) → LLM (reasoning/generati…
Word Error Rate — the standard accuracy metric for STT systems. Calculated as (Substitutions + Deletions + Insertions) /…