Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More
1 contribution to Voice AI Alliance
Speech to speech models
I have some questions regarding speech to speech (voice to voice) models. Has anyone deployed speech-to-speech models like GPT Realtime in voice agents for companies (receptionist,etc.)? Does it handle tool calling the same as traditional stt-llm-tts pipeline? My other concern is STT accuracy for smaller languages like Czech. With a traditional STT → LLM → TTS stack, I can use ElevenLabs for STT, which works alright for Czech. With speech-to-speech, the recognition is built into the model, so I assume you lose that flexibility. How well do S2S models handle Czech or other less common languages in production, especially names, numbers, addresses, etc.? Also, can you use some s2s model inside Elevenagents platform. I asumme not, so the best option is probably livekit right? Also curious if anyone has tried a hybrid setup: S2S for normal conversation, but dedicated STT / traditional pipeline for critical data and tool calls. Thank you for any answers on this matter.🙏
1 like • 14h
You’re asking the exact right questions before committing to an architecture. Having tested both setups extensively in production, here is the ground reality on each point: 1. S2S vs. Traditional Tool Calling GPT Realtime does support function calling, but the execution pattern is different. In a traditional pipeline, you stream text tokens and trigger tools cleanly before speech synthesis begins. In S2S, you’re dealing with asynchronous client/server events over WebSockets or WebRTC. If a tool call takes 800ms–1.5s (e.g., checking a booking calendar), the model either stays silent or needs filler audio ("Let me look that up..."). It works, but state management and interrupt handling require tighter guardrails. 2. Accuracy in Czech (Names, Numbers, Addresses) This is the biggest hurdle for S2S right now. Native S2S models handle conversational flow in non-English languages reasonably well, but they frequently stumble on localized Czech phonetics, inflection, inflectional surnames, and exact street addresses. In a modular stack, you have full control: you can feed domain-specific phonetic corrections or vocabulary dictionaries into a dedicated STT engine (like Whisper or Deepgram) and normalize text with an LLM before synthesis. With S2S, the transcription is baked in, so when it mishears a Czech name or number, the error compounds directly into the reasoning layer. 3. Platform & Frameworks ElevenLabs' Conversational AI platform is built around the modular pipeline (STT → LLM → ElevenLabs TTS). If you want to deploy S2S models with full programmatic control, LiveKit (via livekit-agents) or Pipecat is definitely the way to go. LiveKit gives you the WebRTC infrastructure, voice activity detection (VAD), and turn-detection hooks necessary to make Realtime models usable in production. 4. Is a Hybrid Setup Viable? Hot-swapping mid-call (e.g., switching from S2S to modular just for an address capture) sounds good in theory, but in practice, managing audio streams, session states, and handoffs mid-call introduces latency spikes and edge cases that often break the user experience.
1-1 of 1
John Bernard
1
4 points to level up
@john-bernard-5493
xploring n8n, AI agents, APIs, and automation. Open to learning, collaborating, and connecting with professionals building smarter workflows.

Active 5h ago
Joined Oct 1, 2026
Powered by