You’re asking the exact right questions before committing to an architecture. Having tested both setups extensively in production, here is the ground reality on each point: 1. S2S vs. Traditional Tool Calling GPT Realtime does support function calling, but the execution pattern is different. In a traditional pipeline, you stream text tokens and trigger tools cleanly before speech synthesis begins. In S2S, you’re dealing with asynchronous client/server events over WebSockets or WebRTC. If a tool call takes 800ms–1.5s (e.g., checking a booking calendar), the model either stays silent or needs filler audio ("Let me look that up..."). It works, but state management and interrupt handling require tighter guardrails. 2. Accuracy in Czech (Names, Numbers, Addresses) This is the biggest hurdle for S2S right now. Native S2S models handle conversational flow in non-English languages reasonably well, but they frequently stumble on localized Czech phonetics, inflection, inflectional surnames, and exact street addresses. In a modular stack, you have full control: you can feed domain-specific phonetic corrections or vocabulary dictionaries into a dedicated STT engine (like Whisper or Deepgram) and normalize text with an LLM before synthesis. With S2S, the transcription is baked in, so when it mishears a Czech name or number, the error compounds directly into the reasoning layer. 3. Platform & Frameworks ElevenLabs' Conversational AI platform is built around the modular pipeline (STT → LLM → ElevenLabs TTS). If you want to deploy S2S models with full programmatic control, LiveKit (via livekit-agents) or Pipecat is definitely the way to go. LiveKit gives you the WebRTC infrastructure, voice activity detection (VAD), and turn-detection hooks necessary to make Realtime models usable in production. 4. Is a Hybrid Setup Viable? Hot-swapping mid-call (e.g., switching from S2S to modular just for an address capture) sounds good in theory, but in practice, managing audio streams, session states, and handoffs mid-call introduces latency spikes and edge cases that often break the user experience.