I have some questions regarding speech to speech (voice to voice) models.
Has anyone deployed speech-to-speech models like GPT Realtime in voice agents for companies (receptionist,etc.)?
Does it handle tool calling the same as traditional stt-llm-tts pipeline?
My other concern is STT accuracy for smaller languages like Czech. With a traditional STT → LLM → TTS stack, I can use ElevenLabs for STT, which works alright for Czech. With speech-to-speech, the recognition is built into the model, so I assume you lose that flexibility.
How well do S2S models handle Czech or other less common languages in production, especially names, numbers, addresses, etc.?
Also, can you use some s2s model inside Elevenagents platform. I asumme not, so the best option is probably livekit right?
Also curious if anyone has tried a hybrid setup: S2S for normal conversation, but dedicated STT / traditional pipeline for critical data and tool calls.
Thank you for any answers on this matter.🙏