Hey guys! So glad to be in here. I have been following nour's youtube videos for a while and am here to learn and provide value. Just wanted to share something thats helped me improve my voice AI pipeline a lot!
I spent a lot of time trying to optimize my pipeline by swapping models. While that wasnt the best fix i found what worked was measuring each stage separately first - STT, LLM, TTS, and network - before touching anything. mine turned out to be a stage I'd never suspected.
Also surprisingly logging timestamps at every stage boundary and looking at real numbers .
And streaming everything , so not waiting for full utterance to hand off to the next stage. Most of my "model latency" was just waiting
After lots of testing, my current stack sits around 290ms end to end.. Curious what everyone else is measuring - and where your bottleneck actually turned out to be?