STT Lab

Omli's kids VAD segments your mic; each segment goes to Omli's STT, Gemini, Sarvam REST, or all three — same audio, same boundaries. Latency is not comparable across the three. Omli streams and reports its own inference time. Gemini and Sarvam each get a finished segment, so their numbers cover connect + upload + inference for the whole utterance.

idle

One socket does both jobs. Omli's stt_vad_backend socket supplies the VAD boundaries and its transcript, transcribing as the child talks. The extra vad_enabled=true socket that used to re-send each finished segment has been removed: it existed because kids_v5 under-segmented badly enough that one 38s segment came back as two words, and kids_v6 cuts tightly enough that the workaround costs more than it buys. Gemini and Sarvam are still fed from our own buffer, streamed from vad start, so their latency is the wait after speech ends; Omli reports its own inference time. max_segment_ms 12000 is enforced here, against a buffer we hold — Omli's own stays pinned at 60000, because when its cap fires it discards everything between the flush and the true speech-end, a full spoken sentence, measured. hangover 1000, not Omli's 2000: at 2000 no segment may close before it is 2s old, so a child's half-second "yes" cannot end a turn and gets glued to the adult's next question — measured, one Omli segment ran 49s. 1000 keeps short answers separate while still absorbing most coughs and chair creaks. confidence has NO effect (0.05 and 0.95 are byte-identical — reported).

Seq Timeline Segment Omli STT Gemini Sarvam REST