Murf Adds Word Timestamps for WebSocket Streaming
By Steven Van ·
Falcon 2 word timestamps work across every locale except Chinese, Japanese and Korean.
Voice agents using Murf AI can now track how much of a response has played when a user interrupts. Falcon 2’s WebSocket API returns start and end times for each spoken word.
Enable timestamps for a connection
Set add_word_timestamps to true at the top level of a frame, alongside text. The setting applies to the whole WebSocket connection. Settings placed inside voice_config are ignored without an error or warning.
Timestamps arrive in separate JSON frames alongside the audio, with start_s and end_s values in seconds. Each context’s timing starts at zero, and later sentences continue from the previous sentence’s last word. The returned words reflect spoken text, so “$50” becomes “fifty dollars”.
Handle interruptions at the playback position
Audio is generated faster than real time, so timestamps can arrive before the corresponding words play. To determine what a listener heard, compare each word’s start_s with the current playback position.
When a user interrupts, stop playback and send a clear frame with that turn’s context_id to stop queued audio. Discard received audio beyond the playback position and timestamps whose start_s falls after it.
Supported locales
Word timestamps are available in every locale except Chinese, Japanese and Korean. The Advanced Settings documentation covers the response format and timing behaviour.