Runway details how Characters hits a 1.75s response time
By Steven Van ·
A technical breakdown of the real-time conversational video agent, built on GWM-1, now live via the Runway API and apps.
Runway published a technical breakdown of Runway Characters, a real-time conversational video agent that turns a single reference image, photorealistic, cartoon or fantasy, into an expressive character with no fine-tuning required. Built on Runway's GWM-1 world model, it generates natural lip-sync, facial expressions and head motion at 24fps in HD.
The system runs at an effective 37 milliseconds of model time per frame and a 1.75 second server-side turn-around from when a user stops speaking to when the character starts responding. Runway credits this to autoregressive frame-by-frame generation, with the diffusion transformer and VAE decoder running concurrently so decode time mostly disappears from the critical path, plus parallelization across devices, KV-cache management, CUDA Graphs and tuned kernels.
Alongside the model, Runway built out a product surface for deploying characters:
- Vision: characters can see a user's webcam or shared screen during a session.
- Custom Voice: text-to-voice design from a prompt, or instant voice cloning from an audio sample.
- Tool Calling: characters can trigger defined UI actions or backend calls during a conversation.
- Knowledge base: text and Markdown documents can be attached so a character answers from a company's own material.
- Embeddable widget: a real-time character can be dropped into a web app with one line of code.
- Meeting integration: characters can join Zoom, Google Meet or Teams calls and respond in real time.
Runway Characters is available now via the Runway API and the Runway web and mobile apps.
