StepAudio 3 Realtime Report Published
- •StepFun researchers publish StepAudio 3 Realtime Technical Report for realtime spoken interaction
- •StepAudio 3 reaches 73.0 macro average on StepAudio Chat in reasoning mode
- •Model scores 90.6 on MMSU and 98.9 Overall on Artificial Analysis Full-Duplex Bench
StepFun researchers published the StepAudio 3 Realtime Technical Report on Sep 12, with the paper submitted to Hugging Face Papers by chao yan on Sep 16 and listed as the #3 Paper of the day. The report describes StepAudio 3 Realtime, an audio-language foundation model designed for realtime spoken interaction through a continuous listen-converse-think-act loop.
The model uses Deep Perception to capture acoustic cues for interpreting user intent, according to the abstract. It also uses Seamless Duplex, which models synchronized audio streams so the system can handle pauses, backchannels, and interruptions during conversation. Think-While-Speaking lets the model run private reasoning while it is already delivering speech, aiming to reduce the tension between deep deliberation and latency.
StepAudio 3 reaches a 73.0 macro average on StepAudio Chat in reasoning mode. With Think-While-Speaking, the authors say it achieves dialogue and reasoning performance comparable to dedicated reasoning models while speaking in realtime. An integrated Voice Agent also handles asynchronous tool execution without disrupting dialogue flow.
The report claims top-tier performance across several benchmarks: 90.6 on MMSU, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ-Voice. Hugging Face shows 90 upvotes, +82 near the paper listing, 0 models citing the paper, 0 datasets citing it, 0 Spaces citing it, and 0 collections including it.