← All articles

StepAudio 3 Gen — one model for TTS, voice design, music, and sound effects

· Source: original

⚡ StepFun rolled out StepAudio 3 Gen — universal audio generation without diffusion

StepFun released a general-purpose audio generation model and brought all audio tasks together into a single framework. StepAudio 3 Gen covers zero-shot text-to-speech, voice design, vocal generation, sound effects, music, vibe speech, and hybrid mixes of several audio types — all with one model, rather than a set of separate tools for each type of sound.

Under the hood is a discrete autoregressive generator that models audio directly over RVQ tokens (residual vector quantization). This is a departure from the diffusion Transformer-based continuous generation paradigm that most modern general audio models rely on: instead of continuous diffusion — discrete autoregressive.

Audio compression is handled by the StepAudio Tokenizer: a shared audio representation at 12.5 Hz in a shared 16 × 2048 space.

All of this is described in the StepAudio 3 Gen Technical Report on HuggingFace Papers — huggingface.co/papers/2609.12945

🎬 Work with video and multimedia?

Prompts for video and multimedia scenarios — on topic:

🔗 The entire prompt library · "Video" category

Ready-made product on the topic: AI Image Prompt Pack — grab it and apply it right away.

AIAutomation

🎁 Забери бесплатный набор AI-промптов

6 отобранных промптов для бизнеса, кода и контента + доступ к библиотеке 2000+. Без оплаты.

✈️ Get the kit on Telegram

Need ready-made automations for your business?

Browse products