Product narration
Write a launch script in StepAudio 3 TTS and export a spoken take for demos and onboarding.
Model: stepaudio-3-tts
Click a card to load the prompt into the studio. Generation stays gated until credentials are enabled.
StepAudio 3 gives creators one workspace for StepFun audio. Start with text-to-speech on StepAudio 3 TTS, switch to ASR for recordings, then use Realtime, Gen, and Music when those routes are enabled. Credits are quoted before you submit.
Write a launch script in StepAudio 3 TTS and export a spoken take for demos and onboarding.
Drop a recording into ASR and keep an editable transcript instead of replaying the whole call.
Use Gen to mix speech, room tone, and cues into one clip for a storyboard or trailer.
Give Music a style caption and lyrics to hear a first chorus before a longer session.
One studio for speech, recognition, realtime agents, unified audio, and music — with credits shown before you submit.
Write a script, choose a system voice, and generate downloadable speech with an upfront credit quote.
Switch to ASR for recordings that should become editable transcripts instead of new audio.
The same studio hosts Realtime voice, unified audio generation, and music creation as those routes are enabled.
Signed-in generations stay in your account so you can replay, download, and iterate without losing the last good take.
Start on TTS by default, or switch mode when you need ASR, Realtime, Gen, or Music.
Paste text for speech, upload audio for transcription, or give a creative brief for Gen and Music.
Confirm the model and cost before submission. Credits are reserved before the provider request starts.
Listen in the browser, download the file, and find completed work in your account history.
See the credit cost before every generation, then choose a subscription or a one-time top-up when you need more.
Try StepAudio 3 TTS after sign-in.
For regular individual creative work.
For campaigns and higher-volume production.
For larger workloads, with adjustable credit capacity.
Plans are shown for launch preparation. Checkout opens after production checks are complete.
StepAudio 3 is this site's product name for a StepFun audio studio covering TTS, ASR, Realtime, Gen, and Music. The homepage keyword and primary CTA center on StepAudio 3 TTS.
Text-to-speech uses StepFun model stepaudio-3-tts through the official audio speech API.
No. StepAudio 3 is an independent product that calls StepFun APIs. StepFun and its model names belong to their respective owner.
Generation stays disabled until API credentials, cost gates, and billing checks pass the launch SOP. The studio UI can still be reviewed while live generation is off.
StepAudio 3 is a focused studio for StepFun audio models. Generate speech with StepAudio 3 TTS, transcribe with ASR, try Realtime voice, create audio with Gen, or compose with Music.
Open the studio