StepAudio 3

Guide · Text to speech

StepAudio 3 TTS: turn a script into natural speech

StepAudio 3 TTS reads your script aloud in a chosen voice and follows a short style instruction such as "warm, confident narrator, steady pace". This guide covers voices, languages, writing the instruction, and preparing a script that sounds right on the first take.

Generate speech

In short

  • Text-to-speech runs on StepFun's stepaudio-3-tts model through the official speech API.
  • Choose one of the system voices, then add a one-line style instruction.
  • English and Chinese are supported; Japanese, Korean, French and Spanish are in preview.
  • Up to 1,000 characters per generation, with the credit cost shown before you submit.

Choose a voice

The studio lists the available system voices, for example an elegant gentle female voice, a magnetic English male voice and a confident male voice. Pick the one closest to the brand or character, then use the style instruction to adjust delivery rather than switching voices repeatedly.

Write a style instruction

The instruction describes how the line is delivered. Keep it to one sentence: role, tone, pace.

  • "Warm, confident product narrator. Steady pace."
  • "Calm meditation guide, slow and soft, long pauses between sentences."
  • "Upbeat explainer for a short social video, energetic but clear."

Prepare the script

Write for the ear. Short sentences, spelled-out numbers and punctuation where you want pauses give more natural results. Split long narration into several generations of up to 1,000 characters so you can redo one paragraph without regenerating everything.

Example script
Meet the new trail bottle. It keeps water cold for twenty-four hours, fits in any side pocket, and weighs less than an apple. Fill it, clip it, and go.

Languages

Select the output language to match the script. English and Chinese are fully supported; Japanese, Korean, French and Spanish are marked as preview, so listen closely to names and uncommon words before publishing.

Frequently asked questions

Which model powers StepAudio 3 TTS here?

StepFun's stepaudio-3-tts model, called through the official audio speech API.

How long can the text be?

Up to 1,000 characters per generation.

Can I download the audio?

Yes. Play it in the browser, download it, and find it again in your account history.

Is this site run by StepFun?

No. It is an independent product that calls StepFun APIs; StepFun and its model names belong to their owner.

Updated 2026-10-04. StepAudio 3 is an independent product and is not affiliated with or endorsed by StepFun. Use only text and audio you have the right to submit.