StepAudio 3

Guide · Speech to text

Transcribe audio with StepAudio 3 ASR

The Transcribe mode turns a recording into editable text using StepFun's ASR model. Upload a file, choose the spoken language, and review the transcript in the studio. Here is how to get the cleanest result.

Transcribe a recording

In short

  • Upload MP3, WAV, M4A or OGG audio.
  • Choose the spoken language before you transcribe.
  • The credit cost is quoted before the request starts.
  • Transcripts stay in your account history.

Prepare a recording that transcribes well

  • One speaker at a time where possible; crosstalk lowers accuracy.
  • Close microphone, little background music or noise.
  • Trim long silences and unrelated sections before uploading.
  • Only upload audio you are allowed to process.

Transcribe in three steps

  1. Switch to Transcribe in the StepAudio 3 studio.
  2. Upload the file and pick the language that is actually spoken in it.
  3. Check the quote and transcribe, then copy or edit the text in the result panel.

From transcript to audio again

A common workflow is to transcribe a rough recording, tidy the wording, then switch to Speak (TTS) to produce a clean voice-over from the edited script — all in the same studio.

Frequently asked questions

Which audio formats are supported?

MP3, WAV, M4A and OGG.

Does it add speaker labels?

The Transcribe mode returns the recognised text. Keep one speaker per recording when you need clean attribution.

What else can the studio do?

Besides TTS and ASR, it hosts Realtime voice conversations, Soundscape audio generation and Music.

Updated 2026-10-04. StepAudio 3 is an independent product and is not affiliated with or endorsed by StepFun. Use only text and audio you have the right to submit.