Type
Speech recognition (ASR)
Start free — $2 credit
StepFun API model
StepAudio 2.5 ASR — a 4B MTP streaming transcription model for fast Chinese-English recognition, ITN normalization, subtitles, meetings, and voice agents.
Specs
Transparent USD pricing. No mainland-China account or phone required. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately.
Speech recognition (ASR)
Audio input, text output
Audio, Speech-to-Text, Streaming, Chinese and English, ITN
$0.0208 per audio hour
/v1/audio/transcriptions
StepAudio 2.5 ASR — a 4B MTP streaming transcription model for fast Chinese-English recognition, ITN normalization, subtitles, meetings, and voice agents.
Pricing
$0.0208 per audio hour. Paid in USD. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately. View live pricing
Quickstart
curl -X POST https://api.chinaapi-ru.com/v1/audio/transcriptions \
-H "Authorization: Bearer sk-..." \
-F "model=stepaudio-2.5-asr" \
-F "file=@meeting.wav" \
-F "response_format=json"
Internal links
Use stepaudio-2.5-asr through ChinaAPI with an API key and the OpenAI-compatible endpoint shown above.
stepaudio-2.5-asr is $0.0208 per audio hour, paid in USD. The live pricing page is authoritative. The displayed media rate may include a service margin covering provider input/output billing, payment processing, chargeback exposure, and operations. Any margin is included in the displayed rate and is not added separately.
No. ChinaAPI provides access without a mainland-China account or phone number and includes a $2 free trial. Check live pricing for the displayed USD rate.