最好的中文语音转文字模型是什么?
中文 ASR 榜单要么是英文测试, 要么干净得像朗诵棚. WER (中文里是 CER) 只看每个字 对不对, 但每个错误在意思层面的影响可以非常不一样. 「我喜欢」和「我不喜欢」只差 1 个字, 意思整个反了. 走样把错误分成 4 大类: 意思翻转扣 35 分, 实体认错扣 10 分, 口语填充扣 1 分. 每个模型除了名次, 还有一份错误的具体分布.
Most Chinese-ASR leaderboards are English-only or studio-clean, and WER (CER in Chinese) just counts characters without asking whether the meaning held up. Zouyang sorts errors into 4 groups: a flipped negation costs 35 points, a mangled entity 10, a filler word 1. Every model gets a breakdown of where its errors fall, not just a rank.
榜单 · Leaderboard (专用 ASR / dedicated ASR)
| # | 模型 Model | 厂商 Vendor | 类型 | 保真分Fidelity | 意思翻转Meaning flips | 实体认错Entities | 文化语境Cultural | 转写保真Fidelity | 单价 Price |
|---|---|---|---|---|---|---|---|---|---|
| 1 | MAI-Transcribe-2mai_transcribe_2 | Microsoft | API | 90.4 | -2.9 | -2.8 | -0.2 | -3.8 | $0.0017/min |
| 2 | MiMo-V2.5-ASRmimo | Xiaomi | OSS | 87.5 | -1.4 | -4.7 | -0.5 | -5.8 | $0.0012/min |
| 3 | Qwen3-ASR-Flashflash | Alibaba | API | 84.5 | -3.5 | -4.9 | -0.9 | -6.2 | $0.0021/min |
| 4 | Qwen-Audio-3.1-ASRqwen_audio_3_1_asr | Alibaba | API | 84.4 | -4.6 | -4.0 | -0.9 | -6.1 | - |
| 5 | Scribe v2elevenlabs | ElevenLabs | API | 83.9 | -3.2 | -4.1 | -2.0 | -6.8 | $0.0037/min |
| 6 | Doubao 2.0doubao | ByteDance | API | 82.1 | -4.8 | -6.4 | -0.5 | -6.3 | - |
| 7 | Confucius4-R2T2r2t2 | NetEase Youdao | OSS | 81.9 | -4.2 | -5.9 | -1.1 | -7.0 | - |
| 8 | Soniox stt-async-v5soniox | Soniox | API | 81.0 | -0.9 | -6.2 | -2.9 | -9.0 | $0.0017/min |
| 9 | GPT Transcribegpt_transcribe | OpenAI | API | 80.9 | -2.9 | -4.3 | -1.2 | -10.7 | $0.0045/min |
| 10 | VibeVoice-ASRvibevoice | Microsoft | OSS | 80.4 | -5.3 | -5.7 | -1.2 | -7.4 | - |
| 11 | GPT-4o-mini-transcribewhisper | OpenAI | API | 76.6 | -3.8 | -6.5 | -3.3 | -9.8 | $0.003/min |
| 12 | Qwen3-ASR-1.7Bqwen3_oss | Alibaba OSS | OSS | 76.1 | -5.4 | -7.6 | -1.9 | -9.0 | - |
| 13 | Gemini 3.5 Transcribegemini_3_5_transcribe | API | 74.3 | -4.5 | -6.1 | -2.1 | -12.9 | $0.005/min | |
| 14 | GLM-ASR-Nanoglm_asr_nano | Zhipu OSS | OSS | 73.1 | -7.8 | -7.2 | -1.1 | -8.6 | - |
| 15 | Step-Audio 2.5 ASRstepfun | StepFun | API | 71.3 | -8.4 | -9.1 | -1.1 | -10.1 | - |
| 16 | FireRedASR2-AEDfirered_aed | Xiaohongshu OSS | OSS | 69.3 | -4.3 | -14.2 | -2.1 | -10.1 | 自建 self-host |
| 17 | Fun-ASR-Nanofunasr_nano | Alibaba/DAMO OSS | OSS | 63.5 | -8.3 | -7.6 | -1.1 | -8.3 | - |
| 18 | Whisper-Large-v3-turbowhisper_v3_turbo | OpenAI OSS | OSS | 58.8 | -8.2 | -11.3 | -5.6 | -16.1 | 自建 self-host |
| 19 | Apple SpeechAnalyzer (macOS 26)apple_speechanalyzer | Apple | LOCAL | 48.6 | -4.9 | -23.9 | -6.5 | -16.1 | 自建 self-host |
| 20 | FireRedASR2-CTCfirered | Xiaohongshu OSS | OSS | 35.7 | -4.4 | -31.8 | -6.8 | -21.2 | 自建 self-host |
页面默认只显示基准线 (Qwen3-ASR-1.7B) 以上的 12 个模型, 其余行在这里 一并列出, 顺序和分数与交互版一致. The live board hides rows below the qwen3_oss baseline by default; all of them are listed here.
多模态 LLM 转写 · LLM-as-ASR (reference)
| # | 模型 Model | 厂商 Vendor | 类型 | 保真分Fidelity | 意思翻转Meaning flips | 实体认错Entities | 文化语境Cultural | 转写保真Fidelity | 单价 Price |
|---|---|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.8 Flash (LLM-as-ASR · thinkingLevel=low)gemini_3_8_flash_stt | API | 84.2 | -3.7 | -1.9 | -0.9 | -9.3 | - | |
| 2 | Gemini 3 Pro (LLM-as-ASR · thinkingLevel=low)gemini_pro_stt | API | 82.9 | -3.1 | -3.7 | -0.2 | -10.1 | - | |
| 3 | MiMo V2.5 (LLM-as-ASR)mimo_stt | Xiaomi | API | 79.6 | -5.0 | -4.5 | -1.2 | -9.8 | - |
| 4 | Gemini 3.7 Flash (LLM-as-ASR · thinkingLevel=low)gemini_3_7_flash_stt | API | 78.9 | -5.7 | -2.6 | -1.3 | -11.6 | - | |
| 5 | Gemini 3 Flash (LLM-as-ASR)gemini_flash_stt | API | 76.6 | -5.0 | -3.7 | -1.3 | -13.4 | - | |
| 6 | Gemini 3.5 Flash (LLM-as-ASR)gemini_3_5_flash_stt | API | 76.4 | -3.4 | -4.7 | -2.5 | -13.0 | - | |
| 7 | Gemini 3.6 Flash (LLM-as-ASR)gemini_3_6_flash_stt | API | 69.9 | -6.1 | -4.7 | -2.3 | -17.1 | - | |
| 8 | Gemini 3.1 Flash Lite (LLM-as-ASR)gemini_stt | API | 63.8 | -5.6 | -9.4 | -3.7 | -17.5 | - | |
| 9 | Voxtral-Mini-4B-Realtimevoxtral | Mistral OSS | OSS | 55.7 | -11.2 | -9.5 | -5.4 | -18.3 | - |
| 10 | MOSS-Audio-8B-Instruct (LLM-as-ASR)moss_audio | OpenMOSS / MOSI.AI OSS | OSS | 46.0 | -21.5 | -7.8 | -2.7 | -19.8 | - |
| 11 | Gemini 3.5 Flash-Lite (LLM-as-ASR)gemini_3_5_flash_lite_stt | API | 47.8 | -10.9 | -10.0 | -4.8 | -26.4 | - |
多模态 LLM 不是专用 ASR, 默认不进主榜, 单独作为参照. Multimodal LLMs are excluded from the main board by default and shown here for reference only.
四个维度 · The four groups
意思翻转 · Meaning flips — polarity, semantic-inversion, hallucination, numeric
实体认错 · Entities — brand, tech, proper-noun
文化语境 · Cultural — culture, dialect
转写保真 · Fidelity — mishear-severe, content-drop, paraphrase, mishear-minor, filler, format
数据 · Data
完整聚合数据: eval_aggregate.json (scores, per-category deductions, pricing, per-sample results). 不含音频与参考转写文本. No audio or reference transcripts are published. llms.txt