Fast 4B MTP streaming transcription for Chinese-English audio with ITN normalization and an OpenAI-compatible transcription route.
Model update wall
China AI model release timeline.
A data-driven wall of Chinese AI model releases served by ChinaAPI, plus a small starter set of notable releases not yet on the gateway. 3 estimated dates and 0 unknown dates are flagged in the source data.
Newest first
Release wall.
Cost-efficient multilingual speech synthesis with official and cloned voices, pronunciation mapping, and style controls.
Expressive speech synthesis with official and cloned voices, pronunciation mapping, style controls, and OpenAI-compatible output.
Speed-focused speech synthesis with natural output, emotion controls, language enhancement, and multiple audio formats.
High-fidelity speech synthesis with emotional delivery, voice controls, language enhancement, and multiple audio formats.
Low-latency multilingual speech synthesis with mixed-language input, built-in voices, and character-based usage.
Multilingual file transcription with emotion recognition, ITN, and OpenAI-compatible audio input.
Context-aware expressive speech synthesis with streaming output and controls for voice, speed, and volume.
Multilingual streaming transcription with Chinese dialect coverage, hotwords, and an OpenAI-compatible audio transcription route.
API launch of Kimi's 2.8T-parameter flagship with 1M context, always-on reasoning, and native image and video understanding. The announced open-source weights are due by July 27, 2026.
Context-aware text-to-speech with global and inline natural-language direction, expressive delivery, and direct OpenAI-compatible /v1/audio/speech output.
Tencent Hunyuan 3 gateway model with 256K context and fast, low, and deep thinking modes.
Open MoE reasoning model with 1M context for coding agents and long-horizon tool use.
Lower-latency Doubao Seed 2.1 model for multimodal reasoning, coding, and production agent workflows.
Flagship Doubao Seed 2.1 model for multimodal reasoning, coding, tool use, and long-horizon agent tasks.
HappyHorse 1.1 text-to-video gateway model with improved semantic understanding and camera control.
HappyHorse 1.1 reference-to-video gateway model for multi-reference subject, scene, and style consistency.
HappyHorse 1.1 image-to-video gateway model with stronger texture and identity consistency.
Open flagship GLM with 1M context and stronger long-horizon coding and agent workflows.
Fast Seedance 2.0 mini video model for lower-cost short-form generation.
High-speed Kimi K2.7 gateway variant for lower-latency coding and agent tasks.
Coding-focused Kimi gateway model for long-horizon repository work with vision support.
MiMo speech recognition for Chinese-English code-switching, Cantonese, Wu, Minnan and Sichuan dialects, noisy and far-field recordings, overlapping speakers, and lyrics.
Multimodal Qwen gateway model for visual analysis, planning, files, and tool use.
Native multimodal M-series model with 1M context, MSA attention, coding, agent, image, and video input support.
Fast StepFun gateway model with 256K context, vision input, reasoning, and tool use.
Audio-driven avatar video technical release focused on lip sync, identity consistency, and long-video generation.
Flagship Qwen gateway model for demanding analysis, coding, and production agent workflows.
Qwen 3.6 fast vision-language model for visual reasoning, files, and agent workflows.
1.02T-parameter MiMo agent model with 1M context for demanding coding and long-horizon tasks.
Native omnimodal MiMo model with 1M context and text, image, video, and audio understanding.
Wan 2.7 image-to-video with first/last-frame control, continuation, driving audio, and 720P/1080P output.
Flagship DeepSeek V4 variant with 1.6T total parameters and 1M-token context.
Fast DeepSeek V4 variant for efficient chat, coding assistance, and agent loops.
MiMo expressive text-to-speech with premium built-in voices and natural-language control over pace, emotion, tone, dialect, and character performance.
Open native multimodal Kimi model for long-horizon coding, visual design, and agent-swarm workflows.
A 200K-context multimodal coding foundation model for image, video, file, GUI-agent, and tool-driven workflows.
Wan 2.7 natural-language video editing model for local/global edits and element replacement.
Wan 2.7 text-to-video model for prompt-driven video generation.
Wan 2.7 reference-to-video model for controlled generation from visual references.
GLM update for longer-horizon tasks, coding agents, and hybrid reasoning.
Wan 2.7 Pro image generation and editing with brand-color control, multi-reference consistency, and up to 4096×4096 output.
Wan 2.7 image model for text-to-image, editing, multi-reference generation, and stronger text rendering.
High-speed MiniMax variant for low-latency coding and agent workflows.
Open-weight MiniMax model for chat, coding, office work, and agentic tasks.
Native multimodal-agent Qwen release that preceded the served Qwen 3.6 and 3.7 gateway models.
Efficient GLM-5 turbo variant for faster reasoning, coding, and agent work.
GLM-5 release for coding, reasoning, and agent workflows.
Kling 3.0 Omni variant for reference-to-video and editing workflows.
Kling 3.0 generation with text-to-video, image-to-video, audio, and 720p to 4K output tiers.
Fast Kling 3.0 Turbo path for lower-latency video generation with audio.
Open StepFun flash model with 256K context and tool-use support.
Seedream 5.0-lite image model with controllable creation, retrieval, and stronger subject consistency.
Fast Seedance 2.0 video generation option for cost-efficient clips.
Seedance 2.0 flagship video generation model billed by output resolution.
Open native multimodal Kimi agent model with 256K context and vision-language training.
Seedance 1.5 Pro video model at a lower price point for silent video generation.
Small image-generation model focused on data quality over model size.
Seedream 4.5 image model with multi-image stable fusion and strong editing consistency.
Fast Hailuo 2.3 path for lower-cost image-to-video generation.
Hailuo 2.3 improves complex motion, physical actions, stylization, and micro-expression performance.
Fast Seedance 1.0 Pro option for quick, low-cost video clips.
Budget Hailuo 02 video model for text-to-video and image-to-video generation from 512p.
Seedance 1.0 Pro video model for general text-to-video and image-to-video generation.