Models
25 models across video, image, audio, and chat — one API, one bill, pay only for what you run.
HOTSeedance 2.0
Seedance 2.0. Image-to-video and Text-to-video with synchronized audio.
HOTHappyhorse 1.0
Alibaba's top-of-arena text-to-video and image to video model. Currently leads the Arena Elo leaderboard.

Qwen Image 2512
Qwen image generation. Bilingual prompt support; strong on Chinese typography.
HOTNano Banana 2
Gemini Flash Image v2 · text-to-image & edit. Sharper photoreal than nano-banana.
HOTGPT Image 2
Next-gen image generation & editing with stronger compositional control and text rendering.
HOTFlux 2
Next-gen flagship with sharp subject fidelity and tight prompt adherence, in base and highest-fidelity tiers.

Nano Banana
Gemini 2.5 Flash Image. Persisted to Infer S3.

Ideogram 4
Ideogram 4 — best-in-class typography and graphic design, in Turbo, Default, and Quality tiers.

Uni 1
Strong photoreal generation & editing with up to 8 image refs, available in standard and higher-fidelity tiers.

Qwen Image Max Preview
Qwen Image Max preview. Higher-fidelity Qwen image variant; bilingual prompt support.

Wan 2.7
Alibaba's Wan 2.7 flagship image generation, available in pro and fast tiers.

Seedream 4
ByteDance's flagship image generator with strong photoreal output, for generation & editing.

Seedream 4.5
Newer Seedream revision with refined photoreal output, for generation & editing.

SAM 3
Open-vocabulary detection. Returns bounding boxes for any prompt.

Kling V3
Kuaishou's Kling 3.0 image-to-video and text-to-video, in Standard, Pro, and 4K tiers, plus native-audio variants.

Veo 3.1
Google's Veo 3.1 text-to-video with synchronized audio, in full and fast tiers.

Happyhorse 1.1
Alibaba's latest HappyHorse — top-of-arena text-to-video and image-to-video, now on the 1.1 release.

Wan 2.2
Alibaba's Wan 2.2 Flash — fast image-to-video for rapid iteration.

LTX 2.3
Lightricks' LTX 2.3 image-to-video and text-to-video across 1080p, 1440p, and 4K, in fast and higher-fidelity tiers.
HOTEleven v3
ElevenLabs Multilingual v3. Natural prosody, emotional range, 70+ languages, inline audio tags for laughter / whisper / sigh.

Qwen 3.6
Chat completions — strong reasoning, long context, multilingual — in flagship, mid-tier, and fast low-latency tiers.
HOTNano Banana 2 Lite
Gemini 3.1 Flash-Lite Image · text-to-image & edit. Fast, low-cost 1K image generation.

Seedream 5.0 Pro
Latest flagship Seedream with 2K photoreal output, for generation & editing.

Qwen 3.7 Max Preview
Chat completions — strong reasoning, long context, multilingual.

MiniMax-M3
Multimodal long-context chat from MiniMax with adaptive reasoning, served OpenAI-compatible.