Every major model on one key
Claude for reasoning, DeepSeek for code, Qwen for value, Kimi and GLM for multilingual workflows. Claude & GPT billed at 66% of official rates; other major models at competitive discounts. Live list per key via GET /v1/models.
GPT-5.5 flagship reasoning and Codex-tuned coding routes for the OpenAI agent ecosystem — Responses and Chat formats.
Opus 5: "frontier intelligence of Fable 5 at half the price." Sonnet 5: "the most agentic Sonnet yet." Haiku 4.5: fastest & cheapest Claude — SWE-bench 73.3%.
"The world's first open 3T-class model" — 2.8T params, native vision, 1M-token window, designed for long-horizon coding & knowledge work.
GLM-5.3: "every gain from post-training" — coding +50% vs 5.2, open-weights SOTA. GLM-5.2: MIT open flagship on true 1M context.
V4-Pro (1.6T / 49B active): "performance rivaling the world's top closed-source models." V4-Flash (284B / 13B): fast, efficient and economical. Native 1M context.
3.8-Max: 2.4T-parameter open MoE that "codes and delivers complete projects spanning 10+ days." 3.8-Flash: 6B active per token with 1M context. VL for vision.
Multi-provider roadmap — tell us what you need on Telegram.
How usage is billed
Usage-based, metered per call. No subscription, credits never expire.
Every call is itemised in the console — input, output, cache and thinking tokens per request. Verify live rates with GET /v1/models.