Home / Models

Model catalogue

Claude & GPT at 66% of official rates · DeepSeek & Kimi K3 at 88% · GLM from 70% · Qwen at official rates. Discounts are a policy, not a promo — verify live models with GET /v1/models.

Model catalogue

Every major model on one key

Claude for reasoning, DeepSeek for code, Qwen for value, Kimi and GLM for multilingual workflows. Claude & GPT billed at 66% of official rates; other major models at competitive discounts. Live list per key via GET /v1/models.

OpenAI GPT
GPT-5.5 · Codex routes

GPT-5.5 flagship reasoning and Codex-tuned coding routes for the OpenAI agent ecosystem — Responses and Chat formats.

GPT-5.5 · flagship
Codex · coding routes
LiveResponses API
Claude
Anthropic · frontier trio

Opus 5: "frontier intelligence of Fable 5 at half the price." Sonnet 5: "the most agentic Sonnet yet." Haiku 4.5: fastest & cheapest Claude — SWE-bench 73.3%.

Opus 5 · 1M ctx
Fable 5 · 1M ctx
Sonnet 5 · 1M ctx
Haiku 4.5 · 200K
Online
Kimi
Moonshot · long context

"The world's first open 3T-class model" — 2.8T params, native vision, 1M-token window, designed for long-horizon coding & knowledge work.

Kimi K3
Online
GLM
Zhipu · multilingual

GLM-5.3: "every gain from post-training" — coding +50% vs 5.2, open-weights SOTA. GLM-5.2: MIT open flagship on true 1M context.

GLM-5.3
GLM-5.3 Flash
GLM-5.2
Online
DeepSeek
DeepSeek · value

V4-Pro (1.6T / 49B active): "performance rivaling the world's top closed-source models." V4-Flash (284B / 13B): fast, efficient and economical. Native 1M context.

V4 Pro / V4 Flash
Online
Qwen
Alibaba · value & vision

3.8-Max: 2.4T-parameter open MoE that "codes and delivers complete projects spanning 10+ days." 3.8-Flash: 6B active per token with 1M context. VL for vision.

3.8 Max / Flash
3-VL Plus · vision
Online
Your stack next?
GPT-5.5 / Codex ready

Multi-provider roadmap — tell us what you need on Telegram.

How usage is billed

Usage-based, metered per call. No subscription, credits never expire.

Input tokens 1×
Prompt and context sent to the model. Billed at each model's published rate — Claude & GPT at 66% of official rates, other major models at official rates.
Output tokens 5× input
Generated text. Most models bill output at 5× the input rate (DeepSeek 2×). Streaming replies count the same as non-streaming.
Cache hit (read) 0.1× input
Reused context from the provider's prompt cache — charged at 10% of the input rate. Keep your system prompt stable to hit the cache; this is the single biggest cost saver.
Cache write 1.25× input
Storing context for reuse (Claude models). Charged once at 125% of the input rate, then reads are 10%.
Thinking tokens = output
Reasoning tokens from extended-thinking models are billed exactly like output tokens — no hidden multiplier.
Discounts standing
Discounts are policy, not promotions: Claude & GPT at 66% of official rates; DeepSeek · Kimi · GLM · Qwen at official rates. They do not expire and are not tied to a promotion window.

Every call is itemised in the console — input, output, cache and thinking tokens per request. Verify live rates with GET /v1/models.