Astrolabe
ProductModelsPricingDocs
Sign inRun first request
Astrolabe

One API chooses the lowest-cost model capable of the job.

PricingModelsProductDocumentationTermsPrivacyRefunds
© 2026 Astrolabe CloudOne API. Automatic routing. Lower cost.
Astrolabe
ProductModelsPricingDocs
Sign inRun first request
Universal model directoryCatalog 2026-08-02

Every model. One gateway.

Explore every billable text and embedding model available through Astrolabe. Compare benchmark quality, capabilities, context, token rates, and estimated cost for real workloads.

Models
345
Providers
57
Benchmarked
101

Filters

Provider

Capabilities

Availability
Token price
Context
Benchmark score

Artificial Analysis index. Unrated models are excluded when a minimum is set.

Price per action

345 of 345 models

A

Astrolabe Auto

astrolabe/auto

Managed routing that selects the lowest-cost model capable of satisfying each request.

managed routeautomatic
Tool callingStructured outputVisionReasoning

Benchmarks

Int—Cod—Age—

Context

—

Input / output

Varies / Varies

Coding task

Dynamic

estimated / action

O

GPT-5.6 Sol

openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int58.9Cod77.4Age54.0

Context

1.1M

Input / output

$5 / $30

Coding task

$0.140

estimated / action

X

Grok 4.5

x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int53.8Cod72.4Age45.7

Context

500K

Input / output

$2 / $6

Coding task

$0.038

estimated / action

O

GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int51.2Cod71.4Age45.6

Context

1.1M

Input / output

$0.100 / $0.600

Coding task

$0.0028

estimated / action

G

Gemini 3.5 Flash

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int50.2Cod70.1Age37.4

Context

1.0M

Input / output

$1.5 / $9

Coding task

$0.042

estimated / action

A

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int47.2Cod63.0Age40.8

Context

1M

Input / output

$3 / $15

Coding task

$0.075

estimated / action

M

Kimi K2.6

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int44.2Cod61.8Age30.3

Context

262K

Input / output

$0.600 / $3.41

Coding task

$0.016

estimated / action

D

DeepSeek V4 Pro

deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int44.3Cod59.4Age36.4

Context

1.0M

Input / output

$0.435 / $0.870

Coding task

$0.0070

estimated / action

M

MiniMax M3

minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int44.4Cod58.6Age35.4

Context

1.0M

Input / output

$0.300 / $1.2

Coding task

$0.0066

estimated / action

X

MiMo-V2.5-Pro

xiaomi/mimo-v2.5-pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int42.2Cod60.2Age29.1

Context

1.1M

Input / output

$0.435 / $0.870

Coding task

$0.0070

estimated / action

Q

Qwen3.7 Plus

qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int39.0Cod55.9Age20.8

Context

1M

Input / output

$0.320 / $1.28

Coding task

$0.0070

estimated / action

M

MiniMax M2.7

minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int38.1Cod52.6Age25.6

Context

205K

Input / output

$0.250 / $1

Coding task

$0.0055

estimated / action

X

Grok 4.3

x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int37.6Cod42.2Age24.1

Context

1M

Input / output

$1.25 / $2.5

Coding task

$0.020

estimated / action

G

Gemma 4 31B

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int29.4Cod43.4Age14.4

Context

262K

Input / output

$0.100 / $0.340

Coding task

$0.0020

estimated / action

A

Claude Fable 5

anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int59.9Cod76.5Age52.8

Context

1M

Input / output

$10 / $50

Coding task

$0.250

estimated / action

A

Claude Opus 4.8

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int55.7Cod74.3Age47.2

Context

1M

Input / output

$5 / $25

Coding task

$0.125

estimated / action

O

GPT-5.5

openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int54.8Cod74.9Age44.9

Context

1.1M

Input / output

$5 / $30

Coding task

$0.140

estimated / action

ZA

GLM 5.2

z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int51.1Cod68.8Age43.1

Context

1.0M

Input / output

$0.351 / $1.1

Coding task

$0.0068

estimated / action

O

GPT-5.4

openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int51.4Cod71.1Age41.1

Context

1.1M

Input / output

$2.5 / $15

Coding task

$0.070

estimated / action

M

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int41.9Cod60.8Age29.6

Context

262K

Input / output

$0.730 / $3.5

Coding task

$0.018

estimated / action

D

DeepSeek V4 Flash

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int40.3Cod56.2Age31.1

Context

1.0M

Input / output

$0.140 / $0.280

Coding task

$0.0022

estimated / action

ZA

GLM 5.1

z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

textauto-routed
Tool callingStructured outputReasoning

Benchmarks

Int40.2Cod55.8Age29.9

Context

205K

Input / output

$0.966 / $3.04

Coding task

$0.019

estimated / action

O

GPT-5.4 Mini

openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int40.0Cod56.1Age30.2

Context

400K

Input / output

$0.750 / $4.5

Coding task

$0.021

estimated / action

O

GPT-5.4 Nano

openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int38.2Cod56.1Age27.5

Context

400K

Input / output

$0.200 / $1.25

Coding task

$0.0057

estimated / action

O

GPT-5.6 Luna Pro

openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int—Cod—Age—

Context

1.1M

Input / output

$0.100 / $0.600

Coding task

$0.0028

estimated / action

O

GPT-5.6 Sol Pro

openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int—Cod—Age—

Context

1.1M

Input / output

$5 / $30

Coding task

$0.140

estimated / action

A

Claude Opus 5

anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int60.7Cod78.0Age55.3

Context

1M

Input / output

$5 / $25

Coding task

$0.125

estimated / action

M

Mistral Small 4

mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

textauto-routed
Tool callingStructured outputVisionReasoning

Benchmarks

Int19.6Cod26.6Age4.7

Context

262K

Input / output

$0.150 / $0.600

Coding task

$0.0033

estimated / action

M

Kimi K3

moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int57.1Cod76.2Age50.1

Context

1.0M

Input / output

$3 / $15

Coding task

$0.075

estimated / action

O

GPT-5.6 Terra

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int55.0Cod76.7Age47.4

Context

1.1M

Input / output

$1 / $6

Coding task

$0.028

estimated / action

A

Claude Sonnet 5

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int53.4Cod71.5Age46.7

Context

1M

Input / output

$2 / $10

Coding task

$0.050

estimated / action

A

Claude Opus 4.7

anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int53.5Cod73.6Age44.4

Context

1M

Input / output

$5 / $25

Coding task

$0.125

estimated / action

M

Muse Spark 1.1

meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int50.6Cod71.3Age37.5

Context

1.0M

Input / output

$1.25 / $4.25

Coding task

$0.025

estimated / action

D

DeepSeek V4 Flash 0731

deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

textdirect
Tool callingStructured outputReasoning

Benchmarks

Int49.9Cod69.1Age45.7

Context

1.0M

Input / output

$0.090 / $0.180

Coding task

$0.0014

estimated / action

G

Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int50.1Cod69.2Age38.7

Context

1.0M

Input / output

$1.5 / $7.5

Coding task

$0.037

estimated / action

G

Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int46.5Cod68.8Age21.4

Context

1.0M

Input / output

$2 / $12

Coding task

$0.056

estimated / action

Q

Qwen3.7 Max

qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

textdirect
Tool callingStructured outputReasoning

Benchmarks

Int46.0Cod66.0Age30.6

Context

1M

Input / output

$1.48 / $4.43

Coding task

$0.028

estimated / action

T

Hy3 preview

tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

textdirect
Tool callingReasoning

Benchmarks

Int41.2Cod58.8Age30.7

Context

262K

Input / output

$0.063 / $0.210

Coding task

$0.0013

estimated / action

NA

Nex-N2-Pro

nex-agi/nex-n2-pro

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

textdirect
Tool callingVisionReasoning

Benchmarks

Int41.0Cod59.1Age31.0

Context

262K

Input / output

$0.250 / $1

Coding task

$0.0055

estimated / action

T

Inkling

thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

textdirect
Tool callingVisionReasoning

Benchmarks

Int40.7Cod52.1Age32.3

Context

1.0M

Input / output

$1 / $4.05

Coding task

$0.022

estimated / action

T

Inkling Small

thinkingmachines/inkling-small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is positioned as the smaller, more efficient member of...

textdirect
Tool callingVisionReasoning

Benchmarks

Int40.2Cod52.9Age30.8

Context

524K

Input / output

$0.500 / $1.2

Coding task

$0.0086

estimated / action

Q

Qwen3.6 Plus

qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int39.6Cod54.5Age27.6

Context

1M

Input / output

$0.325 / $1.95

Coding task

$0.0091

estimated / action

X

Grok Build 0.1

x-ai/grok-build-0.1

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int39.8Cod51.5Age28.0

Context

256K

Input / output

$1 / $2

Coding task

$0.016

estimated / action

X

MiMo-V2.5

xiaomi/mimo-v2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int37.2Cod56.8Age23.7

Context

1.1M

Input / output

$0.140 / $0.280

Coding task

$0.0022

estimated / action

Q

Qwen3.6 27B

qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int37.1Cod53.7Age27.0

Context

262K

Input / output

$0.300 / $2

Coding task

$0.0090

estimated / action

N

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

textdirect
Tool callingStructured outputReasoning

Benchmarks

Int37.8Cod49.3Age27.4

Context

512K

Input / output

$0.600 / $3.6

Coding task

$0.017

estimated / action

G

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int36.5Cod49.3Age26.8

Context

1.0M

Input / output

$0.300 / $2.5

Coding task

$0.010

estimated / action

K

KAT-Coder-Pro V2

kwaipilot/kat-coder-pro-v2

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...

textdirect
Tool callingStructured output

Benchmarks

Int33.7Cod59.5Age15.5

Context

262K

Input / output

$0.300 / $1.2

Coding task

$0.0066

estimated / action

O

GPT-5.1

openai/gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int36.9Cod49.4Age21.0

Context

400K

Input / output

$1.25 / $10

Coding task

$0.042

estimated / action

A

Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

textdirect
Tool callingStructured outputVisionReasoning

Benchmarks

Int36.4Cod52.1Age24.6

Context

1M

Input / output

$3 / $15

Coding task

$0.075

estimated / action

One model ID

Or let Astrolabe choose.

Use astrolabe/auto to route each request to the lowest-cost capable model in this catalog.

Run first request
Astrolabe

One API chooses the lowest-cost model capable of the job.

PricingModelsProductDocumentationTermsPrivacyRefunds
© 2026 Astrolabe CloudOne API. Automatic routing. Lower cost.