Model
Model explorer

Inkling-Medium

OPEN
Thinking Machines Lab · Inkling family · released Sep 10, 2026

Thinking Machines Lab's September 10, 2026 middle tier — the Inkling line's first scale-up beyond July's Small (276B/12B). 320B/22B MoE, Apache 2.0 weights, and Tinker-metered API pricing at $1.20/$3.60 per Mtok (the lab's fine-tuning-served inference stack, not a conventional model API). The launch table is the lab's usual narrow but careful set: every shared row lands 2-4 points over Inkling-Small with a first-time Terminal-Bench 2.1 and BrowseComp filing, and GPQA/HLE rows that put it within reach of Sonnet 5's published independent numbers.

ReasoningCodingVisionFunction callingTool useAgentic
2424.1
Elo · rank #44
Parameters
320B
Active params
22B (MoE)
Context
1M tokens
Architecture
320B-total/22B-active sparse MoE trained with Tinker's fine-tuning stack, Apache 2.0 open weights; 1M-token context, natively audio+vision
License
Apache 2.0
Languages
—
API price (in/out)
$1.2 / $3.6
Modalities
text · vision · audio
Benchmark results
Bar shows position within the tracked field; marker = field best
BrowseCompAgents79.4%#25
best: Kimi K3.1 · 92.4%
CharXivVision84.6%#8
best: Qwen3.8-Flash-Next · 90.6%
GPQA DiamondReasoning88.4%#47
best: GPT-6 Astra · 96.0%
Humanity's Last ExamReasoning41.7%#29
best: Claude Opus 5 · 64.7%
IFBenchReasoning85.3%#1
best: this model · 85.3%
MMMU-ProVision77.9%#25
best: Claude Opus 4.7 · 85.5%
SWE-bench ProCoding59.7%#22
best: Claude Fable 5.1 · 81.2%
SWE-bench VerifiedCoding82.6%#9
best: Claude Opus 5 · 96.0%
Terminal-Bench 2.0Coding63.2%#49
best: Gemini 3.8 Flash · 89.4%
Run it locally
VRAM @ Q4
—
VRAM @ FP16
—
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
—
Fine-tune it
Permissive
QLoRA217.6 GB2× H200 141GB
LoRA681.6 GB4× B200 192GB
Full fine-tune5136.0 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $11.38 (2× H200 141GB)
API price $1.2/$3.6 · each benchmark row carries its own source badge (see methodology)