Model
Model explorer

Inkling-Small

OPEN
Thinking Machines Lab · Inkling family · released Jul 30, 2026

Thinking Machines Lab's second open-weight release, two weeks after flagship Inkling: a quarter-sized MoE (276B/12B active vs. 975B/41B) that the model card says matches Inkling on many benchmarks at lower cost and latency, aimed at cheaper multimodal-agent deployment.

ReasoningCodingVisionFunction callingTool useAgentic
2413.8
Elo · rank #31
Parameters
276B
Active params
12B (MoE)
Context
1M tokens
Architecture
Sparse MoE, 276B total / 12B active (roughly a quarter of Inkling's size), same text/vision/audio multimodal design and adjustable inference-time thinking effort as its larger sibling
License
Apache 2.0
Languages
API price (in/out)
$0.58 / $1.44
Modalities
text · vision · audio
Benchmark results
Bar shows position within the tracked field; marker = field best
CharXivVision81.3%#8
best: Muse Spark 1.1 (xhigh) · 88.4%
IFBenchReasoning82.2%#3
best: MiniMax M3 · 83.0%
MMMU-ProVision74.0%#29
best: Claude Opus 4.7 · 85.5%
SWE-bench ProCoding55.9%#29
best: Claude Fable 5 · 80.0%
SWE-bench VerifiedCoding80.2%#18
best: Claude Opus 5 · 96.0%
Run it locally
VRAM @ Q4
VRAM @ FP16
Fits on (Q4)
Multi-node cluster required
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA187.7 GB1× B200 192GB
LoRA587.9 GB4× B200 192GB
Full fine-tune4429.8 GBbeyond 8× B200
QLoRA SFT on ~10k samples ≈ $4.29 (1× B200 192GB)
Inkling family
Elo progression across releases
API price $0.58/$1.44 · each benchmark row carries its own source badge (see methodology)