Model
Model explorer

Granite 4.2 8B

OPEN
IBM · Granite 4.2 family · released Aug 25, 2026

IBM's August 25, 2026 Granite refresh, released Apache 2.0 in 3B, 8B and 30B sizes and aimed squarely at enterprise agentic workflows — tool selection and sequencing, instruction following, coding and native reasoning — rather than leaderboard position. Distributed through Hugging Face, Ollama, watsonx, OpenRouter, Replicate, LM Studio, DeepInfra and CoreWeave. Two Granite Speech 5.0 Turbo CTC models (470M, one non-commercial) shipped the same day with no LLM backbone, reaching ~12,600 RTFx on a single H200 against the Open ASR leaderboard's ~6,000 speed leaders; they are out of scope for this catalog. Only the 8B is entered here — IBM's blog publishes competitor charts as images without a machine-readable table, and no per-benchmark numerals could be sourced to a primary IBM page for any size, so the entry ships spec-only and stays unranked rather than importing a third-party aggregator's row. The 3B and 30B siblings are deliberately deferred until IBM's Hugging Face technical write-up can be transcribed.

ReasoningCodingVisionFunction callingTool useAgentic
Elo · unrated
Parameters
8B
Active params
Undisclosed
Context
128K tokens
Architecture
8B hybrid Mamba-2/transformer Granite 4 architecture with native reasoning and a speculative-decoding layer; trained through foundational RL, an agentic RL phase for enterprise tasks and RLHF, on 1T tokens of CodeAlchemy synthetic code plus a mid-training reasoning stage
License
Apache 2.0
Languages
API price (in/out)
No hosted API
Modalities
text
Benchmark results
Bar shows position within the tracked field; marker = field best
Run it locally
VRAM @ Q4
6 GB
VRAM @ FP16
17 GB
Fits on (Q4)
RTX 3060 12GBRTX 4070 Ti 16GBRTX 3090 24GBRTX 4090 24GBRTX 5090 32GBM4 Pro 48GBM3 Max 128GBM3 Ultra 512GBA100 80GBH100 80GBH200 141GBB200 192GB
Throughput data unavailable.
Quantizations
Fine-tune it
Permissive
QLoRA7.0 GB1× RTX 3060 12GB
LoRA18.6 GB1× RTX 3090 24GB
Full fine-tune130.0 GB1× H200 141GB
QLoRA SFT on ~10k samples ≈ $4.68 (1× RTX 3060 12GB)
Granite 4.2 family
Elo progression across releases
API price weights · each benchmark row carries its own source badge (see methodology)