#AI models
9 articles tagged AI models. Past halfway to the twelve-article threshold.
Timeline
-
What model distillation is — how big AI models teach small ones to cut costs
-
Meta Muse Spark 1.3 pricing — the 21x discount is paid for with your prompts
-
Anthropic's Claude Fable 5.1 — same sticker price, cache reads cut 75%
-
Gemini 3.8 Flash launches — the fourth Flash model in four months, at $0.75 per million tokens
All articles
-
Tech · 3 min readWhat model distillation is — how big AI models teach small ones to cut costs
Definition — a small student model learns to imitate a large teacher, including the teacher's full probability distribution
-
Tech · 3 min readMeta Muse Spark 1.3 pricing — the 21x discount is paid for with your prompts
Price — standard 1.25/4.25 dollars per million tokens; Contributor 0.10/0.20. A 21x gap on output
-
Tech · 4 min readAnthropic's Claude Fable 5.1 — same sticker price, cache reads cut 75%
Price — input $10 and output $50 unchanged; cache reads $1.00 → $0.25 (−75%)
-
Tech · 2 min readGemini 3.8 Flash launches — the fourth Flash model in four months, at $0.75 per million tokens
Cadence — the fourth Flash model in under four months, three weeks after the last one
-
Tech · 3 min readNvidia pays Poolside $6bn — the company selling the chips will now build the models
Size — $6bn license plus $1bn investment. Poolside was valued at $12bn pre-money
-
Tech · 3 min readGPT-5.6 Sol API price cut (August 21) — output falls from $30 to $20 per million tokens
Price — input $5→$4, output $30→$20, cached input $0.50→$0.40 per million tokens. The 33% cut to output is the largest
-
Tech · 2 min readGLM-5.3 claims a 50% coding jump — but there is no API price
Release — August 14, 2026, by Z.ai of China
-
Tech · 4 min readDeepSeek V4-Pro launches — 1.6 trillion parameters, 49 billion switched on
V4-Pro-0813: 1.6tn total parameters, 49bn active per token, 1M-token context, 384K max output
-
Tech · 4 min readWhat is a context window — a model does not read all million tokens
Context window = total tokens visible in one request, input and output combined