Guides & Explainers
In-depth written guides on AI concepts, from beginner to advanced
Large Language ModelsDeepSeek V4-Flash: The Cheapest Model on the Roster
V4-Flash costs $0.22/$0.66 per million tokens off-peak with a 1M context window - the lowest first-party vendor API price we cover, and no longer in beta.
Nemotron 3.5 Lightning: A 30B Agentic Model on One GPU
NVIDIA's Mamba-MoE hybrid runs 30B parameters with 3B active on a single H100, under a commercially-approved open licence. Specs, the 256k context caveat, and who it suits.
Qwen3.8-Max: 2.4 Trillion Parameters, and Cheaper Than the Model It Replaces
Alibaba's largest model runs 2.4T parameters with 95B active. The weights shipped 12 August under a custom licence, and despite near-universal reporting to the contrary it is text-only - the multimodal sibling is Qwen3.8-27B.
Grok 4.6: Same Foundation, and a Pricing Cliff at 200k Tokens
xAI's Grok 4.6 reuses Grok 4.5's 1.5T foundation. Its two-tier pricing doubles input and output above 200k tokens of prompt - here is what that costs in practice.
Gemini 3.7 Flash: Half the Price, With an Expiry Date
Gemini 3.7 Flash costs $0.75/$3.75 per million tokens until 31 December 2026, then doubles. Full specs, what the Pro-line delay means, and how it compares on price.
GLM-5.3 Explained: Same Base Model, Better Training
Z.ai's GLM-5.3 reuses the GLM-5.2 base and takes every gain from post-training. What the vendor-reported benchmarks show, where it trails Mythos 5, and why the weights matter.
How Large Language Models Work: The Technology Behind Modern AI
How large language models actually process text and generate responses, from tokenization to the transformer architecture to training.
Showing 13–19 of 19 guides