TOPIC / MODELS
Models
Tracking how frontier and open models change capability, cost, reliability and what teams can actually build with them.
01 / FEATURED
NEWS · MODELS
Anthropic expands tool-use controls for long-running Claude workflows
New controls focus on safer, more bounded long-running tool use.
Read the story →WHY THIS LEADS
New controls focus on safer, more bounded long-running tool use.02 / LATEST MODELS
Latest stories
News leads this topic. Analysis and experiments appear when they add practical context.
Anthropic expands tool-use controls for long-running Claude workflows
New controls focus on safer, more bounded long-running tool use.
Anthropic expands long-context controls for production Claude workloads
New controls focus on when deeper reasoning should be invoked and how teams bound cost.
Google DeepMind publishes a new method for evaluating multi-step reasoning reliability
The work shifts attention from single-answer accuracy toward repeatable reasoning behavior.
Meta updates its open model family with stronger tool-use and coding behavior
The release narrows the gap for teams that need deployable weights and controllable infrastructure.
Hugging Face launches a reproducibility suite for open model benchmarks
The toolkit records prompts, environments and evaluator settings alongside leaderboard results.
Microsoft adds model routing policies for mixed frontier and small-model deployments
Enterprise teams can route requests by latency, cost and governance constraints.
OpenAI’s new reasoning model changes the cost curve for production AI
The headline is not the benchmark score. It is what happens when stronger reasoning becomes cheap enough to sit inside everyday products.
03 / TOPIC MAP
Inside Models
The recurring adjacent topics shaping this desk.
04 / FROM THE LAB & DESK
Beyond the news cycle
Practical experiments and deeper analysis when the topic needs more than a headline.
EXPERIMENT / MODELS
Testing local models for structured tool calling under latency pressure
Same tool schema, same evaluation set, different model sizes and serving setups. We measured reliability before speed.
See the experiment →DEEP DIVE / MODELS
The inference stack is becoming the product
Why routing, caching, observability and evaluation are moving from supporting infrastructure into the core AI experience.
10 SEPT 2026DEEP DIVE / MODELS
Benchmarks are easy. Reliable agent evaluation is not.
What breaks when evaluation moves from static model scores into long-running agent behavior.
09 SEPT 2026Models groups reporting, experiments and analysis by editorial relevance. Stories are connected through reviewed topic relationships rather than display-only tags.
LAST REVIEWED · 11 SEPT 2026