TOPIC / AGENTS
Agents
Agent systems, tool use, browser automation and long-running workflows.
01 / FEATURED
NEWS · AGENTS
Anthropic expands tool-use controls for long-running Claude workflows
New controls focus on safer, more bounded long-running tool use.
Read the story →WHY THIS LEADS
New controls focus on safer, more bounded long-running tool use.02 / LATEST AGENTS
Latest stories
News leads this topic. Analysis and experiments appear when they add practical context.
Anthropic expands tool-use controls for long-running Claude workflows
New controls focus on safer, more bounded long-running tool use.
Hugging Face releases an open toolkit for reproducible browser-agent benchmarks
The toolkit focuses on repeatable environments and comparable browser-agent runs.
03 / TOPIC MAP
Inside Agents
The recurring adjacent topics shaping this desk.
04 / FROM THE LAB & DESK
Beyond the news cycle
Practical experiments and deeper analysis when the topic needs more than a headline.
EXPERIMENT / AGENTS
We gave three coding agents the same broken production task
Same repository, same acceptance criteria, same time box. The difference was not where we expected it.
See the experiment →DEEP DIVE / AGENTS
Benchmarks are easy. Reliable agent evaluation is not.
What breaks when evaluation moves from static model scores into long-running agent behavior.
09 SEPT 2026Agents groups reporting, experiments and analysis by editorial relevance. Stories are connected through reviewed topic relationships rather than display-only tags.
LAST REVIEWED · 11 SEPT 2026