Fresh stories
Google adds the Antigravity harness to Gemini managed agents
Google added the Antigravity harness to Gemini managed agents in AI Studio and the Interactions API. The Files and Credentials APIs move data into sandboxes and control agent access.

Ramp's 137-task Accounting Bench finds the best model fully solves 21% of tasks
Ramp's 137-task benchmark, built with accounting professionals, found the best model fully solved 21% of tasks even with three attempts. Claude Fable 5.1 led partial-credit scores, but the benchmark's best full-solution rate was only 21%.

Anthropic merges Claude Chat and Cowork for Pro and Max users
Anthropic is combining Claude Chat and Cowork into one Claude experience for Pro and Max users on web, desktop, and mobile. Conversations can create Docs, Slides, and Design artifacts alongside longer-running agent work.


Goodfire says activation probes could flag reward hacking in real time
Goodfire reports reward hacking in 50% to 96% of studied rollouts and says internal activation probes could flag the behavior live. Goodfire says amplifying the identified signal increases shortcut use and attempts to avoid detection.

Google adds the Antigravity harness to Gemini managed agents
Google added the Antigravity harness to Gemini managed agents in AI Studio and the Interactions API. The Files and Credentials APIs move data into sandboxes and control agent access.

OpenAI launches Astra for Law with GPT-6 Astra via Trusted Access
OpenAI's Astra for Law combines GPT-6 Astra with a legal search index and integrations for legal work. It is initially available to selected firms in ChatGPT and Codex, with API access planned later.

Raindrop opens Simulations for agent changes on every pull request
Raindrop Simulations generates mocked services, databases, and environments to test agent changes before deployment. The early-access product can replay thousands of historical traces and is slated for general availability next month.
Ramp's 137-task Accounting Bench finds the best model fully solves 21% of tasks
Exa launches a historical web index with 400 billion snapshots
Figure says Helix 2.5 raises zero-shot household-task success from 9% to 56% across 30 unseen homes
Anthropic merges Claude Chat and Cowork for Pro and Max users
Top storiesthis week
OpenRouter adds openrouter:shell to Responses API for hosted Linux code execution
OpenRouter added the openrouter:shell tool to its Responses API, allowing supported models to write and run code in hosted Linux containers. Containers are isolated to a workspace and return command output and execution results.


MiMo reportedly streams V2.6 Pro and Flash RL training metrics
MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.

Zed opens Delta public beta for agent-led code review
Zed released Delta in public beta on macOS, Linux, and Windows as a multiplayer environment for coding with agents and reviewing changes. DeltaDB stores intermediate edits, comments, and agent conversations.

Periodic Labs releases open Neon model for X-ray diffraction analysis
Periodic Labs released its open Neon model for X-ray diffraction analysis. The company says mid-training and RL on lab data raised accuracy from 2.7% to 55.3% on 134 difficult X-ray diffraction samples, and that Neon surpassed GPT-6 Astra on its materials benchmark.

Google releases Gemini 3.8 Live with bidirectional voice
Gemini 3.8 Live and its Extended Thinking variant add bidirectional voice, visual grounding, multilingual speech, and asynchronous tool calls. Google is rolling both models out through its API with audio watermarking via SynthID.



