Fresh stories
OpenRouter adds openrouter:shell to Responses API for hosted Linux code execution
OpenRouter added the openrouter:shell tool to its Responses API, allowing supported models to write and run code in hosted Linux containers. Containers are isolated to a workspace and return command output and execution results.

MiMo reportedly streams V2.6 Pro and Flash RL training metrics
MiMo is reportedly livestreaming RL training for its V2.6 Pro and Flash models, publishing batch data, harness composition, reward curves, and infrastructure metrics. Reported cost figures list the trillion-parameter Pro run at about $493,000.

Google releases Gemini 3.8 Live with bidirectional voice
Gemini 3.8 Live and its Extended Thinking variant add bidirectional voice, visual grounding, multilingual speech, and asynchronous tool calls. Google is rolling both models out through its API with audio watermarking via SynthID.


Coding-agent harness benchmarks report about 2x completion-time spread on GPT-6 Astra
Two evaluations found harness choice had a modest effect on coding-task success but a much larger effect on execution efficiency. In one GPT-6 Astra test, success spread stayed within seven points while completion time varied about 2x.

OpenRouter adds openrouter:shell to Responses API for hosted Linux code execution
OpenRouter added the openrouter:shell tool to its Responses API, allowing supported models to write and run code in hosted Linux containers. Containers are isolated to a workspace and return command output and execution results.

Anthropic merges Claude Chat and Cowork for Pro and Max users
Anthropic is combining Claude Chat and Cowork into one Claude experience for Pro and Max users on web, desktop, and mobile. Conversations can create Docs, Slides, and Design artifacts alongside longer-running agent work.

OpenAI releases model-misalignment disclosure criteria and timelines
OpenAI published criteria and timelines for tracking, investigating, and publicly disclosing model-misalignment incidents. The report covers unresolved cases and describes six recent examples, including an Astra model carrying jailbreaks.
MiMo reportedly streams V2.6 Pro and Flash RL training metrics
Zed opens Delta public beta for agent-led code review
Devin adds cloud Mac sandboxes for Xcode and iOS app testing
Google releases Gemini 3.8 Live with bidirectional voice
Top storiesthis week
TypeSafe AI launches Jev for predefined choices, scores, and probabilities
Jev returns predefined choices, scores, and probabilities instead of free-form text for bounded software decisions. TypeSafe claims roughly 150 ms responses and lower inference costs for those tasks.


Perplexity reports coding agents helped build CobbleDB, its DynamoDB replacement
Perplexity says two engineers and hundreds of persistent coding agents built CobbleDB, an internal key-value database for its search stack. It is optimized for repeated batch reads of prepared page records and is not offered externally.

ElevenLabs MCP server adds speech, music, image, and video generation
ElevenLabs’ MCP server now lets supported assistants generate speech, music, images, and video. It also exposes transcription, dubbing, sound effects, and tools for refining generated work.

Perplexity launches Portable Computer for Windows RTX PCs
Perplexity launched Portable Computer in its Windows app for supported NVIDIA RTX systems. Local inference requires at least 24GB of VRAM, while the update also adds local MCP support and scheduled tasks.

Polylane says one agent improved quality while cutting latency and cost
Polylane says it replaced role-specific sub-agents with one main agent and improved quality while reducing latency and cost. The report is a practitioner case study, not a general benchmark.



