Fresh stories
Aleph Alpha releases Kolibri with a 1M-token context window
Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.


Aleph Alpha releases Kolibri with a 1M-token context window
Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.

Claude Code 2.1.289 fixes Read-deny bypasses via file references and symlinks
Claude Code 2.1.289 fixes file references and symlinks bypassing Read deny rules. It also addresses managed approval issues and adds teammate agent spawning.

T3 Code ships Orchestrator V2 with cross-harness agent delegation
T3 Code's nightly build ships Orchestrator V2 with official-registry ACP providers and built-in MCP delegation across harnesses and models. Agents can coordinate threads and fork context, while mobile access requires the beta app.

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.
Kevin Kern delegates coding tasks from Opus 5.5 to GPT-6.1 Sol
Kevin Kern uses Opus 5.5 for orchestration and UI work, with GPT-6.1 Sol handling delegated coding. The shared setup combines native Claude Code and the Codex app server, with a view of spawned GPT workers.
SemiAnalysis reports configuration failures in Vultr's ClusterMAX test
SemiAnalysis found vulnerable libraries, login-pod setup errors and storage configuration problems that left pods stuck in Vultr's ClusterMAX test. It also reported strong Lustre throughput alongside networking reliability problems.
Pi Durable prototype resumes Android agent work after worker restarts
Mario Zechner's Pi Durable prototype runs an agent runtime on Android while calling remote models. Shared sessions and subagents are supported, and work resumes after worker restarts.
Top storiesthis week
Cua launches Spaces for agent-controlled desktops
Cua Spaces provides agent-controlled desktops on local or remote machines. The free, source-available release includes approved app-session transfers, a locked local Keyvault, and separate human and agent cursors.


T3 Code adds cross-provider child agents in its upcoming nightly overhaul
T3 Code's upcoming nightly overhaul adds child agents across providers, Pi support, and MCP thread controls. Its creator warns of instability as model switching, queueing, and automatic resumption are introduced.

NerfBench says Claude Opus 5.5's 94.2% score is within normal variance
NerfBench reports Claude Opus 5.5 at 94.2% of its launch baseline but says the drop remains within normal variance. Theo disputes the degradation framing and offers to fund independently audited testing.

xAI releases an experimental TypeScript SDK for Grok
xAI's experimental TypeScript SDK supports Grok text, voice, image, and video models, plus server-side tools. Code execution runs in isolated, stateless Python sandboxes without network or filesystem access.

OpenAI describes Dot coordinating Codex tasks across apps
OpenAI says Dot retains context across apps, coordinates Codex tasks, and flags work needing attention. Practitioner reports describe using it to reduce the overhead of managing coding agents.







