Fresh stories

Claude Code Mods walkthrough says plugins run unsandboxed with user permissions
A Claude Code Mods walkthrough says plugins retain state and run with the user's permissions. Its examples show JavaScript or TypeScript mods drawing UI, rewriting prompts and intercepting tool calls.
Teknium says Hermes checks nearly 500 plugins for malware, not security hardening
Teknium says Hermes lists nearly 500 plugins, but only eight are officially tested. Community submissions get malware and guideline checks, not broader security hardening.

Matt Pocock releases skills v1.3 with /retro transcript reviews
Matt Pocock's skills v1.3 adds /retro to find workflow improvements in old agent transcripts. Its migration prompt compares installed skills, renames CONTEXT.md to GLOSSARY.md and reviews recent skill usage.


Claude Code Mods walkthrough says plugins run unsandboxed with user permissions
A Claude Code Mods walkthrough says plugins retain state and run with the user's permissions. Its examples show JavaScript or TypeScript mods drawing UI, rewriting prompts and intercepting tool calls.

Teknium says Hermes checks nearly 500 plugins for malware, not security hardening
Teknium says Hermes lists nearly 500 plugins, but only eight are officially tested. Community submissions get malware and guideline checks, not broader security hardening.

Pi Durable Android prototype runs without Termux
Mario Zechner demonstrates a native Android app built on his phone with Pi Durable in two days. The prototype also adds paged transcript loading, but he describes it as a demo rather than a product.

Pi Durable agent resumes after redeploying itself with SQLite memory
Raunak demonstrates a Pi Durable agent that resumes after redeploying itself. Its Durable Object setup stores memory and chat history in SQLite, preserving both across deployments.
Peter Gostev reports Opus gains roughly 250 Elo in agent chess test
Geoffrey Huntley demos an SBCL kernel that modifies itself while running
NerfBench publishes method for testing Claude Opus 5.5 performance changes
Matt Pocock releases skills v1.3 with /retro transcript reviews
Top storiesthis week
Claude Code 2.1.289 fixes Read-deny bypasses via file references and symlinks
Claude Code 2.1.289 fixes file references and symlinks bypassing Read deny rules. It also addresses managed approval issues and adds teammate agent spawning.


T3 Code ships Orchestrator V2 with cross-harness agent delegation
T3 Code's nightly build ships Orchestrator V2 with official-registry ACP providers and built-in MCP delegation across harnesses and models. Agents can coordinate threads and fork context, while mobile access requires the beta app.

Aleph Alpha releases Kolibri with a 1M-token context window
Aleph Alpha released Kolibri, an Apache 2.0 open-weight MoE with a 1M-token context window. The model has 78B total parameters and 3.46B active parameters, plus reasoning, tool calling and German-focused training.

Astra's Elo falls in Peter Gostev's repeated-play agent chess test
Astra's Elo fell during Peter Gostev's chess test, where agents can take notes and choose opponents. Later results suggest Opus improved, and GPT-6.1 Sol and Fable were added to the benchmark.

Kevin Kern delegates coding tasks from Opus 5.5 to GPT-6.1 Sol
Kevin Kern uses Opus 5.5 for orchestration and UI work, with GPT-6.1 Sol handling delegated coding. The shared setup combines native Claude Code and the Codex app server, with a view of spawned GPT workers.





