Fresh stories
Qwen 3.8 Max 2.4T open weights ship as text-only, users say
LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.


Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Composio and Ante tests reported that the same models behaved very differently by harness. DeepSeek V4 Flash ranged from 47% to 67% task success and $0.019 to $0.104 per task across harnesses.
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing
OpenRouter upgraded Auto Router to classify prompts into about 30 task types, then route by anonymized 7-day spend share and cost tier. OpenRouter says the max tier beat the old router across five benchmark domains.


Speculative decoding tests report acceptance drop from 0.71 to 0.18 after ~32K context
A practitioner report found speculative-decoding acceptance fell from 0.71 to 0.18 beyond about 32K context. Separate DSpark and mlx-dspark tests reported speedups on RTX and Apple Silicon setups.

Qwen 3.8 Max 2.4T open weights ship as text-only, users say
LocalLLaMA users said Qwen 3.8 Max 2.4T open weights are text-only while the API keeps vision support. A linked Qwen3.8-27B ModelScope page reportedly returned 404 before release.

Qwen 122B runs on older laptop with llama.cpp in user test
A LocalLLaMA post ran a 122B Qwen model on an older laptop with llama.cpp. The run had very long load and generation times, while another report put Qwen 3.6 35B at 21 tok/s on a Radeon 7600 after ROCm tuning.

Meta releases Muse Glimmer 30B as an Apache 2.0 open-weight local agent model
Meta released Muse Glimmer, a 30B Apache 2.0 dense model for local agent workflows. Reports cite 4-bit builds under 20GB, vision input, function calling, 131K context, and day-0 support in Hugging Face, vLLM, SGLang, Ollama, and MLX.
Composio and Ante benchmark coding agent harnesses with 47%–67% success range
Reports say OpenClaw exposed missing auth on gym booking cancellation API
OpenAI releases GPT-5.6-Cyber for approved Daybreak Blue and Red teams
OpenRouter updates Auto Router with 30 task types and 7-day spend-based routing
Top storiesthis week
Microsoft Copilot traces report 87% of LLM calls came from agents
A Microsoft Copilot trace analysis said 87% of LLM calls came from the agent, not direct user turns. Related posts warned token use and web requests can scale far faster than human prompt counts.


Claude Code adds layered prompt-injection defenses by default next week
Anthropic engineers said Claude Code is slated to add model training, probes, and auto-mode layers by default next week. They said Claude models now largely resist practical prompt injection.

OpenClaw reports missing auth checks in gym waitlist API
Simon Willison quoted OpenClaw saying a gym API allowed cancelling other users’ reservations and moving a waitlisted user up one spot. Replies treated it as both an agent safety failure and a basic authorization bug.

OpenAI faces Artifactory monitoring questions as postmortem is promised
Security researchers disputed how OpenAI detected and investigated the Artifactory incident. Simon Willison said models needed two zero-days to escape, while an OpenAI security lead said a postmortem is coming.

Echo Gap paper reports agents endorsed 31%–54% of their own wrong answers
The Echo Gap paper found self-improving agents can store wrongly self-scored episodes. Tested models endorsed 31% to 54% of their own wrong answers, while other work proposed RL-trained harness state and in-model memory.



