Fresh stories
Qwen3.8-Max launches on OpenRouter with 1M-token context
Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

Hermes Agent ships Herald with Ironproxy secrets lockdown
Teknium shipped the Herald release of Hermes Agent as a desktop-agent and runtime update. The release adds voice chats, desktop plugins, Agent2Agent/webhooks, productivity skills, grounded research, Ironproxy secrets lockdown, and trace-driven token-efficiency work from 250,000 conversations.

MiniMax releases H3 open weights with day-zero vLLM-Omni support
MiniMax released H3 weights on Hugging Face for text-to-video, image-to-video, reference-to-video, and editing workflows. vLLM-Omni, ComfyUI, SGLang Diffusion, and fal added support at launch.


OpenAI updates GPT-Live with continuous audio and one-round-trip WebRTC startup
OpenAI says GPT-Live can listen while speaking using continuous audio, async reasoning and tool use, one-round-trip WebRTC startup, and async context compaction. Staff said the rebuilt stack removes a separate turn detector.

Qwen3.8-Max launches on OpenRouter with 1M-token context
Alibaba's Qwen3.8-Max is live on Venice and OpenRouter while open weights are still described as coming soon. Reports cite a 2.4T-parameter model with strong Vals and vision benchmark results.

DeepSeek V4 Flash adds Baseten and Together AI serving with 1M-token context
Baseten and Together AI added DeepSeek V4 Flash with a 1M-token context window, reasoning-effort controls, and DSpark decoding. ValsAI ranked it the cheapest model above 60 on its index, and Nous promoted a short 90% discount.

Agent studies trace reliability failures to harness design and verifier quality
A Renmin survey, Cline's SDK notes, and AutomationBench trajectory reviews point beyond context windows to harness design and verifier quality. A failure-mode paper maps issues across models, memory, tools, users, and environment.
Hermes Agent ships Herald with Ironproxy secrets lockdown
MiniMax H3 releases Hugging Face weights and fal video endpoints
Cursor users report hard-to-audit coding-agent runs and hidden routing
MiniMax releases H3 open weights with day-zero vLLM-Omni support
Top storiesthis week
Codex users route tasks across GPT-5.6 Sol, Terra, and Luna to cut token cost
Practitioners reported better Codex multi-agent runs by raising concurrency and splitting work across Sol, Terra, and Luna. One workflow sends deploy tasks to Luna Max to preserve Sol tokens.


Cline raises free DeepSeek Flash quota 3x for coding agents
Cline tripled its free quota while Nous discounted Flash 0731 by 90% for a week and OpenHands offered free cloud use. OpenCode reported 8T Flash tokens on Aug. 1.

Google DeepMind introduces SkillSmith for KV-cache skill composition in Gemma 3 4B
Google DeepMind introduced SkillSmith, a method that treats prefix weights or KV-cache states as an input modality so a frozen Gemma 3 4B model can synthesize new skill prefixes at inference time. Reported Composite-SNI Elo improved when cache composition was combined with text descriptions, making it a research artifact rather than a deployable runtime.

Vercel says internal @v agent routes finance, docs, and engineering workflows
Vercel said it consolidated dozens of internal agents into @v, an agent/router used across finance, docs, marketing, engineering, analytics, and Slack workflows. The posts describe skills, subagents, per-user memory, and schedules rather than a public product.

Agent builders compare thin harnesses with large skill files for coding agents
Practitioner threads argued coding agents work better with task-specific harnesses, compact runbooks, evals, and deterministic checks. Examples included Sentry MCP traces and state-machine guardrails.



