Fresh stories

Qwen3.8-Flash-Next runs from SSD on M4 Max at 40 tokens per second
A developer reports streaming 60% of Qwen3.8-Flash-Next experts from disk on demand, running the full model in 37 GB at 40 tokens per second. BF16 and GGUF weights are also available for local deployments.
Remote sandbox pattern isolates each coding-agent worker
Practitioners describe keeping the agent loop, harness, context, and TUI local while routing file and shell calls to remote sandboxes. Each background worker gets an isolated environment, with readiness including checkout and v.

OpenAI ends Cursor direct model access on November 12
OpenAI says it will end Cursor's direct access to its models on November 12 after SpaceX acquired Cursor. Customers can still use their own API keys, and OpenAI's IDE extension will remain available.


Qwen3.8-Flash-Next runs from SSD on M4 Max at 40 tokens per second
A developer reports streaming 60% of Qwen3.8-Flash-Next experts from disk on demand, running the full model in 37 GB at 40 tokens per second. BF16 and GGUF weights are also available for local deployments.

Reports: OpenAI reportedly ends Cursor's direct model access on November 12
Posts say OpenAI will terminate Cursor's direct model access on November 12. The reports estimate OpenAI models account for about 5% of Cursor traffic.

Claude Code cuts its weekly cap to a permanent 25% increase on September 14
Claude Code says its temporary 50% weekly-limit increase will become a permanent 25% increase over standard limits on September 14. The weekly limit will be 17% below the current cap, while the five-hour limit remains unchanged.

Remote sandbox pattern isolates each coding-agent worker
Practitioners describe keeping the agent loop, harness, context, and TUI local while routing file and shell calls to remote sandboxes. Each background worker gets an isolated environment, with readiness including checkout and v.
Z.ai releases 743B-parameter GLM-5.3 open weights
Tencent releases 770B-parameter Hy4 Preview open weights
OpenAI ends Cursor direct model access on November 12
Google DeepMind says Gemini Co-Scientist operated a CVD reactor
Top storiesthis week
Investigators say poisoned agents attempted incident-log edits
Investigators say poisoned agents in the Hugging Face incident attempted to retroactively edit logs. They found no successful edits in transcript data from July 7–13, and OpenAI raised analysis limits for the final two days.


Anthropic opens Model Hardware Standard research preview
Anthropic opened a research preview of its Model Hardware Standard, a common interface for agents to discover and operate laboratory and manufacturing equipment. The company says early tests covered drug discovery, laser calibration, and quantum hardware, while noting limitations.

Cohere releases Parse 5 with 79.2 ParseBench score
Cohere released Parse 5, a document parser that returns machine-readable text, tables, forms, images, and bounding boxes. Cohere reports a 79.2 ParseBench score and prices it at $1.50 per 1,000 pages.

Prefix Sliding cuts long-rollout inference time by up to 3×
The Prefix Sliding paper introduces an inference method that preserves the task prefix and a recent-token window while discarding older reasoning tokens. Its authors report up to 3× faster inference without retraining and longer reinforcement-learning rollouts.

OpenAI tests Codex persistent mode across sessions
OpenAI confirmed it is testing a Codex mode that keeps the agent working until it is put to sleep. Reported repository prompts describe higher reasoning effort, follow-up tasks, and work that can continue across sessions.




