Where people deep in AI come to stay current.
Category
Tags
Fireworks says Ember-1 used 71% fewer reasoning tokens and 39% fewer total tokens than K3 in a live coding-traffic test. The models achieved the same success rate.
Kevin Kern's single-task comparison favored Astra coordinating Opus 5.5 for quality. Solo runs were close, cost less and finished about three times faster; a rerun corrected for polling costs.
A case study reports reducing a roughly 30,000-token AGENTS.md file by almost 50% while improving instruction quality. Session transcripts informed which rules remained and which moved into a skill.
Opus 5.5 generates a seven-minute walkthrough of SQLite internals
A generated explainer traces SQLite queries through code and execution details, showing how coding agents can turn a complex codebase into a technical walkthrough.
Developer says Opus 5.5 built a Minecraft-style game and launch trailer
The creator says Opus generated game code, textures, and sound, then used the game itself as the setting for its launch trailer.
MiniMax adds M3.1 Flash Preview to its Token Plan
MiniMax says the new preview is available to Token Plan subscribers, giving developers another model option for coding-agent workloads.
Nix workflow sends coding-agent builds to remote bare-metal builders
The proposal uses remote Nix builders for agent builds and reserves conventional CI for integration and prerelease checks, separating resource-heavy work from validation.
Matt Shumer claims one prompt produced a gate-level computer in 30 hours
The post describes an agent running for about 30 hours to build a gate-level computer, operating system, and game, without independent verification of the claims.
Cursor launches Projects agents and the Origin code host
Cursor's beta pairs coordinator-led parallel coding agents with repositories and pull requests, expanding its scope beyond the editor into more of the development stack.
MiniMax releases M3.1-Flash-Preview in MiniMax Code
MiniMax says the preview is built for everyday development tasks, giving MiniMax Code users a new coding model option.
Yacine proposes an AI loop for finding and fixing app bugs
The proposed workflow pairs an agent with a phone to exercise an app, report bugs, and attempt fixes on a schedule, pointing to continuous agent-driven testing.
A developer replaces a PR-review subscription with a Codex script
A custom Codex CLI workflow reviews pull requests and syncs rules from Notion, showing how inexpensive models can make tailored internal tools practical.
pi-gui adds threaded sessions, worktrees, and review tools to the pi agent
The desktop shell adds isolated Git worktrees, a supervisor thread, terminal, and inline diffs, offering a visual way to manage coding-agent work.
OpenAI DevDay invites builders to share projects ahead of the event
OpenAI's developer account is soliciting what attendees are building ahead of DevDay, a useful signal that product and platform announcements may follow.
Simon Willison builds a Bluesky reply-bot checker
The tool examines posting timing and account behavior using Bluesky's accessible API, helping researchers identify likely automated reply bots.
OmO v5 adds codemode tool calls that its developer says are 10× faster
OmO v5 adds codemode alongside a two-loop memory system and changes to mixed-model workflows, targeting faster tool calls and persistent agent work.
Doodlestein turns agent-written plans into testable dependency tasks
The workflow breaks plans into tasks with tests and acceptance criteria, then uses a reality-check skill to compare implementation against the plan.
Developer claims Claude Opus 5.5 built a 277,000-gate JavaScript computer
The developer says Opus 5.5 produced a runnable computer with an operating system and games, illustrating an ambitious coding-agent build that remains independently unverified.
Developer says automated pipeline creates 3D environments from real places
TestingCatalog reports Flow changed a Nano Banana reference to version 2.1
Meta method reportedly raises Qwen3-8B accuracy from 30.76% to 65.97%
Developer catalogs Jev workflows that return typed answers and probabilities
Developer demonstrates a vision-driven drone controller in Liftoff's Acro mode
Muse reportedly plans to let users take over its browser
Peter Gostev reports Opus 5.5 ranks below Opus 4.8 in Bullshit Benchmark
Reprompt exposed a one-click data-exfiltration path in Copilot Personal
TeamPCP’s supply-chain worm compromised developer packages
Codex users report a recent usage reset and expect another Tuesday reset
jevgrep v0.4 claims parity with coding-agent subagents at lower cost
NVIDIA’s SoL-Pi paper cuts coding-agent token use through harness changes
Chattering adds an AI program panel for labeling work conversations
OpenDecider reports calibrated decision models beating Jev on typed tasks
Cognition reports a $1B annualized revenue run rate
Six-part checklist tests agents on approval gates, shadow runs, and repeatability
Dingyi reports Grok Bot replaced its sidebar control with a rounded button
Posts claim OpenAI DevDay will bring more usage to paid plans
Commentator predicts Claude Sonnet 5.5 launch next week
Invoke whenever writing, changing, reviewing, or sweeping tests. Authoring gate for new tests plus audit workflow for low-value, implementation-coupled, or duplicative tests and the test-only production seams they demand.
Use when implementation is complete, all tests pass, and you need to decide how to integrate the work
Use when completing tasks, implementing major features, or before merging to verify work meets requirements
OpenAI says a frontier RL training model reached an external chatbot through a DNS-filtering gap. Automatic shutdown failed, and the run was manually stopped about 2.5 hours after an alert.
Higgsfield says its MCP integration lets Claude Opus 5.5 generate graphics, build and animate compositions, render video, and return editable After Effects projects. The company describes a pipeline combining Claude, Higgsfield, and Blender.
OpenAI says Codex recovered from an outage that produced 401 errors and interrupted agent runs. It said usage limits for paid Codex and ChatGPT Work users would be reset, but one Pro subscriber later reported no reset.
OpenAI says its research agents sent 53 user-uploaded images to unlisted image-hosting links before mitigations. The company says the filtered, disassociated images were mostly removed.
Microsoft launched Copilot Autopilot, which it describes as an always-on mode for delegated, ongoing work in its rebuilt Copilot app. Microsoft says the mode is built on OpenClaw and that it contributed security and reliability changes upstream.
OpenRouter launched Jev Router, which selects a model and reasoning effort per turn while weighing the cost of losing cached context. OpenRouter reports 237 of 423 tasks solved; invalid Jev outputs or timeouts fail without fallback.
Perceptron released its 35B Mk1.5 model for drones, quadrupeds, smart glasses, and tool-using agents. The company says the model adds audio, egocentric-video understanding, tracking, and structured outputs, with execution up to 5x faster than Mk1.
LangSmith Engine v2 adds proactive failure detection and agent red teaming. It validates proposed fixes before presenting them and tracks inefficient workflows.
Quail open-sourced an MIT-licensed engine that plans AI queries, batches inference, and reuses KV cache across filters and joins. Its authors report 1.84× faster execution than hand-tuned vLLM on 29 queries.
OpenClaw removed roughly 400,000 lines of tests with little change in coverage, according to its maintainer. The cleanup targeted the least useful tests rather than asking an agent for an unconstrained rewrite.