/anthropic-postmortem-xla-topk-precision-bug/ Claude Degraded for a Month. The Cause Was a Workaround Nobody Fully Understood. Sep 17
/hermes-ai-agent-mac-mini/ I Put Hermes on a Mac Mini. The Login Failures Were the Most Useful Part Sep 15
/real-swe-benchmark-frontier-coding-agents-enterprise-codebases/ Real-SWE: The Benchmark Where Coding Agents Meet an Actual Enterprise Codebase Sep 14
/sakana-ai-fugu-max-ultra-v2-orchestration/ Sakana AI's Fugu Ultra v2: When Orchestration Beats a Bigger Model Sep 13
/anthropic-claude-distillation-attacks-china-labs/ Anthropic's Distillation Report: How 200 Million Conversations Became a Threat Report Sep 12
/openai-agents-api-public-beta-codex-harness/ OpenAI's Agents API: Renting the Codex Harness Instead of Building Your Own Sep 11
/mercury-2-5-diffusion-llm-parallel-token-generation/ Mercury 2.5: When an LLM Stops Writing One Token at a Time Sep 10
/openai-alien-mind-chain-of-thought-monitoring/ OpenAI's Chief Scientist Says Its Best Safety Tool Is Losing Its Grip Sep 8
/openai-wiki-incident-misalignment-disclosure/ OpenAI's 'Wiki Incident' Shows Disclosure Rules Weren't Built for Agents Sep 7
/mai-transcribe-2-cheapest-fastest-speech-model/ Microsoft's MAI-Transcribe-2 Undercuts Every Rival on Price and Speed Sep 5
/gpt-6-astra-critical-cybersecurity-threshold/ GPT-6 Astra Is OpenAI's First Model Rated Critical for Cybersecurity Sep 4
/claude-fable-5-1-mythos-5-1-launch/ Claude Fable 5.1 Cuts Cache Reads 75% While Doubling a Key Agent Benchmark Sep 3
/anthropic-enterprise-frontier-safeguards-data-retention/ Anthropic Rebuilds Its Data Retention Deal Around Your Own Cloud Sep 2
/claude-in-chrome-general-availability/ Claude in Chrome Goes GA, and Now It Acts Without Asking First Aug 31
/github-copilot-thinking-effort-dial/ GitHub Copilot Turns Reasoning Depth Into a Dial You Control Aug 30
/claude-code-restricted-mode/ Claude Code's Restricted Mode Closes the Config Escape Hatch on Sandboxed Agents Aug 29
/anthropic-model-hardware-standard-mhs/ Anthropic's Model Hardware Standard Is MCP for Lab Robots and Factory Floors Aug 28
/grok-bot-video-output-agent-context-pricing/ Grok Bot Learns to Explain Itself on Video, and the Context Bill Comes Due at 200K Aug 21
/glm-5-3-coding-agentic-api/ GLM-5.3 Lands: Z.ai's Coding Model Drops the Off Switch on Thinking Aug 20
/google-imagen-4-shutdown-gemini-migration/ Google Kills Imagen 4 Today: The generate_images() Call You Have Is Now Broken Aug 17
/openai-ultrafast-gpt-5-6-sol-cerebras/ OpenAI's Ultrafast Tier Runs GPT-5.6 Sol at 750 Tokens a Second, and It Changes What 'Real-Time AI' Means Aug 15
/deepseek-v4-pro-peak-pricing/ DeepSeek V4 Pro Goes GA, and LLM APIs Get Their First Peak-Hour Pricing Aug 14
/gemini-managed-agents-hooks-spending-ceiling/ Gemini's Managed Agents Get a Kill Switch: Hooks, Budgets, and a Free Tier Aug 13
/mcp-2026-07-28-stateless-spec/ MCP Just Went Stateless, and Your Server Can Now Live on the Edge Aug 12
/meta-muse-glimmer-30b-local-agentic-model/ Meta's Muse Glimmer Puts a 30B Agentic Model on a Single Consumer GPU Aug 11
/claude-code-auto-mode-default-rollout/ Claude Code Turns Auto Mode On by Default, Because You Approve 97% of Prompts Anyway Aug 10
/claude-code-gateway-spend-limits-workspace-trust/ Claude Code Adds Spend-Limit Warnings and a Trust Prompt for Autonomous Agents Aug 9
/claude-code-self-hosted-environments-cross-session-messaging/ Claude Code Sessions Can Now Run on Your Own Servers and Talk to Each Other Aug 8
/chi-autoresearch-harness-llm-coding-agents/ Building Chi (ฯ): An Autoresearch Harness for LLM Coding Agents Aug 1