AI News

Latest AI news, breakthroughs, and industry trends.

Prime Agent Scores 95.5% on ARC-AGI-3: The RLM Harness Debate and the Case for Model-Harness Co-Evolution

Prime Agent Scores 95.5% on ARC-AGI-3: The RLM Harness Debate and the Case for Model-Harness Co-Evolution

An open-source agent framework just beat the human baseline on ARC-AGI-3 — and got publicly dismantled a day later. Critics flag RLM_MAX_DEPTH=1 and a public-eval 95.5%; the RLM paper's first author pushes back. Beyond the fight lies the real signal: Prime Agent's self-rewriting Continual Harness (and its Factorio cheating incident) mark the shift from training smarter brains to co-evolving models with their harnesses. What agent engineers should steal — and what to guard against.

Read more →
First Open-Source 100B MoE Video Model: MAGI-2 Preview Runs 114B Parameters at 6B Active Cost

First Open-Source 100B MoE Video Model: MAGI-2 Preview Runs 114B Parameters at 6B Active Cost

Sand.ai just open-sourced MAGI-2 Preview, the first hundred-billion-parameter MoE video generation model: 114B total parameters with only 6B activated, a 10-second 1080p clip for about 7 US cents, and a sixth-place leaderboard finish that beats closed models activating 100B+ per token. Inside the three-layer engineering answer — single-stream audio-video architecture, 3,072 fine-grained experts per layer, and Head Parallel communication that turns routing jitter into a constant — plus why video scaling resisted the LLM playbook for two years.

Read more →
AI Agent Paper Audits: 99.2% of Top Conference Papers Flagged, Only 8 ICML 2026 Orals Reproduce

AI Agent Paper Audits: 99.2% of Top Conference Papers Flagged, Only 8 ICML 2026 Orals Reproduce

AI agents just completed the largest reproduction audit in ML conference history — and the numbers are brutal. Of 168 ICML 2026 oral papers, agents could reproduce over 80% of verifiable claims in only 8; 99.2% of scanned top-conference papers carry at least one objective error, and per-paper error counts are up 55% since 2021. The root cause is scale, not malice: peer review never had the resources to verify anything. With verification costs collapsed, publication is becoming the starting line for machine auditing — and the researchers who embrace that shift inherit an almost unclaimed frontier.

Read more →
Claude Code Auto Mode Goes Default: Five Permission Fixes to Make Before the Switch

Claude Code Auto Mode Goes Default: Five Permission Fixes to Make Before the Switch

Claude Code is making auto mode the default for Pro, Max, and Team sessions — every tool call filtered through a risk classifier while hard-deny rules stay non-negotiable. Human testers caught just 13.6% of dangerous commands; the classifier caught 89%. Before the switch lands: pin your team default, audit the Bash allow-rules that 43% of users have left effectively wide open, extend hard-deny coverage, and learn the Shift+Tab escape hatch. Cloud channels get a one-month grace window before the same default arrives on AWS, Google Cloud, and Azure deployments.

Read more →
Hidden Reasoning Chains Broken: Lightweight Models Decrypt the Private Thinking of OpenAI, Anthropic, and Gemini

Hidden Reasoning Chains Broken: Lightweight Models Decrypt the Private Thinking of OpenAI, Anthropic, and Gemini

A cryptographic bypass across all three major LLM vendors lets lightweight models act as decryptors for flagship hidden reasoning traces — and the anti-distillation moat just collapsed. Inside the three-layer flaw, the $720 economics, and what vendors must fix now.

Read more →
Filter Migration: Pocket FM and the $400M Volume Game Reshaping Content

Filter Migration: Pocket FM and the $400M Volume Game Reshaping Content

My Vampire System runs 4,192 episodes. It launched four years ago and has been streamed more than 1.5 billion times. To hear all of it, you either burn through 30 minutes of daily free listening for four years, or you pay per episode — anywhere from a few cents to over $3 each. Some users have spent hundreds, even thousands, of dollars on a single series. Behind the show stands no profess …

Read more →
AI Daily Briefing – Sep 25: Personal agents go from chatting to cutting your bills

AI Daily Briefing – Sep 25: Personal agents go from chatting to cutting your bills

Meta's personal agent Muse went viral, Google refined its voice AI stack, and Lovable crossed another revenue milestone — today's clearest thread is AI agents moving from "can chat" to "can act and monetize." 🔥 Top 3 Stories 1. Meta's Muse pulls off viral "agent haggling" stunts, tops the App Store and sends shares up 11% Muse, the personal agent from Meta Superintelligence Labs, topped …

Read more →
PCIe GPU Inference: How Software Unlocks 7x More AI Throughput

PCIe GPU Inference: How Software Unlocks 7x More AI Throughput

A graphics card running at one-seventh of its speed Eight ordinary PCIe GPUs running DeepSeek-V4.1-Flash produced an input throughput of 1,932 tokens per second on the community's Day 0 baseline. The same eight cards, the same model, with not a single weight changed, later delivered 13,274 tokens per second — a 6.87x jump. What happened in between? No new algorithm. No new chip. Just the …

Read more →
AI Content Farms: Why Hit Rate No Longer Matters

AI Content Farms: Why Hit Rate No Longer Matters

One Indian company uploads 200,000 hours of new content to its platform every month and pulls in roughly $400 million a year. Its single biggest hit — an AI-narrated audio drama with 4,192 episodes — has generated close to $90 million. Pocket FM is not a story about an entertainment company adopting AI. It is a signal about what happens when the marginal cost of content collapses: the yar …

Read more →
DeepSeek DSec: Inside the Sandbox Infrastructure Behind Its Agent Training

DeepSeek DSec: Inside the Sandbox Infrastructure Behind Its Agent Training

DeepSeek just published the floor plan of a building most of its rivals didn't know existed. On September 23, a 31-page systems paper appeared on arXiv (number 2609.22978), titled "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale." More than 130 authors signed it, with Tsinghua University listed as a collaborator and founder Liang Wenfeng, …

Read more →
DeepSeek DSec: Agent RL Training's Bottleneck Moves Off the GPU

DeepSeek DSec: Agent RL Training's Bottleneck Moves Off the GPU

DeepSeek just published a 31-page systems paper, and the most interesting thing about it is what it does not discuss. "DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale" — posted to arXiv on September 19 with more than 130 authors including founder Wenfeng Liang — contains almost nothing about attention, Mixture of Experts, or reinforcement …

Read more →
AI Daily Briefing – Sep 24: Muse goes multimodal and Meta ships camera-free AI glasses

AI Daily Briefing – Sep 24: Muse goes multimodal and Meta ships camera-free AI glasses

AI Daily Briefing for Sep 24: Meta Connect 2026 steals the night — the Muse agent gets its own email address and video calling, is headed to smart glasses, and Meta shipped a lighter, camera-free pair of AI glasses; Anthropic's wet lab reports Claude autonomously discovered a CRISPR-like enzyme system; DeepSeek published a Liang Wenfeng–authored paper laying bare the infrastructure that s …

Read more →
Scroll to Top