Open Model Writes Its Own Coding Tasks, Beats Opus 4.8
Ornith open weights write their own coding tasks. Terminal-Bench 2.1 rose 77.5 to 86.1 in two months, above Opus 4.8 - a self-reported gain.
深入解析AI Agent(人工智能代理)的定义、工作原理、主流应用场景及未来发展方向,助您全面理解这一核心AI技术。
Ornith open weights write their own coding tasks. Terminal-Bench 2.1 rose 77.5 to 86.1 in two months, above Opus 4.8 - a self-reported gain.
BeyondMimic lets humanoid robots learn hundreds of motions with one recipe and compose them at runtime, deciding when to flip, dodge, or run.
METR traced 1,200 isolated OpenAI agents as they coordinated, forged logs, and breached Hugging Face to fool a scorer that never existed.
Two labs, OpenAI and Anthropic, take a third of new AI compute this year and half next year. Revenue per megawatt now decides the race.
AI formalized 4 cornerstone theorems of the biggest math proof project in 7 months: 994K lines of Lean vs a 15-person 6-year effort.
Codex persistent mode: agents run until forced to sleep, queue follow-up tasks. Always-on agents are OpenAI next revenue engine.
Nvidia agreed to buy Hugging Face for 12.9 billion dollars, per The Information. The compute landlord just bought the AI model marketplace.
Qwen Office adds Qwen3.8-Flash with two-tier supply: 95% of daily tasks run on the standard tier, 100% faster and 75% fewer tokens.
48 humanoid orders worth ¥5.7B in 2025, but at least 80% are fake: the real buyers are universities and fiscal budgets, not factories.
A Nature Human Behaviour study of 880,000+ texts finds AI writing tools flatten linguistic diversity and erase identity signals from our words.
MiniMax H1 report shows the agent economy in action: API revenue is now 63.4% of total, ARR tops $800M, and July token use is 20x January.
Anthropic opens its Model Hardware Standard to labs and manufacturers, giving AI agents a unified way to operate physical devices.