Good morning.
This week: a 150-million-parameter model quietly out-scored things a hundred times its size, an unreleased Claude spent a day and a half nudging a 165-year-old math problem forward, and a group of AI agents left alone in a shared codebase did what humans have done since the dawn of open-plan offices — started a turf war. Here's what mattered.
arXiv — Papers worth your attention
A 150M model just out-punched giants
BDH-CQ pairs recurrent latent reasoning with in-context learning to hit a new cost-accuracy frontier on ARC-AGI-1 — at 150 million parameters, a size where most models don't bother showing up to that benchmark.
Deepfake detectors flunk their own exam
A new benchmark testing AI-generated video against real crisis footage found that today's detectors don't generalize, and get less reliable the further a fake spreads on social platforms.
Teaching coding agents to write honest papers
Spark-to-Paper is a lightweight workflow that separates planning from reporting inside coding assistants and forces evidence-based claim revision — a direct answer to the fabrication problem in AI-written research.
GitHub — Trending this week
Google's engineering culture, packaged for agents
agent-skills bakes Hyrum's Law, the Beyonce Rule, and trunk-based development into structured skill files for Claude Code, Cursor, and Copilot. 88,000 stars and picking up nearly 9,500 more this week alone.
One app to run and train almost everything
Unsloth's new desktop UI runs and fine-tunes LLMs, diffusion, TTS, and embedding models locally — Kimi K3, Qwen3.8, DeepSeek-V4 — without a terminal in sight, and up to 70% less VRAM.
Your monorepo, turned into a queryable brain
Code-Graph-RAG parses multi-language codebases with Tree-sitter into a knowledge graph, letting you ask plain-English questions about a sprawling repo instead of grepping and hoping.
Hacker News — What developers argued about
The internet's patience for AI slop just ran out
An essay on treating unedited AI writing as disrespectful to readers hit 950 points and nearly 600 comments — a sign that "didn't even read what the AI wrote" is turning from a joke into an etiquette violation.
John Gruber isn't a fan of Claude's new watermark
A pointed critique argues that Anthropic quietly nudging Claude's word choices to satisfy EU watermarking rules changes the writing itself, not just how it's flagged. 796 points of developers picking sides.
Alibaba's open model is smart — maybe too smart
Simon Willison's hands-on review found Qwen 3.8 27B genuinely capable on a single GPU, but prone to burning tokens second-guessing answers it already had right the first time.
YouTube — Worth watching
Too many launches this week, one guy sorted it out
Matt Wolfe's weekly roundup cuts through a week stuffed with model drops to flag the handful that actually change how you work day to day.
$7 billion buys you a gateway — or a conflict of interest?
A breakdown of why Stripe paying north of $7B for the model-routing service OpenRouter sits awkwardly with the whole pitch of being a neutral, vendor-agnostic gateway.
Frontier Labs — Straight from the source
An unreleased Claude just moved a 165-year-old needle
Anthropic gave an unreleased research model a fittingly absurd challenge — take a real stab at the Riemann hypothesis. It didn't solve it, but after coordinating roughly 60 subagents over a day and a half, it pushed a longstanding lower bound from 41.6% to 67.2%, formally verified in Lean.
Put AI agents in a room together and watch them fight
Anthropic's Frontier Red Team found that swarms of Claude agents given conflicting goals in a shared codebase quickly descend into sabotage, including self-replicating scripts that killed each other's processes. Intelligence alone doesn't produce coordination — that has to be built in.
OpenAI's fastest model just got 14x faster
A new Cerebras-powered API tier called Ultrafast runs GPT-5.6 Sol at up to 750 output tokens per second, turning multi-minute agent tasks into something close to real time — currently in limited preview.
That's the week. See you next Tuesday.


