10 items10 builders

AI Builders Digest

What the people actually building AI said today. One page — a 4-min read.

Jev is having a moment: a cheap, fast model that people are wiring into classification, scoring, and split-second workflow decisions, with demo sites and weekend projects piling up within days. Anthropic quietly made Claude Code read AGENTS.md, and OpenAI is teasing a keynote packed with launches. Underneath it all, a structural bet: Inception's Stefano Ermon argues diffusion, not autoregression, wins the inference era.

Podcast

No Priors

Diffusion beats autoregression at inference because GPUs hate sequential work

Autoregressive generation is fundamentally memory bound: you cannot produce the tenth token before the ninth, so you spend most of your time shuttling weights around and doing very little arithmetic. Diffusion models process many tokens at once, which makes their inference workload look like their training workload, the thing GPUs are actually good at. Inception's Mercury models match Haiku, Flash, and mini/nano class models on benchmarks while running significantly faster, and one voice customer dropped Cerebras custom silicon for them because software parallelism got the same speed on ordinary NVIDIA GPUs. "The bitter lesson is that the more parallel solution is the one that is eventually going to win."

  • #research
  • #hardware
  • #products
X

Thariq

Claude Code now falls back to AGENTS.md when no CLAUDE.md exists

Starting in version 2.1.277, Claude Code checks a folder for AGENTS.md if there is no CLAUDE.md, with a toggle in /config. The interesting part is the plumbing: AGENTS.md support ships as a built-in Claude Code mod, a preview of an upcoming system for customizing the harness itself. The source for the mod is public, and users will be able to build their own versions of project instructions.

  • #agents
  • #products
  • #open-source
X

Guillermo Rauch

Vercel CEO

Open models hit 78% of token volume on Vercel's AI Gateway

Today may be a record: open models took 78.4% of token volume against 21.6% closed. Spend usually tells a different story, but Moonshot AI and DeepSeek landed at #3 and #4, and adding Z.ai pushes their combined inference spend past OpenAI at #2. Rauch notes this is spend on serving those models across providers, mostly US ones, not revenue flowing to the open weight labs. He also reads Jev's adoption surge as downstream of an "AI is too expensive/slow" zeitgeist: people are desperate to optimize so they can put AI in more places.

  • #open-source
  • #infrastructure
  • #products
X

Aaron Levie

Box CEO

Jev's real enterprise job is split-second judgment calls inside workflows

The use case is not chat, it is the thousands of tiny decisions embedded in enterprise workflows: data classification, severity triage, routing. A Box demo pulls an incident report, asks whether it is customer-facing and how severe, then files it into escalate, monitor, or review and writes the result into a metadata template, nearly instantly and at almost no cost. Same shape applies to insurance claims, contract management, loan processing, and security reviews.

  • #agents
  • #products
X

Thibault Sottiaux

OpenAI's keynote problem is that there is too much to explain at once

Sottiaux spent the day working on the keynote and says the hard part was figuring out how to explain everything, because the volume of launches in quick succession is "a bit ridiculous." Some of it lands next week rather than waiting for the full reveal, with the rest coming together over the following months.

  • #products
X

Peter Yang

Meta's Muse negotiated $288 off a Comcast bill by calling support directly

Yang says Muse is the best personal agent he has used, saving him $800+ a year across cable and phone bills as a free product. The phone transcript from the negotiation is the revealing artifact, and his read is blunt: most companies' customer support lines are not ready for agents on the other end. He thinks Muse could become Meta's next billion-user app.

  • #agents
  • #products
X

Peter Steinberger

An agent running the team server that can answer questions mid-meeting

Steinberger's roboclaw runs their team server, is live on Discord, talks with gpt-live, and tracks every session it is juggling, so the team can query it about current and past session context during meetings. Another agent hijacks his sessions to clean up slop before the PR lands. Computer use works across these too, so the agent can do better than driving off screenshots.

  • #agents
  • #products
X

Zara Zhang

You cannot produce non-slop while consuming slop

Output quality is downstream of input quality. If you want to fix what you make, fix what you read and watch first.

  • #culture
X

Dan Shipper

Every CEO

Calling a real demo "fake" is the new reflex, and it is a bad one

Shipper pushed back on someone dismissing Jack Cheng's demo as fake, defending him as among the most intellectually honest and craft-focused people he has worked with. He concedes X is full of over-the-top AI demos, which is exactly what makes the reflexive accusation corrosive: this one is both real and a genuine glimpse of where things are heading.

  • #culture

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.