10 items10 builders

AI Builders Digest

What the people actually building AI said today. One page — a 4-min read.

Evals are now treated as central to building agents. Box's Aaron Levie and Meta's Madhu Guru both argue that agent quality depends on testing against real work environments and on treating evals as the spec. Guillermo Rauch argued that agents thrive under strict compile-time constraints that humans found too tedious, while Garry Tan and Thariq both focused on how agent harnesses spend tokens.

X

Guillermo Rauch

Vercel CEO

Vercel's gdp-ts makes the typechecker demand proof of an auth check

Guillermo Rauch released gdp-ts (Ghosts of Departed Proofs for TypeScript), a library, linter and AI skill. With it, sensitive functions only accept a call that carries a 'proof' that an authorization check ran, and the typechecker verifies this at compile time. He says the pattern stayed niche in Haskell because humans had to review it and pay its syntactic overhead. Now agents write more code than anyone can review, and they thrive in tight loops with hard constraints. The README models a real Vercel rule: changing a Project's password requires proof of a certain role plus a certain entitlement.

  • #agents
  • #security
  • #dev-tools
X

Aaron Levie

Box CEO

Levie: agent adoption is bottlenecked on testing in realistic work environments

Aaron Levie says a lot of agent deployment is held back by the inability to test, tune and optimize agents in real work environments: the files, CRM, email and other systems they actually touch. Without that, you can't know how an agent does on your evals today or how it will hold up after a model change or workflow upgrade. Right now every enterprise builds this one by one. His prediction is that every enterprise will have someone managing evals and the infrastructure for simulated environments, and he calls it a huge space.

  • #agents
  • #evals
  • #enterprise
X

Madhu Guru

Meta Sr Director, AI

Your evals are your product spec, not a QA step after the agent is built

Madhu Guru, who previously led Gemini, Veo and Nano Banana at Google, says the most common mistake teams make is treating evals as an extra QA step added once the agent is built. AI products are fundamentally different, he argues: the evals are the product spec.

  • #evals
  • #agents
X

Garry Tan

Y Combinator President & CEO

Garry Tan: lab harnesses have an incentive to burn tokens, startups can fix that

Garry Tan argues that agent harnesses built by the labs are incentivized to burn tokens, which gives harnesses from startups real utility. His example is Grep, which watches how an agent is used and automatically replaces that token burn with deterministic, tested code that repeats the same way every time. He also predicted that 'AGI Science Loops' are coming and said Halmos is building in that space.

  • #agents
  • #startups
X

Thariq

'Local hands': Claude runs in the cloud but reaches your local files

Thariq from Claude Code says his favorite name for one pattern is 'local hands': Claude runs in the cloud but can access your files locally, and the pattern is also coming to Cowork. Separately, he said the planning approach he showed is much more token efficient than raw HTML. The model doesn't have to rebuild components or logic for common things like state machines, diagrams and code snippets.

  • #agents
  • #products
X

Ryo Lu

Ryo Lu on Steve Jobs: tech is stuck refining instead of making things with heart

Remembering Steve Jobs, Ryo Lu says he got into interface design by making his beige PC look like the iMac G4 he couldn't afford. Since Jobs died, he argues, the industry has been stuck in a loop of constant refinement, better supply chains and more ways to keep people scrolling, while people have grown more disconnected and lonely. "We spend so much energy trying to beat humans. I wish we spent more time learning from people's lives and making them better." He also shipped ryOS Subtitles, a Chrome extension for watching Netflix with subtitles in two languages at once, plus pronunciation guides for Japanese, Chinese and Korean.

  • #design
  • #products
X

Aditya Agarwal

SPC General Partner

Aditya Agarwal: a local Muse/Dot computer could beat locked-down macOS for agents

Aditya Agarwal says it would be interesting to run the Muse/Dot 'computer' locally, since that feels like a better agent environment than an increasingly locked-down macOS. SPC is also hosting Sergey Levine, Berkeley EECS faculty since 2016 and cofounder of Physical Intelligence. Levine builds the algorithms that let robots learn from experience by linking what they see to how they move.

  • #agents
  • #robotics
X

Peter Yang

Build a voice-call Japanese tutor with Gemini Live and Nano Banana

Peter Yang published a tutorial on the app he built that teaches him Japanese through live voice calls. It walks through using his spec skill to create the key designs, setting up Gemini Live APIs for voice, and making diorama art with Nano Banana. The goal is an app that teaches you 100 phrases in any language before a trip. He also pointed to Gemini 3.8 Live.

  • #products
  • #voice

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.