12 items12 builders
Archive

AI Builders Digest

What the people actually building AI said today. One page — a 5-min read.

Cost benchmarks moved from tokens to tasks today, with Vercel publishing numbers that put Grok 4.5 at the top of the cybersecurity price-performance chart and Kimi's agent sandbox paper arguing containers are not a real security boundary. Underneath that, the more interesting thread was about interfaces: Granola's CEO and Dan Shipper both landed on the same idea, that agents and humans need to manipulate the same UI at the same time. And Aaron Levie kept pushing back on the AI jobs apocalypse with what he is hearing from enterprises.

X

Guillermo Rauch

Vercel CEO

Grok 4.5 tops cybersecurity price-performance; containers are not a safe agent boundary

Vercel's latest benchmarks put Grok 4.5 as the best cybersecurity model on price-performance: 10x cheaper than Sol, 5.7x cheaper than Opus 5, and 2.2x cheaper than Kimi K3, while landing at roughly Kimi-level performance. Sol still holds the frontier outright, ahead of Opus 5. Separately, Rauch flagged Kimi's paper on where agents should be allowed to run: container-level isolation is not enough, because in their experiments agents crashed the underlying machine via kernel panics. His read is that Firecracker microVMs, as used in Vercel Sandbox, are the boundary that actually holds.

  • #evals
  • #security
  • #agents
Podcast

AI & I by Every

Granola's CEO: meeting notes are not the prize, the work interface is

Chris Pedregal thinks the thing everyone is fighting over right now, AI meeting notes, is not where the value ends up. Notion, OpenAI, and Zoom all shipped competing versions and it changed Granola's growth rate not at all, which reinforced his view that the real contest is over what interface we use for work in an AI-native world. His most concrete bet is pre-generation: Granola generates millions of pre-meeting briefs knowing only a small fraction get opened, because the useful moment is the fifteen seconds when you are late and cannot remember who you are about to talk to. He describes the product as a handrail, invisible until you trip. On the founder side: "I thought startups were just really hard when they weren't working. Turns out they're really hard even when they're working as well." The team only started tracking token spend last week, after someone burned a couple thousand dollars in a day on Fable.

  • #products
  • #agents
X

Swyx

Dollars per token stopped being a real cost measure a year ago

Swyx argues that price per input and output token died as a meaningful comparison sometime last year, and that anyone still plotting cost that way instead of dollars per task is not worth taking seriously. He also turned the knife on his own agent lab thesis: Claude Code got effectively open sourced this year and roughly nothing happened to its roadmap or its competitors', which he names as the strongest argument against the position he is known for.

  • #evals
  • #agents
X

Aaron Levie

Box CEO

Enterprises are still hiring, just for different roles

Levie says the predicted AI jobs collapse keeps not showing up in the enterprises he talks to. They are hiring engineers to attack problems that were previously out of reach, hiring in sales because AI lets them go deeper with clients, and hiring internal forward-deployed engineers to actually get AI shipped, which he reads as Jevons paradox playing out. His sharper claim: companies using AI purely to cut costs get outcompeted by companies using it to serve customers better. He allows the trend could reverse.

  • #policy
  • #products
X

Peter Yang

Codex edited and shipped a launch video from a bike ride

Jason Liu, who works on DevEx at OpenAI, got asked to fix a launch video while out on a bike. He connected to Codex from his phone, had it use computer use to edit the video, export it, and post it back to the Slack thread, then told it to check the thread every thirty minutes and export V2, V3, and V4 off the feedback. The video was greenlit by the time he got home. It is a clean example of an agent running an async revision loop against a human review channel with nobody at a keyboard.

  • #agents
  • #products
X

Madhu Guru

Meta Sr Director, AI

Product reviews should simulate the market, not report status

Madhu Guru's argument is that a good product review compresses months of learning into an hour by simulating how the market will react to your ideas, and that only works if the people in the room genuinely know the space, have taste, hold strong opinions, and are usually right. Most reviews have drifted into status updates, leadership visibility, and cross-functional alignment instead. That drift is why people resent them: the meeting stops being learning and becomes overhead.

  • #products
X

Amjad Masad

Replit CEO

The next frontier to map is the computational universe

Masad frames the current moment as a new age of exploration. Previous generations mapped the Earth and then went to space; this one gets the space of algorithms, programs, proofs, and designs that AI agents can search through. It is a framing bet on agents as search engines over possibility rather than as productivity tools.

  • #agents
X

Matt Turck

Under 40% of VCs have ever had a successful investment

Turck points at a new study finding that fewer than 40% of venture capitalists have a single successful investment to their name, then delivers the punchline himself: every VC reading it is quietly certain they are in the 40%.

  • #funding
X

Peter Steinberger

One agent filed the bug, another agent fixed it overnight

Steinberger's agent reported a bug and Jarred Sumner's robobun setup fixed it, all in the same night with no human in either loop. He calls that setup the future. He also notes the frustration of doing serious security work with strong partner teams and still having people assume the product is unsafe.

  • #agents
  • #security
X

Nan Yu

Linear head of product

If your smart people are so smart, have them improve your own product

Yu's pitch against building your own internal version of a tool: a lot of very smart people work very hard to make a product good and pleasant to use, so if you have very smart people, point them at your own product and pay a nominal fee for someone else's. It is the build-versus-buy argument compressed into a jab.

  • #products

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.