16 items16 builders

AI Builders Digest

What the people actually building AI said today. One page — a 6-min read.

The plumbing for agent-run work got a lot of attention today: OpenAI stacked five Astra launches into one pre-DevDay week while publishing an unusually specific postmortem on its quality complaints, and Coinbase laid out why card rails cannot carry machine-to-machine payments. Running underneath it all was a shared worry about measurement, with three separate builders arguing that benchmark scores no longer tell you whether a model works on your actual job. A few people also stepped away from product talk to mark the 9/11 anniversary.

X

Thibault Sottiaux

OpenAI ships five Astra products in a week, names three causes of quality drop

Images 2.5, GPT-Live-1, an Agents API, Data Agent, and ChatGPT for Financial Services all landed before DevDay, with more planned next week. The quality complaints got a real postmortem rather than a shrug: skills written for older models were triggering too often and stopping the model from checking its own work, an opt-in context management experiment caused early stops and replies to older messages for an estimated 4-5k users and has been disabled, and some badly configured engines degrading a long tail of traffic were removed. OpenAI also brought on Aidan and Sasha from the Git AI team, whose open-source tool measures how coding agents actually contribute to a codebase, and committed to keeping it open source.

  • #products
  • #agents
Podcast

No Priors

76% of agent commerce payments are under 30 cents, below what cards can carry

Card rails charge roughly 30 cents flat plus a percentage, so they break entirely for the transactions agents actually make. Brian Armstrong's framing: "We don't want the AIs to be unbanked." Coinbase now hands any agent a self custodial wallet from a single pasted prompt, with no KYC, and routes payments over x402, a protocol it incubated and donated to the Linux Foundation with Google, Cloudflare, and AWS participating. Internally the company is chasing recursive self improvement by keeping a per repo and per team "brain" of every incident, financial control, AB test, and accepted or rejected PR, and enforcing that any human correction to an agent's work gets written back into it. Also disclosed: 88% of revenue is now non Bitcoin trading, prediction markets hit a $100M run rate within months of launch, and tokenized stocks are live outside the US as a real security held one to one in custody.

  • #agents
  • #payments
X

Madhu Guru

Meta Sr Director of AI

Enterprise AI fails when a central team builds for people instead of with them

Three failure patterns: a CEO taps a trusted lieutenant who pulls in trusted colleagues, importing a team structure and launch playbook built for fifteen years of incremental improvement into work that demands experimentation and invention. Second, chronic under investment in evals, where most companies do not even know what good looks like. Third, a central AI team declaring itself a platform, shipping tools disconnected from the workflows and judgment of the people doing the work, and getting begrudging adoption with no productivity gain. The fix is to embed your best AI builders inside finance, sales, and support rather than building at them from outside.

  • #evals
  • #enterprise
X

Dan Shipper

Every CEO

Every turns three years of model vibe checks into personal benchmarks

Better benchmark scores say almost nothing about how a model handles your real work, which is why Every has published hands on long form reviews of each new model for three years instead. Now those vibe checks are getting quantitative: Hammer and Nityesh built an internal platform where every person on the team constructs a personal benchmark from their own day to day work.

  • #evals
X

Peter Yang

Software factories do not work yet, and nobody has produced a counterexample

Outside of verification and testing, AI cannot self improve a product or build a new feature end to end without a human in the loop. The failure mode is specific: loop something overnight and a single wrong assumption turns the entire run into wasted tokens. The open challenge is worth taking seriously, since no one has named a product or feature built end to end with no human defining requirements or checking the work. Separately, a clean split on tooling: local scheduled tasks stay in Codex, cloud tasks move to Grok Bot.

  • #agents
X

Thariq

Claude Code adds plugin evals so skills survive model upgrades

Running `claude plugin eval init` in a plugin folder answers the question of whether your skills still work after a new model ships. The deeper point is that raw pass/fail scores have become nearly uninterpretable: many benchmark failures trace to overly strict hidden tests, and in some cases the model's answer is more sensible than the expected result.

  • #evals
  • #agents
X

Aaron Levie

Box CEO

Box now mounts directly into agent sandboxes for file reads and writes

An agent can read and write files on its own computer with Box mounted into the sandbox. The argument behind it: as agents take over critical enterprise workflows, they need the same primitives people have always had, starting with a real filesystem.

  • #agents
  • #products
X

Guillermo Rauch

Vercel CEO

Tailscale's model router runs on Vercel AI Gateway

AI gateways are becoming what CDNs were: you can go direct to origin, but it is brittle, and rolling your own is painful and costly. Tailscale's model router is now built on Vercel AI Gateway as its underlying infrastructure.

  • #products
X

Nikunj Kothari

FPV Ventures Partner

Expect VC track records to get quietly rewritten over the next two years

Large exits and markups are genuinely rare, and they are what you raise the next fund on, so successful VCs fight hard over attribution for hot deals. The prediction is concrete: names will get scrubbed or conveniently omitted as people revise history to claim credit for winners. Emerging GPs take the worst of it, since track record is the only thing they can bring to future LPs. The defense is to hold your founders close, because they are the real reference checks.

  • #funding
X

Zara Zhang

The one-person company is overrated because building alone is lonely

AI genuinely lets one person do far more, but that misses what a cofounder is for. Building something new is a profoundly lonely experience, and you need someone to brainstorm with, suffer with, and celebrate with. Motivation collapses easily when you have not tied yourself to the mast alongside somebody else.

  • #startups
X

Peter Steinberger

Astra plays Doom in a cloud session through computer use

Running Astra on OpenClaw in a cloud session with CUA driving the machine, the model played Doom. Peter Steinberger's assessment: not AGI, but probably beats a fly brain. Getting there required submitting a patch to trycua to make keys work reliably under Linux, which he still rates as a solid framework overall.

  • #agents
  • #open-source
X

Matt Turck

FirstMark Capital VC

The 53rd floor instead of the 104th, and a startup that survived 9/11

TripleHop, a 25 person enterprise search startup doing roughly what Glean does now, took an office in the North Tower because the World Trade Center was courting young dot-com companies with preferential terms. They chose the cheaper 53rd floor over a space near the 104th, where the views were unbelievable but the elevator ride took forever. Only one employee, Kenton Beerman, was in the office at 8:46am, and he got out unscathed; the CTO's backups had the business running again by the end of the day. Visiting the memorial for the first time a few weeks ago, the thought was how little it would have taken: an hour later, a few floors lower, or the other office.

  • #startups
X

Garry Tan

Y Combinator President & CEO

Scoring 1600 should unlock a harder test, not end the measurement

A perfect SAT score caps out the instrument exactly where it stops being informative, so there should be a second harder test stacking another score on top of the 1600. Instead the test gets banned, admissions turns into a random lottery, and excellence becomes impossible to recognize.

  • #policy
X

Aditya Agarwal

SPC General Partner

SPC opens a fall series on tech regulation and a two-party California

The questions on the table are how to encourage real innovation in California, how to enable a two-party state, and how technology should be regulated, with decision makers joining through the fall and Steve Hilton first up on September 23rd. Separately, a 9/11 reflection: hearing the news in a CMU lecture in Wean Hall 7500 as people panicked about a plane near Pittsburgh, and hoping he would have had the courage of the passengers on United 93.

  • #policy

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.