12 items12 builders

AI Builders Digest

What the people actually building AI said today. One page — a 6-min read.

Two big arguments today about where the ceiling is. Jerry Tworek and Rohan Anil say it is the transformer itself, which cannot learn after it leaves the lab, while Aaron Levie says the ceiling is verifiability and that the hardest work automates first precisely because you can test it. Meanwhile Vercel put an internal agent at the center of its own operations, and one investor called the current funding market fully divorced from fundamentals for at least another year.

Podcast

Training Data

Core Automation's founders: the transformer is the bottleneck, not the compute

Jerry Tworek scaled RL at OpenAI believing it would finish the job, and says 2025 was supposed to be the year everything got solved. Benchmarks kept climbing while real world tasks stayed messy, because the evals and the training data are two sides of the same coin and neither matches deployment. His conclusion is that models must learn at test time, which transformers structurally cannot do, so Core Automation is hunting for a replacement architecture and defining AGI as a model that improves itself with no human in the loop. As he puts it, the human LLM hybrid is really successful right now, but LLMs without humans, not so much. Rohan Anil adds the practical wall: a QR kernel competition they ran needed roughly three people on earth plus $100,000 of coding agents over four weeks to beat cuSolver by 60x, and no frontier model comes close to solving it.

  • #research
  • #agents
  • #hardware
X

Guillermo Rauch

Vercel CEO

Vercel now runs its internal operations through a single in-house agent

Guillermo Rauch says every day-to-day job at Vercel now involves an internal agent called @v, growing exponentially in both daily interactions and token use, with expertise spanning finance, comms, docs, marketing, engineering and analytics. It keeps per-user memories, workflows and schedules, and it is the basis for the design of Eve. The routing point is the interesting one: Vercel had built dozens of separate agents, which he compares to giving every agent its own domain name, so @v became both agent and router with sub-agents, skills and delegation. His argument against just wiring up a vendor Slack integration is ownership, since if agents become synonymous with modern companies you want control from source to runtime to data to token.

  • #agents
  • #products
X

Aaron Levie

Box CEO

The hardest work automates first because it is the easiest to verify

Aaron Levie's framing is that math, cyber and code get automated early not despite being hard but because correctness is objectively testable, which gives training clearer reward signals and lets you confirm at scale that deployment is working. Legal terms, marketing campaigns, sales messaging and budget setting have no single right answer, depend on operator risk tolerance, and often cannot be graded until long after the model produces output. His implication cuts against pure model scaling: even with exponential capability gains, most of the value gets built at the applied AI layer, and the underlying business processes themselves will have to change. He suggests we may need entirely new ways to test knowledge work, the way software already has.

  • #evals
  • #agents
X

Dan Shipper

Every CEO

Dan Shipper maps the three stages of an "agency rupture"

Watching a model do something that used to require you at every step reads as a kind of death, because we equate ourselves with our outputs and nobody asked our permission. Shipper says the reaction follows a predictable arc: first you see only the AI, so you say "AI solved that Erdos problem"; then you notice the human scaffolding and say "my job is just to babysit Claude"; finally you rebuild agency around the scaffolding itself, which now looks like the real work, and you go back to saying "I did this" with the AI implied. His bet is that the ability to metabolize these ruptures into playfulness and curiosity predicts who thrives in the AI economy, and that every capability jump creates fresh ruptures for populations, like mathematicians, that had not been touched yet.

  • #agents
  • #culture
X

Nikunj Kothari

FPV Ventures Partner

"VC has effectively fully become vibes capital"

Nikunj Kothari says early and mid stage pricing is now completely divorced from fundamentals, with some companies raising wild rounds on nothing while apparent sure things struggle. Being in the in-vogue sector is what determines outcomes, and rather than the correction most people expect, he thinks dry powder and the AI tailwind extend this for at least 12 to 18 months. Public markets are running on the same fuel, with trillion dollar stocks swinging more than 5% on vibes and model releases, memory and the KOSPI being recent examples, and rotations getting faster. His advice is that profitability and a clean cap table still win eventually, but founders should understand the current vortex before touching the capital markets.

  • #funding
  • #markets
X

Thariq

Jevons paradox is already visible in mathematics, and demand is rising

Thariq argues that AI in math is playing out like chess did. More is happening, it is easier to understand, and mathematicians have more time to discuss it with outsiders at higher levels of abstraction. His conclusion is the opposite of the displacement story: demand for people who think and know about math goes up, not down.

  • #research
  • #culture
X

Peter Yang

How Hermes keeps its self-built skills from turning into slop

Hermes builds its own skills to get work done, which raises the obvious question of what stops that library from rotting. Nous Research co-founder karan4d told Peter Yang the answer is Hermes Curator, a background task running inside your agent on a schedule, asking where the slop is and where things can be made more efficient as it cleans up skills and memory. Because it is open source, you can hand it your own definition of slop and it rewrites its own cleanup loop to match.

  • #agents
  • #open-source
X

Swyx

A computer-use agent worked a support chat and won the argument

Swyx is collecting Codex computer-use moments ahead of a podcast on the topic, and the current standout is the agent handling a support chat on his behalf to escalate for faster resolution. The support rep tried to pin the problem on his side, and the agent came back with complete receipts. His note: the humans have no idea they are talking to a bot.

  • #agents
  • #products
X

Garry Tan

Y Combinator President & CEO

Garry Tan: the outcome is the territory, everything else is the map

Tan's take on meritocracy is that people keep mistaking the map for the territory, and that in markets the territory is simply whether you made something people want. Alongside it he calls AI-driven economic growth the best white pill available, and notes that the public sense of wonder vanished at exactly the moment the amount of wonder went parabolic.

  • #culture
  • #markets
X

Ryo Lu

If we are leaving apps behind, what part of software stays visible?

Ryo Lu credits Rdio, Mailbox and Apple as his early mentors, apps that invented patterns which made software feel simpler and intuitive to touch. With that era ending, he is asking what parts of software remain visible at all, and how those parts should feel.

  • #design
  • #products
X

Andrej Karpathy

Karpathy publishes a playable, forkable version of the pelican bicycle test

Following up on Simon Willison's latest on the pelican on a bicycle test, Karpathy uploaded the source so it runs in the browser and anyone can fork it. His aside: watch for GTA Hobbiton to drop before GTA VI.

  • #evals
  • #open-source

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.