8 items8 builders
Archive

AI Builders Digest

What the people actually building AI said today. One page — a 4-min read.

Multi-model routing was the day's loudest idea: a frontier model planning and handing work to cheaper models is now the core pattern for cutting agent costs, and token economics is the question everyone from Box to Booking is chewing on. The other thread was moats, or the lack of them, with two CEOs arguing durable advantage comes from insight and relentless building rather than scale or capital. Benchmark gaming and how to reshape companies and hiring around coding agents rounded things out.

X

Aaron Levie

Box CEO

Frontier planner plus cheap workhorse model cuts agent costs 15X

Multi-model agentic systems are becoming the core design pattern for complex agents. Citing new Cursor research, a frontier model handles the parts that genuinely need intelligence, the original decomposition, design decisions, and certain trade-offs, then collapses the ambiguity into explicit instructions a cheaper workhorse model simply follows, cutting total token cost 15X. The applied layer will differentiate here: companies that know a domain and can route across model tiers will win larger coding, finance, legal, and healthcare workloads that were otherwise too expensive to deploy.

  • #agents
  • #products
Podcast

No Priors

Booking CEO: there is no such thing as a moat

Glenn Fogel, who joined Priceline when it was worth a few hundred million and helped grow it past $130B, insists durable advantage comes only from constantly building new services, never from a moat that innovation can't erode. Priceline's agentic assistant Penny has doubled adoption every month, lifting conversion and cutting cancellations, though it's still tiny against $186B in annual travel, and the open questions are token economics and which model to route where. He's blunt on AI and jobs: society always adapts, but the speed of change and the 50-year-old truck driver who can't retrain worry him more than the technology itself.

  • #agents
  • #products
  • #policy
X

Madhu Guru

The road to AGI is paved with economically valuable tasks

Enterprise AI is one of the most important frontiers because that's where the economically valuable tasks live, and those tasks are the real path toward AGI. The tokenomics debate that actually matters now isn't crypto's, it's open versus closed weights, inference costs, and model routing. It's the greatest time ever to have product sense.

  • #agents
  • #products
X

Swyx

How labs quietly game benchmarks by training on test lookalikes

A trajectory-comparison writeup buried in the RLM paper from Alex Zhang and Omar Khattab pokes at an open secret of frontier training: even without training on the test set, you can train on lookalikes to goalseek almost any benchmark number, with plausible deniability because open-weight releases rarely ship the datasets or RL environments that would expose it. Their preliminary fix applies standard NLP distance metrics to hidden trajectories. There's no ultimate solution, but the exploration also supports the finding that reasoning language models generalize to unseen tasks sharing latent structure with what they saw in training.

  • #evals
  • #open-source
X

Zara Zhang

Two kinds of companies now: built before coding agents, and after

Companies founded after coding agents are structurally different from day one: teams under ten because they genuinely don't need more, work organized by projects instead of departments, each person closing their own loop, and almost no internal meetings. Everyone built before is scrambling to retrofit. Her matching hiring design: one in-person round with no AI allowed to test raw domain expertise, then a project that's impossible to finish without AI, where the candidate is judged on the agent chat transcript as much as the result.

  • #agents
  • #hiring
X

Peter Yang

Use a separate agent with a rubric to review the first agent's work

For subjective outputs like whether a video short is any good, don't let the model grade itself. Thariq explains you should run a separate verification agent that reads a rubric, reviews the output, and gives feedback, because of self-preferential bias: when a model prefers its own output it goes easy when checking it. Yang also argues that banning Chinese open models would be the same self-own as banning Chinese EVs.

  • #agents
  • #evals
X

Nikunj Kothari

No moats in AI doesn't make scale and capital your moat

Founders of the last 18 months are about to relearn that being well-capitalized and at scale is no substitute for a unique insight. The graveyard is full of companies that were structurally and financially strong and still got beaten or collapsed under their own weight: Webvan, Groupon, MySpace, Yahoo, AltaVista, Blockbuster, Nokia. The needle to thread is finding an insight worth a 10-plus-year journey while staying prudent enough not to let capital and scale stand in for it.

  • #funding
  • #products
X

Guillermo Rauch

Vercel CEO

The big lesson from AI: everything is code

A slide deck is code, design is code, that cool promo video is code, Excel automation is code, and probably the universe too. Rauch's framing for why AI collapses so many creative and business artifacts into the same malleable medium.

  • #products

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.