14 items14 builders

AI Builders Digest

What the people actually building AI said today. One page — a 6-min read.

The harness, not the model, was the day's obsession. Between Vercel and Linear both describing the same Issue → Agent → PR → Release loop, Box's Aaron Levie calling the harness the variable that matters most as tasks scale into hundreds of millions of tokens, and Claude Code shipping shareable artifacts, the theme was infrastructure around models rather than the models themselves. Meanwhile an xAI co-founder made the case that proprietary labs are getting squeezed from both sides.

Blog

Claude Blog

Claude Code sessions can now publish themselves as live, shareable web pages

Claude Code can turn a session into an artifact: a PR walkthrough, an incident timeline, a filterable dashboard, a release checklist that fills itself in as work proceeds. The page is built from the session's full context, including codebase, connectors, and conversation, so an incident page can pull the failing test, the error spike from a connected monitoring tool, and the root-cause reasoning into one view without wiring up data sources. Every publish is a new version at the same URL with history and restore. Artifacts are private to the author by default, viewable only by authenticated org members, never public, and the feature is in beta for Team and Enterprise.

  • #products
  • #agents
Podcast

Unsupervised Learning

xAI co-founder: proprietary model builders are getting squeezed from both directions

Igor Babushkin argues closed labs are in a genuinely hard spot. Make the model too capable and regulation or your own caution stops you releasing it; hold back and open weights close the gap every month. He thinks training returns are hitting the wall, since you cannot cover the earth with GPUs, and the escape route is innovation rather than scale. His new company River AI makes three bets: an RL and fine-tuning API, models that personalize to the individual instead of the average user, and local hardware that runs a frontier model in a box in your home. On the enterprise side his warning is blunt: "the worst place you can be in is if all of your knowledge is on the internet." On what actually blocks progress now, it is rollout length and non-verifiable rewards, not data. If an agent trajectory takes twenty four hours to roll out, every training step takes very long. He is also unworried about backdoors in Chinese open weights today, since we would have seen an example in the wild by now, but wants the best open model in the world trained in the US.

  • #open-source
  • #policy
  • #hardware
X

Aaron Levie

Box CEO

Levie: the harness becomes the second most important variable after model capability

The harness barely mattered when tasks consumed hundreds of thousands or a few million tokens. At tens or hundreds of millions, how work gets broken down and routed to the right model at the right time becomes the main lever on both accuracy and cost. Levie is careful that he does not know whether the specific numbers he saw generalize across tasks, but calls the direction clear and the opportunity huge.

  • #agents
  • #products
X

Nan Yu

Linear Head of Product

30% of Linear's bugs go Issue to Agent to PR to Release with no human in the middle

The most common automation Linear customers write is that exact loop, and about 30% of Linear's own bugs make it all the way through it. The nuances matter more than the loop: instruct the agent to research root cause extensively and pull evidence from Datadog and Sentry MCPs, and tell it to attempt a fix only at high certainty, otherwise you burn a lot of tokens. If it needs repro steps, it should comment on the issue asking the reporter. As Yu puts it, agents need to be told to follow good practices, just like people.

  • #agents
  • #products
X

Guillermo Rauch

Vercel CEO

Vercel ships per-key AI budgets and tells customers to quit token-maxing

AI Gateway now does budgets per key, team, and project, alongside failover, model and provider choice, and realtime observability. The framing is pointed: if you are still in a token-maxing fever dream, wake up. Rauch also expects Issue → Agent → PR → Release to become the norm as projects turn into agentic software factories, with the maintainer's job shifting to designing the loop that yields the highest quality product and setting the criteria for what gets worked on at all.

  • #products
  • #agents
X

Amjad Masad

Replit CEO

An 8B model hits ~1500 Elo at chess and beats GPT 5.6 with high reasoning

The small model consistently beats frontier models and Stockfish level 0, and does it spending one to two seconds per move against roughly 30 seconds for GPT 5.6 with high reasoning and response chaining. Masad shared a link to play it.

  • #evals
  • #open-source
X

Zara Zhang

65% of PRs from Anthropic product and eng teams now come from Claude Tag

Zhang flagged that figure from an interview on how Claude Tag changed the way work gets done inside Anthropic. Her read: for non-engineering teams the ultimate agent interface is wherever they already work, like Slack. She tracks her own interface moving that direction over six months, from the terminal in January, to the desktop app in March, to her work collaboration tool by June, each step closer to how humans naturally communicate. The agent should meet the user where they are. Separately, on posting: what feels totally native to you is brand new to someone outside your circle, and the voice saying "duh, isn't that obvious?" preceded every one of her viral posts.

  • #agents
  • #products
X

Swyx

Swyx says people quit /loop and /goal too early in the g5.6 and c5 era

He is in the minority among AI leaders in still actively using both, and thinks everyone who stopped is wrong, not forever, just premature. The case for them is narrower than it used to be: when you want the right mix of steerability and autonomy, and when you want an open-ended loop-that-generates-loops end state without deeply specifying the path there. He cites a long action-reasoning turn where having a goal set saved him. Separately, he notes the pejorative edge on "vibe coding" has fully vanished now that everyone from nontechnical to supertechnical is doing it.

  • #agents
X

Sam Altman

Altman's favorite ChatGPT Work use case is a morning podcast about your own kids

Connect the family calendars, explain each kid's interests, and have it generate a podcast for the school drive covering one kid's soccer game that afternoon, another's upcoming birthday, and some news. He also responded to a Moore's law comparison with "i see your moore's law and i raise you 20x."

  • #products
X

Dan Shipper

Every CEO

Shipper: momentum has been shifting to OpenAI since early spring

Quoted in a WSJ piece on OpenAI versus Anthropic, he stands behind the call and frames it as a comeback story. He also floated what programmer interviews might look like in 2027: describe the last three unresolved mathematical conjectures you solved and share your prompts, describe the last cyber felony your agent unintentionally committed and how you mitigated it.

  • #policy
X

Nikunj Kothari

FPV Ventures Partner

New essay questions venture's belief that the best founders are running from pain

There is a quiet conviction in venture that great founders are driven by a troubled childhood, a chip on the shoulder, something that keeps the foot on the gas. Kothari's essay asks where the drive actually comes from and points at a gear most people never reach.

  • #funding

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.