10 items10 builders

AI Builders Digest

What the people actually building AI said today. One page — a 5-min read.

Two frontier labs cut prices on the same day: GPT-6 Sol and Luna arrived with a permanent 50% API cut, and Opus 5.5 became the default across Claude's products with rate limits stretching 25% further. Vercel's fresh Next.js evals then put Opus 5.5, GPT-6 Sol, and Fable 5.1 in a dead heat at 97%. The most useful reading came from Box, which measured what cheaper tokens actually buy: the same tasks, higher accuracy, a fraction of the spend.

X

Thibault Sottiaux

GPT-6 Sol and Luna ship with a permanent 50% API price cut

GPT-6 Sol and Luna are out, and OpenAI is halving API prices permanently rather than as a launch promo, which opens both models to a class of use cases that could not carry the old token cost. The improvement is described as across the board, with writing and general feel being where you notice it first. Plus, Pro, and Business accounts also get a banked usage reset. The framing is efficiency funded by capability: you only get to make everything below the frontier cheap if you have incredible models at the top.

  • #products
  • #pricing
X

Cat Wu

Opus 5.5 becomes the default in Claude Code and the Claude app

Opus 5.5 is now the default for Pro, Max, and Team across Claude Code, the Claude app, and Cowork. The default effort setting is medium, described as comparable to Fable 5.1 on intelligence but faster, and rate limits go 25% further than they did on Opus 5. The invitation to users is to hand it an ambitious task rather than a small one.

  • #products
  • #agents
X

Aaron Levie

Box CEO

Box measured Opus 5.5 using 63% fewer tokens than Opus 5

Box tested Opus 5.5 on complex enterprise knowledge work and saw 63% fewer tokens, 42% less verbosity, and 30% faster completion than Opus 5, alongside accuracy gains of 39% on financial due diligence and 65% on cloud cost analysis. On a clinical diagnostics task the model caught that two groups' standard deviations differed more than 100-fold, re-ran the comparison correctly, and found the dry-season difference did not actually hold. Levie reads the simultaneous price cuts from both labs as Jevons paradox applied to agents: cost per task in AI falls faster than in any technology before it, and each drop unlocks another tier of deployment, from scanning all your code for security issues to reading every line of log data.

  • #pricing
  • #agents
  • #enterprise
X

Guillermo Rauch

Vercel CEO

Vercel's Next.js evals end in a three-way tie at 97%

Fresh Next.js evals put Opus 5.5, GPT-6 Sol, and Fable 5.1 all at 97%, with Grok 4.7 close behind at 94% and 2x to 7x cheaper than the leaders. The broader claim is that cheap generation changes what software is: liked Google Reader? Generate and deploy your own, and keep it forever. Rauch also argues the case for headless on the web has finally arrived, since any page can now take whatever whimsical shape you want and there is no excuse left for not pushing the design frontier.

  • #evals
  • #pricing
X

Boris Cherny

16 PRs of race condition fixes from pointing Opus 5.5 at Lean

Formally verifying the Claude Agent SDK in Lean took a couple of short prompts and produced 16 PRs fixing bugs and race conditions, written by someone who does not know Lean well. TLA+ works too, and combining the two is useful for data flow, concurrency, and state management problems a human reviewer would probably never spot. Separately, Opus 5.5 and Fable 5.1 each ported HAProxy from C to Rust and both passed nearly all of HAProxy's tests, but Opus 5.5 finished in 9.5 hours against 12 and at 51% less cost. The open question: is formal verification where bug finding ends up?

  • #coding
  • #research
X

Alex Albert

One prompt rebuilds 1906 Market Street in Blender from archive sources

Better 3D modeling and vision in Opus 5.5 mean a single prompt can produce an entire world: Market Street in San Francisco on the afternoon before the 1906 earthquake, from the Ferry Building up to Fifth, including the Palace Hotel, the Call Building and Lotta's Fountain. The prompt bans downloaded meshes, textures and HDRIs, and requires reusable generators for Victorian facades, mansard roofs, cable cars and gas lamps, assembled from a source file built out of Sanborn fire insurance maps, the Miles Brothers film shot that April, and period photo archives, with every fact carrying its source and a confidence level.

  • #coding
  • #products
X

Peter Steinberger

A ChatGPT crash on macOS 27 traced back to a 14 year old libuv bug

ChatGPT started crashing intermittently after the macOS 27 update, and Astra tracked the cause down to a bug that has been sitting in libuv for roughly 14 years. The OS update is what surfaced it; the defect itself predates most of the code depending on it.

  • #open-source
  • #agents
X

Nikunj Kothari

FPV Ventures partner

Assume a big round headline is mostly SPV money

The volume of SPVs being organized by even the best investors is large enough that when you read a headline fundraise number, you can reasonably assume a lot of it came from SPVs regardless of what the announcement says. Stack tranched valuations and revenue figures that do not mean what they appear to mean on top of that, and the conclusion is blunt: stop believing the headlines. The proof is in the pudding and the pudding is not publicly visible anywhere.

  • #funding
X

Aditya Agarwal

SPC General Partner

Waymo's real achievement is the eval infrastructure, not the driving

Today's safety and alignment argument is about models, but the original machine learning safety debate was autonomous vehicles, and that field had to settle it with evidence. The striking part of hosting Dmitri Dolgov at SPC was the extensive eval and testing infrastructure Waymo built to earn the confidence to release 2-ton robots moving at 30mph through cities. A first Waymo ride still lands as a religious experience: technology that actually works on real roads. Full video coming.

  • #evals
  • #hardware
X

Thariq

Better models should buy prototypes, not 10x more shipped features

The wrong response to a capability jump is pushing ten times as many features to production. The right one is spending the surplus on understanding your users, running experiments, building prototypes, and learning about things you do not understand, so that what you ship actually works. Applied to games and 3D generation: it is great for imagining your game into existence, but you still have to find a satisfying game loop first.

  • #products
  • #agents

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.