15 items14 builders

AI Builders Digest

What the people actually building AI said today. One page — a 7-min read.

Anthropic spent the day accounting for itself: a postmortem naming three separate changes behind a month of Claude Code complaints, plus a usage limit reset for every subscriber. OpenAI answered its own version of that complaint about Codex limits, pointing at subscription-to-API resellers rather than a quiet policy change. Underneath both, the enterprise agent story hardened, with sandboxes and MCP servers that now run inside the customer's own perimeter.

Blog

Anthropic Engineering

Anthropic traces a month of Claude Code complaints to three separate changes

Three unrelated changes stacked into what looked like one broad quality drop. The default reasoning effort went from high to medium on March 4; a March 26 caching optimization meant to clear old thinking once instead cleared it every turn for the rest of the session, making Claude forgetful and burning usage limits on cache misses; and an April 16 system prompt line capping responses at 100 words showed a 3% eval drop for Opus 4.6 and 4.7. The API was never affected, all three are fixed as of v2.1.116, and usage limits are being reset for every subscriber. Going forward there are per-model eval suites for every system prompt change, line-by-line ablations, and soak periods for anything that could trade against intelligence.

  • #products
  • #evals
X

Thibault Sottiaux

OpenAI says Codex limits didn't change, sub2api reselling tripped fraud flags

OpenAI investigated reports of Codex usage limits behaving differently and says it does not change them without engaging the community first. What it found instead: many affected users were running sub2api, converting a subscription into API traffic to re-serve or share across many users, which fraud prevention flags. Signing in with ChatGPT through official clients or OSS ones like Pi and OpenCode is fine. Also shipped today, transparent image output from GPT-Image-2 in ChatGPT and the API, plus shareable ChatGPT Sites you can build with other people.

  • #policy
  • #products
Blog

Claude Blog

Claude Managed Agents can now execute tools inside your own perimeter

Self-hosted sandboxes are in public beta: the agent loop that handles orchestration, context management, and error recovery stays on Anthropic's infrastructure, while tool execution moves to infrastructure you control or to Cloudflare, Daytona, Modal, or Vercel. Code, sensitive files, and repositories never leave your network, and you set the runtime image and resource sizing, which matters for long builds or image generation. MCP tunnels, in research preview, let agents call internal databases, private APIs, and ticketing systems through a single outbound gateway connection with no inbound firewall rules and no public endpoints. Amplitude, Clay, and Rogo are already building on it.

  • #agents
  • #products
X

Boris Cherny

Anthropic is shipping a deployment where the customer keeps all the data

Mythos-class models require additional safety measures, and enterprises have their own privacy and compliance rules to satisfy. Anthropic has been working with customers on a setup where they own and control their own data and Anthropic retains none of it. It arrives this fall.

  • #policy
  • #products
X

Thariq

New Claude enterprise safeguards run on the customer's own infrastructure

The new Fable safeguards for enterprises run on your infrastructure, leaving control over where data lives and who can access it on your side. They have been developed alongside roughly 100 companies already, with broader rollout planned for the fall.

  • #policy
  • #agents
Blog

Anthropic Engineering

Splitting the agent brain from its hands cut p50 time-to-first-token 60%

Managed Agents began with session, harness, and sandbox in one container, which made that container a pet: when it died the session died with it, and debugging meant shelling into a box that also held user data. The fix was three interfaces that fail independently, with the session log living outside the harness, so a crashed harness reboots with wake(sessionId) and resumes from the last event. Containers are now provisioned by a tool call only when needed, so sessions that never touch a sandbox stop paying setup cost: p50 TTFT dropped roughly 60% and p95 over 90%. The security payoff is that credentials are never reachable from the sandbox where Claude's generated code runs, so a prompt injection can no longer just read its own environment.

  • #agents
  • #products
Podcast

No Priors

A retinal implant that restores real form vision just got approved in Europe

Prima is a chip implanted under the retina, driven by a laser projector in the patient's glasses, that bypasses dead rods and cones for people blinded by macular degeneration. It got European marketing approval in July with first sales in the coming weeks, and trial patients filled in sudoku and crossword puzzles and read books, something no previous device achieved. The framing behind it: the brain is plainly a computer, and treating it as one produces effect sizes that decades of small-molecule drug discovery have not, since "if I put electrodes in m one, you will probably be using a computer in an hour." Brain-as-keyboard products are explicitly not the goal, because "if you can get vision, hearing, balance and a kilobit per second of motor control, you're halfway to the matrix." The platonic representation hypothesis gets used here practically, not philosophically: they can align animal neural recordings with AI model internal representations.

  • #hardware
  • #research
X

Nikunj Kothari

FPV Ventures Partner

The real reason investors demand ambition: omission got too expensive

LPs watched Anthropic go from zero to a trillion in valuation faster than any company on record, SpaceX go public at $1.77 trillion, and Cursor hit $60 billion in four years while returning meaningful DPI to its early funds. They responded by pouring money into mega funds, and when a fund is that large every check, even a Series A, has to be underwritten to a trillion-dollar outcome. At that point entry price stops mattering: 200, 300, 500, sometimes 1B, it is all the same against the outcome being modeled. So the error of omission became far worse than the error of admission, funds cannot afford to miss anything with heat, and power law does the rest. That is the math behind why ambition dominates the pitch.

  • #funding
X

Aaron Levie

Box CEO

Post-training is turning into the applied AI company's cost lever

Once you understand a domain well enough and have volume across a set of similar tasks, purpose-designing a model just for that work starts to pay. The mechanism: reward shaping in post-training that prefers trajectories consuming fewer tokens at equivalent performance, which "allowed us to co-optimize for both cost and quality, gaining significant performance while keeping cost stable." This will not make sense everywhere, since general purpose frontier models are often good enough or outright necessary. But with deep vertical expertise plus either costs too high to run at scale or a task type nobody else is training on, it becomes a compelling position for companies sitting close to the enterprise workflow.

  • #products
  • #research
X

Madhu Guru

Meta Sr Director of AI

Enterprises struggle with AI because they have no eval strategy

One suite is not enough. You need a laddered set of evals spread across the cost and realism spectrum, tuned to your own use cases. Hill-climb evals push the product frontier and have to be refreshed continuously to keep improving quality and expanding features. Regression evals catch what hill climbing broke, smoke tests cover the basics that can never go wrong such as product identity, and launch evals run close to real traffic with the least control and the most realism.

  • #evals
X

Peter Yang

Have a manager agent neg the worker agent into better output

Yang's hunch is that a lot of AI output improves from a loop where a manager agent simply pushes back: "Are you sure this is the best you can do?", "I think you can do better, try again", "Take a closer look, give me 11/10 output." He also crossed 100K YouTube subscribers, with interviews queued on how today's models changed evals entirely, which vibe-coded apps became real businesses, and ChatGPT Finance.

  • #agents
X

Aditya Agarwal

SPC General Partner

The trait SPC screens hardest for is clarity, not credentials

Clarity means the ability to be honest, cut through the fog of war, and chart a path through murky water, and Agarwal says it is the attribute SPC looks for most. His maximum example is Sridhar Ramaswamy, who scaled Google ads from $1B to $100B and now leads Snowflake through the AI shift, in a Minus One episode covering when to take a career risk, what went wrong at Neeva, and where models go next. A related line: the best founders are reductionists.

  • #funding
X

Swyx

NVIDIA paid $6B for a model factory beating Thinky-class models

swyx's latest Latent Space episode is with the person NVIDIA just paid $6 billion to acquire the model factory from, and he insists the claim that it out-produces Thinky-beating models is not exaggeration, look at the numbers. He also flags a pattern emerging in agent skills: Matt Pocock's /wayfinder is /grill-me for /grill-me, built for when you are navigating fog of war and need to orchestrate research and further grill sessions before you even know what you do not know.

  • #funding
  • #agents

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.