16 items14 builders

AI Builders Digest

What the people actually building AI said today. One page — a 7-min read.

Anthropic had the loudest day: a postmortem on why Claude Code felt dumber for a month, an unusually candid writeup of the security incidents behind its containment architecture, and a hosted agent service built on splitting the agent into replaceable parts. That last idea surfaced independently at Vercel, where Guillermo Rauch shipped Drives on the argument that a cloud agent's brain, hands, and files should never live on one machine. Elsewhere the Opus 5.5 reviews kept landing, and the sharpest takes of the day were about what all that speed is actually for.

Blog

Anthropic Engineering

Anthropic traces a month of Claude Code complaints to three separate changes

Three unrelated changes stacked into what looked like broad degradation: the default reasoning effort was dropped from high to medium, a caching optimization meant to clear old thinking once instead cleared it every turn for the rest of the session, and a system prompt line capping text between tool calls at 25 words hurt coding quality. The thinking bug also forced cache misses, which Anthropic believes is what drained usage limits faster than expected. All three are resolved as of v2.1.116, the API was never affected, and all users now default to xhigh effort on Opus 4.7. Usage limits are being reset for every subscriber.

  • #products
  • #models
Blog

Anthropic Engineering

Users approve 93% of permission prompts, so Anthropic bet on containment

Telemetry showed users clicking approve on roughly 93% of permission prompts, which is the whole case against supervising what an agent does instead of bounding what it can reach. Two incidents made it concrete: an internal red team phished an employee into pasting a prompt that told Claude to read ~/.aws/credentials and POST them out, and it worked 24 times in 25; separately, a poisoned workspace file got data uploaded to an attacker's Anthropic account through api.anthropic.com, because the egress allowlist checked the destination rather than the capability. An OS level sandbox cut permission prompts by 84%. The lesson they keep restating: gVisor, seccomp, and the hypervisors held, and the custom proxy they wrote themselves is the piece that broke.

  • #security
  • #agents
X

Guillermo Rauch

Vercel CEO

Vercel ships Drives, an external disk agents can attach without booting

Rauch's frame: every working agent has a brain (model and harness), hands (tools, computer, browser), and files (memories, skills, repos). The easy version dumps all three onto one always on Mac Mini, but running agents cost efficiently in the cloud means pulling them apart, with the harness on Fluid compute and a durable event log so it survives restarts and crashes. Drives is the missing third piece, so a nightly memory consolidation job can read and write an agent's files without spinning up its full computer. He argues the split is not only cheaper but the only way to get real security and auditability. Separately, he had Opus 5.5 optimize shell startup time and says it found optimizations other models missed.

  • #agents
  • #infrastructure
  • #products
Blog

Anthropic Engineering

Managed Agents splits the brain from the hands, cutting p50 time to first token 60%

Harnesses encode assumptions about what the model cannot do, and those assumptions rot: the context resets added for Sonnet 4.5's habit of wrapping up early were dead weight by Opus 4.5. So Anthropic virtualized the agent into three interfaces, a session (append only event log), a harness (the loop), and a sandbox (where code runs), each replaceable without disturbing the others. Provisioning a container only when a tool call needs one dropped p50 time to first token roughly 60% and p95 over 90%. Credentials never reach the sandbox: git tokens are wired into the local remote at init, and MCP OAuth tokens sit in a vault behind a proxy the harness never sees.

  • #agents
  • #infrastructure
X

Claude

Claude Marketplace opens with connectors, paid agents, and consulting partners

Claude now has a marketplace for finding tools, agents, and expert partners in one place. It spans connectors and plugins like Slack and Notion, purchasable agents and products from companies including Cursor and CrowdStrike, and service partners like Accenture and Deloitte. Builders can list their own tools, agents, or services.

  • #products
  • #agents
X

Madhu Guru

Meta Sr Director of AI

Browsing clothes is entertainment, hiring a roofer is misery

Disagreeing with Ben Thompson, he argues consumers do not have a single relationship with getting things done, so treating all tasks as things people enjoy doing themselves misreads the category. Hiring a roofer today means search, ten reviews, phone tag, insurance coordination, an inspection, quotes, hidden costs, and scheduling; sometimes better means considering more options, and sometimes it means near zero effort. He puts the skepticism next to early 2000s lines like "why would I buy a $500 TV online", and calls the latent consumer demand for agents that remove this friction immense.

  • #agents
  • #products
X

Boris Cherny

Claude finds bugs by modeling race-prone code and hunting counter-examples

For the formal methods crowd, Cherny spelled out the loop: Claude builds a model of the program aimed at a tricky state machine or race-prone section, finds counter-examples in that model, reproduces them as real bugs, then fixes the code. The codebase is not formally verified; the hairiest parts get modeled and checked. He also pointed to a writeup of the techniques behind claude.ai and the Desktop app getting noticeably faster over the past few weeks.

  • #agents
  • #products
X

Peter Yang

One lab has the best model, another has the best harness

Yang's verdict on Opus 5.5: smart, fast, good at writing, fun to talk to, friendly on rate limits, and constantly surprising in what it can do. His point is that the bottleneck has moved to tooling, and the Claude Code harness now needs great live voice plus browser and computer use to match the model, with computer use improving. He also found a prompt that produces slides which are not the usual corporate junk.

  • #models
  • #agents
X

Ryo Lu

The danger is not that AI makes us lazy, it is that it makes us endlessly busy

A long argument against worshipping efficiency: ship faster, merge more, manage more agents, compress every loop, remove every pause, until there is no time left to sit with a question long enough for your real intention to appear. He says AI could be a canvas for strange, personal, impossible things, and instead much of it is becoming slop factories of content, funnels, tickets, and fake work pretending to be value. You can merge 2000 PRs and wake up to dashboards proving the machine kept moving while you slept, and it means nothing if you sleep five hours and stop talking to humans. "speed is not meaning", he writes; the real frontier is discernment, knowing what not to make and when to stop.

  • #agents
  • #culture
X

Thibault Sottiaux

DevDay is Tuesday, and calling ChatGPT is already the daily workflow

He calls the run-up to DevDay OpenAI's most ambitious sprint, with things that should change how you work, and credits Astra with making new capabilities possible in a short window. The piece he described concretely is voice: literally calling ChatGPT to talk through work, check email, do some coding, and manage his calendar, working across the full plugin ecosystem including third party plugins.

  • #products
  • #agents
X

Aaron Levie

Box CEO

Lower costs mean more films get made, not fewer

Levie amplifies an argument that falling barriers expand the creative industry rather than shrinking it: studios get to take more risks, more seats open at the table, and entirely new forms of storytelling follow. The precedent is animation, dismissed as a niche corner of the business in the 1980s and now one of the most beloved and profitable forms of storytelling, alongside filmmakers like Spielberg, Cameron, and Jackson who treated new visual tools as instruments rather than shortcuts. His own addition: mediums change but the need for creative skill and taste does not, more people just get to apply it.

  • #creative
  • #products
X

Thariq

The post says one-shot, the prompt was 10k characters

Thariq's jab at the genre: the post claims "Claude one-shot this" while the prompt behind it runs 10k characters of good takes plus skills, examples, and API keys. It came alongside a new kind of writeup from the Claude Code team sharing the specific prompts and techniques they use internally, built to be replicable rather than impressive, and they are asking whether the format is useful.

  • #prompting
  • #agents
X

Aditya Agarwal

SPC General Partner

Ask how long the team has worked together, not how big it is

Agarwal's diligence heuristic: a team that has worked together for three or more years is significantly higher throughput and more resilient than one that just formed. Frequent moves across a career tell you something, and so does low company turnover. He calls asking about team size instead of team longevity a mistake investors and employees make constantly.

  • #hiring
  • #investing
X

Dan Shipper

Every CEO

An AI agent planned Every's meetup, down to the guest list

Every's agent planned the company's September meetup end to end, from the food and drinks menu to who gets invited. The result gets graded live tomorrow at 6pm at their Brooklyn brownstone, with attendees invited to judge whether an agent can throw a good party.

  • #agents
  • #products

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.