7 items7 builders

AI Builders Digest

What the people actually building AI said today. One page — a 4-min read.

Anthropic shipped Claude in Chrome to general availability and published the prompt injection numbers that made it possible, which is the first time a browser agent vendor has put per-model attack success rates next to a GA announcement. The rest of the day was builders converging on the same thesis from different angles: agents are about to become the primary consumer of software and the primary channel for commerce, and the businesses built around humans looking at screens have not priced that in. On the tooling side, the interesting argument was against throwing more context at agents rather than less.

Blog

Claude Blog

Claude in Chrome hits GA and can act in your browser without per-click approval

Claude in Chrome is now generally available on every paid plan, and it approves its own safe actions by default instead of asking you each time, using the same mechanism as auto mode in Claude Code. The pitch is the long tail that has no integration: internal dashboards, legacy systems, vendor portals, all reached through your existing logins. Anthropic published the red-team numbers behind the decision: on its current evaluation, attacks that reached the model succeeded 17.6% of the time against Opus 4.5 and 3.8% against Opus 5 before safeguards, and with probes plus the action classifier no attacks succeeded against Sonnet 5 or Opus 5, with 0.3% against Fable 5. You still need the desktop app for local files, and it does not run on other Chromium browsers or mobile.

  • #agents
  • #products
  • #security
X

Aaron Levie

Box CEO

Agents that transact on your behalf become the commerce layer

Levie argues the monetization story for personal agents is bigger than assistant subscriptions: once an agent can handle an arbitrarily complex task end to end, a substantial share of commerce inevitably routes through it, and lowering friction means people spend more than they did before. That creates value at two layers at once, the agent providers and whatever the agents transact against in commerce, local, and b2b services. His second point is the enterprise mirror of it: agents will use software 100 times more than people ever did, and the core primitives matter more, not less, when an agent can take destructive actions or work off the wrong context. He puts the opportunity with whoever becomes the security layer, data manager, and workflow orchestrator for agents.

  • #agents
  • #products
X

Peter Yang

Agents browsing sites for you could gut the targeted display ad business

Yang's prediction: a rude awakening is coming for ad markets, because if your business is showing targeted display ads to humans, agents that browse your site and complete the job without a human ever seeing the page break the model outright. He frames his own behavior as the leading indicator, saying he no longer lives in email or text but in the chat with his agents. He also relays a ChatGPT Finances use case from its product lead Ethan, who was billed for a hotel he had booked with points and only caught it because ChatGPT flagged the charge and worked with customer service to get him reimbursed.

  • #agents
  • #products
X

Nikunj Kothari

FPV Ventures Partner

Tokenmaxxing correlates with worse products, not better ones

With few exceptions, Kothari says the companies boasting about tokenmaxxing ship the worst product experiences, because good products come from curation and gardening rather than throwing the kitchen sink at an agent and hoping it sorts things out. Less is more has never been more apt. He separately rates Codex on Mac as undefeated for Computer Use work and suggests taking a workflow you do by hand and asking it to one-shot the whole thing, while noting that Instinct and Muse are the ones showing what a good agent feels like for the masses with zero setup.

  • #agents
  • #products
X

Garry Tan

Y Combinator President and CEO

Capy is outrunning Codex and Claude Code on large multi-step PRs

Tan calls capy.ai his favorite agentic coding weapon of the past week, saying it tracks multi-step workflows and lands large PRs faster than Codex or Claude Code on their own, though he admits he does not know how it does it. His example is an ambitious bug fix wave on GBrain with clear task delineation, automatic parallelization, and clean GitHub PR and CI workflow. He also defends Cluely on the merits, arguing that a realtime thought helper and semi-adversarial assistant carrying ongoing context is still a good idea.

  • #agents
  • #products
X

Peter Steinberger

Meta built its own agent inspired by OpenClaw, not on top of it

Steinberger corrects the circulating story that Meta uses OpenClaw: Nat and his team built their own agent, inspired by it, and he gives them credit for the work. He also points to a security review of OpenClaw that turned up nothing critical, saying they did their homework. His third note is the structural argument for self-hosting an agent: if you run the claw yourself, nobody can block you.

  • #open-source
  • #agents
  • #security

Get this in your inbox

One email a day. Unsubscribe in one click.

Where this comes from

Source data comes from the open-source project follow-builders by zarazhangrui, released under the MIT license. Summaries are generated by an LLM from that project's public feeds, and the summarization prompts are adapted from it. Every item above links to its original source.

Summaries generated automatically. Read the original before relying on any claim.