Anthropic Engineering
Anthropic traces a month of Claude Code complaints to three separate changes
Three unrelated changes stacked into what looked like one broad quality drop. The default reasoning effort went from high to medium on March 4; a March 26 caching optimization meant to clear old thinking once instead cleared it every turn for the rest of the session, making Claude forgetful and burning usage limits on cache misses; and an April 16 system prompt line capping responses at 100 words showed a 3% eval drop for Opus 4.6 and 4.7. The API was never affected, all three are fixed as of v2.1.116, and usage limits are being reset for every subscriber. Going forward there are per-model eval suites for every system prompt change, line-by-line ablations, and soak periods for anything that could trade against intelligence.
- #products
- #evals
