Training Data
Chai Discovery took antibody design from a 0.1% hit rate to about 15%
When the company started, state of the art antibody design bound about one molecule in a thousand, weak enough that most targets returned no hits at all and you could never see drug-like properties in the results. Chai-2 gets roughly 15%, so screening a thousand candidates returns about 150 and you can finally read real statistics off the output. The approach is imported straight from LLM work: refuse to add the twenty-fourth submodule, keep the architecture simple enough that scaling laws are findable, and treat antibodies, mini proteins, and everything else as different prompts to the same model. Their contrarian claim is that biology is more verifiable than code, since binding and manufacturability are objective readouts while code taste is not, and the discipline that demands is blunt: "You can fool yourself so easily in biology."
- #research
- #evals
- #startups
