Anthropic has demonstrated Claude operating autonomously for 48 hours without any human intervention in scientific decision-making. Given a list of 16 target proteins, the agent independently researched the biological context, spun up the required environment, coordinated a two-tier team of subagents, and linked ten specialized generators—including RFdiffusion, PXDesign, and Genie—into 24 operational workflows.

Wet-lab validation confirmed the results: out of 1,320 tested designs, 354 proteins successfully bound to 14 of the 15 targets (one was excluded due to intrinsic protein aggregation). The overall binding success rate hit 26.8%, compared to the industry standard baseline of 10–15%. On the RBX1 target, Claude's best design achieved a binding affinity of 3.9 nanomolar, easily outperforming the 45 nM mark set by a recent open-competition winner. However, limitations emerged: binding rates dropped to near zero on the synthetic protein BBF-14 and bacterial MBP due to biases in the training datasets of the underlying generators.

The strategic takeaway for the industry is clear: LLMs moving from text generation to the autonomous orchestration of specialized ML models fundamentally disrupts the economics of early-stage drug discovery. Instead of relying on expensive bioinformaticians to spend weeks manually stitching together tools and infrastructure, the agent shoulders the computational grind—compressing the primary screening cycle down to just two days.

AnthropicAI AgentsAI in HealthcareMachine LearningAutomation