Bacterial antimicrobial resistance is quietly mounting a catastrophic toll: roughly five million deaths were tied to drug-resistant infections in 2021, a figure projected to double by 2050. Compounding this crisis, the pharma pipeline has been functionally dry for half a century, with zero novel classes of antibiotics brought to market over the past 50 years, as noted by bioengineer César de la Fuente. Traditional wet-lab drug discovery remains a grueling, multi-year brute-force exercise of isolating and screening physical compounds from nature.

To break this economic and operational bottleneck, de la Fuente's lab at the University of Pennsylvania shifts early-stage discovery into software, querying the genomes of both living and extinct organisms for antimicrobial properties. By training deep-learning models to detect structural and sequence patterns across massive genomic and proteomic libraries, the lab replaces slow manual pipetting with computational screening. This approach collapses the initial timeline for identifying viable drug candidates from several years down to mere hours.

Integrating Language Models into the Discovery Pipeline

Beyond specialized biological sequence classifiers, the group embeds off-the-shelf generative tools directly into their scientific infrastructure. Operating at the intersection of computational biology, chemistry, and software engineering requires rapid cross-disciplinary synthesis and heavy tooling overhead.

The team deploys OpenAI's Codex and ChatGPT to draft computational pipelines, generate analysis scripts, process genomic datasets, and test research hypotheses. Converting early R&D from physical molecular trial-and-error into in-silico screening fundamentally alters the unit economics of biotech: compute-validated candidates enter wet-lab synthesis with pre-vetted statistical confidence, lowering the barrier to clinical entry and resetting early discovery timelines from years to weeks.

Artificial IntelligenceGenerative AIMachine LearningAI in HealthcareOpenAI