On this page

Propel

Propel is a Broad initiative to expand what scientists can discover with a new generation of AI research agents. It gives scientists the power to pursue biological questions previously beyond their reach, with the ambition of changing how we understand and treat disease.

A new instrument for discovery

AI systems now prove theorems, design molecules that work in the wet lab, and carry out analyses that once took experts weeks, and they improve every few months.1 The same capability reaches into every part of scientific work, from the literature to data, code, and the tools of the laboratory.2 Propel puts scientists in charge of that new capability, directing it toward questions that matter to their own work. Their experience helps them recognize promising results and judge whether they hold up, so they can build and direct work that was previously beyond their capacity.3

Propel begins with a small group given two months to pursue that work. Nothing here waits on a future breakthrough. Models already follow instructions and use scientific tools, and they improve faster than the institutions around them can adjust. What the first scientists learn will guide what gets built next and let the scientists who follow pursue ambitious questions of their own.

A confluence at the Broad

At the Broad, scientists from MIT, Harvard, and Harvard's teaching hospitals work together to tackle some of the hardest problems in biomedicine. Propel's first participants include inventors of the methods they use, who are now skilled at directing research agents. They know their datasets and can draw on experimental collaborators to test the ideas that emerge. That combination of scientific experience and experimental access is difficult to reproduce elsewhere. For Broad scientists already working with research agents, results like these are becoming familiar.

A Broad scientist spent two years studying melanoma with spatial genomics. He wanted to understand how neighboring cells help cancer cells evade immune attack.4 He gave an agent his notes and asked it to distill the analysis procedures and choices from that work. On a new dataset from 25 newly collected colorectal cancer samples, it prepared the data and ran the starting analyses in about a day.

A computational biologist leading a regulatory genomics lab has made AI agents part of how his lab uses and develops models that predict what a stretch of DNA does in a cell. With collaborators, he built an open source framework that lets agents run leading models from plain-language questions. Working hands-on with Claude, he also took a new model through implementation, computational testing, and manuscript submission in about five months. This single model predicts regulatory activity across cell types, designs DNA for a chosen activity, and makes the DNA patterns behind both inspectable.5

A postdoc coming from outside medicinal chemistry co-authored the study introducing Fleming, an agent for designing tuberculosis antibiotics. It combined molecular models with literature searches to design and refine candidates, weighing predicted activity, toxicity, and other drug properties. His collaborators synthesized six compounds from its designs, and all six inhibited the pathogen in the wet lab.6

Broad scientists were pursuing these projects months ago, with models less capable than those available today. If that is the new floor, where does the real work begin? The first Propel projects will explore that question by tackling biological puzzles and testing how much more of the work scientists can entrust to agents.

The first questions

  • Spatial genomics. How do tumors recruit immune cells, and can patterns in spatial data lead to a mechanism we can test?7Jackson Weir
  • Liver toxicity. Which cellular phenotypes signal liver injury, and can we detect them reliably in images?8Shantanu Singh
  • Regulatory DNA across cell types. Which DNA patterns control gene activity across cell types, and can we use them to design sequences for experimental testing?9Luca Pinello
  • Optical pooled screening. Can an agent learn from an expert's records to automate the analysis of an optical pooled screen?10Matteo Di Bernardo
  • Rapid experimental loops. Can a scientist train an agent to carry out wet-lab work, using experimental results to pursue a question over repeated rounds?11Yasha Ektefaie

Provisions for the crossing

Participants and projects. A small group of researchers will spend two months pursuing biological questions or developing ways to carry out scientific work that still requires sustained expert involvement. Both start from problems in the researchers' own work. Participants must already be highly skilled with research agents, including coding agents such as Claude Code and Codex equipped with scientific tools. The two-month window is deliberately aggressive: research agents make it plausible to attempt projects that would normally take six to eight months.

The proposal. Each researcher will explain the biological question or scientific workflow problem and how they plan to begin investigating it. They will also state what they need to get started. They should identify a result they could obtain in two months and how they would test it.

The work. The Broad will provide compute, storage, tokens, and a reasonable budget for early pilot experiments. Researchers will decide how to organize and carry out the work, meeting only when useful. They should be free to change direction as they learn. The process need not be clean; the goal is to make discoveries and advance the science.

Research-agent sessions will capture much of the work automatically. The logs record the agent's actions alongside the information it received. They preserve its plans and the results of its commands. When an attempt fails, the record shows how the agent tries again or changes direction, including any corrections from the scientist.12 For benchwork, participants will keep a lightweight log of key decisions and results. Each project will maintain a GitHub repository for its code, notebooks, and outputs.

The outcome. After two months, each researcher will report the findings and explain how they would pursue the remaining scientific questions. The report and session logs should let others follow the work, including unsuccessful approaches. The larger scientific question may still be open.

For those who follow

Scientific work rarely follows a path laid out in advance. It is more like an ocean crossing: only afterward do we see its wake as "a brief, glimmering trace on the waters."13

Research-agent logs let us reconstruct more of that scientific journey.

With each project lead's permission, a small analysis group, which may include the scientists themselves, will examine completed work during and after the two months. It will use research agents to read the logs alongside the results, examining where a scientist's judgment changed an investigation's direction and what helped an agent work through a difficulty.14 Project leads will review the interpretations, and the group will distinguish individual experiences from patterns that recur across projects.

The group will use research agents to build tools and guidance from what it learns, so other scientists can draw on that experience.1516 It will test when tools and guidance are enough to help with a new problem and when the fuller record of an earlier investigation is more useful.17 Some of that work may involve making existing scientific resources easier for agents to use.18

Perhaps the greatest lesson is simply knowing that a difficult thing can be done.19 Scientists are using research agents to pursue ideas they once set aside for lack of time or specialist help. Propel will show what they accomplish and how they do it, so other scientists can judge which of their own ideas they can now pursue.

The findings may reveal how collaboration and support needs change as scientists take on questions that research agents put within reach. The Broad can draw on that understanding to adapt as scientific practice evolves.20

Propel will turn each crossing's wake into a chart for those who follow.21

We invite philanthropic partners to join us in expanding what scientists can discover with research agents. To discuss Propel, contact Shantanu Singh at the Broad Institute.


  1. The pace now concerns the people building these systems, and the concern predates their essays. In May 2023, Hinton, Bengio, and the heads of OpenAI, Anthropic, and Google DeepMind signed a one-sentence statement that mitigating the risk of extinction from AI should be a global priority alongside pandemics and nuclear war. In April 2025, AI 2027 described AI research automating itself into superhuman systems within about two years. That September, Yudkowsky and Soares argued that superhuman AI should not be built at all. In January 2026, Anthropic's Dario Amodei set out the risks of powerful AI, from misuse to loss of control. On September 12, 2026, he wrote that AI "has been advancing drastically faster, driven primarily by AI's growing ability to build the next generation of AI," and that left unchecked this "could outrun our ability to understand and control these systems." The same essay expects AI to cure most major diseases within a decade. He asked the industry to pace the frontier, and OpenAI's Sam Altman agreed the same day. We take the risk and the promise seriously. This document asks a narrower question, which is what scientific work these systems already make possible and how scientists should direct it. Center for AI Safety, Statement on AI Risk (2023); Kokotajlo et al., AI 2027 (2025); Yudkowsky and Soares, If Anyone Builds It, Everyone Dies (2025); Amodei, The Adolescence of Technology (2026) and We Must Pace the Frontier (2026).↩︎

  2. At the Broad, we often think of these systems as apprentices: collaborators that can take on increasingly complex scientific work across literature, data, code, and experimental tools while scientific judgment remains with the scientist. Coding agents such as Earendil's Pi, OpenAI's Codex, and Anthropic's Claude Code can serve as research agents when equipped with scientific tools and access to data. Other examples include Anthropic's Claude Science and more autonomous systems such as Edison Scientific's Kosmos and Google DeepMind's Co-Scientist. Scientific examples also include Coscientist, which carried out chemistry experiments through an automated laboratory, and Virtual Lab, which designed nanobodies under human guidance. Bunne and Regev discuss a broader role for research agents in connecting models of cells with experimental investigation: "Beyond Representation: AI in Cellular Discovery", Daedalus (2026). The apprentice framing draws on Licklider's vision of people and computers developing ideas together, with people setting goals and evaluating the results: "Man-Computer Symbiosis", IRE Transactions on Human Factors in Electronics (1960). Autor and Manyika make the case for designing AI to work with human expertise and strengthen it: "A Better Way to Think About AI", The Atlantic (2025).↩︎

  3. Greka emphasizes the scientific intuition needed to frame meaningful biological questions and validate AI-assisted findings: "Building the Drug Discovery Engine of the Future with AI-Empowered Nodal Biology", Daedalus (2026). Koller argues that faster agent-driven experimentation depends on choosing objectives that reflect the biology of human disease: "Drug Discovery Has No Magic Wands", Deep Phenotype (2026).↩︎

  4. Weir et al., "A bone repair external state program drives cancer cell dedifferentiation and immune evasion in melanoma", Cancer Research (2026) [conference abstract].↩︎

  5. Chorus. Sequence-to-function models predict molecular activity, such as chromatin accessibility, transcription factor binding, or gene expression, directly from DNA sequence. They help scientists investigate the effects of genetic variants, but using several models together often requires substantial computational expertise. Chorus brings leading models, including AlphaGenome, Enformer, Borzoi, and ChromBPNet, behind one interface that an AI agent can operate. A researcher asks a question in plain language, such as what a disease-associated variant might do in liver cells. The agent selects appropriate models, runs them, and compares their predictions. Worked examples show how analyses that previously required days of assembling separate pipelines can be completed in an afternoon through conversation. The predictions provide a starting point for experimental validation and depend on the cell types and assays the models cover. Penzar et al., "Chorus: chatting with genomic oracles", Genomics x AI (2026).↩︎

  6. Wei et al., "Fleming: An AI Agent for Antibiotic Design for Mycobacterium tuberculosis", bioRxiv (2026).↩︎

  7. Spatial genomics. Slide-tags profiles nuclei in tissue without losing their spatial context, allowing scientists to study how neighboring cells shape one another. Jackson Weir co-invented the technology and is now using it to understand how tumors recruit immune cells into the tumor microenvironment. An agent recently prepared a new dataset and ran the starting analyses in about a day. Can they now go from those analyses to a biological mechanism they can test? He aims to build a discovery system that identifies local patterns associated with immune cell recruitment, then proposes a mechanism and helps him begin an experiment to test it. Russell et al., "Slide-tags enables single-nucleus barcoding for multimodal spatial genomics", Nature (2024).↩︎

  8. Liver toxicity. OASIS combines Cell Painting, transcriptomics, and proteomics to develop alternatives to animal testing. Its first study detected bioactivity at lower doses than standard assays, but the broad morphology profiles were difficult to link to specific functional changes or mechanisms of toxicity. Shantanu Singh built Cell Painting's computational foundations and now co-leads OASIS. He wants to understand which cellular phenotypes signal liver injury and whether we can reliably detect them in images. Research agents will help generate and test hypotheses, aiming to identify a first set of interpretable phenotypes. Rouquié et al., "The OASIS Consortium: integrating multi-omics technologies to transform chemical safety assessment", Toxicological Sciences (2025). Ewald et al., "Cell Painting for cytotoxicity and mode-of-action analysis in primary human hepatocytes", Cell Systems (2026).↩︎

  9. Regulatory DNA across cell types. Luca Pinello's lab developed CRISPResso2 to measure which DNA changes genome editing produces and how often. With DNA-Diffusion, the team designed short regulatory sequences and showed experimentally that they could drive gene activity preferentially in selected cell types. The new model learns reusable combinations of short DNA patterns recognized by transcription factors. It uses those patterns to predict regulatory activity and design DNA sequences of 200 base pairs for a desired activity in a chosen cell type. It currently covers four cell types. The next questions are whether the learned patterns can help predict regulatory activity in additional cell types, and whether the designed sequences perform as predicted in experiments. If validated, these designs could help scientists control which cell types express a gene introduced through gene therapy. Clement et al., "CRISPResso2 provides accurate and rapid genome editing sequence analysis", Nature Biotechnology (2019). DaSilva et al., "Designing synthetic regulatory elements using the generative AI framework DNA-Diffusion", Nature Genetics (2026).↩︎

  10. Optical pooled screening. A largely automated system recently completed a genome-wide Cell Painting screen of more than 5 million cells across 21,732 gene knockouts in eight days, with human intervention at each sequencing round. The analysis still required an expert to make dozens of decisions about the data and how to process them. Matteo Di Bernardo built that pipeline and has used it on several datasets. Can an agent learn from its session logs to make the analytical decisions that still require expert judgment? The project will test how far that experience can take an agent toward running the analysis independently. Kirby et al., "A one-week automated genome-wide optical pooled screen using OttoSeq", Genome Biology (2026). Di Bernardo et al., "Brieflow: An Integrated Computational Pipeline for High-Throughput Analysis of Optical Pooled Screening Data", bioRxiv (2025).↩︎

  11. Rapid experimental loops. Yasha Ektefaie co-authored the study introducing Fleming, the antibiotic-design agent described above. He is building a lab where agents use experimental results to design and run further experiments. He wants to do this with off-the-shelf equipment and downloadable tools. The work includes MelanomaBox, a wet-lab sandbox for testing agent-proposed drug combinations in living melanoma cells. Recently, a small team spent two days connecting agents to a robot to test that idea. The agents proposed drug combinations to kill melanoma cells, which the team tested using the robot with some human assistance. The results fed back to the agents for another round. He is now building a prototype in which a scientist trains a research agent to carry out wet-lab work with a robot arm. The instrumentation is itself part of the scientific question: what could an agent explore with access to the bench? Rao, "Measuring the frontier of scientific capability in the age of AI", Substack (2026).↩︎

  12. Logs may also help improve how agents work. Meta-Harness uses prior execution traces and evaluation results to revise an agent's harness, the code that manages its context and interaction with tools. See Lee et al., "Meta-Harness: End-to-End Optimization of Model Harnesses", arXiv (2026). Zhang and Khattab argue that harnesses can help models tackle unfamiliar problems by breaking them into familiar parts: "Language Model Harnesses Are Compositional Generalizers" (2026). These approaches suggest ways to turn project experience into better tools; their value for Propel's scientific work would need to be tested.↩︎

  13. Whyte, "Ambition," Consolations (2015): "A life's work is not a series of stepping stones onto which we calmly place our feet, but more like an ocean crossing, where there is no path, only a heading, a direction, in conversation with the elements. Looking back, we see the wake we have left as only a brief, glimmering trace on the waters." Whyte contrasts ambition fixed on a destination with a calling that changes us through the work. He closes with generosity, passing on to others "the core artistry that made the journey a journey."↩︎

  14. Engelbart, Augmenting Human Intellect: A Conceptual Framework (1962). Deming, The New Economics for Industry, Government, Education (1993). This feedback loop has precedents in Engelbart's bootstrapping (using tools to improve the tools themselves) and Deming's Plan-Do-Study-Act cycle for iterative improvement.↩︎

  15. One simple example comes from planning experiments for MelanomaBox, discussed above. A researcher needed to choose a concentration range for each drug. He pointed Codex to an earlier expert discussion about whether reported potency values described effects on a drug's target or on cell survival. Codex built a script that combines evidence from public databases and papers to propose concentration ranges, preserving what each assay measured. The saved session logs record the questions and corrections that shaped the analysis. Another scientist could give Codex those logs and the script alongside a new compound list and experimental design, then ask which parts apply and what needs to change. Propel could compare the quality of the experimental choices and the effort required with those of the scientist's usual approach. Propel would study this pattern of adapting earlier work as questions and evidence change, including in more complex investigations.↩︎

  16. Some difficulties encountered in these projects could become science sandboxes: controlled settings where agents investigate a system through repeated experiments, recording how their explanations change with the evidence. These would let us test whether better tools or guidance help agents overcome a particular difficulty. See Rao et al., "Science sandboxes measure the scientific capability of AI agents", arXiv (2026), and the accompanying blog post, Rao et al., "Measuring the frontier of scientific capability in the age of AI", Substack (2026).↩︎

  17. Recent agent benchmarks find that episodic records can preserve useful information lost when agents distill experience into general guidance. See Zhang et al., "Useful Memories Become Faulty When Continuously Updated by LLMs", arXiv (2026), and Lei et al., "SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills", arXiv (2026). Their results do not yet establish which will work best for scientific investigations.↩︎

  18. The projects may reveal where agents need better ways to access and use the Broad's existing data resources. A recent evaluation of NCBI Virus shows how much a better interface can change agent performance: Luebbert, "Paving the way for agents in biology", Anthropic (2026). Through the human interface, agent accuracy ranged from 16.9% to 91.3% and varied across runs; with gget virus, every evaluated agent exceeded 90%, peaking at 99.7%, and run-to-run variability largely disappeared. Illustrative targets, to be finalized with resource owners: gnomAD (GraphQL), the Connectivity Map / L1000 (CLUE REST), DepMap (Breadbox API), the Broad-hosted GeneGenie results API (REST; includes FinnGen R13, UK Biobank, MVP, and other cohorts), the GTEx Portal (REST), and file-based resources such as the JUMP Cell Painting data, MSigDB, and the PROSPECT tuberculosis chemical-genetics data. Depending on the resource, the right interface may be an API, a lightweight data file, or a reusable skill.↩︎

  19. Roger Bannister ran the first sub-four-minute mile on 6 May 1954. John Landy surpassed his record 46 days later. World Athletics, "Fab five: mile races" (2019).↩︎

  20. Repenning and Kieffer's dynamic work design offers a precedent for improving how work is organized through feedback from the people doing it. Their account includes experience in the Broad's sequencing lab. How far such ideas extend to open-ended scientific discovery remains a question for Propel. Repenning and Kieffer, There's Got to Be a Better Way: How to Deliver Results and Get Rid of the Stuff That Gets in the Way of Real Work (2025); excerpt.↩︎

  21. Acknowledgments. Participants: Jackson Weir, Yasha Ektefaie, Shantanu Singh, Matteo Di Bernardo, Luca Pinello. On the framing: Nodar Gogoberidze, Alan Munoz, Arya Rao, Eric Lander, Johan Haslum, Jess Ewald, David Dao, Jackson Weir, Yasha Ektefaie, and Zeming Lin. For thinking out loud: Aleksandr Zimin, Alex Kalinin, Ank Kumar, Anna Greka, Anne Carpenter, Aviv Regev, Beth Cimini, Blake Lash, Brendan Saloner, Bryan Hsu, Carolina Wahlby, Daphne Koller, David Autor, Deb Hung, Enoch Huang, Jason Swedlow, Jan De Boer, Jesse Boehm, Jessika Baral, John Arevalo, Jonne Rietdijk, Iain Cheeseman, Luca Pinello, Matteo Di Bernardo, Misha Belkin, Nabiha Saklayen, Nelson Repenning, Pardis Sabeti, Paul Blainey, Puneet Batra, Ramnik Xavier, Roby Bhattacharya, Sandeep Kambhampati, Soumya Raychaudhuri, Tim Treis, Wei Ouyang, Wolfgang Huber, and Yajit Jain.↩︎