AI and Precision Medicine: Expanding the Genomic Landscape in Oncology

Introduction
Artificial intelligence has rapidly become a cornerstone of precision oncology, not only for optimizing drug discovery, but for elucidating mechanistic layers that govern therapeutic response. Within the life sciences, AI is increasingly being used to enhance drug design and to match patients with therapies based upon their personal data and how the treatment works. Successful outcomes critically depend on the quality of training data and allowing models access to as much systematically structured empirical data as possible, particularly around different facets of drug target and efficacy. Doing so can address long-standing translational biological questions, such as what DNA sequences and features are most vulnerable to damage and mutation in a cancer’s genome upon therapeutic targeting, evolution, and resistance? And given this, how can we better formulate a therapeutic strategy to achieve the best outcome for the patient?
As an example of AI’s expanding clinical reach, a machine-learning method called SCORPIO was developed using data from routine blood tests and clinical records to estimate the probability of benefit from immune checkpoint inhibitors and the resulting survival outcomes.7 Predicted outcomes exceeded those of conventional biomarker assays such as tumor mutation burden and PD-L1 immunostaining, illustrating how AI can extract predictive signal from inexpensive, broadly available data.
The intersection of AI and precision oncology represents a paradigm shift in how we approach cancer treatment. By leveraging computational power to analyze vast biological datasets, researchers can now identify patterns and susceptibilities that were previously invisible to traditional analytical methods. This capability becomes particularly powerful when integrated with multi-omic data, including comprehensive genome-wide analysis platforms that extend beyond the mere 2% of the genome that codes for proteins.
From the Human Genome Project to AI: A Brief History of Precision Oncology
To appreciate where AI is taking precision oncology, it helps to recall how far the field has already traveled. The modern era of molecularly targeted cancer therapy began at the turn of the millennium with two landmark approvals: trastuzumab in 1998 for HER2-positive breast cancer, and imatinib in 2001 for chronic myeloid leukemia (CML), which targets the BCR-ABL fusion protein generated by the Philadelphia chromosome.1 Imatinib, in particular, became the springboard for an entire generation of small-molecule inhibitors, demonstrating that a drug rationally designed against a defined molecular driver could transform a once-fatal disease into a manageable condition. In its first randomized Phase III study, imatinib produced complete cytogenetic responses in 76.2% of patients compared with 14.5% for the prior standard of care.1
The completion of the Human Genome Project in 2003, which sequenced roughly 92% of the human genome, supplied the foundational map that made systematic molecular profiling of tumors conceivable.2 Over the subsequent two decades, plummeting sequencing costs, advances in next-generation sequencing (NGS), and creation of large-scale variant databases turned the genomic characterization of tumors from a research luxury into a clinical expectation. Histology-agnostic, biomarker-defined approvals followed, where the molecular alteration, rather than the tissue of origin, defines the indication for treatment.
Yet this progress exposed a persistent bottleneck. Even for a cancer with a single, well-understood therapeutic target like CML, the journey from discovery to standard-of-care clinical use historically took more than thirty years.3 Physicians and molecular tumor boards became increasingly overwhelmed by the sheer volume and complexity of molecular data, struggling to organize, cross-reference, and translate it into patient benefit.3 It is precisely this gap—between the explosion of genomic data and our capacity to interpret it—that artificial intelligence is now poised to close.
AI-Driven Genomic Analysis: Beyond the Coding Genome
In genomics, AI has already been deployed to probe the vast information stored in a cell’s DNA. GROVER is a large language model developed by Dr. Anna Poetsch’s team at the Biotechnology Center of Dresden University of Technology that has been trained on the full human genomic sequence to understand with new precision the complex mechanistic traits present in our genome.4 GROVER stands for “Genome Rules Obtained Via Extracted Representations,” and extracts information beyond genic regions to delve into the “dark matter” that exists in the noncoding parts of the genome.
In this model, human DNA is viewed as a long sequence of text within which specific patterns of ‘words’ exist. Through sequence patterns and contextual extrapolation, this AI model can accurately predict the DNA sequences that follow as well as the function of the DNA sequence itself, whether it serves as a promoter, an enhancer, binds to proteins, or more. The significance of GROVER’s approach lies in its ability to learn the “grammar” of DNA—understanding not just individual nucleotides but the contextual relationships between sequences that determine biological function.4
The way the GROVER team built the model is itself a notable advance. Instead of slicing DNA into fixed-length chunks, they applied byte-pair encoding—a compression technique borrowed from natural-language processing—over hundreds of iterative cycles, letting the most frequently recurring multi-letter combinations surface as the model’s vocabulary. The result is a dictionary of genomic “words” tuned to the statistical structure of the genome itself. When the researchers probed what the model had internalized, they found its representations captured properties such as how often a token appears, what it is composed of, and how long it is; a handful of tokens clustered tightly within repetitive elements, while the majority ranged broadly across the genome.4 Tellingly, functional annotations could be recovered from sequence context alone, meaning the model inferred biological meaning it was never explicitly taught. For oncology, this kind of genomic literacy offers a sharper lens on how a drug target actually behaves within its surrounding sequence.
GROVER is part of a rapidly expanding class of genomic foundation models that increasingly read the non-coding genome with unprecedented fidelity. In 2025, Google DeepMind introduced AlphaGenome, a model that takes one million base pairs of DNA sequence input and extrapolates thousands of functional genomic features—gene expression, chromatin accessibility, transcription factor binding, splicing, and three-dimensional chromatin contacts—at single-base-pair resolution, matching or exceeding the best specialized models on 25 out of 26 variant-effect-prediction benchmarks.5 Critically, AlphaGenome’s developers emphasize that more than 98% of human genetic variation is non-coding, but the model was still able to correctly identify the functions of clinically relevant variants near the TAL1 oncogene.5 In parallel, the Arc Institute’s Evo 2, a foundation model trained on roughly 9 trillion DNA base pairs bridging all domains of life, learned to understand and reveal the functional and pathogenic consequence of sequence variation across both the non-coding and coding regions in the absence of focused task adjustments.6
Together, these models signal a decisive shift: AI is learning to interpret the functional consequences of variation across the entire genome, not just its protein-coding fraction.
Precision Oncology and Drug Target Refinement
Precision oncology centers on tailoring treatment to defined patient subgroups using diagnostic and prognostic biomarkers that track disease progression, gauge treatment response, and expose the molecular roots of drug resistance. Computational drug prioritization has become central to this effort. Rather than leaning on any single data stream, current approaches weave together biological network analysis, curated knowledge bases, and predictive modeling to rank candidate therapies, turning sprawling molecular datasets into ranked, decision-ready options for the clinic.8
Much of this capability rests on machine-learning models trained to connect a tumor’s molecular profile to how it is likely to respond to a given drug. Large pharmacogenomic resources—chief among them the Genomics of Drug Sensitivity in Cancer (GDSC) and the Cancer Cell Line Encyclopedia (CCLE)—supply the labeled examples these models learn from, pairing gene expression and mutational features with measured drug sensitivities. Deep-learning architectures have proven especially relevant to the problem, because they can fuse several layers of omics data at once and capture the nonlinear relationships among genes, pathways, and compounds. Tools in this vein, including DeepDR, DeepSynergy, and GraphDRP, predict drug response or effective drug combinations, surfacing resistance-linked patterns and nominating rational combination strategies that might otherwise stay hidden.8,9
Layering multiple omics measurements together yields a fuller picture of tumor biology than any one modality alone. Integrative platforms such as PANOPLY and MOalmanac merge genomic and transcriptomic data to nominate and rank therapeutic targets for individual tumors.10,11 Applied across large genomic and epigenomic datasets, the same AI methods can flag candidates for drug repurposing, identify compounds aimed at specific epigenetic modifications, and build models that tie epigenomic and clinical features to patient outcomes—broadening the target space well beyond protein-coding mutations.10,11
Refining Mechanism of Action and Improving Drug Development Success
Beyond identifying targets, AI is beginning to reshape the economics and probability of success of drug development itself—a domain where the stakes could hardly be higher. Bringing a single drug to market traditionally takes over a decade and more than $2 billion, with roughly 90% of candidates failing somewhere along the way.12 Much of this attrition stems from an incomplete understanding of a candidate’s true mechanism of action and from the late discovery of toxicity or lack of efficacy.
Early evidence suggests AI can meaningfully shift these odds. A 2024 analysis of AI-native biotechnology companies found that AI-discovered molecules achieved an 80–90% success rate in Phase I trials, considerably higher than the industry average of roughly 40–65%.13 This implies that AI is decidedly adept at generating and screening molecules with favorable drug-like properties and safety profiles. Tellingly, the same analysis found that Phase II success rates—where the bar shifts from safety to efficacy—remained near the historical average of about 40%.13 The lesson is instructive: AI has become proficient at the chemistry and pharmacology of making a viable molecule, but efficacy still hinges on choosing the right target and correctly understanding how a drug engages the disease biology. The field is also maturing realistically; candidates such as Recursion’s REC-994 have been discontinued when long-term data failed to confirm early signals, even as AI-originated drugs like Insilico Medicine’s ISM001-055 have produced positive Phase IIa results in idiopathic pulmonary fibrosis.14
This is where a comprehensive mechanistic understanding of how a drug acts on the cancer, such as where it triggers the DNA in cancer cells to break enough to cause cell death, becomes invaluable. If AI’s current blind spot is efficacy—knowing whether a well-made molecule will truly disable the right vulnerability in the right patient—then richer, empirical data on a drug’s mechanism of action directly addresses the weakest link in the development chain. Refining the mechanism of action is not merely an academic exercise; it is the most leveraged point at which to improve the probability that a candidate survives Phase II and beyond.
DNA Damage Response Inhibitors and Precision Medicine
The DNA damage response (DDR) has become one of oncology’s most actively pursued target spaces, and the logic is rooted in a strategic vulnerability. Many tumors already carry defects in their DDR machinery or labor under chronic replication stress, leaving them precariously dependent on whatever repair capacity remains. Inhibiting that residual capacity tips the balance, allowing unrepaired damage to accumulate until the cancer cell dies—the principle of synthetic lethality.15,16 The clearest validation of this strategy came from poly(ADP-ribose) polymerase (PARP) inhibitors, which have delivered meaningful benefit to patients whose tumors carry BRCA mutations or broader homologous recombination deficiency (HRD).
Applying DDR inhibitors precisely means matching each drug to the patients whose tumor biology predicts a response. Specific mutations can forecast sensitivity, but the field is still actively mapping the full landscape of exploitable weaknesses, the routes by which tumors escape, and the ways different repair pathways reinforce or compensate for one another.15 A growing roster of agents now reaches beyond PARP to inhibit other nodes of the network—DNA-PK, ATM, ATR, CHK1, CHK2, and WEE1 among them—and rationally designed combinations are showing early promise against resistance.17,18 Yet one question still persists at scale: across the roughly three billion base pairs of a tumor’s genome, where exactly does a given therapy land its lethal blows—and does that map look different in cells that respond versus those that survive?
BreakSight: A Unique Angle on the Genome’s Hidden 98%
New and pioneering efforts to better characterize the entire genome of a diseased cell, and not just its genes, could enhance the way we derive insights about disease and subsequent therapeutic success. Given all this incredible headway, how can we further utilize our evolving understanding of the way DNA plays a role in disease and therapy to gain further traction in precision medicine? Information on the precise sites of breakage across the cancer’s DNA that result from targeted cancer drugs, such as DDR inhibitors, can be integrated into AI-backed analyses to more holistically capture critical characteristics about the cancer that can better predict therapeutic efficacy.
This is the gap BreakSight sets out to close. Today’s DNA biomarkers are overwhelmingly mutations within expressed genes—the 2% of the genome that codes for proteins. BreakSight’s aim is to extend therapeutically meaningful coverage into the other 98%.19 What sets the approach apart is what it chooses to measure. Instead of cataloguing which genes are mutated, it records where treatment physically fractures the DNA, producing an empirical, functional readout of a drug’s mechanism inscribed directly on the genome. As one of the first commercial assays to pinpoint sequences tied to site-specific DNA damage, BreakSight captures something most genomic tests do not: the genome’s response to therapy at sequence resolution, rather than its fixed starting state.
At the platform’s core is DDsite, which tags, recovers, and sequences double-strand break sites genome-wide—a workflow built for cells that have been exposed to DNA-damaging agents. The service runs the full pipeline from break labeling within permeabilized cells, DNA break site retrieval, and library preparation through next-generation sequencing at 150-bp paired-end resolution.19 Conceptually, while similar, it deviates from in situ break-labeling methods such as BLESS (Breaks Labeling, Enrichment on Streptavidin, and next-generation Sequencing), which ligates biotinylated oligonucleotides to break ends, by avoiding fixation-induced artefactual breaks and extending captured DNA lengths to allow for more accurate break site resolution at repetitive regions.20 A companion assay, DDsite-Cyto, isolates and reads cytoplasmic breaks apart from nuclear ones—a distinction that matters for immuno-oncology, given that cytoplasmic DNA is a potent trigger of innate immune signaling, and the sequences it flags may help anticipate which tumors will synergize with immunotherapy.19
Once breaks are captured and sequenced, the DDinsight analysis suite takes over, mapping sequenced reads across the genome, distinguishing statistically significant, reproducible, and treatment-specific break sites from background, and comparing break patterns across conditions. Its extended version, DDinsight+, digs deeper—surfacing enriched repeats and sequence motifs and anchoring break sites to annotated genome features through public resources such as ENCODE.19 For drug developers, the practical payoff is a focused set of readouts: recurring breaks in compound-sensitized cells, treated-versus-resistant comparisons that expose emerging resistance mechanisms and fresh targets, the combined genomic footprint of drug combinations, and complementary mechanisms that inform how to dose and schedule for maximum effect with minimal off-target liability.21 An optional RNA-Seq add-on ties these break signatures back to the pathways switched on or off by treatment or tumor biology, clarifying how shifts in protein expression and network activity render particular sequences more prone to breakage or repair.19
Taken together, BreakSight reframes the question of drug mechanism. Instead of inferring how a therapy works from indirect markers, it offers a direct, genome-wide map of the molecular wounds a drug leaves behind across coding and, crucially, non-coding regions alike. This is the empirical substrate that AI has, until now, largely lacked.
Integrating DNA Break Signatures with AI-Driven Precision Oncology
The convergence of comprehensive DNA break mapping with AI-powered genomic analysis creates unprecedented opportunities for refining drug target mechanisms of action. When DNA break signature data from a platform like BreakSight’s is paired with genomic foundation models like GROVER, AlphaGenome, or Evo 2, researchers can develop a holistic understanding of how therapeutic agents interact with the cancer genome at multiple levels—from individual nucleotide contexts to large-scale genomic vulnerability patterns. The relationship is naturally complementary: foundation models predict what a sequence does and how a variant might alter its function, while DNA break-mapping provides the observable measurement of where therapy acts. One supplies hypotheses, the other supplies evidence.
This integration enables several critical advances. First, by mapping DNA break sites across the entire genome rather than just coding regions, researchers can identify previously unknown mechanisms by which DDR inhibitors and other targeted therapies exert their effects. These break signatures may reveal that a drug’s primary action involves disrupting regulatory elements, repetitive sequences, or structural features in non-coding regions—mechanisms that would be invisible to traditional gene-focused analyses, yet precisely the territory that models like AlphaGenome are now learning to interpret.
Second, AI models trained on both genomic sequence context and empirical break-site data can predict which genomic regions in a patient’s tumor are most likely to be vulnerable to specific therapeutic agents. This predictive capability could guide treatment selection by matching a patient’s unique genomic landscape with drugs most likely to induce catastrophic DNA damage in their specific cancer while sparing normal tissues. The AI can learn sequence motifs, chromatin features, and contextual patterns that correlate with therapeutic sensitivity, creating actionable biomarkers that extend well beyond simple gene mutations.
Third, quantitative break data integrated with AI analysis can refine dosing strategies and combination therapies. By understanding the dose-dependent relationship between treatment and genome-wide break patterns, and by identifying synergistic break signatures when multiple agents are combined, oncologists and developers can optimize regimens for maximum efficacy with minimal toxicity. Because efficacy is precisely where AI-discovered drugs still falter in Phase II, feeding models with empirical mechanism-of-action data of this kind targets the single greatest source of late-stage attrition in drug development.
Clinical Applications and Future Directions
In addition to preclinical discoveries, AI and machine learning have been proposed and utilized for clinical data management. Phase III trials often contain millions of data points, so harnessing the advantages of AI in this sphere carries enormous benefits. AI can further expedite IND and other regulatory document preparation, protocol generation, adverse-event prediction, and more. Regulatory agencies worldwide are increasingly supportive of these approaches, and the broader clinical-trial pipeline is beginning to adapt to keep pace with an accelerating discovery engine.12
Advanced AI can additionally generate and monitor digital twins of patients, providing real-time insights for the most personalized decision-making possible. Multi-modal data collected from tumor biopsies—including omics, digital pathology, and multiplexed immunofluorescence—can be integrated to develop nuanced digital facsimiles of individual patients. Incorporating DNA damage-sensitive sequence data into these digital twin models would add another critical dimension, enabling simulation of how different therapeutic strategies would affect a patient’s specific genomic vulnerabilities.22
While single-biomarker assays augment patient selection for different treatment regimens, integrating various types of biological data—including whole-genome sequence context, gene expression analyses, and multi-omics profiles—and using AI-driven algorithms to ascertain the best treatment option given all the identifiable features of a particular cancer has the potential to significantly improve clinical outcomes for any given patient. This comprehensive approach represents the future of precision oncology, where treatment decisions rest not on isolated genetic or protein markers but on a systems-level understanding of each tumor’s unique biology.
For these tools to reach the bedside, sequencing analysis and drug-prioritization algorithms will have to live inside clinical decision-support systems—and that, in turn, demands a level of interoperability most institutions have yet to achieve. Data, metadata, analysis software, and computing infrastructure must all speak to one another, which calls for consistent nomenclature, genomic datasets carefully annotated and linked to clinical and pathological records, and dependable approaches for sharing data across sites.9 Any platform that pairs DNA break mapping with AI analysis will face the same bar: it must slot cleanly into existing clinical workflows and electronic health record systems to deliver on its promise.
Conclusion
The convergence of AI-powered multidimensional analysis, comprehensive biological annotation, and precision therapeutic strategies represents a transformative moment in cancer treatment. The field has traveled from the first targeted therapies and the Human Genome Project to a present in which foundation models can read the regulatory grammar of DNA at single-base resolution without prior validation. By extending our analytical capabilities beyond the coding genome to encompass the full genomic landscape, and by grounding sophisticated AI models in empirical biological data, such as how drug targeting leads to specific DNA damage landscapes within different pathological contexts, we can develop a far more nuanced understanding of cancer vulnerabilities and therapeutic mechanisms.
Platforms like BreakSight, which map site-specific DNA breaks across the genome, combined with AI models like GROVER, AlphaGenome, and Evo 2, which decode the sequence context and functional meaning of genomic regions, provide complementary tools that together can reveal patterns invisible to either approach alone. This pairing enables researchers to identify novel biomarkers in non-coding regions, predict patient-specific therapeutic responses, optimize combination strategies, and refine drug mechanisms of action with a precision that directly addresses the efficacy bottleneck still limiting drug development success.
As these technologies mature and integrate into clinical practice, they promise to move precision oncology beyond simple genetic profiling toward true genome-wide, multi-omics, and mechanism-informed treatment selection. The future of cancer care lies not in treating tumors based solely on their mutational status, but in understanding and exploiting the complex interplay between signaling pathway dynamics, DNA damage vulnerabilities, and therapeutic interventions across the entire genome.
As AI models evolve toward context-aware genomic reasoning, pairing them with empirical maps of DNA damage and repair offers a path to more mechanistically grounded precision oncology in the DDR space. By bridging machine learning and molecular measurement, platforms like BreakSight may help the field navigate from descriptive biomarkers toward predictive, mechanism-based frameworks that illuminate how therapeutic pressure reshapes the cancer genome—and, in doing so, help more of tomorrow’s promising candidates survive the journey from discovery to the patients who need them.
References
- Stuart DD, et al. Precision oncology comes of age: designing best-in-class small molecules by integrating two decades of advances in chemistry, target biology, and data science. Cancer Discov. 2023;13(10):2131-2149. https://doi.org/10.1158/2159-8290.CD-23-0280
- Parums DV. Editorial: Twenty Years On from Sequencing the Human Genome, Personalized/Precision Oncology Prepares to Meet the Challenges of Checkpoint Inhibitor Therapy. Med Sci Monit. 2023;29:e940911. https://doi:10.12659/MSM.940911
- Klement GL, Arkun K, Valik D, et al. Future paradigms for precision oncology. Oncotarget. 2016;7(29):46813-46831. https://doi:10.18632/oncotarget.9488
- Sanabria M, Hirsch J, Joubert PM, Poetsch AR. DNA language model GROVER learns sequence context in the human genome. Nat Mach Intell. 2024;6:911-923. https://doi.org/10.1038/s42256-024-00872-0
- Avsec Ž, Latysheva N, Cheng J, et al. Advancing regulatory variant effect prediction with AlphaGenome. Nature. 2026;649(8099):1206-1218. https://doi:10.1038/s41586-025-10014-0
- Brixi G, Durrant MG, Ku J, et al. Genome modelling and design across all domains of life with Evo 2. Nature. 2026;652(8112):1349-1361. https://doi:10.1038/s41586-026-10176-5
- Yoo SK, Fitzgerald CW, Cho BA, et al. Prediction of checkpoint inhibitor immunotherapy efficacy for cancer using routine blood tests and clinical data. Nat Med. 2025;31:869-880. https://doi.org/10.1038/s41591-024-03398-5
- Srivastava R. Advancing precision oncology with AI-powered genomic analysis. Front Pharmacol. 2025;16:1591696. https://doi:10.3389/fphar.2025.1591696
- Dhar, S. Precision oncology: current status of targeted therapies and prospective advances. Premier Journal of Science. 2024;1:100039. https://doi.org/10.70389/ PJS.100039
- Mani DR, Maynard M, Kothadia R, et al. PANOPLY: a cloud-based platform for automated and reproducible proteogenomic data analysis. Nat Methods. 2021;18(6):580-582. doi:10.1038/s41592-021-01176-6
- Reardon B, Moore ND, Moore NS, et al. Integrating molecular profiles into clinical frameworks through the Molecular Oncology Almanac to prospectively guide precision oncology. Nat Cancer. 2021;2(10):1102-1112. doi:10.1038/s43018-021-00243-3
- Clinical trials gain intelligence. Nat Biotechnol. 2025;43:1017-1018. https://doi:10.1038/s41587-025-02754-1
- Kp Jayatunga M, Ayers M, Bruens L, Jayanth D, Meier C. How successful are AI-discovered drugs in clinical trials? A first analysis and emerging lessons. Drug Discov Today. 2024;29(6):104009. https://doi:10.1016/j.drudis.2024.104009
- Dharmasivam M, Kaya B, Akinware A, Azad MG, Richardson DR. Leading artificial intelligence-driven drug discovery platforms: 2025 landscape and global outlook. Pharmacol Rev. 2026;78(1):100102. https://doi:10.1016/j.pharmr.2025.100102
- Li Q., Qian W., Zhang Y. et al. A new wave of innovations within the DNA damage response. Sig Transduct Target Ther. 2023;8:338. https://doi.org/10.1038/s41392-023-01548-8
- Li J, Jia Z, Dong L, et al. DNA damage response in breast cancer and its significant role in guiding novel precise therapies. Biomark Res. 2024;12(1):111. https://doi:10.1186/s40364-024-00653-2
- Qian J, Liao G, Chen M, et al. Advancing cancer therapy: new frontiers in targeting DNA damage response. Front Pharmacol. 2024;15:1474337. https://doi:10.3389/fphar.2024.1474337
- Fontenot R, Biyani N, Bhatia K, Ewesuedo R, Chamberlain M, Sharma P. Clinical outcomes of DNA-damaging agents and DNA damage response inhibitors combinations in cancer: a data-driven review. Front Oncol. 2025;15:1577468. https://doi:10.3389/fonc.2025.1577468
- BreakSight, Inc. Technology and platform overview. https://www.breaksightdna.com/
- Crosetto N, Mitra A, Silva MJ, et al. Nucleotide-resolution DNA double-strand break mapping by next-generation sequencing. Nat Methods. 2013;10(4):361-365. https://doi:10.1038/nmeth.2408
- BreakSight, Inc. Drug development applications. https://www.breaksightdna.com/drug-development/
- Mao Y, Shangguan D, Huang Q, et al. Emerging artificial intelligence-driven precision therapies in tumor drug resistance: recent advances, opportunities, and challenges. Mol Cancer. 2025;24(1):123. https://doi:10.1186/s12943-025-02321-x