Jump to content

Why AI Alone Will Not Solve Oligonucleotide Discovery

July 2, 2026
Justyna Lisowska, Ming Wang

The promise of AI in oligonucleotide discovery is clear: faster sequence design, smarter experiments, and more efficient paths to promising candidates — from antisense oligonucleotides and siRNA to mRNA-based medicines. But its impact still depends on the quality of the data behind it.

To truly advance oligonucleotide discovery, organizations must look beyond computational approaches alone. The focus must also shift to the quality, consistency, and completeness of the experimental data these models are trained on.

AI Brings Speed, but Not Complete Understanding

Oligonucleotide therapeutics are well-suited to AI-driven approaches. Their design involves exploring vast combinatorial possibilities — different nucleotide sequences, chemical modifications, and delivery strategies. AI helps narrow this search space by identifying promising candidates more efficiently than traditional trial-and-error methods.

This is especially valuable in rare diseases, where accelerated drug development is critical to addressing urgent unmet medical needs. By guiding experiment design, AI can significantly reduce the number of iterations required to identify viable candidates.1

However, faster predictions do not necessarily mean deeper biological insight. AI models learn from patterns in existing experimental results, but in oligonucleotide research, that data is often incomplete, inconsistent, or too diverse to capture the full complexity of biological systems.

The Data Gap in Oligonucleotide Research

Developing predictive models for oligonucleotide discovery typically begins with collecting experimental data on sequences. Each sequence is described by a set of features, and algorithms are trained to associate these features with outcomes such as efficacy or safety. 

In practice, however, building such datasets is challenging. Compared to other areas of drug discovery, oligonucleotide research still suffers from limited data availability.

To compensate, researchers frequently rely on publicly available data extracted from scientific literature and patents. While useful, these sources introduce several limitations:

  • Experimental conditions vary widely across studies, reducing comparability
  • Sequence and chemistry coverage are often narrow
  • Failed experiments are rarely reported, leading to a lack of negative examples
  • Safety-related data, including toxicity and off-target effects, is often incomplete

As a result, AI models trained on such data struggle to generalize. Reported predictive performance indicates only moderate alignment between computational predictions and experimental results, underscoring the difficulty of translating models into reliable decision support.

Data Quality: The Often-Overlooked Constraint

Even when data is available, its quality can vary significantly. In many cases, data is aggregated through automated extraction tools that interpret information from publications or databases. While this enables scaling, it also increases the risk of errors.

Inconsistent, incomplete, or incorrect metadata can all degrade model performance. These issues are particularly critical in oligonucleotide discovery, where subtle differences in sequence or chemistry can have major biological consequences.

This leads to an important conclusion: larger datasets do not automatically lead to better predictions. Without rigorous quality control, additional inputs may introduce noise rather than improve insight.

Designing Better Data, Not Just Better Models

Improving outcomes in AI-driven oligonucleotide discovery requires a different focus. Instead of prioritizing algorithmic complexity, organizations must invest in intentional data generation and curation. 

High-value datasets share several defining characteristics: 

  • They cover a broad and diverse chemical and sequence space
  • They include both successful and unsuccessful outcomes
  • They capture critical safety and specificity parameters
  • They are generated under controlled and reproducible conditions  

Achieving this level of quality typically requires internally generated data rather than reliance on external sources. Controlled experimental design ensures that variables are systematically explored and that results can be meaningfully compared. 

Systematic Screening to Improve Predictability 

Large-scale screening campaigns play a central role in building such datasets. By testing thousands of oligonucleotide candidates with carefully designed variations, researchers can generate dense and structured data that reveals meaningful relationships between sequence, chemistry, and biological response. 

These datasets are essential for training AI models that go beyond superficial correlations and instead capture deeper mechanistic insights. They also provide the foundation for improved predictability in key areas such as efficacy, safety, and off-target effects. 

In this sense, experimental design becomes a critical enabler of AI performance. Without well-constructed screening strategies, even advanced models remain constrained by insufficient input data. 

Designing Better Data, Not Just Better Models

Connecting Data Across the Discovery Workflow 

Generating results is only the first step. How those results are managed and integrated also determines their usability. 

In many organizations, experimental data is distributed across multiple systems and teams, often stored in incompatible formats. This fragmentation limits the ability to reuse data and reduces its value for AI applications. 

To overcome this, organizations must adopt approaches that: 

  • Enable seamless access across workflows and teams
  • Standardize data formats and annotations
  • Preserve context, including experimental conditions and metadata 

When data is accessible, standardized, and context-rich, AI systems can use it more effectively. 

Reframing the Role of AI 

The key takeaway is not that AI is insufficient, but that its role must be clearly understood. 

AI is a powerful tool for identifying patterns and guiding decisions, but it does not replace the need for high-quality experimental science. Its performance is inherently tied to the data it learns from and the context in which it operates. 

 In oligonucleotide discovery, its greatest value comes when computational insight is paired with a reliable experimental context. 

Toward a Data-Centric Future 

Oligonucleotide therapeutics are opening new possibilities in precision medicine, and AI will remain a critical enabler of this progress. 

But meaningful impact will come from more than better algorithms. It will depend on the quality of the experimental data behind them. Organizations that invest in high-quality experimental datasets, structured data environments, and integrated workflows will be better positioned to turn AI from a promising tool to a reliable engine for oligonucleotide R&D. 

                                                
                                                                                 Simplify Your Oligonucleotide R&D

FAQs

AI can accelerate oligonucleotide sequence design and help identify promising candidates, but its predictions are only as reliable as the experimental data behind them. Incomplete, inconsistent, or poorly annotated data can limit model performance and reduce confidence in AI-driven decisions.

AI models benefit from high-quality datasets that cover diverse sequences and chemistries, include both successful and unsuccessful experimental outcomes, capture safety and off-target effects, and are generated under controlled, reproducible conditions.

Genedata helps biopharma organizations build the structured, annotated, and traceable data foundation needed for reliable AI. By connecting experimental workflows, standardizing data capture, and preserving scientific context, Genedata enables teams to turn complex oligonucleotide data into AI-ready datasets that support more confident discovery decisions.