Data mining for aptamer discovery
| Title: |
Data mining for aptamer discovery |
| DNr: |
Berzelius-2026-224 |
| Project Type: |
LiU Berzelius |
| Principal Investigator: |
Erik Benson <erik.benson@ki.se> |
| Affiliation: |
Karolinska Institutet |
| Duration: |
2026-08-18 – 2027-03-01 |
| Classification: |
10203 |
| Keywords: |
|
Abstract
Aptamers are short, single-stranded DNA or RNA molecules that fold into three-dimensional structures capable of binding specific molecular targets with high affinity and selectivity. Their chemical stability, relatively low production cost, ease of modification, and potential for reproducible synthesis make them attractive alternatives to antibodies in diagnostics, therapeutics, biosensing, drug delivery, and environmental monitoring. However, conventional aptamer discovery remains labor-intensive and is commonly based on iterative experimental selection procedures, such as systematic evolution of ligands by exponential enrichment. Improving the reuse of previously generated aptamer data could reduce experimental costs, accelerate candidate identification, and expand the range of targets for which functional aptamers can be developed.
The aptamer field has already accumulated substantial knowledge in the form of sequence–target pairs, experimentally measured binding properties, structural information, selection conditions, and application-specific annotations. Several databases published before 2023 organize portions of these data, while a rapidly growing body of literature published since 2023 contains additional aptamer sequences and target associations. Nevertheless, much of the newer information remains embedded in unstructured publication text, tables, figures, supplementary files, and inconsistent sequence formats. As a result, current resources are incomplete, difficult to update, and insufficiently standardized for large-scale computational analysis. Valuable information is therefore fragmented across databases and publications, limiting its reuse for systematic characterization and predictive modeling.
The goal of the Data Mining for Aptamer Discovery project is to develop an integrated computational platform that automatically extracts, harmonizes, and models aptamer sequence–target knowledge. The project will first train a literature-mining model using existing publications and curated sequence–target pairs from established aptamer databases. The model will identify aptamer sequences, molecular targets, binding measurements, experimental context, and relationships between these entities in newly published articles. Extracted records will be standardized, quality-controlled, assigned confidence scores, and integrated with existing datasets to create an updatable representation of the field.
The resulting dataset will then support the development of predictive models for aptamer function. Sequence-derived, structural, physicochemical, motif-based, and target-related features will be evaluated through feature-selection procedures to create optimized data embeddings. These embeddings will be used to characterize known aptamer families, quantify similarities between sequences, and train models that estimate whether a previously unseen sequence is likely to exhibit aptamer-like binding behavior and which target classes or specific targets it may recognize. The system will also report confidence levels and identify the nearest experimentally validated sequences.
The final output will be a searchable database and user-facing analysis platform. Users will be able to query validated aptamers by target, sequence, experimental property, or publication; submit new sequences to retrieve nearest neighbors; and obtain model-based predictions of potential binding targets and aptamer suitability. By transforming fragmented literature into structured, continuously updated, and computationally actionable knowledge, this project is expected to improve data accessibility, reduce duplication of experimental effort, support evidence-based candidate prioritization, and accelerate the discovery of aptamers for biomedical, biotechnological, and environmental applications.