IIMI InstituteMünchen

DNA × AI

Medicine starts with data.

Genetics, artificial intelligence and personalised medicine. The IIMI Institute brings these fields together with many years of medical experience.

Medical experienceAI at the coreMunich, Germany

Built on understanding

From sequence to meaning.

DNA

The biological basis of individuality.

Learn more

DNA stores genetic information. To understand how it is expressed, researchers take into account cellular activity and other biological data.

Artificial intelligence

Finding connections in complex data.

Learn more

AI helps to cross-reference data, spot patterns and propose ideas for testing. Its capabilities are evaluated on specific tasks.

Personalised medicine

More attention to what makes each person different.

Learn more

For IIMI, a personalised approach begins with the differences between people. Medical experience helps keep technology connected to the individual.

At the IIMI Institute, artificial intelligence is part of everyday work. The institute combines many years of medical experience with the development of its own technologies.

Transition

Layer 1Sequence
Layer 2Cellular context
Layer 3Data structure

Depth in connections

What turns data into understanding?

Data about a person arrives in layers: hereditary factors, laboratory measurements, a history of observations. On its own, each layer says little.

Three translucent curved layers joined into one three-dimensional object: an image of data being brought together.

More understanding.

Understanding emerges when the layers are cross-referenced: hereditary factors alongside measurements, measurements alongside history. This is how the IIMI Institute approaches data: not source by source, but through the connections between them.

Related research

R03 · AlphaGenome · Nature, 2026. Open

R04 · GET · Nature, 2025. Open

R06 · Synthetic enhancers, Drosophila · Nature, 2023. Open

Scientific context

The change is already visible.

AI helps to read DNA, cross-reference data and test new ideas. Studies show how the field’s capabilities are expanding.

  1. Scale
  2. Models
  3. Research
  4. Horizons
01 / 09
Measured data · 2002–2025

DNA data keeps growing.

Open archives show the scale of the genetic data that modern biology works with.

GenBank traditionalWGS
10⁹10¹⁰10¹¹10¹²10¹³10¹⁴2002201020182025nucleotide bases · log
GenBank: number of nucleotide bases in two parts of the archive. December snapshots; logarithmic scale. Data retrieved on 11 Sep 2026.

The archive holds sequences from many organisms, not the genomes of individual people.

Source and notes

Traditional GenBank and WGS (whole-genome shotgun) records are shown separately. Archive size at a given release is not the number of new studies in that year. December was chosen for a consistent annual series; the August 2026 snapshot is available separately.

R01 · GenBank and WGS Statistics · NCBI / NLM / NIH. View source

View data
SnapshottraditionalWGS
Dec 200228,507,990,1666,702,372,564
Dec 200336,553,368,48514,523,454,868
Dec 200444,575,745,17635,009,256,228
Dec 200556,037,734,46259,638,900,034
Dec 200669,019,290,70581,611,376,856
Dec 200783,874,179,730106,505,691,578
Dec 200899,116,431,942141,374,971,004
Dec 2009110,118,557,163158,317,168,385
Dec 2010122,082,812,719177,385,297,156
Dec 2011135,117,731,375239,868,309,609
Dec 2012148,390,863,904356,002,922,838
Dec 2013156,230,531,562556,764,321,498
Dec 2014184,938,063,614848,977,922,022
Dec 2015203,939,111,0711,297,865,618,365
Dec 2016224,973,060,4331,817,189,565,845
Dec 2017249,722,163,5942,466,098,053,327
Dec 2018285,688,542,1863,656,719,423,096
Dec 2019388,417,258,0096,277,551,200,690
Dec 2020723,003,822,00711,830,842,428,018
Dec 20211,053,275,115,03014,922,033,922,302
Dec 20221,635,594,138,49319,086,596,616,569
Dec 20232,570,711,588,04424,651,580,464,335
Dec 20245,085,904,976,33832,983,029,087,303
Dec 20256,651,459,875,40842,125,323,988,215
Aug 20268,236,878,868,45050,829,714,144,609
Measured data · 2015–2025

Science studies the shape of molecules.

Biology is more than sequence. Three-dimensional structures help researchers study how molecules are built.

20159,242
201610,801
201711,056
201811,157
201911,467
202013,976
202112,567
202214,247
202314,445
202415,279
202517,566
010,00020,000
entries released per year
Structure entries publicly released by the wwPDB each year. Full years; table as of 1 September 2026.

These are archive entries, not unique proteins. The chart does not say why the archive is growing.

Source and notes

The source is the ‘Number of Structures Released per year’ table, not the deposition statistics. The partial year 2026 is excluded. Historical values may be revised when the status of entries changes.

R02 · wwPDB · as of 1 Sep 2026. View source

View data
YearEntries released
20159,242
201610,801
201711,056
201811,157
201911,467
202013,976
202112,567
202214,247
202314,445
202415,279
202517,566
Research model · Nature, 2026

AI helps to read DNA more deeply.

AlphaGenome links DNA sequence to predictions of gene activity and other molecular processes.

InputDNA segment · about 1 million base pairs
ModelResearch model
TranscriptionRNA-seq · CAGE · PRO-cap
SplicingSites · usage · junctions
Accessibility and regulatory marksDNase · ATAC · histone · TF
Spatial contactsChromatin contact maps
About 1 million base pairs in; 11 signal types out, shown in the diagram as four groups.
Source and notes

Groups: transcription, splicing, accessibility and regulatory marks, spatial contacts. Resolution and evaluation tasks differ between outputs. The illustration does not show a real sequence or an IIMI result.

R03 · AlphaGenome · Nature · 28 Jan 2026. View source

Research model · Nature, 2025

What DNA means depends on the cell.

The GET model combines information about DNA with measurements in the cell to predict gene activity.

Input 1DNA motifs
Input 2 · measurementAccessibility of DNA regions in the cell
ModelGET
OutputPredicted gene activity
SeparatelyComparison with measured gene activity
Two inputs: DNA motifs and the accessibility of DNA regions measured in the cell (chromatin accessibility). Output: a prediction of gene activity.

The model looks at more than DNA: measurements from the cell itself are supplied separately.

Source and notes

The paper describes 213 cell types for pre-training and 153 cell types with paired data for later stages. These are neither patient numbers nor a performance percentage. What matters in the diagram is the structure of the inputs, not a contest between models.

R04 · GET · Nature · 8 Jan 2025. View source

Developers’ benchmark · technical preprint, 2026

Documents become data.

The study compared how two versions of an AI model extract information from laboratory documents.

MedGemma 1 4BMedGemma 1.5 4B
EHR Dataset 2 · private
7891
EHR Dataset 3 · private
5071
EHR Dataset 4 · private, synthetic
2564
Mendeley Clinical Laboratory Test Reports · public
85
050100
Macro-F1, scale 0–100
Field extraction from laboratory documents. Macro-F1 on a 0–100 scale. The table used reports no confidence intervals.

Three datasets are private, one of them synthetic; the fourth is public.

Source and notes

The published score pairs for the four datasets are 78→91, 50→71, 25→64 and 85→85. The two versions can be compared within each dataset; averaging the four datasets into an overall ‘AI accuracy’ is not valid.

R05 · MedGemma 1.5 Technical Report · arXiv v2 · 1 May 2026. View source

View data
Dataset1 4B1.5 4BAccess
EHR Dataset 27891private
EHR Dataset 35071private
EHR Dataset 42564private, synthetic
Mendeley Clinical Laboratory Test Reports8585public
Drosophila experiment · Nature, 2023/2024

Ideas from AI are tested by experiment.

DNA segments designed with a model were tested in Drosophila. The researchers measured whether the intended activity appeared in the selected cells.

Kenyon cells
10 of 13
Perineurial glia
4 of 6
a filled symbol means the intended activity was detected; an open one means it was not
The target activity was detected in 10 of 13 selected constructs for Kenyon cells and in 4 of 6 for perineurial glia. These are two separate series from one study.

Sample sizes are small, and this is an experiment in Drosophila, not in humans.

Source and notes

Three constructs without the target GFP signal in the first series also lacked the cell marker, so the open symbols cannot automatically be labelled ‘AI errors’.

R06 · Cell-type-directed design of synthetic enhancers · Nature · 12 Dec 2023. View source

View data
SeriesDetectedTested
Kenyon cells1013
Perineurial glia46
Research agent · Science, 2026

AI helps to carry out research.

Research agents can plan tasks, use tools and run code. Biomni is one example of this approach.

01Question
02Plan
03Tools
04Code / analysis
05Verifiable result
Checks feed back into the plan
Conceptual diagram: question → plan → tools and code → verifiable result.
Source and notes

Based on the description in the final publication. Speed, success rates and other figures are not quoted here: they would need to be checked against the full text rather than carried over from an early preprint.

R07 · Biomni · Science · 20 Aug 2026. View source

Expert scenario · 2035 horizon

AI and experiment work more closely together.

One expert scenario: models help to study cells, and measurements make it possible to test the models.

2035
Scenario

Model and experiment converge: automation of research steps, interpretation of new measurements, testable biological models.

Data qualityVerifiabilityInteroperability
A possible direction: more automated research and testable models of biology.
Source and notes

Conditions for progress: high-quality data, reproducible experiments, interoperable tools and accessible research infrastructure.

R08–R10 · Sanger Genomics Futures Series · Wellcome Sanger Institute. View source

Expert scenario · 2050 horizon

More complex digital models of biology.

Expert discussions consider more complex models and wider access to research technologies.

2050
Scenario

Modelling complex systems, interaction between AI and researchers, international access to research tools. The structure is left open: this is a horizon, not a measure of how much precision is still missing.

ValidationSafetyAccess
Future possibilities: modelling complex systems, interaction between AI and researchers, international access to tools.
Source and notes

Questions of data quality, model validation, safety, trust and equal access remain open. These conditions cannot be reduced to a single rising curve.

R08–R10 · Wellcome Sanger Institute, 2025 horizons. View source

Not just more data. More understanding.
The IIMI Institute approach

For IIMI, what matters is the connection between medical experience, data and the individual person.

International focus

Based in Germany. International in focus.

The IIMI Institute is headquartered in Munich. The institute is expanding its international connections in medicine, genetics and artificial intelligence.

BaseMunich, Germany
ApproachMedical experience × AI
OutlookInternational cooperation

The conversation starts here

Start with a question.

Share what interests you. A conversation will help to clarify which possibilities are worth discussing.

Nikolai Savtchenko
Nikolai Savtchenko

Director, IIMI Institute