More attention to what makes each person different.
Learn more
For IIMI, a personalised approach begins with the differences between people. Medical experience helps keep technology connected to the individual.
At the IIMI Institute, artificial intelligence is part of everyday work. The institute combines many years of medical experience with the development of its own technologies.
Transition
Layer 1Sequence
Layer 2Cellular context
Layer 3Data structure
Depth in connections
What turns data into understanding?
Data about a person arrives in layers: hereditary factors, laboratory measurements, a history of observations. On its own, each layer says little.
More understanding.
Understanding emerges when the layers are cross-referenced: hereditary factors alongside measurements, measurements alongside history. This is how the IIMI Institute approaches data: not source by source, but through the connections between them.
R06 · Synthetic enhancers, Drosophila · Nature, 2023. Open
Scientific context
The change is already visible.
AI helps to read DNA, cross-reference data and test new ideas. Studies show how the field’s capabilities are expanding.
Scale
Models
Research
Horizons
01 / 09
Measured data · 2002–2025
DNA data keeps growing.
Open archives show the scale of the genetic data that modern biology works with.
GenBank traditionalWGS
GenBank: number of nucleotide bases in two parts of the archive. December snapshots; logarithmic scale. Data retrieved on 11 Sep 2026.
The archive holds sequences from many organisms, not the genomes of individual people.
Source and notes
Traditional GenBank and WGS (whole-genome shotgun) records are shown separately. Archive size at a given release is not the number of new studies in that year. December was chosen for a consistent annual series; the August 2026 snapshot is available separately.
Biology is more than sequence. Three-dimensional structures help researchers study how molecules are built.
20159,242
201610,801
201711,056
201811,157
201911,467
202013,976
202112,567
202214,247
202314,445
202415,279
202517,566
010,00020,000
entries released per year
Structure entries publicly released by the wwPDB each year. Full years; table as of 1 September 2026.
These are archive entries, not unique proteins. The chart does not say why the archive is growing.
Source and notes
The source is the ‘Number of Structures Released per year’ table, not the deposition statistics. The partial year 2026 is excluded. Historical values may be revised when the status of entries changes.
AlphaGenome links DNA sequence to predictions of gene activity and other molecular processes.
InputDNA segment · about 1 million base pairs
ModelResearch model
TranscriptionRNA-seq · CAGE · PRO-cap
SplicingSites · usage · junctions
Accessibility and regulatory marksDNase · ATAC · histone · TF
Spatial contactsChromatin contact maps
About 1 million base pairs in; 11 signal types out, shown in the diagram as four groups.Source and notes
Groups: transcription, splicing, accessibility and regulatory marks, spatial contacts. Resolution and evaluation tasks differ between outputs. The illustration does not show a real sequence or an IIMI result.
The GET model combines information about DNA with measurements in the cell to predict gene activity.
Input 1DNA motifs
Input 2 · measurementAccessibility of DNA regions in the cell
ModelGET
OutputPredicted gene activity
SeparatelyComparison with measured gene activity
Two inputs: DNA motifs and the accessibility of DNA regions measured in the cell (chromatin accessibility). Output: a prediction of gene activity.
The model looks at more than DNA: measurements from the cell itself are supplied separately.
Source and notes
The paper describes 213 cell types for pre-training and 153 cell types with paired data for later stages. These are neither patient numbers nor a performance percentage. What matters in the diagram is the structure of the inputs, not a contest between models.
The study compared how two versions of an AI model extract information from laboratory documents.
MedGemma 1 4BMedGemma 1.5 4B
EHR Dataset 2 · private
7891
EHR Dataset 3 · private
5071
EHR Dataset 4 · private, synthetic
2564
Mendeley Clinical Laboratory Test Reports · public
85
050100
Macro-F1, scale 0–100
Field extraction from laboratory documents. Macro-F1 on a 0–100 scale. The table used reports no confidence intervals.
Three datasets are private, one of them synthetic; the fourth is public.
Source and notes
The published score pairs for the four datasets are 78→91, 50→71, 25→64 and 85→85. The two versions can be compared within each dataset; averaging the four datasets into an overall ‘AI accuracy’ is not valid.
DNA segments designed with a model were tested in Drosophila. The researchers measured whether the intended activity appeared in the selected cells.
Kenyon cells
10 of 13
Perineurial glia
4 of 6
a filled symbol means the intended activity was detected; an open one means it was not
The target activity was detected in 10 of 13 selected constructs for Kenyon cells and in 4 of 6 for perineurial glia. These are two separate series from one study.
Sample sizes are small, and this is an experiment in Drosophila, not in humans.
Source and notes
Three constructs without the target GFP signal in the first series also lacked the cell marker, so the open symbols cannot automatically be labelled ‘AI errors’.
R06 · Cell-type-directed design of synthetic enhancers · Nature · 12 Dec 2023. View source
View data
Series
Detected
Tested
Kenyon cells
10
13
Perineurial glia
4
6
Research agent · Science, 2026
AI helps to carry out research.
Research agents can plan tasks, use tools and run code. Biomni is one example of this approach.
01Question
02Plan
03Tools
04Code / analysis
05Verifiable result
Checks feed back into the plan
Conceptual diagram: question → plan → tools and code → verifiable result.Source and notes
Based on the description in the final publication. Speed, success rates and other figures are not quoted here: they would need to be checked against the full text rather than carried over from an early preprint.
One expert scenario: models help to study cells, and measurements make it possible to test the models.
2035
Scenario
Model and experiment converge: automation of research steps, interpretation of new measurements, testable biological models.
Data qualityVerifiabilityInteroperability
A possible direction: more automated research and testable models of biology.Source and notes
Conditions for progress: high-quality data, reproducible experiments, interoperable tools and accessible research infrastructure.
R08–R10 · Sanger Genomics Futures Series · Wellcome Sanger Institute. View source
Expert scenario · 2050 horizon
More complex digital models of biology.
Expert discussions consider more complex models and wider access to research technologies.
2050
Scenario
Modelling complex systems, interaction between AI and researchers, international access to research tools. The structure is left open: this is a horizon, not a measure of how much precision is still missing.
ValidationSafetyAccess
Future possibilities: modelling complex systems, interaction between AI and researchers, international access to tools.Source and notes
Questions of data quality, model validation, safety, trust and equal access remain open. These conditions cannot be reduced to a single rising curve.
R08–R10 · Wellcome Sanger Institute, 2025 horizons. View source
Not just more data. More understanding.
The IIMI Institute approach
For IIMI, what matters is the connection between medical experience, data and the individual person.
The IIMI Institute is headquartered in Munich. The institute is expanding its international connections in medicine, genetics and artificial intelligence.
BaseMunich, Germany
ApproachMedical experience × AI
OutlookInternational cooperation
The conversation starts here
Start with a question.
Share what interests you. A conversation will help to clarify which possibilities are worth discussing.