Why I am building PopGenLM Bench
The first public step toward a population-genetic evaluation layer for genomic language models.
An authentic record of learning to build scientific AI - grounded in evolutionary biology, shaped by the history of biological data, and tested through open, reproducible work.
This is not a collection of polished conclusions. It is a record of questions, design choices, failures, revisions, and evidence.
Biology has always been an information science. Taxonomists organised visible diversity; geneticists inferred inheritance from crosses; population geneticists turned allele frequencies into models of drift, selection, migration, and demography. The computational scale changed gradually and then dramatically. Handwritten observations became databases. Individual loci became whole genomes. Sanger sequencing gave way to high-throughput short reads, long reads, chromosome-scale assemblies, population resequencing, metagenomes, single-cell measurements, spatial assays, and continuous environmental data.
The bottleneck therefore moved. Generating data remains demanding, but modern biology is increasingly constrained by our ability to organise, integrate, evaluate, and interpret information that is too large and interconnected for manual reasoning alone. Bioinformatics emerged to manage sequence and annotation; data science added scalable statistics, visualisation, and reproducible computation; big-data infrastructure made distributed storage and processing routine. Scientific AI now adds another layer: models that can learn representations from sequence, structure, images, text, and multimodal evidence.
This progression does not make biological knowledge less important. It makes domain knowledge more consequential. A model can identify statistical regularities at a scale no individual can inspect, but it cannot by itself guarantee the correct genome build, distinguish demographic history from selection, define a meaningful negative control, or decide whether a result supports a biological claim. Those responsibilities remain scientific.
From the reading notebook · Chapter 1: What does evolution need in order to work? returns to the Darwinian minimum—inheritance, heritable variation and selection—and asks why all three are required before populations can accumulate adaptation, diverge and begin to encode information about their environments.
From the reading notebook · Chapter 3: What can a genome remember? follows one question through measurement, evolutionary change, and relationships between sites—with a pencil sketch from the margins.
From the reading notebook · Chapter 4: Can we replay evolution? follows Dallinger’s warming cultures, Lenski’s frozen E. coli and Avida to ask how selection, chance and earlier history shape what can happen next.
From the reading notebook · Chapter 5: When does more become complexity? asks what we measure when we call a genome, network or evolutionary path complex—and whether an apparent upward trend reflects selection, chance or the way we chose to measure it.
From the reading notebook · Chapter 6: Can evolution become robust to chance? asks how mutation rate, population size and epistasis can make very different genetic neighbourhoods evolutionarily robust—and why the highest-fitness genotype is not always the safest place to be.
Genomic language models are a natural meeting point between my past work and the engineering skills I am developing. They treat DNA as structured information and learn context-dependent sequence patterns that can be used to score variants. Their promise is substantial: prioritising candidate variants, identifying disrupted regulatory grammar, comparing constraint across sequence contexts, and extending annotation beyond a list of previously characterised genes.
But the scale and apparent precision of model output can create false confidence. Millions of scores can be produced before anyone notices an allele mismatch, reference incompatibility, hidden version change, unexamined calibration problem, or unsupported evolutionary interpretation. My aim with PopGenLM Bench is to make that transition from score to evidence explicit. The project connects model engineering with the population-genetic questions that determine whether a prediction is scientifically coherent.
Not a smooth software demo alone. We need reference-safe inputs, reproducible inference, independent evidence, uncertainty estimates, negative controls, sensitivity analysis, and language that does not outrun the data.
Scientific AI becomes useful when model capability and scientific validation develop together.
PyTorch, model inference, packaging, testing, APIs, deployment, and engineering patterns explained through the biological problems that required them.
Each substantial entry should connect a scientific question to code, a reproducible result, a benchmark, or a concrete software decision.
What was checked, what could fail, what the result supports, which alternatives remain, and where confidence should stop.
How an experienced computational biologist develops the engineering depth needed to design, release, and maintain reliable AI for life sciences.
The most valuable applications will not be defined by novelty alone. They will emerge where AI helps researchers interrogate complex biological systems while preserving provenance, uncertainty, and experimental accountability.
Combine sequence representations with population frequency, conservation, regulatory context, phenotype, and experimental evidence to prioritise variants transparently.
Analyse genomic change through time in natural populations, pathogens, crops, and conservation programmes while separating signal from drift and sampling noise.
Connect genomic variation with climate, traits, and demographic vulnerability to support testable decisions rather than opaque predictions.
Link sequence, structure, expression, environment, images, and literature into representations that can be evaluated across scales.
Assist with quality control, provenance, documentation, and reproducible analysis while keeping researchers responsible for assumptions and conclusions.
Build teaching systems that adapt practice and feedback to the learner without replacing approved sources, instructor judgement, or intellectual effort.
Five fortnightly notes follow my reading of Christoph Adami's The Evolution of Biological Information. They move from the history of genetic information to probability, relationships between sites, entropy, and the difficult question of what a gene can tell us about its biological context.
These are not substitutes for the chapter. They are a personal record of what I understood, where I remain cautious, and which questions I want to pursue next.
Reading notes, technical articles, build decisions, failures, and release retrospectives appear below. The newest entry is shown first.