<?xml version="1.0" encoding="UTF-8"?>
<rss  xmlns:atom="http://www.w3.org/2005/Atom" 
      xmlns:media="http://search.yahoo.com/mrss/" 
      xmlns:content="http://purl.org/rss/1.0/modules/content/" 
      xmlns:dc="http://purl.org/dc/elements/1.1/" 
      version="2.0">
<channel>
<title>Dr. Tahir Ali</title>
<link>https://tahirali-biomics.github.io/genomes-to-ai.html</link>
<atom:link href="https://tahirali-biomics.github.io/genomes-to-ai.xml" rel="self" type="application/rss+xml"/>
<description>Genomes to AI - Dr Tahir Ali&#39;s working journal of becoming a scientific AI engineer.</description>
<image>
<url>https://tahirali-biomics.github.io/assets/writings/genomes-to-ai-journey.png</url>
<title>Dr. Tahir Ali</title>
<link>https://tahirali-biomics.github.io/genomes-to-ai.html</link>
<height>75</height>
<width>144</width>
</image>
<generator>quarto-1.8.24</generator>
<lastBuildDate>Sun, 30 Aug 2026 22:00:00 GMT</lastBuildDate>
<item>
  <title>Why I am building PopGenLM Bench</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-08-31-popgenlm-bench/</link>
  <description><![CDATA[ 




<p>Genomic language models can score millions of variants, but a completed model run does not by itself establish biological validity. Before interpreting a score, I want to know whether the alleles matched the intended reference, whether model and software versions were recorded, whether sequence orientation was treated consistently, and whether the output behaves sensibly against independent evolutionary evidence.</p>
<p>That is the reason for <strong>PopGenLM Bench</strong>, my first open-source scientific AI engineering project.</p>
<section id="information-about-what" class="level2">
<h2 class="anchored" data-anchor-id="information-about-what">Information about what?</h2>
<p>Reading Chapter 2 of Christoph Adami’s <em>The Evolution of Biological Information</em> sharpened the question behind this project. In information theory, a sequence or score is not informative in isolation. Information is a relationship: knowing one variable must reduce uncertainty about a specified target across an appropriate ensemble.</p>
<p>For a genomic language-model score, I therefore need to ask: information about what? A score may reflect sequence plausibility, evolutionary conservation, regulatory context, population frequency or a technical property of the training data. These possibilities are related, but they are not equivalent. Producing a confident number does not establish which relationship has been learned.</p>
<p>This is why PopGenLM Bench is designed as an evaluation layer rather than a score viewer. Its task is to state the target, define the comparison set, measure whether the model adds predictive value beyond simpler baselines, and expose where that relationship fails.</p>
</section>
<section id="the-first-deliberately-small-milestone" class="level2">
<h2 class="anchored" data-anchor-id="the-first-deliberately-small-milestone">The first deliberately small milestone</h2>
<p>Version 0.1 does one thing carefully: it scores exactly 100 deterministic SNVs with the published GPN Brassicales model. The variants occupy a safe internal section of the public <em>Arabidopsis thaliana</em> region used in GPN’s official example, ensuring complete sequence context for every site.</p>
<p>The run uses pinned upstream code and model revisions, validates every REF allele against the FASTA, invokes GPN’s maintained variant-effect command, and records input and output hashes, package versions, device, runtime, and the exact command.</p>
<div class="code-copy-outer-scaffold"><div class="sourceCode" id="cb1" style="background: #f1f3f5;"><pre class="sourceCode bash code-with-copy"><code class="sourceCode bash"><span id="cb1-1"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">uv</span> python install 3.13</span>
<span id="cb1-2"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">uv</span> sync <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--extra</span> gpn <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--group</span> dev</span>
<span id="cb1-3"><span class="ex" style="color: null;
background-color: null;
font-style: inherit;">uv</span> run popgenlm demo <span class="at" style="color: #657422;
background-color: null;
font-style: inherit;">--project-root</span> .</span></code></pre></div></div>
</section>
<section id="verified-result" class="level2">
<h2 class="anchored" data-anchor-id="verified-result">Verified result</h2>
<p>The model returned all <strong>100 expected scores</strong>. The median alternate-minus-reference score was <strong>−1.0309</strong>, with <strong>77%</strong> of the deterministic alternate alleles scoring below their reference allele. The complete CPU run took <strong>80.7 seconds</strong>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/popgenlm/benchmark-overview.png" class="img-fluid figure-img" alt="Four-panel summary of GPN scores showing the score distribution, positional profile, ranked profile, and bootstrap interval for the median."></p>
<figcaption>Verified PopGenLM benchmark overview</figcaption>
</figure>
</div>
</section>
<section id="what-this-result-does-not-mean" class="level2">
<h2 class="anchored" data-anchor-id="what-this-result-does-not-mean">What this result does not mean</h2>
<p>The distribution cannot be used to infer natural selection or biological information content. These variants were generated deterministically from one public sequence region; they were not sampled from a natural population and do not form an evolutionary ensemble suitable for that claim.</p>
<p>This distinction is central to the project. A technically successful run demonstrates that the inference and validation machinery works. Biological support requires an independent benchmark with observed population variants, allele-frequency information, uncertainty estimates, and appropriate null comparisons.</p>
</section>
<section id="what-comes-next" class="level2">
<h2 class="anchored" data-anchor-id="what-comes-next">What comes next</h2>
<p>The next milestone will use a documented public <em>A. thaliana</em> population dataset. It will test whether GPN scores reduce uncertainty about allele-frequency classes beyond appropriate sequence baselines, using effect sizes, block-bootstrap confidence intervals, permutation tests and sensitivity analyses. Later milestones will incorporate conservation scores, genomic annotation classes and controls for genomic and phylogenetic dependence.</p>
<p>The guiding principle is now clearer to me: a model output becomes scientific evidence only when its predictive relationship to biology is explicit, reproducible and tested outside the model that produced it.</p>
<p>The evolving project overview, verified outputs, and roadmap are available on the <a href="../../popgenlm.html">PopGenLM Bench page</a>. The complete statistical workflow can also be rerun in the <a href="https://colab.research.google.com/github/tahirali-biomics/tahirali-biomics.github.io/blob/main/notebooks/popgenlm-bench-demo.ipynb">verified Colab notebook</a>. The public package repository will become active with the first reviewed release.</p>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Genomic language models</category>
  <category>Population genomics</category>
  <category>Open source</category>
  <category>Build note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-08-31-popgenlm-bench/</guid>
  <pubDate>Sun, 30 Aug 2026 22:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/popgenlm/benchmark-overview.png" medium="image" type="image/png" height="92" width="144"/>
</item>
<item>
  <title>Can evolution become robust to chance?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-04-18-can-evolution-become-robust-to-chance/</link>
  <description><![CDATA[ 




<p><em>Reading Christoph Adami, <strong>The Evolution of Biological Information</strong>, Chapter 6.</em></p>
<p><img src="https://tahirali-biomics.github.io/assets/writings/09-can-evolution-become-robust-to-chance.webp" class="img-fluid" alt="Graphite notebook sketch of neutral networks, fitness peaks, mutation, genetic drift and kinetoplast RNA editing, surrounded by handwritten questions."></p>
<p>I thought I knew what <em>robustness</em> meant before starting this chapter.</p>
<p>Something changes. The organism absorbs the disturbance. Life goes on.</p>
<p>Simple enough.</p>
<p>But fairly early in Chapter 6 I realised that Adami was asking something much stranger.</p>
<p>What if evolution can change not only the organism, but also <strong>what happens when the organism is changed</strong>?</p>
<p>I kept returning to that thought.</p>
<p>A genome is constantly being copied. Mutations appear. Alleles are lost by chance. Population size changes how efficiently selection can distinguish one effect from another.</p>
<p>So perhaps surviving well is not only about occupying a good genotype.</p>
<p>Perhaps it is also about living in a <strong>good neighbourhood of possible genotypes</strong>.</p>
<p>That was the point where this chapter opened up for me.</p>
<section id="neutral-but-neutral-in-what-sense" class="level2">
<h2 class="anchored" data-anchor-id="neutral-but-neutral-in-what-sense">Neutral — but neutral in what sense?</h2>
<p>The chapter begins with Kimura and neutral evolution.</p>
<p>The result itself is one I already knew from population genetics: a new neutral mutation has a fixation probability of roughly (1/N), while roughly (N) neutral mutations arise, so the population-level substitution rate becomes approximately ().</p>
<p>Population size disappears from the final expression.</p>
<p>Beautiful.</p>
<p>But this time another question bothered me more:</p>
<p><strong>If a mutation does not change fitness, does that really mean it does nothing?</strong></p>
<p>My instinctive answer had always been: for selection, essentially yes.</p>
<p>Then I realised how much is hidden inside the words <em>does not change fitness</em>.</p>
<p>We have measured the organism carrying the mutation.</p>
<p>We have not yet asked what happens to its descendants when <strong>another mutation</strong> occurs.</p>
<p>That is a very different question.</p>
<p>A change can therefore be neutral with respect to the organism standing in front of us and still alter the consequences of mutations that have not happened yet.</p>
<p>I wrote beside this:</p>
<blockquote class="blockquote">
<p><strong>neutral now ≠ irrelevant later</strong></p>
</blockquote>
<p>And suddenly neutrality no longer looked like empty evolutionary space.</p>
<p>It looked like hidden architecture.</p>
</section>
<section id="i-had-been-drawing-genomes-as-points" class="level2">
<h2 class="anchored" data-anchor-id="i-had-been-drawing-genomes-as-points">I had been drawing genomes as points</h2>
<p>This was probably my strongest shift while reading the chapter.</p>
<p>We routinely draw fitness landscapes as peaks and valleys and place a genotype somewhere on them.</p>
<p>One genotype. One point. One fitness.</p>
<p>But a replicating genome never really exists as an isolated point.</p>
<p>Every replication creates the possibility of moving somewhere nearby.</p>
<p>One nucleotide away.</p>
<p>One mutation away.</p>
<p>Then another.</p>
<p>So I started thinking less about the height of the point and more about the shape of the ground around it.</p>
<p>Imagine two genotypes with almost the same present fitness.</p>
<p>Around the first, most single mutations are damaging.</p>
<p>Around the second, many mutations produce sequences that still work.</p>
<p>If I only measure the two organisms today, they may appear equally successful.</p>
<p>But their futures are not equivalent.</p>
<p>One stands on brittle ground.</p>
<p>The other stands on something more forgiving.</p>
<p>That felt like an important change in perspective:</p>
<p><strong>perhaps selection sometimes acts on the neighbourhood, not merely the point.</strong></p>
</section>
<section id="then-came-the-phrase-that-sounds-almost-wrong" class="level2">
<h2 class="anchored" data-anchor-id="then-came-the-phrase-that-sounds-almost-wrong">Then came the phrase that sounds almost wrong</h2>
<p><strong>Survival of the flattest.</strong></p>
<p>The first time you encounter it, it almost feels like a contradiction inserted deliberately into evolutionary biology.</p>
<p>Surely the highest fitness peak should win.</p>
<p>That is practically the picture we teach.</p>
<p>But now imagine that the peak is very narrow.</p>
<p>Its summit is excellent, but almost every mutational step away from it is disastrous.</p>
<p>Nearby is another peak.</p>
<p>Lower.</p>
<p>Less impressive.</p>
<p>But broad.</p>
<p>Mutations around it usually still leave functioning descendants.</p>
<p>At a sufficiently high mutation rate, I can no longer judge these populations only by the fitness of their best genotype.</p>
<p>The population constantly leaks into neighbouring sequence space.</p>
<p>And once I saw it that way, the contradiction disappeared.</p>
<p>The lower peak can win because evolution is not comparing two isolated sequences.</p>
<p>It is comparing two <strong>mutating populations</strong>.</p>
<p>That was one of those moments where a familiar diagram suddenly means something different.</p>
<p>I could almost hear myself objecting:</p>
<p><em>But the other genotype is fitter.</em></p>
<p>And the answer coming back:</p>
<p><strong>Fitter as what? A sequence—or as a lineage that must repeatedly reproduce under mutation?</strong></p>
<p>That distinction stayed with me.</p>
</section>
<section id="robustness-does-not-even-mean-making-every-mutation-harmless" class="level2">
<h2 class="anchored" data-anchor-id="robustness-does-not-even-mean-making-every-mutation-harmless">Robustness does not even mean making every mutation harmless</h2>
<p>Then the chapter did something I did not expect.</p>
<p>My intuition was that mutational robustness should always mean turning damaging mutations into milder ones.</p>
<p>Bad mutation → less bad mutation → neutral mutation.</p>
<p>That seems obvious.</p>
<p>Except it is not the only possibility.</p>
<p>A mildly deleterious mutation can survive.</p>
<p>It can reproduce.</p>
<p>It can become somebody else’s starting point.</p>
<p>A lethal mutation cannot.</p>
<p>So under some circumstances, if a deleterious change cannot be made harmless, making it <strong>lethal rather than weakly damaging</strong> can actually protect the lineage from carrying that damage forward.</p>
<p>I stopped at this because it felt almost backwards.</p>
<p>Robustness through lethality?</p>
<p>Yet the logic is remarkably clean.</p>
<p>The important question is not simply:</p>
<p><strong>How severe is the mutation?</strong></p>
<p>It is also:</p>
<p><strong>Can its consequences propagate?</strong></p>
<p>That changed what the word <em>robustness</em> meant for me.</p>
<p>Robustness is not necessarily gentleness.</p>
<p>Sometimes it is containment.</p>
</section>
<section id="and-then-population-size-changed-the-answer-again" class="level2">
<h2 class="anchored" data-anchor-id="and-then-population-size-changed-the-answer-again">And then population size changed the answer again</h2>
<p>At this point I thought I had the chapter’s logic.</p>
<p>High mutation pressure can favour flatter regions of a fitness landscape.</p>
<p>Good.</p>
<p>Then came drift robustness.</p>
<p>And my neat picture broke again.</p>
<p>In a small population, selection loses resolution.</p>
<p>A mutation with a tiny deleterious effect may technically reduce fitness, but if that effect is smaller than the noise created by genetic drift, selection may be unable to remove it reliably.</p>
<p>Such mutations can become effectively neutral.</p>
<p>One can fix.</p>
<p>Then another.</p>
<p>Then another.</p>
<p>The problem is no longer simply mutation.</p>
<p>It is the gradual, stochastic loss of information that selection is too weak to prevent.</p>
<p>I found myself asking:</p>
<p><strong>If flatness protects against mutation, shouldn’t flatness also protect a small population?</strong></p>
<p>No.</p>
<p>And this was probably my favourite conceptual reversal in the chapter.</p>
<p>For drift, a shallow fitness effect can be precisely the problem.</p>
<p>If a mutation is only slightly harmful, drift can carry it across the population.</p>
<p>But if epistasis makes that mutation strongly deleterious, selection can suddenly see it again.</p>
<p>So protection from drift can require something closer to a <strong>steeper</strong> local landscape.</p>
<p>I drew the two possibilities beside each other:</p>
<p><strong>high mutation rate → flatter can be safer</strong></p>
<p><strong>small population → steeper can be safer</strong></p>
<p>For a moment they look contradictory.</p>
<p>Then the real question appears:</p>
<p><strong>Robust to what?</strong></p>
<p>That, for me, is the centre of Chapter 6.</p>
<p>There is no universally robust genome.</p>
<p>There is only robustness relative to a particular source of evolutionary danger.</p>
</section>
<section id="same-genome.-different-evolutionary-physics." class="level2">
<h2 class="anchored" data-anchor-id="same-genome.-different-evolutionary-physics.">Same genome. Different evolutionary physics.</h2>
<p>Mutation and drift are both sources of change.</p>
<p>But they threaten populations differently.</p>
<p>Mutation continually generates new variants.</p>
<p>Drift determines which variants may survive or disappear simply because populations are finite.</p>
<p>So the architecture that protects a lineage against one kind of uncertainty need not protect it against the other.</p>
<p>That sounds obvious after saying it.</p>
<p>It did not feel obvious before reaching this chapter.</p>
<p>And it made me think again about something population geneticists routinely do: we talk about a variant’s effect as though that effect were a complete evolutionary description.</p>
<p>But the fate of that effect depends on mutation rate, population size, genetic background, epistasis and the surrounding fitness landscape.</p>
<p>A mutation does not arrive in a vacuum.</p>
<p>Neither does selection.</p>
</section>
<section id="then-adami-takes-us-to-one-of-biologys-strangest-genomes" class="level2">
<h2 class="anchored" data-anchor-id="then-adami-takes-us-to-one-of-biologys-strangest-genomes">Then Adami takes us to one of biology’s strangest genomes</h2>
<p>The <em>Trypanosoma</em> section made the abstract argument feel much less abstract.</p>
<p>Kinetoplastid mitochondria are extraordinary.</p>
<p>Their mitochondrial DNA is divided between maxicircles and enormous numbers of minicircles. Many mitochondrial transcripts are not ready to translate directly from the encoded DNA. Guide RNAs direct extensive RNA editing before functional messages emerge.</p>
<p>The system looks, at first sight, unnecessarily complicated.</p>
<p>Why would evolution tolerate such machinery?</p>
<p>This is precisely where the robustness perspective becomes provocative.</p>
<p>If mutations in important mitochondrial sequences disrupt the editing cascade, many damaging changes can become effectively lethal rather than being allowed to accumulate gradually.</p>
<p>Other sequences can gain protection through overlap or multifunctionality.</p>
<p>And the guide-RNA population itself can maintain a relatively stable collective distribution even though individual sequence variants turn over.</p>
<p>I found myself staring again at the same question:</p>
<p><strong>Is this complexity merely something evolution became stuck with—or is part of it doing evolutionary work that is invisible if I inspect only present-day fitness?</strong></p>
<p>Adami contrasts this interpretation with constructive neutral evolution.</p>
<p>That distinction matters.</p>
<p>One explanation asks how elaborate machinery can accumulate without being directly favoured.</p>
<p>The robustness argument asks something different:</p>
<p><strong>What does that machinery do to the future consequences of mutation and drift?</strong></p>
<p>It may be almost invisible when nothing goes wrong.</p>
<p>Its importance appears when perturbation arrives.</p>
<p>And that thought immediately reminded me of engineering.</p>
<p>The value of redundancy, error correction or a backup system is often invisible while everything is functioning normally.</p>
<p>You discover what it was doing only when the system is challenged.</p>
</section>
<section id="this-is-where-i-suddenly-thought-about-genomic-ai" class="level2">
<h2 class="anchored" data-anchor-id="this-is-where-i-suddenly-thought-about-genomic-ai">This is where I suddenly thought about genomic AI</h2>
<p>Until this point I had been reading Chapter 6 as an evolutionary biologist.</p>
<p>Then I started seeing language-model scores everywhere.</p>
<p>A genomic language model gives us something very tempting:</p>
<p>one sequence → one score.</p>
<p>Change one nucleotide → another score.</p>
<p>We can ask whether the alternative allele looks more or less plausible than the reference.</p>
<p>Useful.</p>
<p>But after this chapter, that feels incomplete.</p>
<p>Suppose two sequences receive almost identical model scores.</p>
<p>Now mutate every position around each sequence.</p>
<p>For the first sequence, almost every neighbour receives a much worse score.</p>
<p>For the second, dozens of neighbouring sequences remain perfectly plausible.</p>
<p>Are those two sequences really equivalent?</p>
<p>The original scores say yes.</p>
<p>Their <strong>local landscapes</strong> say no.</p>
<p>And suddenly the connection to robustness becomes hard to ignore.</p>
<p>Perhaps genomic models should not only be evaluated by asking:</p>
<p><strong>Did the model score this variant correctly?</strong></p>
<p>Perhaps we should also ask:</p>
<p><strong>What landscape has the model learned around this sequence?</strong></p>
<p>Is it sharp?</p>
<p>Flat?</p>
<p>Asymmetric?</p>
<p>Are there connected neutral-like regions?</p>
<p>Do the model’s tolerated sequence neighbourhoods correspond to evolutionary conservation, allele frequencies or experimentally measured mutational tolerance?</p>
<p>And perhaps most importantly:</p>
<p><strong>does the biological meaning of that landscape change with population context?</strong></p>
<p>A model may identify molecular constraint beautifully and still know nothing about effective population size.</p>
<p>It may recognize a highly constrained nucleotide without knowing whether a weakly deleterious allele is visible to selection in the population where it occurs.</p>
<p>That is not necessarily a failure of the model.</p>
<p>But it is a limitation of what its score means.</p>
<p>And distinguishing those two things matters.</p>
</section>
<section id="maybe-a-variant-is-the-wrong-unit-of-thought" class="level2">
<h2 class="anchored" data-anchor-id="maybe-a-variant-is-the-wrong-unit-of-thought">Maybe a variant is the wrong unit of thought</h2>
<p>This chapter left me wondering whether our obsession with individual variant scores is partly inherited from the structure of our datasets.</p>
<p>Variant.</p>
<p>Position.</p>
<p>Reference.</p>
<p>Alternative.</p>
<p>Score.</p>
<p>Next row.</p>
<p>But evolution does not read a VCF one row at a time.</p>
<p>A mutation appears inside a genome.</p>
<p>That genome has neighbours.</p>
<p>Those neighbours have neighbours.</p>
<p>Their effects interact.</p>
<p>And the whole system is filtered through mutation, selection, drift and population history.</p>
<p>So perhaps a genuinely evolutionary genomic AI should eventually tell us something larger than:</p>
<blockquote class="blockquote">
<p>this mutation looks bad.</p>
</blockquote>
<p>I would want to know:</p>
<blockquote class="blockquote">
<p><strong>What kind of evolutionary neighbourhood am I standing in?</strong></p>
</blockquote>
<p>That feels like a much richer question.</p>
<p>And perhaps a much harder benchmark.</p>
</section>
<section id="the-sentence-i-am-carrying-out-of-chapter-6" class="level2">
<h2 class="anchored" data-anchor-id="the-sentence-i-am-carrying-out-of-chapter-6">The sentence I am carrying out of Chapter 6</h2>
<p>If Chapter 5 made me suspicious of the word <em>complexity</em>, Chapter 6 has made me suspicious of the word <em>fitness</em> when it appears alone.</p>
<p>A genotype can have high fitness and still be fragile.</p>
<p>A genotype can sacrifice some immediate fitness and make its lineage safer under mutation.</p>
<p>A mutation can be neutral now and profoundly alter the effects of mutations that come later.</p>
<p>And a small population can favour a genetic architecture very different from the one favoured under high mutation pressure.</p>
<p>So evolution is doing something subtler than climbing toward better sequences.</p>
<p>It is also reshaping the consequences of where the next step might land.</p>
<p>That is the idea I did not have when I began the chapter.</p>
<p>And it may be the most useful one I am taking with me toward genomic AI:</p>
<p><strong>a genome is not just a sequence. It occupies a neighbourhood of possibilities.</strong></p>
<p>If our models truly understand something about biological sequence, perhaps we should eventually expect them to recover not only the score of the sequence in front of us—</p>
<p>but the shape of that neighbourhood around it.</p>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Genomes to AI</category>
  <category>Evolution</category>
  <category>Population genetics</category>
  <category>Scientific AI</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-04-18-can-evolution-become-robust-to-chance/</guid>
  <pubDate>Fri, 17 Apr 2026 22:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/09-can-evolution-become-robust-to-chance.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>When does more become complexity?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-04-04-when-does-more-become-complexity/</link>
  <description><![CDATA[ 




<p><em>A Genomes to AI reading note on Chapter 5 of Christoph Adami’s</em> The Evolution of Biological Information.</p>
<p>In <a href="../../posts/2026-03-23-can-we-replay-evolution/index.html">my previous note</a>, the frozen ancestors in Lenski’s experiment made me ask whether a new evolutionary path could be replayed. Chapter 5 asks a harder question about the paths we see across life: <strong>if evolution produces complexity, what exactly is increasing?</strong></p>
<p>Adami begins with a tree of life. Its branches do not form a ladder. Bacteria and archaea have not been replaced by animals; they are still here, thriving in their own ways. A path from the root to a human can make it look as if evolution has been climbing toward us. Look at the whole tree, and the story changes. Before drawing an arrow marked <em>more complex</em>, I need to say what I counted and where I looked.</p>
<section id="could-i-count-the-parts" class="level2">
<h2 class="anchored" data-anchor-id="could-i-count-the-parts">Could I count the parts?</h2>
<p>Different cell types, tissues and levels of organization give us one way to describe structure. But what level do I choose? One cell can contain intricate molecular machinery; a multicellular body can be divided into parts in several ways. Count the parts alone, and I miss their connections. Add the connections, and I still do not know what they <em>do</em>.</p>
<p>Perhaps genome length would give me a simpler answer. Every organism builds itself, in part, from genetic instructions. Yet the C-value paradox interrupts that thought: large genomes do not line up neatly with our sense of organismal complexity. Some organisms carry much more DNA than humans without fitting a simple ranking by structure or function. The extra sequence is still real; its length simply answers a different question.</p>
<p>I tried another word in the margin: <em>compressibility</em>. A repetitive sequence can be described briefly, while a random one resists compression. Kolmogorov complexity formalizes that distinction through the length of the shortest program that could produce a string. It is an elegant idea, but a random string can score highly without doing anything for an organism. Also, finding the shortest possible program is not generally computable. <strong>Hard to compress is not the same as useful to a cell.</strong></p>
<p>Adami’s move is to make the world part of the question. How much of a sequence is informative <em>given the environment in which it evolved</em>? In his account, physical or informational complexity is tied to what a sequence records about that environment. This is not a property I can read from a solitary DNA string. I need an ensemble of sequences, a specified environment and a way to compare them. And even then, the link between the theoretical construction and a measurable quantity requires assumptions.</p>
<p>The RNA aptamer experiments give that abstract idea something to hold onto. In one set of laboratory-evolved GTP-binding RNAs, stronger binding was associated with greater sequence information and, in the studied structures, more elaborate folding. This is a relationship within a defined molecular task, not a conversion rule that lets me rank every organism on one scale. A different task, or an environment that varies across a life cycle, changes what information may be useful. Adami even sketches an extension to multiple environments, while admitting how difficult it would be to measure in practice.</p>
<p>I wrote <strong>“what does it do <em>here</em>?”</strong> beside my first attempt at a universal ruler.</p>
</section>
<section id="does-the-wiring-tell-me-what-the-system-does" class="level2">
<h2 class="anchored" data-anchor-id="does-the-wiring-tell-me-what-the-system-does">Does the wiring tell me what the system does?</h2>
<p>Chapter 5 then moves from sequences to networks. Genes help specify proteins and regulation; proteins participate in metabolic systems; neurons connect into circuits. A graph can show which parts are connected, but two graphs with the same bare pattern may have different meanings if the parts have different roles.</p>
<p>Adami uses an artificial cell model to make this testable. A small coded genome begins with a few metabolic functions and evolves larger networks. In the model, information measured in the genome grows alongside metabolic organization. The changing environment matters: a cell in a predictable world can rely on available precursors, whereas a cell facing fluctuating supplies may need machinery to make them. Bigger networks and higher information can emerge together in this model; it does not follow that every additional node in a living network improves function.</p>
<p>The worm <em>C. elegans</em> makes the limitation of a wiring diagram vivid. Count small patterns of connections between its neurons without identifying the neurons, and some patterns look unremarkable. Mark which nodes are sensory neurons, interneurons or motor neurons, and a sensory-to-interneuron-to-motor arrangement stands out against shuffled assignments. The connections stayed the same; the <strong>roles</strong> made a relationship visible. Adami also examines network modules and motif frequencies, while warning that neither modularity nor a motif count supplies a universal measure of complexity. How we label a node is itself a scientific choice.</p>
<p>That is what I wanted the two small networks in my notebook to remind me: <strong>same wires, different roles.</strong></p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/08-when-does-more-become-complexity.webp" class="reading-sketch img-fluid figure-img" alt="Graphite notebook sketch on ruled paper. A central tree retains low and high branches; two marginal sketches are labelled passive spread from a floor and driven upward bias. Side questions ask which world sequence information concerns, why node roles matter in a network, and how A and B act on the same background."></p>
<figcaption>A wide ruled notebook page in dark graphite. A branching tree keeps small and elaborate lineages together, while annotations distinguish passive spread from a driven trend and ask what genome length, network roles and paired variants can reveal.</figcaption>
</figure>
</div>
</section>
<section id="if-the-upper-branches-rise-is-evolution-pushing-them-upward" class="level2">
<h2 class="anchored" data-anchor-id="if-the-upper-branches-rise-is-evolution-pushing-them-upward">If the upper branches rise, is evolution pushing them upward?</h2>
<p>The tree led me to Adami’s distinction between <em>passive</em> and <em>driven</em> trends. Imagine many lineages whose trait values wander up or down, but cannot go below a lower bound. As they branch and spread, the maximum can rise even when changes in either direction are equally likely. A driven trend requires evidence that increases are favored over decreases. <strong>An increasing maximum by itself cannot tell me which process produced it.</strong></p>
<p>That distinction turns a vague claim about progress into a test. In fossil marine animals, body size has increased in ways that fit a size-biased model better than an unbiased one in the study Adami discusses. But body size is not complexity, and a result for one trait cannot be carried across the entire tree of life. He turns to proteins too: in a comparison grouped by inferred gene age, younger proteins tend to change faster, while older proteins are on average longer and more structurally ordered. Those are useful population-level patterns. They do not mean that every protein grows longer as it ages, or that any individual lineage must grow more complex.</p>
<p>Here I noticed a trap in my own question. <em>Does complexity increase?</em> has been quietly shifting between at least three things: the maximum among branches, a mean across groups, and the path taken by one lineage. These can behave differently even when all the observations are correct.</p>
</section>
<section id="what-happens-along-one-lineage" class="level2">
<h2 class="anchored" data-anchor-id="what-happens-along-one-lineage">What happens along one lineage?</h2>
<p>To study shorter paths, Adami returns to fitness. In a deliberately restricted model with fixed fitness values, no mutation and selection acting on existing types, mean fitness rises as the fitter types spread. Fisher’s theorem makes that result precise. Let mutations enter, and the simple guarantee no longer follows; Price’s equation separates the part associated with selection from changes transmitted across generations. In a finite population, chance can also complicate the tidy upward path.</p>
<p>Lenski’s evolving <em>E. coli</em> populations give this section an experimental anchor. Fitness improved substantially over tens of thousands of generations, yet the gains slowed. Adami examines the fitted trajectory and asks whether diminishing returns mean that adaptation is approaching a final summit. They do not, by themselves, establish one. An apparent peak in a two-dimensional landscape may conceal other accessible routes when many genes interact.</p>
<p>The NK model makes those interactions adjustable. Change how many positions affect one another, and the landscape becomes more or less rugged. Evolutionary runs can then take different routes; with some settings, a lineage can even pass through a lower-fitness state on its way to another peak. The model is intentionally small and controlled. It helps isolate what genetic interaction can do, without being a literal map of an <em>E. coli</em> genome.</p>
<p>This is where my question changes again. If mutation A helps on one background, will it help after mutation B? Epistasis means that the combined effect need not equal the effects of A and B treated independently. Even <em>diminishing returns</em> requires care: a slowing fitness trajectory and a particular sign of pairwise epistasis are not interchangeable statements. The route through sequence space changes what a later step can mean.</p>
</section>
<section id="from-genomes-to-ai" class="level2">
<h2 class="anchored" data-anchor-id="from-genomes-to-ai">From genomes to AI</h2>
<p>I came to Chapter 5 with a genomic-model version of the same tempting shortcut: reduce a complex system to one number and compare the numbers. A language model can give each variant a precise score. But what would a score be a measure <em>of</em>? Sequence regularity, molecular function, long-term constraint, or a fitness effect in a particular environment? Those are distinct targets. A model trained to predict sequence can be excellent at that task without settling which variants matter for an organism.</p>
<p>This chapter suggests several questions for <a href="../../popgenlm.html">PopGenLM Bench</a>. Does model performance depend on the genomic context or the functional class being tested? Would a prediction change when a variant is placed on another haplotype? Do paired variants behave as the sum of their individual scores suggests? Does agreement with conservation persist in different environments or lineages? I have not established those answers. The benchmark needs independent evidence, appropriate controls and explicit comparison units to test them.</p>
<p>The most honest answer I can take from this chapter is a method of asking. <strong>When does more become complexity?</strong> When I specify <em>more of what</em>, in <em>which system</em>, under <em>which conditions</em>, and show why that measure speaks to the function or evolutionary process I care about. Adami does not close the book on a universal upward trend; his own chapter ends with that larger question open. My sketch therefore leaves the branches unfinished. Before I extend an arrow, I want to know what it measures.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em> (2024), Chapter 5, “Evolution of Complexity,” §§5.1–5.5. The claims about RNA aptamers, artificial metabolic networks, colored neuronal motifs, fossil body size, protein age, the LTEE and NK landscapes follow examples discussed there. The notebook questions and the connection to genomic AI are my interpretations, not results from PopGenLM Bench.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Evolutionary genomics</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-04-04-when-does-more-become-complexity/</guid>
  <pubDate>Fri, 03 Apr 2026 22:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/08-when-does-more-become-complexity.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>Can we replay evolution?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-03-23-can-we-replay-evolution/</link>
  <description><![CDATA[ 




<p><em>A Genomes to AI reading note on Chapter 4 of Christoph Adami’s</em> The Evolution of Biological Information.</p>
<p>In <a href="../../posts/2026-03-15-what-can-a-genome-remember/index.html">my previous note</a>, I came to think of a genome as a record of encounters with the world. But a record shows me what happened, not all the things that <em>could</em> have happened. If I could take an earlier population and let it evolve again, would it arrive at the same answer?</p>
<p>Adami’s Chapter 4 moves that question from a thought experiment to the laboratory. The settings change from warm cultures to frozen bacteria to digital organisms, but one question follows me through them: <strong>What can an experiment tell me about the path, not just the outcome?</strong></p>
<section id="how-slowly-can-a-world-change" class="level2">
<h2 class="anchored" data-anchor-id="how-slowly-can-a-world-change">How slowly can a world change?</h2>
<p>William Henry Dallinger cultivated microorganisms while raising the temperature of their environment over generations. The later populations tolerated heat that had harmed their predecessors. I underlined one word in my notes: <em>gradually</em>.</p>
<p>What if Dallinger had raised the heat to the final level on the first day? The experiment does not answer that counterfactual, but it makes the question hard to ignore. A population must survive each step to encounter the next one. Dallinger could watch adaptation unfold without being able to read the mutations that enabled it. When I see an adapted population today, I too can be tempted to jump from its present state to a tidy story about how it got there. <strong>Was the route as important as the destination?</strong></p>
</section>
<section id="what-does-a-frozen-ancestor-give-us" class="level2">
<h2 class="anchored" data-anchor-id="what-does-a-frozen-ancestor-give-us">What does a frozen ancestor give us?</h2>
<p>Richard Lenski’s long-term <em>E. coli</em> experiment offers an unusual way to ask. Twelve populations began from a common ancestor and evolved separately in the same laboratory environment. Samples were frozen at intervals. Researchers could compare descendants with their predecessors—and later thaw an earlier sample to start a new run.</p>
<p>Citrate was present in the medium throughout the experiment. Yet only one of the original populations evolved the ability to grow aerobically on it. That detail made me stop. If all twelve populations encountered the same resource, <strong>why this one?</strong> Had it simply been lucky, or had earlier changes opened a path that was less accessible to the others?</p>
<p>The freezer allowed researchers to test that question. They thawed clones from different points in the successful lineage’s past and ran evolution again. Citrate use arose in some replays started from later backgrounds, while replays from the original ancestor did not produce it. Earlier genetic changes had altered the chance of a later innovation: <strong>historical contingency</strong> made a difference. Subsequent genomic work described <em>potentiation</em>, <em>actualization</em> and <em>refinement</em> in the origin of citrate use. Those names help reconstruct what happened in this lineage; they do not mean the outcome was guaranteed.</p>
<p>I wrote beside the freezer: <strong>Same environment, different histories.</strong> A replay is a test, not a rewind. It begins with a preserved clone, but new mutations and chance events follow. Even a genetic background that makes citrate use more accessible does not ensure that every replay will find it.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/07-can-we-replay-evolution.webp" class="reading-sketch img-fluid figure-img" alt="Graphite reading notebook on ruled paper. Dallinger's heat experiment, Lenski's common ancestor and archived E. coli lineage, replays from one thawed intermediate clone with differing outcomes including Cit+, and Avida digital replicators. Marginal questions concern selection, chance and history."></p>
<figcaption>A wide ruled notebook page drawn in uneven graphite. Dallinger’s rising-temperature cultures sit in one margin; Lenski’s replicate E. coli lineages lead to a frozen intermediate and several replay outcomes; a small Avida vignette asks whether the steps can be tested.</figcaption>
</figure>
</div>
</section>
<section id="what-can-a-digital-organism-tell-me" class="level2">
<h2 class="anchored" data-anchor-id="what-can-a-digital-organism-tell-me">What can a digital organism tell me?</h2>
<p>Then Adami takes the experiment into a computer. In Avida, self-replicating programs carry instructions that can mutate; their success depends on the rules of a digital environment. I paused over the word <em>organism</em>. These programs have no cells or bacterial metabolism. Why study them alongside Dallinger’s cultures and Lenski’s <em>E. coli</em>?</p>
<p>Because here, the lineage can be inspected instruction by instruction. In Avida experiments, complex logic functions evolved through histories that included simpler functions and other changes. The first program able to perform a complex task could be only a mutation or two from its parent, yet many consequential changes away from its ancestor. Seen only at the moment it appeared, the new function might look sudden. Seen along the lineage, it has a history.</p>
<p>Avida lets us test evolutionary possibilities under explicit rules; the microbial experiments concern living populations with their own biology. I cannot use a digital lineage to explain a bacterium’s exact mutations. But I can use the comparison to ask a better question of both: <strong>Which steps can I actually recover, and which am I only inferring from the end?</strong></p>
</section>
<section id="from-genomes-to-ai" class="level2">
<h2 class="anchored" data-anchor-id="from-genomes-to-ai">From genomes to AI</h2>
<p>I often ask a genomic model to score one variant at a time. Chapter 4 makes me pause before treating that score as a property of the variant alone. Which sequence surrounds it? Which other mutations are already present? Which environment tests it? A change that matters in one background may have a different effect in another.</p>
<p>This gives me questions for <a href="../../popgenlm.html">PopGenLM Bench</a>. Could a model’s prediction change across genetic backgrounds? If closely related sequences appear in both training and evaluation, what does good performance really show? Could experimental lineages provide independent evidence about which predictions hold? These are directions I want to test, not results I already have.</p>
<p>So, <strong>can we replay evolution?</strong> We can return to a preserved starting point and watch another history begin. We cannot demand the same ending. That is what makes the replay informative: when outcomes differ, or become possible only after certain earlier changes, the record of the past becomes something we can question experimentally.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em> (2024), Chapter 4, “Experiments in Evolution,” §§4.1–4.5. For the replay and digital-evolution details, see <a href="https://doi.org/10.1073/pnas.0803151105">Blount, Borland and Lenski (2008)</a>, <a href="https://doi.org/10.1038/nature11514">Blount and colleagues (2012)</a> and <a href="https://doi.org/10.1038/nature01568">Lenski and colleagues (2003)</a>. The notebook questions and the connection to genomic AI are my interpretations, not findings from PopGenLM Bench.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Evolutionary genomics</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-03-23-can-we-replay-evolution/</guid>
  <pubDate>Sun, 22 Mar 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/07-can-we-replay-evolution.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>What can a genome remember?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-03-15-what-can-a-genome-remember/</link>
  <description><![CDATA[ 




<p><em>A Genomes to AI reading note on Chapter 3 of Christoph Adami’s</em> The Evolution of Biological Information.</p>
<p>I began this chapter expecting an account of how evolution adds information to DNA. Adami began with a ruler.</p>
<p>At first I wanted to hurry past it. What could the length of a stick have to do with the history of a genome? Then I realized that I had already made the assumption the example was meant to unsettle: that information simply sits in an object, waiting for us to read it out. A ruler tells me something about a stick because its markings respond to length. It may be precise and still tell me nothing about the stick’s colour. The property, the measuring device and the question have to meet.</p>
<p>My first question in the margin was: <strong>What is a genome measuring?</strong></p>
<p>Adami asks us to imagine a DNA sequence and the environment in which it evolves as two initially uncertain variables. Mutation offers possible changes; survival and reproduction test them in a particular world. If a variant repeatedly succeeds, inheritance can keep its sequence in a population. In the chapter’s idealized example, an adaptive fixation makes a sequence state more predictable given the environment. Mutation alone has not learned anything. The relationship comes from variation, differential success and inheritance.</p>
<p>I had been picturing a genome as a store of answers. Now I pictured a <em>record of encounters</em>: a molecular possibility met a particular world, and descendants carried an outcome forward. “Memory” is a useful metaphor, provided I remember that there is no intention in the molecule and no foresight in selection.</p>
<section id="why-does-the-record-last" class="level2">
<h2 class="anchored" data-anchor-id="why-does-the-record-last">Why does the record last?</h2>
<p>If molecules are physical, their records are vulnerable. Adami’s Maxwell demon makes that problem tangible. The imaginary demon measures the motion of molecules and sorts them through a door. Sorting seems to create order for free until we count the physical record of the measurements and the cost of erasing it. Information has a material history.</p>
<p>Selection performs a different kind of sorting. A favourable change can spread; a damaging change may disappear with the lineages carrying it. Replication makes surviving arrangements available for another round. In Adami’s deliberately clean picture, this can preserve information acquired in an environment against the noise that would otherwise erode it.</p>
<p>But I paused over the word <em>clean</em>. I work with population data; I cannot read a fixed difference as proof that selection discovered a better answer. Drift can fix variants. Populations are finite. Recombination reshuffles combinations. Organisms occupy different niches, and their neighbours evolve too. Adami makes room for these complications in his “leaky natural demon.” The leak matters: it prevents the metaphor from becoming a claim that information must rise with every generation.</p>
<p>It also changes the question. <strong>Information about which world?</strong></p>
</section>
<section id="does-a-conserved-protein-carry-the-same-information-everywhere" class="level2">
<h2 class="anchored" data-anchor-id="does-a-conserved-protein-carry-the-same-information-everywhere">Does a conserved protein carry the same information everywhere?</h2>
<p>Adami compares information profiles of homeodomain and COX2 proteins across branches of life. Their patterns do not trace one tidy staircase toward more information; the lineages separate in different ways. I returned to the ruler. The meaning of sequence constraint depends on what the protein does, where it does it, and which sequences I put in the comparison. An alignment can show me a pattern; it cannot alone tell me what caused it.</p>
<p>The chapter’s ancestral protein example sharpens the distinction. Reconstructed fluorescent proteins from a coral lineage can be made and examined for their colours. Inferring a past sequence from a tree and testing a property of the resulting protein are different steps. A line on an entropy plot can start a question about a change in function; it cannot finish it.</p>
</section>
<section id="can-a-good-answer-become-wrong" class="level2">
<h2 class="anchored" data-anchor-id="can-a-good-answer-become-wrong">Can a good answer become wrong?</h2>
<p>The lineage comparisons span deep time. Adami’s HIV protease example brings environmental change into sharper focus. A protease inhibitor alters the world in which a viral protein must function. Viral populations that survive encounter a new selective pressure and explore new sequence combinations.</p>
<p>In comparisons of patient sequences, drug exposed proteases become more variable at particular residues than proteases from untreated patients. Count each residue separately and it looks as though information has been lost. That is a striking possibility: a population adapting to treatment may look less ordered when inspected one position at a time.</p>
<p>Then Adami asks where the information is being counted. In his time series, estimated single-site information falls in treated proteases while estimated dependencies between pairs of residues rise. Some associations join residues near the active site to others farther away. Adaptation may be distributed across combinations that a list of individual sites would miss.</p>
<p>I kept returning to a different question: <strong>What if the unit I inspect is too small?</strong></p>
<p>This needs care. A correlation between residues does not establish its mechanism, and the chapter could not reliably estimate all higher-order dependencies. A rising sequence statistic is not proof that a virus has solved every problem posed by treatment. Yet the change between the single-site and pairwise views stays with me. At individual positions I might tell a story of loss; in their relationships I can see part of a possible reconstruction.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/06-what-can-a-genome-remember.webp" class="reading-sketch img-fluid figure-img" alt="Graphite notebook drawing with a loop through variation, a changing environment, and inherited sequences. Handwritten notes question measurement, drift, HIV resistance and information between sites."></p>
<figcaption>A ruled notebook page in thick graphite: a changing world cuts across a loop of variation, selection and inheritance; questions about measurement and memory surround it, while paired DNA marks prompt a second look between sites.</figcaption>
</figure>
</div>
</section>
<section id="what-if-the-interesting-message-lies-outside-the-gene" class="level2">
<h2 class="anchored" data-anchor-id="what-if-the-interesting-message-lies-outside-the-gene">What if the interesting message lies outside the gene?</h2>
<p>Adami then moves into regulatory DNA. A transcription factor must find suitable binding sites in a very long sequence. A motif can provide some of the specificity for that search. The chapter links the frequencies of nucleotides in binding sites to binding energy under a model, and asks how much specificity is <em>enough</em> to find targets among many alternatives in a genome.</p>
<p>Why should the strongest possible binding site always be best? Binding has to work within a regulatory system. A consensus motif is not necessarily the sequence we find in an organism. A weight matrix learned from known CRP sites can rank candidates, but a strict threshold misses some known sites while a loose one admits many others. A precise motif score cannot by itself tell me what happens in a cell.</p>
<p>The last example returned me to a thought from my earlier reading note: <a href="../../posts/2026-02-07-meaning-lives-in-relationships/index.html">meaning can live between things</a>. Dorsal binding sites involved in fly development differ depending on whether a Twist site is nearby. Separate the sites by that context and part of their sequence pattern becomes visible; pool them and the distinction blurs. Adami measures a modest association between the Dorsal sequence and the proximity of a Twist motif. For me, the discovery is not that DNA literally “knows” its neighbour. It is that my choice of comparison can hide a relationship useful for prediction.</p>
<p>That is why the little paired marks sit apart from the main loop in my drawing. The loop asks how a record changes through time; the marks ask where I am looking for that record. <strong>The record can be revised, and some of it is in the connections.</strong></p>
</section>
<section id="from-genomes-to-ai" class="level2">
<h2 class="anchored" data-anchor-id="from-genomes-to-ai">From genomes to AI</h2>
<p>This chapter gives me a more demanding way to think about genomic language models. A model can produce a reproducible score for a variant. That score may help predict a measured outcome. But it is not the same quantity as Adami’s evolutionary information, and it cannot tell me on its own which aspect of biology the model has recovered. Does it track long-term constraint, a molecular assay, a change in binding, or fitness in a particular environment? These targets may disagree.</p>
<p>This is one reason I am building <a href="../../popgenlm.html">PopGenLM Bench</a>. I want to put variant scores beside independent evolutionary and functional evidence, state the genomic context, and test where agreement holds or breaks. The chapter also makes me uneasy about treating every variant as an isolated question. If function sometimes depends on combinations, neighbouring binding sites or environmental change, how much can any one variant score tell me without that context? This is a question for the benchmark to test, not a conclusion it has already established.</p>
<p>I closed the chapter with a changed version of my opening question. <strong>What can a genome remember?</strong> Relationships shaped by its history, to the extent that they persist and matter in a given context. <strong>What can an AI model learn from that record?</strong> Perhaps a great deal. But to find out <em>what</em> it has learned, I have to name a biological question and look for evidence beyond the score.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em> (2024), Chapter 3, “Evolution of Information,” §§3.1–3.5. The notebook questions and the connection to genomic AI are my interpretations of the chapter, not claims made by Adami or experimental findings from PopGenLM Bench.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Evolutionary genomics</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-03-15-what-can-a-genome-remember/</guid>
  <pubDate>Sat, 14 Mar 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/06-what-can-a-genome-remember.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>What Does a Gene Know?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-03-07-what-does-a-gene-know/</link>
  <description><![CDATA[ 




<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/05-what-does-a-gene-know.webp" class="reading-sketch img-fluid figure-img" alt="Handwritten pencil notes ask what information is about, emphasise ensembles and testing predictions against biology."></p>
<figcaption>A simple pencil notebook page linking a sequence ensemble to a folded protein and a leaf in its environment.</figcaption>
</figure>
</div>
<p>The final part of the supplied excerpt asks a question that sounds almost poetic: what does a gene know? The chapter uses “know” in a precise, non-mental sense. A biological molecule has information about its world when its state supports successful predictions or interactions in that world.</p>
<p>This question brought the whole reading series together for me.</p>
<section id="information-is-always-about-something" class="level2">
<h2 class="anchored" data-anchor-id="information-is-always-about-something">Information is always about something</h2>
<p>Section 2.3 turns from relationships within sequences to the information content of genes. A molecule such as haemoglobin is adapted to a world containing oxygen, particular chemical structures, cellular partners and physiological demands. Its sequence and folded form embody relationships accumulated through evolution.</p>
<p>Calling this “knowledge” is useful only if we remain precise. The molecule does not think. Its structure carries predictive relationships to an environment: what to bind, where to function, and which interactions to avoid.</p>
<p>This gives biological information a direction. A sequence is not simply full of information in the abstract. It may carry information about a molecular partner, an environment, a developmental state or a functional requirement.</p>
</section>
<section id="why-one-sequence-is-not-enough" class="level2">
<h2 class="anchored" data-anchor-id="why-one-sequence-is-not-enough">Why one sequence is not enough</h2>
<p>Box 2.2 addresses a common difficulty. Shannon information is a correlational property of ensembles, so it cannot normally be measured by inspecting a single sequence. One observed string does not reveal which of its patterns are reliably connected to a target and which occurred by chance.</p>
<p>The distinction between a message and an ensemble is crucial. A particular clue may happen to predict correctly once, but its information can be assessed only across repeated possibilities. For molecular sequences, evolution helps generate the ensemble: mutation explores alternatives and selection changes which alternatives persist.</p>
<p>This is why conservation and covariation become informative. Across many sequences, the evolutionary process reveals which changes are tolerated and which relationships are maintained. The information is not printed visibly inside one string; it appears through structured variation across possible strings and worlds.</p>
</section>
<section id="information-content-and-biological-complexity" class="level2">
<h2 class="anchored" data-anchor-id="information-content-and-biological-complexity">Information content and biological complexity</h2>
<p>The chapter suggests that information content may be a useful proxy for complexity, while acknowledging how difficult it is to estimate. Genes do not act alone. Regulatory systems store substantial context about when, where and how gene products are used. Protein-coding sequence is therefore only part of an organism’s informational architecture.</p>
<p>The excerpt begins the discussion of proteins as amino-acid sequences translated from codons. Even at this early stage, the challenge is clear: counting possible symbols is not the same as measuring functional information. Function depends on folding, interactions, cellular context, environment and the ensemble of viable alternatives.</p>
<section id="notes-from-my-margin" class="level3 pencil-margin-note">
<h3 class="anchored" data-anchor-id="notes-from-my-margin">Notes from my margin</h3>
<ul>
<li>Information must have a target: information <em>about what</em>?</li>
<li>A single sequence can be examined, but an ensemble is needed to measure a reliable relationship.</li>
<li>Evolutionary history is part of the evidence, not background noise.</li>
</ul>
</section>
</section>
<section id="my-interpretation-for-popgenlm-bench" class="level2">
<h2 class="anchored" data-anchor-id="my-interpretation-for-popgenlm-bench">My interpretation for PopGenLM Bench</h2>
<p>This reading changes the language I want to use around genomic AI. A model score is not “biological information” simply because it was produced from DNA. It becomes useful evidence only if it improves prediction of a clearly defined biological target in an appropriate ensemble.</p>
<p>For variant-effect models, possible targets include population frequency, evolutionary constraint, molecular function, expression, phenotype or experimental effect. These are related but not interchangeable. A score that predicts conservation may not predict present-day fitness in a particular population. A score calibrated in one species may not transfer to another. A correlation learned from sequence may reflect shared history rather than a causal mechanism.</p>
<p>PopGenLM Bench should therefore make four questions explicit:</p>
<ol type="1">
<li>What biological quantity is the score expected to predict?</li>
<li>Which ensemble defines the test—sites, individuals, populations or species?</li>
<li>How much uncertainty does the model reduce beyond appropriate baselines?</li>
<li>Under which changes of reference, lineage, annotation and sampling does that relationship remain stable?</li>
</ol>
<p>This is a stronger foundation than asking whether the software ran or whether the plot looks convincing.</p>
</section>
<section id="my-commentary-on-chapter-2" class="level2">
<h2 class="anchored" data-anchor-id="my-commentary-on-chapter-2">My commentary on Chapter 2</h2>
<p>The chapter’s most important contribution, for me, is discipline of language. It moves from the history of genetic information to random variables, probability, entropy, conditional dependence and gene information without allowing “information” to remain a loose synonym for sequence.</p>
<p>I also appreciate that the biological examples expose the limits of the mathematics. An entropy profile depends on sampling. Mutual information detects dependence, not cause. A single sequence does not establish what it predicts. Recent ancestry can resemble constraint. These are not reasons to reject the framework; they are reasons to use it scientifically.</p>
<p>The reading has made my AI journey feel more grounded. The question is not whether AI can process more genomic data. It can. The harder question is whether its outputs carry stable, testable information about the biological world.</p>
</section>
<section id="where-i-want-to-learn-next" class="level2">
<h2 class="anchored" data-anchor-id="where-i-want-to-learn-next">Where I want to learn next</h2>
<p>My next commentary series will return to the chapter topic by topic, with simple comparisons and small reproducible examples: entropy versus conservation, mutual information versus correlation, single sequences versus ensembles, and model likelihood versus biological effect.</p>
<p>I also want to extend the reading into information content in regulatory DNA, phylogeny-aware sequence models, calibration and the design of benchmarks that separate technical capability from biological support.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em>, Chapter 2, Section 2.3, Section 2.3.1 as included in the supplied excerpt, and Box 2.2, “Information in Single Sequences?”</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Genes and environment</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-03-07-what-does-a-gene-know/</guid>
  <pubDate>Fri, 06 Mar 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/05-what-does-a-gene-know.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>When Uncertainty Reveals Structure</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-02-21-uncertainty-reveals-structure/</link>
  <description><![CDATA[ 




<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/04-when-uncertainty-reveals-structure.webp" class="reading-sketch img-fluid figure-img" alt="Handwritten pencil notes say that uncertainty can be measured, shared variation can reveal structure and recent ancestry can mislead."></p>
<figcaption>A simple pencil notebook page showing uncertain dots, a small correlation grid and a rough RNA cloverleaf.</figcaption>
</figure>
</div>
<p>For the last two weeks I have been thinking about a counter-intuitive idea: uncertainty is not merely a defect in data. When it is measured carefully, its pattern can reveal biological organisation.</p>
<p>Sections 2.2 and 2.2.1 introduce entropy and information. The formulas matter, but the conceptual sequence matters more: first measure how uncertain a variable is; then ask how much that uncertainty decreases when another variable becomes known.</p>
<section id="entropy-measuring-what-remains-unpredictable" class="level2">
<h2 class="anchored" data-anchor-id="entropy-measuring-what-remains-unpredictable">Entropy: measuring what remains unpredictable</h2>
<p>Shannon entropy is calculated from the probability distribution of a variable. If all states are equally likely, uncertainty is maximal. If one state occurs with certainty, entropy is zero. The base of the logarithm determines the unit: base two gives bits; other alphabets can use units scaled to their number of possible states.</p>
<p>I find it helpful to think of entropy as <em>potential information</em>. A variable with several plausible states has something left to reveal. A variable already known with certainty does not.</p>
<p>Applied to a sequence alignment, entropy can be measured at each site. Conserved positions have low entropy; highly variable positions have high entropy. In the tRNA alignment, this produces a landscape of constraint and variability across the molecule.</p>
<p>But low entropy is not automatically the same as biological importance. A site can appear conserved because the sampled sequences share a recent ancestor and have not had time to diverge. Sample composition and evolutionary history therefore shape the apparent signal.</p>
</section>
<section id="information-uncertainty-reduced-by-knowledge" class="level2">
<h2 class="anchored" data-anchor-id="information-uncertainty-reduced-by-knowledge">Information: uncertainty reduced by knowledge</h2>
<p>The chapter defines mutual information as the reduction in uncertainty about one variable after another variable is known. This is a much stricter idea than calling a sequence “informative” because it looks complex.</p>
<p>The measure is symmetric: the shared dependence between two sites is the same whichever direction we describe it. Conditional entropy captures what remains unknown; mutual information captures what is shared.</p>
<p>In the tRNA example, calculating mutual information across every pair of sites creates a matrix. Strong off-diagonal patterns correspond to coordinated substitutions produced by base pairing. From sequence variation alone, the analysis begins to reveal RNA secondary structure.</p>
<p>This is beautiful because variability is doing the explanatory work. If paired sites never changed, there would be no covariation to detect. Structure becomes visible because evolution explores alternatives while selection preserves compatible combinations.</p>
</section>
<section id="the-limitation-that-matters" class="level2">
<h2 class="anchored" data-anchor-id="the-limitation-that-matters">The limitation that matters</h2>
<p>The same example contains its own warning. The method needs a sufficiently diverse ensemble. If sequences are too closely related, most sites may look invariant whether or not they are functionally constrained. If the alignment contains uneven phylogenetic representation, shared ancestry can generate correlations that resemble functional coupling.</p>
<p>Information theory measures dependence in the supplied distribution. Biology must still explain how that distribution arose.</p>
<section id="notes-from-my-margin" class="level3 pencil-margin-note">
<h3 class="anchored" data-anchor-id="notes-from-my-margin">Notes from my margin</h3>
<ul>
<li>Entropy is not disorder in a vague sense; it is uncertainty under a stated model.</li>
<li>Variation can reveal constraint when changes are coordinated.</li>
<li>A clean pattern can still reflect sampling history rather than mechanism.</li>
</ul>
</section>
</section>
<section id="my-interpretation-for-scientific-ai" class="level2">
<h2 class="anchored" data-anchor-id="my-interpretation-for-scientific-ai">My interpretation for scientific AI</h2>
<p>This reading helps me distinguish three quantities that are often blurred together: model confidence, statistical dependence and biological validity.</p>
<p>A language model can assign a sharp probability distribution, but calibration determines whether that confidence is trustworthy. A model score can correlate with an annotation, but dependence alone does not show causation. A benchmark can report strong discrimination, but biological validity depends on the target, controls, sampling and domain of use.</p>
<p>For PopGenLM Bench, entropy and information suggest useful evaluation questions. Does adding a model score reduce uncertainty about an independently measured outcome beyond simpler features such as conservation or GC content? Is that reduction stable across chromosomes, annotation classes and species? Does it remain after controlling for relatedness and genomic structure?</p>
<p>These questions are more demanding than plotting a score distribution. They are also more scientifically meaningful.</p>
</section>
<section id="additional-learning-directions" class="level2">
<h2 class="anchored" data-anchor-id="additional-learning-directions">Additional learning directions</h2>
<p>I want to study finite-sample bias in entropy estimates, sequence reweighting, permutation-based null models and methods for higher-order interactions. Pairwise mutual information can detect relationships, but biological systems often involve networks of sites rather than isolated pairs.</p>
<p>I also want to compare mutual information with attention maps and learned representations from genomic foundation models. Similar-looking heat maps can arise from very different mathematical objects; comparison requires precise definitions and tests.</p>
</section>
<section id="my-commentary" class="level2">
<h2 class="anchored" data-anchor-id="my-commentary">My commentary</h2>
<p>This was the section where information theory stopped feeling like an imported engineering language and began to feel biologically native. Evolution generates distributions. Selection, mutation and history shape their uncertainty. Relationships within those distributions can reveal organisation—but only when the ensemble and its history are taken seriously.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em>, Chapter 2, Sections 2.2 and 2.2.1 on entropy, conditional entropy, mutual information and the inference of tRNA structure from sequence covariation.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Entropy</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-02-21-uncertainty-reveals-structure/</guid>
  <pubDate>Fri, 20 Feb 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/04-when-uncertainty-reveals-structure.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>Meaning Lives Between Things</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-02-07-meaning-lives-in-relationships/</link>
  <description><![CDATA[ 




<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/03-meaning-lives-in-relationships.webp" class="reading-sketch img-fluid figure-img" alt="Handwritten pencil notes emphasise relationships between sites, context-dependent probability and the difference between prediction and cause."></p>
<figcaption>A simple pencil notebook page connecting two sites in short DNA sequences, with a rough base-pair sketch.</figcaption>
</figure>
</div>
<p>This fortnight the chapter moved from individual variables to relationships. That shift felt important. A nucleotide can be common, rare, conserved or variable on its own, but biological structure often appears only when I ask how one position changes with another.</p>
<p>The central idea is simple: context changes probability.</p>
<section id="joint-probability-seeing-states-together" class="level2">
<h2 class="anchored" data-anchor-id="joint-probability-seeing-states-together">Joint probability: seeing states together</h2>
<p>Two variables can be combined into a joint variable. For nucleotide sites, this means counting pairs such as AA, AC or GT across an alignment. The joint distribution records how often two states occur together. By summing over one variable, we recover the marginal distribution of the other.</p>
<p>The language can sound abstract, but the biological question is direct: do two positions vary independently, or are some combinations favoured and others avoided?</p>
<p>If two coins are independent, learning the outcome of one does not change what I expect from the other. If their outcomes are linked, that knowledge changes the probability. The same logic applies to two sites in a molecule.</p>
</section>
<section id="conditional-probability-what-changes-after-i-know-more" class="level2">
<h2 class="anchored" data-anchor-id="conditional-probability-what-changes-after-i-know-more">Conditional probability: what changes after I know more?</h2>
<p>Conditional probability asks for the distribution of one variable after the state of another is known. It is the formal version of an “if-then” question. Bayes’ theorem relates alternative directions of that conditioning.</p>
<p>What I found especially useful is the chapter’s caution that conditional dependence does not by itself establish causation. A site can predict another because of physical pairing, shared ancestry, selection, population structure or an unmeasured third factor. Prediction is scientifically valuable, but explanation requires more evidence.</p>
<p>In the tRNA alignment, knowing the nucleotide at one site changes the expected distribution at another. For one pair the prediction is partial. For another pair, the relationship is almost deterministic and reflects complementary base pairing. Sequence statistics therefore reveal something not obvious from either site considered alone: a structural relationship in the molecule.</p>
</section>
<section id="from-letters-to-organisation" class="level2">
<h2 class="anchored" data-anchor-id="from-letters-to-organisation">From letters to organisation</h2>
<p>This example changed the way I read an alignment. I usually see rows of characters, conserved columns and substitutions. Conditional probability asks me to look across columns for coordinated change. Two positions may each appear variable, yet their variation together can be tightly constrained.</p>
<p>That is a powerful idea. Biological function does not require every component to remain unchanged. Sometimes function is preserved because components change together.</p>
<p>It also clarifies why simple conservation scores are incomplete. A position may tolerate several states, but only in combination with compatible states elsewhere. The information is relational.</p>
<section id="notes-from-my-margin" class="level3 pencil-margin-note">
<h3 class="anchored" data-anchor-id="notes-from-my-margin">Notes from my margin</h3>
<ul>
<li>An isolated value may look noisy; a relationship can reveal a rule.</li>
<li>Conditional probability improves prediction, not automatically explanation.</li>
<li>Co-variation can preserve structure while individual sites change.</li>
</ul>
</section>
</section>
<section id="my-interpretation-for-language-models" class="level2">
<h2 class="anchored" data-anchor-id="my-interpretation-for-language-models">My interpretation for language models</h2>
<p>DNA language models are built around context. They estimate how plausible a nucleotide or sequence is given surrounding sequence. In that sense, their power comes from learning many conditional relationships.</p>
<p>But a learned relationship is not automatically a biological mechanism. A model may capture motif grammar, compositional bias, phylogenetic history, annotation patterns or technical artefacts. A strong score tells me that a substitution changes what the model expects. It does not yet tell me why, whether the change affects fitness, or whether the relationship transfers to another lineage.</p>
<p>This gives me a sharper question for PopGenLM Bench: when a model says that one allele is less compatible with its context, does that score improve prediction of independent biological evidence? For example, is it related to allele frequency, conservation, functional annotation or experimentally observed effect after accounting for confounders?</p>
<p>The benchmark should therefore compare conditional predictions with external outcomes, not merely display model confidence.</p>
</section>
<section id="additional-learning-directions" class="level2">
<h2 class="anchored" data-anchor-id="additional-learning-directions">Additional learning directions</h2>
<p>My next steps are to explore:</p>
<ol type="1">
<li>mutual information as a quantitative measure of shared dependence;</li>
<li>phylogenetic correction, because related sequences are not independent samples;</li>
<li>direct-coupling methods that try to distinguish direct from indirect sequence relationships;</li>
<li>causal diagrams for separating prediction from mechanism; and</li>
<li>negative controls that reveal whether a model is using biologically meaningful context or simpler background patterns.</li>
</ol>
</section>
<section id="my-commentary" class="level2">
<h2 class="anchored" data-anchor-id="my-commentary">My commentary</h2>
<p>The phrase I wrote in my notes was “meaning lives between things.” It is not a definition from the chapter; it is my own interpretation. A base, a gene or a model score becomes informative through a relationship to something else. The lesson is both technical and philosophical: connection can carry structure that isolated objects do not show.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em>, Chapter 2, the later part of Section 2.1 on joint, marginal and conditional probabilities, Bayes’ theorem, independence and conditional relationships between tRNA sites.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Conditional probability</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-02-07-meaning-lives-in-relationships/</guid>
  <pubDate>Fri, 06 Feb 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/03-meaning-lives-in-relationships.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>The Model Is Not the World</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-01-24-the-model-is-a-choice/</link>
  <description><![CDATA[ 




<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/02-the-model-is-a-choice.webp" class="reading-sketch img-fluid figure-img" alt="Handwritten pencil notes say that a model is not the world, states are chosen and an ensemble gives context."></p>
<figcaption>A simple pencil notebook page with a coin, four nucleotide states and a small sequence ensemble.</figcaption>
</figure>
</div>
<p>During the last two weeks I worked through Section 2.1 on random variables and probabilities. The mathematics begins with familiar objects—a coin, possible outcomes and their frequencies—but the biological lesson is deeper: before measuring uncertainty, I must decide what counts as a possible state.</p>
<p>That decision is part of the model. It is not given automatically by nature.</p>
<section id="random-variables-are-descriptions" class="level2">
<h2 class="anchored" data-anchor-id="random-variables-are-descriptions">Random variables are descriptions</h2>
<p>A discrete random variable has a finite set of states, often called an alphabet, and a probability for each state. A coin can be represented by two states; a nucleotide site by A, C, G and T; an amino-acid site by twenty possibilities.</p>
<p>The chapter insists on a distinction I find extremely useful: the random variable is not the physical object. A real coin has material, orientation, wear and many possible behaviours. The two-state model keeps only the distinction needed for a particular question. Even “heads” and “tails” are choices made by an observer using a particular measurement.</p>
<p>The same is true in genomics. A site may be represented as four nucleotides, but an analysis may also need a gap, an unresolved base, methylation state, genotype uncertainty or structural context. The alphabet depends on the biological question and the measurement system. If I change that representation, I may change the probabilities and therefore the information I calculate.</p>
<p>This is not a weakness of modelling. It is why transparent modelling matters.</p>
</section>
<section id="probabilities-come-from-finite-observations" class="level2">
<h2 class="anchored" data-anchor-id="probabilities-come-from-finite-observations">Probabilities come from finite observations</h2>
<p>For a physical system, probabilities are usually estimated from repeated observations. With a finite sample, frequency is only an approximation to the underlying probability. This obvious fact becomes easy to forget when software prints many decimal places.</p>
<p>The chapter illustrates several ways to treat a DNA sequence. A complete sequence of length (L) can be one state among (4^L) possibilities. Alternatively, its (L) positions can be treated as repeated observations of one nucleotide variable, allowing an estimate of overall base composition. A third approach treats each position as its own variable and uses an alignment of related sequences to estimate site-specific probabilities.</p>
<p>These views use the same letters but answer different questions.</p>
<p>Overall nucleotide frequencies can reveal GC or AT bias. Such bias may have biochemical consequences, but it does not have a single guaranteed evolutionary cause. A site-wise ensemble adds another layer: strongly conserved positions may indicate functional constraint, whereas variable positions may tolerate change. Yet this interpretation depends on alignment quality, sampling and evolutionary history.</p>
</section>
<section id="why-the-ensemble-matters" class="level2">
<h2 class="anchored" data-anchor-id="why-the-ensemble-matters">Why the ensemble matters</h2>
<p>An ensemble is the collection from which observations are drawn. For a genomic site, an alignment of homologous sequences can provide that collection. Across the alignment, each position becomes a distribution rather than a single letter.</p>
<p>The tRNA example in the chapter makes this concrete. Some positions vary, while others are nearly or completely conserved because the molecule’s structure and function constrain acceptable changes. Gaps and unresolved characters must also be treated explicitly; ignoring how they enter the alphabet can alter the analysis.</p>
<p>For me, this is the bridge from a sequence as an object to a sequence as evidence. One string shows what occurred once. An ensemble begins to show what could vary, what repeatedly remains, and how uncertain those conclusions are.</p>
<section id="notes-from-my-margin" class="level3 pencil-margin-note">
<h3 class="anchored" data-anchor-id="notes-from-my-margin">Notes from my margin</h3>
<ul>
<li>Every dataset has an alphabet, even when it is not written down.</li>
<li>More decimal places do not create more observations.</li>
<li>A single sequence is a record; an ensemble makes variation measurable.</li>
</ul>
</section>
</section>
<section id="my-interpretation-for-genomic-ai" class="level2">
<h2 class="anchored" data-anchor-id="my-interpretation-for-genomic-ai">My interpretation for genomic AI</h2>
<p>This section immediately changes how I think about model inputs. Tokenisation, genome build, strand orientation, reference allele, treatment of missing bases and sampling population are not preliminary details. They define the states the system is allowed to see.</p>
<p>A DNA language model also learns from an ensemble—its training data. If that ensemble represents some species, regions or sequence contexts much better than others, the model’s apparent confidence may exceed the evidence available for a new organism. Before evaluating a score, I therefore need to ask: what alphabet did the model use, what ensemble shaped its probabilities, and how closely does my new dataset match that world?</p>
<p>This is one reason PopGenLM Bench must begin with reference validation and provenance rather than a plot. A biologically incorrect representation can pass through perfectly functioning code.</p>
</section>
<section id="where-i-want-to-learn-next" class="level2">
<h2 class="anchored" data-anchor-id="where-i-want-to-learn-next">Where I want to learn next</h2>
<p>I want to connect this section to practical methods for:</p>
<ul>
<li>quantifying finite-sample uncertainty in nucleotide frequencies;</li>
<li>weighting related sequences so that a large clade is not mistaken for many independent observations;</li>
<li>handling gaps and ambiguous states in alignments;</li>
<li>comparing frequentist estimates with Bayesian probability models; and</li>
<li>measuring how training-set composition affects transfer across species.</li>
</ul>
</section>
<section id="my-commentary" class="level2">
<h2 class="anchored" data-anchor-id="my-commentary">My commentary</h2>
<p>The most thought-provoking lesson was not a formula. It was the reminder that measurement begins with an agreement about distinctions. Biology is richer than any alphabet we impose on it. Good science does not pretend otherwise; it states the simplification, tests its consequences and changes it when the question demands more.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em>, Chapter 2, Section 2.1, from random variables and finite probability estimates through nucleotide composition, site-specific variables and the tRNA alignment.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>Probability</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-01-24-the-model-is-a-choice/</guid>
  <pubDate>Fri, 23 Jan 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/02-the-model-is-a-choice.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>When Life Became Legible</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-01-10-inheritance-becomes-information/</link>
  <description><![CDATA[ 




<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/01-inheritance-becomes-information.webp" class="reading-sketch img-fluid figure-img" alt="Handwritten pencil notes connect inheritance, copying and new scientific questions."></p>
<figcaption>A simple pencil notebook sketch linking a pea pod, chromosome and DNA, with notes about inheritance and information.</figcaption>
</figure>
</div>
<p>This fortnight I began Chapter 2, “Information Theory in Biology,” from Christoph Adami’s <em>The Evolution of Biological Information</em>. I expected a mathematical introduction. What surprised me first was the history. Before information could be measured, biologists had to learn how to see heredity as something separable, locatable, copyable and changeable.</p>
<p>That history made the chapter feel immediately connected to my own journey. Genomics can appear to begin with sequencing machines and large files, but its deeper foundations were built through a sequence of conceptual steps. Each step changed not only what scientists knew, but what they believed could be asked.</p>
<section id="what-i-learned-from-the-opening-section" class="level2">
<h2 class="anchored" data-anchor-id="what-i-learned-from-the-opening-section">What I learned from the opening section</h2>
<p>The chapter begins with a strong reductionist claim: life on Earth is organised through information. This is not simply the familiar metaphor of DNA as a “book.” The important point is operational. DNA can be copied, altered and inherited. Its molecular structure links continuity with variation, so the same mechanism that preserves biological organisation also makes evolution possible.</p>
<p>This changed how I thought about the word <em>information</em>. In ordinary conversation, information is often treated as meaningful content. In this chapter, the more useful starting point is prediction. A biological pattern carries information when knowing it improves what we can predict about another state: a nucleotide, a structure, a function or an environment.</p>
<p>The chapter then asks deceptively simple questions. How much information is present in a genome? How much is shared by two organisms? How much is gained through adaptation, transmitted between generations or lost through extinction? These questions are easy to state and very difficult to answer because they require an explicit comparison, a defined set of possible states and enough observations to estimate probabilities.</p>
</section>
<section id="the-historical-path-from-character-to-code" class="level2">
<h2 class="anchored" data-anchor-id="the-historical-path-from-character-to-code">The historical path: from character to code</h2>
<p>Box 2.1 traces the emergence of genetic information through several distinct advances.</p>
<p>Mendel showed that inherited characters behave as if discrete factors pass between generations. Bateson and Saunders helped establish that these units can occupy alternative states—what we now call alleles. Johannsen separated the inherited unit from its visible expression and gave us the term <em>gene</em>. Morgan, Sturtevant and Bridges then placed genes in a linear order on chromosomes, turning heredity into something that could be mapped.</p>
<p>The next steps connected location with mechanism. Mutations showed that changing a chromosomal position could alter or destroy function. Work on gene action linked hereditary material to protein production. The genetic code then made the relationship between nucleotide sequence and amino-acid sequence experimentally legible. By this point, biological molecules could be treated as messages drawn from a space of possible messages—a perspective that made Shannon’s mathematical theory relevant to biology.</p>
<p>What I find most important is that no single discovery created “biological information.” The idea became possible only after inheritance, physical location, copying, mutation, expression and coding had been separated and then reconnected.</p>
<section id="notes-from-my-margin" class="level3 pencil-margin-note">
<h3 class="anchored" data-anchor-id="notes-from-my-margin">Notes from my margin</h3>
<ul>
<li>A new instrument is valuable when it opens a better question.</li>
<li>Copying explains continuity; mutation makes history possible.</li>
<li>“Information” becomes scientific only after I say what it predicts.</li>
</ul>
</section>
</section>
<section id="my-interpretation" class="level2">
<h2 class="anchored" data-anchor-id="my-interpretation">My interpretation</h2>
<p>The history in this section is a warning against treating scientific AI as a sudden replacement for earlier biology. AI belongs to the same longer progression. Natural history organised variation; genetics formalised inheritance; molecular biology identified a physical code; sequencing made that code observable at scale; bioinformatics made large collections analysable. AI may help us detect structure across those collections, but it does not remove the need to define the biological states, comparisons and evidence.</p>
<p>This matters because an impressive prediction can still be scientifically empty. If I cannot say what a model output is information <em>about</em>, which observations support it, or under what conditions it fails, then I have a number rather than knowledge.</p>
<p>I also see a connection with temporal genomics. A genome sampled at one date is a record, but a series of genomes sampled through time lets us ask what was retained, altered or lost. The value lies not in the sequence alone; it lies in the comparison that makes change visible.</p>
</section>
<section id="where-i-want-to-learn-next" class="level2">
<h2 class="anchored" data-anchor-id="where-i-want-to-learn-next">Where I want to learn next</h2>
<p>My next steps are to understand three transitions more deeply:</p>
<ol type="1">
<li>How probability turns biological variation into a measurable ensemble.</li>
<li>How Shannon entropy separates uncertainty from information.</li>
<li>How evolutionary processes create the correlations that allow genomes to carry evidence about environments.</li>
</ol>
<p>I also want to revisit the history of the genetic code and chromosome mapping—not as a list of discoveries, but as examples of how measurement changes explanation.</p>
</section>
<section id="a-simple-commentary-on-the-chapter-so-far" class="level2">
<h2 class="anchored" data-anchor-id="a-simple-commentary-on-the-chapter-so-far">A simple commentary on the chapter so far</h2>
<p>Adami’s presentation is mathematical, but the opening argument is philosophical in the best scientific sense. It asks what we mean when we say that life contains information, and then refuses to leave the term vague. My main lesson is that biological information is not a substance hidden inside DNA. It is a measurable relationship made visible by a carefully defined question.</p>
<p>That is a demanding standard. It is also a useful one for the AI systems I want to build.</p>
<div class="reading-source">
<p><strong>Reading for this note:</strong> Christoph Adami, <em>The Evolution of Biological Information</em>, Chapter 2, opening discussion and Box 2.1, “The Emergence of the Concept of Genetic Information.” This article is my own synthesis and commentary on the supplied chapter excerpt.</p>
</div>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Biological information</category>
  <category>History of genetics</category>
  <category>Learning note</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-01-10-inheritance-becomes-information/</guid>
  <pubDate>Fri, 09 Jan 2026 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/01-inheritance-becomes-information.webp" medium="image" type="image/webp"/>
</item>
<item>
  <title>What does evolution need in order to work?</title>
  <dc:creator>Dr Tahir Ali</dc:creator>
  <link>https://tahirali-biomics.github.io/posts/2026-01-01-what-does-evolution-need-to-work/</link>
  <description><![CDATA[ 




<p><em>Reading Christoph Adami, <strong>The Evolution of Biological Information</strong>, Chapter 1.</em></p>
<p>I thought Chapter 1 would feel like familiar ground.</p>
<p>Inheritance. Mutation. Selection. Species. Darwin.</p>
<p>The vocabulary is so deeply embedded in biology that I almost expected to move through the chapter quickly, collecting definitions on the way to the more difficult material.</p>
<p>Instead, I got stuck on a deceptively simple question:</p>
<blockquote class="blockquote">
<p><strong>What has to be true before evolution can do anything at all?</strong></p>
</blockquote>
<p>Not before adaptation becomes impressive.</p>
<p>Not before complexity appears.</p>
<p>Before evolution can even begin to accumulate a history.</p>
<p>The answer sounds almost too elementary:</p>
<p><strong>something must be copied, something must vary, and those differences must matter for what gets copied next.</strong></p>
<p>I knew all three pieces.</p>
<p>What I had not been thinking about carefully enough was why none of them is evolution on its own.</p>
<p>That became the thread I followed through the chapter.</p>
<section id="copying-is-more-interesting-than-inheritance" class="level2">
<h2 class="anchored" data-anchor-id="copying-is-more-interesting-than-inheritance">Copying is more interesting than inheritance</h2>
<p>We usually introduce inheritance from the outside.</p>
<p>Offspring resemble parents.</p>
<p>Traits run in families.</p>
<p>Characters persist across generations.</p>
<p>But Adami immediately pushes the idea one level deeper. The mechanistic core of inheritance is <strong>replication</strong>—and, more abstractly still, the copying of information.</p>
<p>That small shift changed the picture for me.</p>
<p>The evolutionary object is not simply a visible trait being passed forward. Something encoded has to survive the transition from one generation to the next with enough fidelity that a lineage has continuity.</p>
<p>So copying is the conservative side of evolution.</p>
<p>It keeps yesterday from disappearing.</p>
<p>And then the first paradox appears.</p>
<p>Perfect copying sounds ideal.</p>
<p>For preservation, it is.</p>
<p>For evolution, it is a dead end.</p>
<p>If every copy were exact forever, the population could preserve what it already knows, but it could never discover anything new.</p>
<p>I wrote in the margin:</p>
<blockquote class="blockquote">
<p><strong>perfect memory, no exploration</strong></p>
</blockquote>
<p>That feels like a useful way to think about inheritance.</p>
<p>Evolution needs memory.</p>
<p>But not perfect memory.</p>
</section>
<section id="mutation-is-not-only-a-mistake" class="level2">
<h2 class="anchored" data-anchor-id="mutation-is-not-only-a-mistake">Mutation is not only a mistake</h2>
<p>The natural next move is to call mutation an error.</p>
<p>That is physically reasonable. Replication happens in a noisy material world. Molecules are copied by imperfect machinery. Bases are substituted. Segments are inserted, deleted, duplicated, inverted or shuffled. Recombination creates still more combinations.</p>
<p>From the perspective of the copied sequence, these are departures from fidelity.</p>
<p>But from the perspective of evolution, something more interesting is happening.</p>
<p>Mutation creates alternatives.</p>
<p>And without alternatives, selection has nothing to compare.</p>
<p>This made me notice an asymmetry that I had stopped seeing because it is so familiar:</p>
<p><strong>replication protects existing information; mutation risks it.</strong></p>
<p>Yet evolution needs both.</p>
<p>Too much instability and the lineage cannot retain successful solutions.</p>
<p>Too much fidelity and it cannot search beyond them.</p>
<p>That is already a much richer picture than “mutation introduces variation”.</p>
<p>The genome is being held between two requirements:</p>
<p><strong>remember enough to persist; change enough to explore.</strong></p>
<p>I suspect I will keep coming back to that tension throughout the book.</p>
</section>
<section id="but-variation-alone-knows-nothing-about-where-to-go" class="level2">
<h2 class="anchored" data-anchor-id="but-variation-alone-knows-nothing-about-where-to-go">But variation alone knows nothing about where to go</h2>
<p>At this point another question appeared.</p>
<p>If mutation creates possibilities, does mutation create adaptation?</p>
<p>No.</p>
<p>Variation is blind to whether a new state is useful.</p>
<p>A mutation can be beneficial, neutral, deleterious or lethal only relative to a biological context in which its consequences are expressed.</p>
<p>So the mutation itself is not the answer.</p>
<p>It is a proposal.</p>
<p>The environment—and the population living in it—determines what happens next.</p>
<p>That makes the familiar phrase <em>variation meets selection</em> feel slightly misleading to me now.</p>
<p>What variation really meets is a <strong>world</strong>.</p>
<p>And only through that encounter does a difference acquire an evolutionary consequence.</p>
</section>
<section id="fitness-is-not-who-happened-to-survive" class="level2">
<h2 class="anchored" data-anchor-id="fitness-is-not-who-happened-to-survive">Fitness is not “who happened to survive”</h2>
<p>The chapter then pauses over fitness, and this was more useful than I expected.</p>
<p>“Survival of the fittest” is one of those phrases that becomes less helpful the more casually it is repeated.</p>
<p>If the survivor is defined as the fittest and the fittest is defined as the survivor, the phrase collapses into a circle.</p>
<p>But that is not how fitness works.</p>
<p>Fitness is an expectation about reproductive success associated with a lineage or genotype—not a guarantee about the fate of one individual.</p>
<p>That distinction matters because chance never disappears.</p>
<p>A member of a high-fitness lineage can die before reproducing.</p>
<p>A member of a lower-fitness lineage can get lucky.</p>
<p>Realized survival is noisy.</p>
<p>Expected reproductive success is the population-level quantity selection acts through.</p>
<p>I found myself thinking of it this way:</p>
<blockquote class="blockquote">
<p><strong>fitness is a tendency; survival is an outcome</strong></p>
</blockquote>
<p>They are related, but they are not identical.</p>
<p>That one distinction already protects us from a surprising amount of sloppy evolutionary reasoning.</p>
<p>It also makes selection less mysterious.</p>
<p>Selection does not need foresight.</p>
<p>It does not inspect organisms and reward the “best”.</p>
<p>If heritable differences systematically alter expected reproductive success, frequencies change.</p>
<p>That is enough.</p>
</section>
<section id="then-the-chapter-performs-the-cleanest-possible-experiment" class="level2">
<h2 class="anchored" data-anchor-id="then-the-chapter-performs-the-cleanest-possible-experiment">Then the chapter performs the cleanest possible experiment</h2>
<p>This was probably my favourite part of Chapter 1.</p>
<p>Adami constructs an intentionally simple population of sequences on an intentionally simple fitness landscape.</p>
<p>The model is not trying to imitate a real organism.</p>
<p>That is exactly why it is useful.</p>
<p>It lets us ask a sharper question:</p>
<blockquote class="blockquote">
<p><strong>Which pieces of the Darwinian mechanism are actually necessary?</strong></p>
</blockquote>
<p>With replication, mutation and selection all operating, the population can move through sequence space toward the defined optimum.</p>
<p>Then one ingredient is removed at a time.</p>
<section id="no-mutation" class="level3">
<h3 class="anchored" data-anchor-id="no-mutation">No mutation</h3>
<p>There is still selection.</p>
<p>There is still reproduction.</p>
<p>The best sequence already present can spread.</p>
<p>Mean fitness rises for a while.</p>
<p>But once the initial variation has been sorted, nothing genuinely new can appear.</p>
<p>The population has climbed as far as its starting material allows.</p>
<p>Then evolution stalls.</p>
<p>Selection can amplify an existing answer.</p>
<p>It cannot invent the next one.</p>
</section>
<section id="no-reproduction" class="level3">
<h3 class="anchored" data-anchor-id="no-reproduction">No reproduction</h3>
<p>Mutation still produces new sequences.</p>
<p>A fitness ranking still exists.</p>
<p>But there is no mechanism by which successful sequences leave more descendants.</p>
<p>The system can wander.</p>
<p>It cannot build a lineage of accumulated improvements.</p>
<p>A discovery that is not preferentially copied is evolutionarily ephemeral.</p>
</section>
<section id="no-selection" class="level3">
<h3 class="anchored" data-anchor-id="no-selection">No selection</h3>
<p>Replication continues.</p>
<p>Mutation continues.</p>
<p>The population still changes.</p>
<p>Some variants can even become common by chance.</p>
<p>But the relationship between sequence and the defined environment no longer directs the change.</p>
<p>There is motion without adaptive direction.</p>
<p>And suddenly the Darwinian triad stops feeling like three items in a textbook list.</p>
<p>It feels like a machine.</p>
<p>Remove one gear and the behaviour changes qualitatively.</p>
<p>I wrote:</p>
<blockquote class="blockquote">
<p><strong>copying preserves<br>
mutation opens<br>
selection filters</strong></p>
</blockquote>
<p>But even that is slightly incomplete.</p>
<p>The real point is that the three only become evolution <strong>together</strong>.</p>
<div class="quarto-figure quarto-figure-center">
<figure class="figure">
<p><img src="https://tahirali-biomics.github.io/assets/writings/00-what-does-evolution-need-to-work.webp" class="img-fluid figure-img" alt="Graphite notebook-style reading notes on lined paper, with hand-drawn DNA, evolutionary branching, questions about measurement, information, fitness, selection, and the distinction between passive and driven evolutionary trends."></p>
<figcaption>My Chapter 1 reading notebook. I kept returning to the same question: what exactly has to meet before evolutionary change can accumulate into a history?</figcaption>
</figure>
</div>
</section>
</section>
<section id="a-strange-detail-in-the-toy-model-points-toward-the-rest-of-the-book" class="level2">
<h2 class="anchored" data-anchor-id="a-strange-detail-in-the-toy-model-points-toward-the-rest-of-the-book">A strange detail in the toy model points toward the rest of the book</h2>
<p>The simulation is deliberately simple enough that mutations contribute independently to fitness.</p>
<p>That means the order in which beneficial changes occur does not matter very much.</p>
<p>Real biology is not like that.</p>
<p>The effect of one mutation can depend on what is already present elsewhere in the genome.</p>
<p>Adami flags this immediately: once mutations interact—once <strong>epistasis</strong> enters—the landscape becomes enormously richer.</p>
<p>I liked that the chapter tells us, in effect:</p>
<p><em>Here is evolution with much of the interesting difficulty temporarily removed.</em></p>
<p>That is good modelling.</p>
<p>Strip the system down until the mechanism becomes visible.</p>
<p>Then put the complexity back.</p>
<p>It also made me think about how easily we confuse a useful simplification with a biological claim.</p>
<p>A model can be intentionally unrealistic and still teach us something true.</p>
<p>The danger begins when we forget which part was deliberately left out.</p>
<p>That thought will matter later when I get to genomic AI.</p>
</section>
<section id="species-are-not-another-ingredient" class="level2">
<h2 class="anchored" data-anchor-id="species-are-not-another-ingredient">Species are not another ingredient</h2>
<p>I had another small conceptual correction when the chapter turned to speciation.</p>
<p>Species are central to Darwin’s title.</p>
<p>But speciation is not a fourth element added beside inheritance, variation and selection.</p>
<p>It is an <strong>outcome</strong> that can emerge when those processes operate through time in structured populations.</p>
<p>That distinction is important.</p>
<p>A population can be split geographically.</p>
<p>Gene flow can be interrupted.</p>
<p>Independent histories accumulate.</p>
<p>Eventually the difference can become biological rather than merely geographic.</p>
<p>Or populations can begin diverging while occupying the same broader region, particularly when ecological specialization and mate choice reduce gene exchange.</p>
<p>For asexual organisms the details change again.</p>
<p>And ecology becomes unavoidable.</p>
<p>If every beneficial mutation in an asexual population simply replaced what came before, why would so many lineages coexist?</p>
<p>Because organisms need not compete for exactly the same resource in exactly the same way.</p>
<p>Different niches can sustain different solutions.</p>
<p>So by the time the chapter reaches species, the neat three-part mechanism has already opened into something larger:</p>
<p><strong>population structure, gene flow, competition, resources, frequency dependence, ecology.</strong></p>
<p>The Darwinian minimum is simple.</p>
<p>Its consequences are not.</p>
</section>
<section id="then-the-chapter-does-something-i-did-not-expect-it-asks-why-this-took-so-long-to-see" class="level2">
<h2 class="anchored" data-anchor-id="then-the-chapter-does-something-i-did-not-expect-it-asks-why-this-took-so-long-to-see">Then the chapter does something I did not expect: it asks why this took so long to see</h2>
<p>The historical half initially felt like a change of subject.</p>
<p>It was not.</p>
<p>In fact, it may be the part that made the first half more vivid.</p>
<p>Today, evolution is so conceptually normal that the Darwinian mechanism can feel almost inevitable.</p>
<p>But that feeling is an illusion produced by hindsight.</p>
<p>Before Darwin’s synthesis could become possible, the intellectual world itself had to change.</p>
<p>Life had to become something that could be classified systematically.</p>
<p>Extinction had to become believable.</p>
<p>The Earth had to become old.</p>
<p>Species had to become mutable.</p>
<p>Fossils had to become records of lost worlds rather than curiosities.</p>
<p>Geology had to show that enormous transformations could emerge from ordinary processes acting for immense periods of time.</p>
<p>Population growth had to be placed against finite resources.</p>
<p>And the old image of nature as a fixed ladder had to break.</p>
<p>That last point stayed with me.</p>
<p>A ladder gives every organism a preassigned place.</p>
<p>A tree gives every lineage a history.</p>
<p>Those are not just different diagrams.</p>
<p>They are different ways of thinking about life.</p>
</section>
<section id="the-people-before-darwin-were-not-simply-wrong" class="level2">
<h2 class="anchored" data-anchor-id="the-people-before-darwin-were-not-simply-wrong">The people before Darwin were not simply “wrong”</h2>
<p>I also liked that the history is messier than the usual compressed story.</p>
<p>Linnaeus helped make biological diversity legible through classification even while working within a largely fixed view of species.</p>
<p>De Maillet, Buffon and Erasmus Darwin entertained transformation and a much more dynamic natural world.</p>
<p>Malthus sharpened the consequence of reproduction under limited resources.</p>
<p>Lamarck was wrong about the inheritance mechanism he emphasized, but he was not wrong to take species change seriously.</p>
<p>Cuvier’s comparative anatomy and fossils made extinction difficult to ignore even though he did not embrace Darwinian transformation.</p>
<p>Lyell supplied deep geological time and the power of ordinary processes.</p>
<p>Wallace independently reached natural selection from observations of variation, geography and resources.</p>
<p>Seen this way, Darwin did not walk into an empty room and switch on the light.</p>
<p>The room had been filling with pieces for a long time.</p>
<p>His achievement was to connect them into a mechanism powerful enough to reorganize biology.</p>
<p>That gave me another margin note:</p>
<blockquote class="blockquote">
<p><strong>sometimes the breakthrough is not a new fact—it is a new relationship among facts already present</strong></p>
</blockquote>
<p>That sentence feels very relevant to scientific AI too.</p>
</section>
<section id="a-deeper-question-what-exactly-is-the-environment-writing-into-the-genome" class="level2">
<h2 class="anchored" data-anchor-id="a-deeper-question-what-exactly-is-the-environment-writing-into-the-genome">A deeper question: what exactly is the environment writing into the genome?</h2>
<p>Near the end of the mechanistic discussion, Adami makes a move that points toward the rest of the book.</p>
<p>A well-adapted organism can be thought of as carrying information about its environment in its genes.</p>
<p>I find that formulation both exciting and dangerous.</p>
<p>Exciting, because it gives adaptation a quantitative direction.</p>
<p>Dangerous, because it is easy to turn “information about the environment” into a metaphor and stop there.</p>
<p>So I found myself asking:</p>
<p><strong>What exactly has been written?</strong></p>
<p>A genome does not contain a literal description of temperature, predators, nutrients or mates.</p>
<p>What it contains are sequence states whose consequences have repeatedly survived encounters with those conditions.</p>
<p>That is a very different kind of knowledge.</p>
<p>It is knowledge encoded through differential persistence.</p>
<p>The genome does not <em>describe</em> the environment.</p>
<p>It bears traces of what has worked in it.</p>
<p>And those traces are always conditional on history.</p>
<p>Change the environment and yesterday’s information can become irrelevant—or harmful.</p>
<p>That makes biological information feel less like a stored encyclopedia and more like a compressed record of successful interactions.</p>
</section>
<section id="this-is-where-i-began-thinking-about-genomic-language-models" class="level2">
<h2 class="anchored" data-anchor-id="this-is-where-i-began-thinking-about-genomic-language-models">This is where I began thinking about genomic language models</h2>
<p>Up to this point I had been reading as an evolutionary biologist.</p>
<p>Then the AI connection became difficult to ignore.</p>
<p>A genomic language model is trained on sequence.</p>
<p>But the sequence corpus itself is not raw nature.</p>
<p>It is already an evolutionary product.</p>
<p>Every extant genome is a survivor of copying, mutation, selection, drift, demography, extinction and historical contingency.</p>
<p>That means the model is not merely learning “DNA”.</p>
<p>It is learning patterns from an <strong>archive produced by evolution</strong>.</p>
<p>That sounds powerful.</p>
<p>It is.</p>
<p>But it creates a difficult interpretive problem.</p>
<p>A pattern can be present in the archive for many reasons.</p>
<p>Constraint.</p>
<p>Mutation bias.</p>
<p>GC-biased gene conversion.</p>
<p>Demographic history.</p>
<p>Linkage.</p>
<p>Recombination.</p>
<p>Phylogenetic relatedness.</p>
<p>Selection.</p>
<p>Chance.</p>
<p>The model may learn the pattern beautifully without knowing which process produced it.</p>
<p>And that led me to a question I want to carry into PopGenLM Bench:</p>
<blockquote class="blockquote">
<p><strong>Can a model learn the evolutionary trace without learning the evolutionary cause?</strong></p>
</blockquote>
<p>Almost certainly, yes.</p>
<p>That is not a criticism.</p>
<p>It is a reminder to be precise about what a sequence model score means.</p>
</section>
<section id="a-high-confidence-prediction-is-not-yet-an-evolutionary-explanation" class="level2">
<h2 class="anchored" data-anchor-id="a-high-confidence-prediction-is-not-yet-an-evolutionary-explanation">A high-confidence prediction is not yet an evolutionary explanation</h2>
<p>Suppose a model strongly prefers one allele over another.</p>
<p>That may be useful.</p>
<p>But what does the preference correspond to?</p>
<p>Is the alternative allele rare because it disrupts molecular function?</p>
<p>Because the local sequence context is unusual?</p>
<p>Because related species share the reference state?</p>
<p>Because the region is mutation-poor?</p>
<p>Because selection has repeatedly removed similar changes?</p>
<p>Because the training data overrepresent one lineage?</p>
<p>The score itself does not answer those questions.</p>
<p>And Chapter 1 gives me a simple reason why.</p>
<p>Evolution is not a property of one sequence in isolation.</p>
<p>It is a process involving <strong>replication, variation and differential propagation through populations and environments</strong>.</p>
<p>A language model can learn the residue left by that process.</p>
<p>Whether it has learned the process itself is a much harder claim.</p>
<p>That distinction feels important enough to state plainly:</p>
<blockquote class="blockquote">
<p><strong>prediction from evolutionary output is not the same as modelling evolutionary dynamics</strong></p>
</blockquote>
<p>The two can be related.</p>
<p>They are not interchangeable.</p>
</section>
<section id="maybe-the-right-comparison-is-not-genome-versus-model-but-process-versus-representation" class="level2">
<h2 class="anchored" data-anchor-id="maybe-the-right-comparison-is-not-genome-versus-model-but-process-versus-representation">Maybe the right comparison is not genome versus model, but process versus representation</h2>
<p>This chapter also made me rethink what I want from an AI model in science.</p>
<p>I do not need a genomic language model to reproduce evolution internally in order for it to be useful.</p>
<p>That would be an unreasonable standard.</p>
<p>But I do need to know which aspects of evolutionary structure its representation preserves.</p>
<p>Does it recover constrained positions?</p>
<p>Does it capture dependencies among sites?</p>
<p>Do its variant scores track allele frequency?</p>
<p>Do its local sequence landscapes correspond to experimentally measured mutational tolerance?</p>
<p>Are its predictions stable across ancestry, species and genomic background?</p>
<p>Where does agreement break down?</p>
<p>Those are empirical questions.</p>
<p>And that is where the analogy with Adami’s toy simulation becomes useful.</p>
<p>The simulation is valuable because its assumptions are explicit and the consequences can be tested by removing components.</p>
<p>A scientific AI model should be treated with the same discipline.</p>
<p>What information was available to it?</p>
<p>What structure did it learn?</p>
<p>What biological evidence does that structure predict?</p>
<p>What disappears when context is removed?</p>
<p>What conclusion would falsify my interpretation of the score?</p>
<p>That is much more interesting to me than simply asking whether the model is “accurate”.</p>
</section>
<section id="perhaps-the-most-important-thing-chapter-1-removes-is-inevitability" class="level2">
<h2 class="anchored" data-anchor-id="perhaps-the-most-important-thing-chapter-1-removes-is-inevitability">Perhaps the most important thing Chapter 1 removes is inevitability</h2>
<p>There is another idea underneath the whole chapter that I nearly missed.</p>
<p>Evolution produces astonishing adaptation.</p>
<p>Once we see the finished organism, it is tempting to read the outcome backward as though it had been waiting to happen.</p>
<p>But the mechanism contains no such promise.</p>
<p>Variation has to arise.</p>
<p>It has to be heritable.</p>
<p>Its effect has to matter in that environment.</p>
<p>Selection has to be strong enough relative to stochastic forces.</p>
<p>The population has to persist long enough.</p>
<p>Historical paths constrain what becomes reachable next.</p>
<p>Species branch.</p>
<p>Some lineages disappear.</p>
<p>Many alternatives are never explored.</p>
<p>So the Darwinian mechanism is powerful without being prophetic.</p>
<p>It can generate extraordinary fit without containing a blueprint of the destination.</p>
<p>That is refreshing because it changes the emotional tone of adaptation.</p>
<p>The marvel is not that evolution knew where to go.</p>
<p>The marvel is that <strong>a process with no foresight can accumulate structure that looks as though it had one</strong>.</p>
</section>
<section id="the-sentence-i-am-carrying-out-of-chapter-1" class="level2">
<h2 class="anchored" data-anchor-id="the-sentence-i-am-carrying-out-of-chapter-1">The sentence I am carrying out of Chapter 1</h2>
<p>I started the chapter with three familiar words:</p>
<p>inheritance, variation, selection.</p>
<p>I am leaving it with a different picture.</p>
<p>A lineage needs enough fidelity to remember.</p>
<p>Enough error to explore.</p>
<p>And a world that makes some inherited differences matter more than others.</p>
<p>From that interaction, populations can accumulate adaptation.</p>
<p>From populations come divergence, ecological specialization and species.</p>
<p>Across deep time, those processes leave structured traces in genomes.</p>
<p>And those genomes are now becoming the training material for our models.</p>
<p>So the question I want to keep at the edge of every later chapter is not simply:</p>
<p><strong>what pattern is in the DNA?</strong></p>
<p>It is:</p>
<blockquote class="blockquote">
<p><strong>What evolutionary process could have written that pattern there—and what part of that history can a model actually recover?</strong></p>
</blockquote>
<p>That feels like a better starting point for <em>Genomes to AI</em>.</p>
<p>Before asking whether an AI can read the genome, Chapter 1 has made me ask something more fundamental:</p>
<p><strong>what had to happen before there was anything in the genome worth reading?</strong></p>


</section>

<a onclick="window.scrollTo(0, 0); return false;" id="quarto-back-to-top"><i class="bi bi-arrow-up"></i> Back to top</a> ]]></description>
  <category>Genomes to AI</category>
  <category>Evolution</category>
  <category>Population genetics</category>
  <category>Scientific AI</category>
  <guid>https://tahirali-biomics.github.io/posts/2026-01-01-what-does-evolution-need-to-work/</guid>
  <pubDate>Wed, 31 Dec 2025 23:00:00 GMT</pubDate>
  <media:content url="https://tahirali-biomics.github.io/assets/writings/00-what-does-evolution-need-to-work.webp" medium="image" type="image/webp"/>
</item>
</channel>
</rss>
