The extension from the Kingman process to the two-locus and continuous genome scenarios is the author’s own reasoning, informed by the existing Griffiths two-locus model and Hudson algorithm, though not intended to exactly replicate their derivations.
Kingman’s coalescence process
Given a sample of N haplotypes, exactly N-1 coalescence events occur over a continuous time interval. It is assumed that, within any given time interval, any pair of lineages has equal probability of coalescing. Under the limit of large N, at most one coalescence event occurs per infinitesimal time interval. Kingman proved that, as N goes to infinity, the density function $f_{Ti}$ (eq. 3.9 [1]) of each inter-coalescence interval is independent of the others. Furthermore, the fewer lineages remaining (i.e., the closer to the common ancestor), the longer the expected waiting time to the next coalescence, as determined by the lineage count in $\mathbf{E}(Ti)$ .
We will use algorithm 1 to build the diagram [2] below.

Algorithm 1
# Kingman process
def density_coal(i: int):
'''theotical density function of coal given number of lineages'''
...
return f
samples = {0, 1, 2, 3, 4}
# number of lineages
i = len(samples)
graph = ... # some proper object
while i > 1:
t = simulator(density_coal(i)) # simulate time need for a coal event, note in `graph`
choose 2, {s1, s2}, make `new_node`, note in `graph`
samples = (samples - {s1, s2}) | {new_node}
i -= 1
Two-locus model
We consider two loci in each sample, instead of one in Kingman’s. The recombination is the major concern.

The schematic diagram illustrates the reconstruction of an ARG, where the true underlying process is unobserved and the two per-locus trees are reconstructed individually.
In each per-locus diagram, the standard Kingman coalescent applies; the tree is therefore strictly bifurcating, as at most one pair of lineages can coalesce per infinitesimal time interval.
The ARG is derived by augmenting the two per-locus diagrams. Paths shared across both loci are denoted by a black line.
Note that: 1) internal nodes (the labelled ones), relative to the true process, appear in the ARG only when indegree and outdegree are unequal, that is, when a coalescence event has occurred$^*$. 2) The true number of recombination events is not guaranteed to be captured in the ARG.
The concept here is straightforward to be extended to a multi-locus scenario.
Continuous genome
So far, we have considered the two-locus case, in which each locus has its own coalescent history. We now generalise: each haplotype in the sample is treated as carrying an unknown number, $n_{i}$, of segments, not necessarily corresponding to discrete loci, each of variable length, $l_{ij}$. Allelic states can then be assigned to those segments across the sample. The parameters would be inferred as discussed later.
Inference
The sequence of tree topologies and branch lengths is generated by the coalescence and recombination process as a prior. SNPs are then applied to compute the likelihood of each tree.
Unlike in standard phylogenetics, the tree is constrained by the recombination rate governing transitions between adjacent locus trees, the effective population size, and related parameters. It is therefore not a random traversal of tree space, as phylogenetic packages typically perform.
Notes:
[1] Coalescent Theory: An Introduction, John Wakeley, 2008.
[2] diagram generated by msprime
$*$ The “Big ARG” formulation of Griffiths & Marjoram (1997) is not adopted here.
© 2026 Mike Lang. Licensed under CC BY-NC-ND 4.0. Cite as: Yapeng Mike Lang, “An algorithmic review of ARG and its inference”, yplang-blog