Identity by Descent: A Scientific Guide for DNA Matches

Identity by Descent: A Scientific Guide for DNA Matches

Identity by descent (IBD) is a shared DNA segment that two individuals inherited from a common ancestor through an unbroken chain of transmission, without intervening recombination breaking that segment apart. The practical takeaway is direct: the longer a shared segment, the more recently the two people share that ancestor; short segments often reflect distant population-level sharing rather than a genealogically meaningful connection.

When you see a match in your DNA results, three things deserve immediate attention:

  • Longest single segment length (in centimorgans, cM): the primary signal of recency
  • Total shared cM across all segments: reflects overall relationship depth
  • Testing platform: SNP-chip density and phasing quality affect which segments are reported at all

Key Takeaways

IBD segment length is the primary signal of genealogical recency: longer segments indicate more recent shared ancestry, while short segments often reflect population-level background sharing rather than a meaningful genealogical connection.

Point Details
Segment length signals recency Longer single segments (above ~15 cM) indicate more recent common ancestry; below 7 cM warrants skepticism.
IBD differs from IBS IBS means identical sequence at observed markers; IBD requires descent from a common ancestor without recombination.
Relationship ranges are probabilistic 185 total cM suggests second cousin to first cousin once removed, but variance, endogamy, and platform density all shift the estimate.
Algorithm choice affects detection BEAGLE, GERMLINE, IMPUTE, and Refined IBD differ in sensitivity, speed, and false-positive rate; platform method shapes which segments are reported.
Legal proof requires a formal test Consumer match reports are not court-admissible; chain-of-custody kinship tests are required for legal, inheritance, or medical proceedings.

Table of Contents

What is identity by descent, and how do IBD segments arise?

IBD segments are inherited haplotype stretches passed down through successive meioses, and recombination is the force that progressively fragments them across generations. Understanding that mechanism is the foundation for interpreting any shared-segment result.

During meiosis, homologous chromosomes exchange material at crossover points. Each crossover event can split a previously intact inherited segment. A grandparent passes a child roughly half of each chromosome, and that child passes roughly half of their own chromosomes to their children. Two first cousins, for example, each received a copy of their shared grandparent's chromosome, but each copy was already fragmented by one round of recombination in the parent generation. By the time the cousins' children compare DNA, another round of recombination has occurred, and the original grandparental segment is shorter still.

Diagram of chromosome crossover on desk

Coalescent theory formalizes this intuition: recombination causes expected IBD segment length to decline with generations, so the expected length of an IBD segment is inversely related to the number of generations separating two individuals from their common ancestor (source). More generations means more opportunities for recombination to cut the segment, so older shared ancestry produces shorter, more fragmented segments. This is why researchers treat segment length as a proxy for time to the most recent common ancestor.

Two important complications qualify the simple pedigree picture. Mutation and genotyping error can introduce apparent mismatches within a true IBD segment, causing detection algorithms to truncate or miss it. Endogamy, where a population has a history of intermarriage within a limited gene pool, inflates observed segment sharing beyond what a simple pedigree would predict, because multiple independent lines of descent converge on the same ancestral chromosomes. Both factors mean that raw segment length is a probabilistic signal, not a deterministic one.


IBD versus IBS: why the distinction shapes every match interpretation

IBS (identity by state) means two individuals carry identical alleles at observed marker positions; IBD means that identical stretch traces back to a specific common ancestor without recombination. The distinction matters because IBS does not require a genealogically recent relationship, while IBD does.

Abstract DNA strands illustrating shared and identical segments

Every IBD segment is also IBS at the genotyped markers within it, but the reverse is not true. A short stretch of identical alleles can arise by chance, particularly when those alleles are common in the population. Marker density, sequencing depth, and population history all affect the probability that an observed IBS alignment is actually IBD. A 3 cM match on a low-density SNP chip in a population with strong linkage disequilibrium (LD) may well be coincidental IBS, while a 25 cM match on a dense panel is generally considered more likely to be true IBD.

Phasing adds another layer. When genotype data is unphased, algorithms cannot always determine which alleles sit on the same chromosome copy, increasing the chance of spurious matches. Phased haplotype data, especially from whole-genome sequencing, is thought to substantially reduce this ambiguity.

Red flags that a reported match may be IBS rather than true recent IBD:

  • Segment length below 7 cM, particularly on a low-density chip
  • No shared matches with the proposed relative on other platforms
  • The matching region corresponds to a known high-LD or high-frequency population segment
  • Phasing uncertainty flagged by the detection software
  • Many small segments with no single dominant long segment

What segment length and count tell you about generational distance

Longer total shared cM and longer individual segments imply more recent common ancestry, but there is substantial variance even among relatives at the same degree. The coefficient of relationship provides the theoretical expected proportion of shared DNA, and the table below translates those expectations into approximate cM ranges used in genetic genealogy.

The pedigree formula for the coefficient of relationship is (1/2)^(2g − 1), where g is the number of generations of descent from the common ancestor. First cousins share approximately 1/8 of their DNA by this formula; second cousins share approximately 1/32.

Statistic note: The ranges are qualitative heuristics from genetic genealogy practice and pedigree theory; actual shared DNA varies due to stochastic recombination and individual inheritance patterns.

Caveats that shift these interpretations:

  • Endogamy: relatives from endogamous populations (Ashkenazi Jewish, Finnish, certain island populations) often share more cM than the pedigree relationship alone would predict, because multiple ancestral lines converge
  • Variance at distant degrees: at third-cousin distance and beyond, the observed cM can be zero even for a documented genealogical relationship
  • Platform SNP density: a lower-density chip may miss short segments, artificially reducing total shared cM
  • Genealogical versus genetic ancestry: a documented ancestor 10–11 generations back may have contributed zero detectable DNA to a given descendant, because genetic inheritance diverges from genealogical ancestry over many generations

How IBD is detected: algorithms, phasing, and key software tools

IBD is detected by statistical algorithms that scan phased haplotypes or unphased genotypes for long shared haplotype stretches; the choice of method directly affects sensitivity and false-positive rate. Recent methodological advances have extended IBD detection from small pedigrees to datasets spanning hundreds of thousands of individuals.

Four algorithm families dominate current practice:

GERMLINE uses a seed-and-extend approach: it identifies short exact haplotype matches across individuals and extends them to find longer IBD segments. The method is fast and scales well to large datasets, making it a common choice for population-scale screens, though it can produce false positives at short segment lengths.

BEAGLE performs phasing and imputation using a hidden Markov model (HMM) built on a Li-Stephens framework. Phased haplotypes from BEAGLE feed downstream IBD detection, and BEAGLE itself includes IBD detection functionality. Its probabilistic framework handles missing data and low-density chips more gracefully than purely deterministic methods.

IMPUTE is primarily an imputation tool that increases effective marker density by inferring unobserved genotypes from a reference panel. Higher post-imputation density improves the ability of downstream IBD callers to distinguish true IBD from coincidental IBS, particularly for shorter segments.

Refined IBD is a post-processing refinement layer, typically applied after an initial BEAGLE IBD call, that trims segment endpoints to reduce false-positive extensions and improves boundary accuracy. It is especially useful when downstream analyses depend on precise segment coordinates.

Method Primary input Speed at scale Sensitivity for short segments Typical use case
GERMLINE Phased haplotypes High Moderate Population-scale pairwise IBD screens
BEAGLE Unphased or phased genotypes Moderate Moderate-high Phasing, imputation, and IBD detection in one pipeline
IMPUTE Unphased genotypes + reference panel Moderate Indirect (via density increase) Pre-processing to boost marker density before IBD calling
Refined IBD BEAGLE IBD output Fast (post-processing) High (boundary refinement) Improving segment-endpoint accuracy after initial IBD call

Phased sequence data and dense marker panels increase the ability to distinguish true IBD from IBS, while low-density chips and unphased data raise uncertainty. Whole-genome sequencing generally offers better resolution for very short segments, though it also introduces more complex phasing requirements.

Common detection limitations:

  • Phasing errors introduce false breaks in true IBD segments or false joins across non-IBD regions
  • Low marker density on older SNP chips reduces sensitivity for detection of segments below roughly 5–7 cM.
  • Population structure can create systematic IBS patterns that mimic IBD in certain genomic regions
  • Endogamy inflates apparent IBD by creating background haplotype sharing across the population

Why researchers and genealogists rely on IBD analysis

IBD serves three distinct application areas: genetic genealogy, IBD mapping for rare variants and disease, and population demographic inference. Each exploits the same underlying signal, shared haplotype length, but at different scales and with different analytical goals.

In genetic genealogy, IBD segment data helps identify candidate common ancestors by narrowing the generational window. A 150 cM match almost certainly represents a second-cousin-or-closer relationship; a 20 cM match could be a third cousin or a more distant relative. Genealogists combine segment data with shared-match networks and documentary records to triangulate the most probable ancestral line. For practitioners focused on identity through ancestry, IBD analysis within approximately the last ten generations is most genealogically informative, because segments from more distant ancestors are often too short to detect reliably.

Genealogy research workspace with DNA test kit elements

IBD mapping uses shared segments to localize rare disease variants in families or founder populations. If multiple affected individuals in a pedigree share a long IBD segment that unaffected relatives do not, the causal variant likely lies within that segment. This approach has been particularly productive in founder populations, where a limited number of founding chromosomes means that many carriers of a rare variant share a common haplotype.

At biobank scale, newer methods enable inference of multi-individual IBD clusters across hundreds of thousands of samples, supporting demographic analyses of migration, admixture, and population bottlenecks. When datasets are large enough, the distribution of IBD segment lengths across a population encodes information about effective population size through time.


How to interpret match-length thresholds and avoid false positives

Testing services set minimum segment-length thresholds to reduce false positives; very small segments are often IBS or population-level background sharing and should be treated with caution rather than as evidence of recent kinship. The ISOGG Wiki summarizes community-practical guidance: small matches are frequently not reliable evidence of a genealogically meaningful relationship, and many services filter out segments below a platform-specific threshold before reporting them.

Practical checklist when evaluating a borderline match:

  • Largest single segment: below 7 cM warrants skepticism; above 15 cM is generally more reliable
  • Total shared cM: a high total built from many tiny segments is less convincing than a similar total from a few long segments
  • Number of segments: many short segments with no dominant long one often indicate population LD rather than recent kinship
  • Shared matches: corroborating matches with the same proposed relative's known family members strengthen the case
  • Population frequency of the matching region: regions with high LD in a given population produce more coincidental IBS matches
  • Full identical regions (FIRs): FIRs indicate both parental chromosomes match and are generally found only among full siblings or very close relationships; multiple FIRs in a purported distant relative suggest endogamy

Pro Tip: When a match sits near your platform's reporting threshold, check whether the same individual appears as a shared match with relatives whose relationship you have already confirmed. Convergent evidence from multiple independent matches is a stronger signal than segment length alone.

When IBD evidence needs to support a legal claim, such as an inheritance dispute or a custody proceeding, a consumer DNA match is not sufficient. Court-admissible results require a formal chain-of-custody test conducted under accredited laboratory protocols; a consumer match report does not meet that standard.


A worked example: reasoning from segment data to a relationship estimate

Given a hypothetical dataset of 185 total shared cM, a longest single segment of 42 cM, and 9 total segments reported on a standard SNP-chip platform, the conservative relationship estimate is second cousin to first cousin once removed, with first cousin twice removed also plausible.

Here is the step-by-step reasoning:

  1. Locate the total cM in the relationship table. 185 cM falls within the first-cousin-once-removed range (~200–650 cM) at its lower end, and overlaps with the upper end of the second-cousin range (~40–360 cM). Neither range is exclusive.

  2. Check the longest single segment. A 42 cM segment is long enough to be confident this is true IBD rather than coincidental IBS. Segments above roughly 15 cM on a well-phased chip are generally reliable; 42 cM points toward a relationship no more distant than second cousin in most non-endogamous populations.

  3. Cross-check with the coefficient-of-relationship formula. 185 cM represents approximately 5.4% of the genome. The expected value for a first cousin once removed is ~6.25% and for a second cousin is ~3.125%. The observed value sits between these two, consistent with either relationship or with variance around either expected value.

  4. Account for variance and endogamy. Stochastic recombination means any two relatives at the same degree can show a wide range of observed cM. If either individual has ancestry from an endogamous population, the total cM may be inflated relative to the pedigree relationship.

  5. List alternative plausible relationships. Half-first cousin (~12.5% expected, but with high variance), first cousin twice removed (~3.125%), and half-second cousin are all consistent with 185 cM depending on population background and chance.

Pro Tip: This reasoning is probabilistic, not deterministic. A 185 cM match cannot be resolved to a single relationship without additional evidence: shared matches, documentary records, or a formal kinship test. If you need a definitive answer for legal, medical, or inheritance purposes, a kinship DNA test from a laboratory following chain-of-custody procedures is the appropriate next step.


Why science-first explanations reduce misinterpretation of DNA matches

The most common mistake people make when reading DNA match results is treating a short shared segment as proof of a specific genealogical relationship. A match just over the minimal reporting threshold does not confirm a fourth cousin; it confirms that two people share a stretch of identical sequence that may or may not trace to a recent common ancestor. The distinction between IBD and IBS is not a technicality for researchers alone; it is the difference between a meaningful genealogical signal and background noise.

Platform differences compound this problem. A match reported on one consumer service may not appear on another, not because the relationship is absent, but because the two platforms use different SNP densities, phasing algorithms, and minimum reporting thresholds. Readers who see a match on one platform and not another often conclude one result is wrong, when in fact both are operating within their respective detection limits.

Practical steps when you see a new match:

  • Validate the match with multiple corroborating shared matches before drawing genealogical conclusions
  • Check the shared-match network to see whether the proposed relative connects to a known ancestral line
  • Consider a formal kinship test when the relationship has legal, medical, or inheritance implications

US Diagnostics Center offers kinship testing options for sibling, grandparent, and avuncular relationships, with results in 2–3 business days of lab processing. For cases that require court-admissible documentation, a legal chain-of-custody test is the appropriate standard, not a consumer match report.


Sources

The sources below represent the most useful starting points for primary literature, software documentation, and community-practical guidance on IBD analysis.


FAQ

What is identity by descent in simple terms?

Identity by descent (IBD) is a segment of DNA that two people share because they both inherited it from the same ancestor, without recombination breaking it apart between that ancestor and them. The longer the shared segment, the more recently that ancestor likely lived.

How far back does 1% shared ancestry reach?

One percent shared DNA qualitatively corresponds roughly to a third-cousin relationship within a genealogical timeframe of a few generations, though this can vary. Beyond that distance, observed shared cM can drop to zero even for documented genealogical relatives, because genetic inheritance diverges from genealogical ancestry over many generations.

What is the difference between ancestry and descent in genetics?

Genealogical ancestry refers to all documented biological relatives in a family tree, while genetic descent refers only to the DNA segments actually transmitted and detectable today. A person may have hundreds of genealogical ancestors ten generations back but carry measurable DNA from only a fraction of them.

What genetics are passed down from mother versus father?

Autosomal DNA is inherited roughly equally from both parents, with each parent contributing approximately 50% of a child's autosomal genome through recombination. Mitochondrial DNA passes exclusively through the maternal line, and the Y chromosome passes exclusively from father to son; neither undergoes the same recombination process as autosomal chromosomes.

How many generations until two relatives share no detectable DNA?

At approximately seven to ten generations of separation, the probability of sharing a detectable IBD segment drops substantially, and many genealogical relatives at that distance share zero measurable DNA. This is why genetic genealogy is most reliable within roughly ten generations, and why a documented ancestor beyond that range may have contributed no detectable genetic material to a given descendant.


This article is part of our Understanding DNA Testing: How It Works and What to Expect guide.

0 comments

Leave a comment

Please note, comments need to be approved before they are published.