What Probability of Paternity Actually Means (and Why It Never Reaches 100 Percent)

What Probability of Paternity Actually Means (and Why It Never Reaches 100 Percent)

If you have opened a paternity test report and seen a number like 99.9999%, you have probably wondered why it isn't just 100%. The alleged father matched at every marker tested. The lab confirmed the result. So why hedge with all those nines?

The short answer is that Probability of Paternity is a statistical estimate, not a physical measurement. It is derived from a formula that, by design, can approach 100% but never actually reach it. Understanding why that is takes a quick tour through Bayesian statistics, the math behind DNA matching, and how courts read these numbers in practice.

The Two Numbers on Your Report: CPI and PP

Most paternity reports show two related figures near the top: the Combined Paternity Index (CPI) and the Probability of Paternity (PP). They are related, but they answer slightly different questions.

The Combined Paternity Index is a likelihood ratio. It compares two hypotheses:

  • Hypothesis 1: The tested man is the biological father.
  • Hypothesis 2: A random, unrelated man from the same population is the biological father.

A CPI of 500,000 means the DNA evidence is 500,000 times more likely to be observed if the tested man is the father than if a random unrelated man is the father. It is a ratio of probabilities, not a probability itself. That distinction matters.

The Probability of Paternity converts that ratio into something that looks more like a percentage. It answers a different question: given the DNA evidence, what is the probability that the tested man is the actual father? That number is bounded between 0 and 100%, and it is the one most people focus on when they open their report.

The bridge between the two is a formula from Bayesian statistics, which is what the rest of this article walks through.

Probability of Paternity Is a Bayesian Posterior

Bayes' theorem is a rule for updating a belief when new evidence comes in. You start with a prior probability (what you thought before the test), you incorporate the evidence (the DNA results, represented by the CPI), and you end up with a posterior probability (your updated belief).

For a paternity test, the formula labs use is:

PP = CPI ÷ (CPI + 1) × 100

That formula assumes a prior probability of 0.5, meaning that before the DNA test, the odds of the tested man being the father are treated as 50/50. Every U.S. lab that follows standard reporting practice, including those guided by the AABB parentage testing standards and the ISFG DNA Commission recommendations, disclose this assumption somewhere on the report.

Why 50/50 as a default? Because the lab knows nothing about the case. It doesn't know whether the tested man had a relationship with the mother, when, how often, or whether there are other possible fathers. Assigning 50/50 is the most neutral starting point possible. If a court or the tested parties want to plug in a different prior based on non-DNA evidence, they can, and the resulting number will change.

What "Prior Probability" Actually Means

The prior probability represents everything you knew before the DNA test. In practice, this is usually a mix of circumstantial evidence: whether the mother and alleged father were in a relationship at the time of conception, whether there were other possible partners, whether anyone has admitted or denied paternity, and so on.

If a court decides the non-DNA evidence points strongly toward the tested man being the father, the prior might be set at 0.7 or higher. If there's strong reason to doubt, the prior might be set at 0.3. The DNA evidence then updates that starting belief.

Here's what a shift in prior does to the final number, holding the CPI constant at 500,000:

Prior Probability CPI Resulting Probability of Paternity
0.10 (unlikely) 500,000 99.99820%
0.50 (neutral, lab default) 500,000 99.99980%
0.90 (likely) 500,000 99.99998%

The strong DNA evidence overwhelms even a low prior, which is one of the reasons the 50/50 default is defensible. When the CPI runs into the hundreds of thousands or millions, the prior barely matters. The math will land in the "practical certainty" range no matter what reasonable prior you plug in.

Why the Number Approaches 100% but Never Reaches It

Look at the formula again: PP = CPI ÷ (CPI + 1) × 100.

No matter how large CPI becomes, the denominator is always CPI + 1, which is slightly larger than the numerator. The ratio is always less than 1, so multiplying by 100 always yields a number less than 100.

You can prove this to yourself with a few values:

  • CPI = 100 → PP = 100/101 × 100 = 99.0099%
  • CPI = 10,000 → PP = 10,000/10,001 × 100 = 99.9900%
  • CPI = 1,000,000 → PP = 1,000,000/1,000,001 × 100 = 99.99990%
  • CPI = 1,000,000,000 → PP = 1,000,000,000/1,000,000,001 × 100 = 99.9999999%

The number gets closer and closer to 100 as CPI grows, but the +1 in the denominator will never disappear. In calculus terms, 100 is the limit as CPI approaches infinity, but the function never actually reaches it. This is not a rounding convention or a legal disclaimer. It is baked into the math.

The Deeper Reason: Two Sources of Irreducible Uncertainty

The formula's asymptote reflects two real-world uncertainties that no amount of testing can fully eliminate.

1. A random, unrelated man could match by coincidence

Autosomal Short Tandem Repeat (STR) markers, the type used in nearly all commercial paternity tests, are chosen because they vary a lot between people. But they still have a finite number of common variants (alleles), which means two unrelated people can share the same allele at any given marker just by chance.

At a single marker, the coincidence probability might be around 5-10%. Across 20 or more markers, the combined probability of a random unrelated man matching at every marker drops to something like 1 in hundreds of millions or 1 in trillions. But it never drops to zero. The NIST STRBase reference database maintains published allele frequencies for common STR markers used in forensic and parentage testing, which is where these probability calculations come from.

Because the underlying probability of coincidental match is not zero, the calculated Probability of Paternity cannot be exactly 100%. There is always a mathematical scenario, however unlikely, where an unrelated man happens to share the same allele set.

2. Identical twins cannot be distinguished by autosomal STR

Identical (monozygotic) twins share essentially the same DNA at all autosomal STR markers because they came from the same fertilized egg. If the alleged father has an identical twin who is also a plausible candidate, standard paternity testing cannot tell them apart. Reports typically include a note about this if it is a possibility. Distinguishing identical twins in a paternity dispute requires very deep whole-genome sequencing looking for rare somatic mutations, which is not a service most parentage labs offer.

The Floor Underneath a "Match": Why It Almost Never Drops Below 99.99%

When a test returns an inclusion (the tested man is not excluded as the father), Probability of Paternity is usually reported at 99.99% or higher. Why does the number have such a high floor when it "counts" as a match?

The answer is that most U.S. labs, following AABB and ISFG guidance, set 99.9% or 99.99% as the reporting threshold for what qualifies as an inclusion. If the calculated PP falls below that threshold, either more markers get tested or the result is reported as "inconclusive" rather than as a match. This threshold-based reporting is intentional; it keeps the number of false inclusions extremely low in the general population.

The floor is also driven by the number of markers tested. When a lab analyzes 20 or more STR markers, the combined statistical power is high enough that a genuine father will almost always produce a CPI in the tens of thousands or higher, which mathematically pushes PP above 99.99%.

A Worked Example: 25 Markers, CPI = 500,000

Suppose an alleged father is tested against a child and matches at every allele across 25 STR markers. Each marker contributes its own Paternity Index (PI), and the Combined Paternity Index is the product of all of them. Let's say the lab calculates a CPI of 500,000. Plug that into the formula:

PP = 500,000 / (500,000 + 1) × 100
PP = 500,000 / 500,001 × 100
PP = 0.99999800004 × 100
PP = 99.99980%

The report will typically round this to 99.9998% or 99.99% depending on the lab's reporting conventions. What this number means in plain English: given the DNA evidence, and starting from a neutral 50/50 assumption, it is 500,000 times more likely that the tested man is the father than a random unrelated man, and the corresponding probability that he is the father is 99.9998%.

What it does not mean: there is a 0.0002% chance he isn't the father. That interpretation confuses the math. The residual probability is the mathematical space left over by the formula's asymptote plus the small residual chance of an unrelated coincidence, not a real-world estimate of doubt in this specific case.

How Courts Read Probability of Paternity

State family courts across the U.S. rely on paternity test results to establish or disprove legal paternity. Many state statutes explicitly define the threshold at which DNA evidence creates a legal presumption of paternity. The National Conference of State Legislatures summarizes state-by-state statutory language on paternity establishment.

The most common threshold is a Probability of Paternity of 99% or higher, which most state statutes treat as creating a rebuttable presumption of paternity. A few examples of typical statutory framing:

  • Some states require a PP of 95% or higher for admissibility, with 99% creating a presumption.
  • Others use 98% or 99% as the presumption threshold.
  • Some statutes reference the AABB standard directly and defer to whatever threshold the accredited lab uses.

What courts will not accept is a Probability of Paternity below the state's statutory threshold. And courts don't require 100% because they know 100% isn't achievable under the formula. The whole system is built around the assumption that "practically certain" is enough for legal purposes.

Here's a rough guide to how the numbers are typically read:

CPI Probability of Paternity How This Is Generally Read
1 to 99 50% to 99% Not sufficient for inclusion. Additional markers or retest typically needed.
100 to 999 99.0% to 99.9% May meet lower state thresholds but generally below the 99.9% AABB reporting bar.
1,000 to 9,999 99.9% to 99.99% Meets most state statutory presumption thresholds. Often the reporting floor for inclusion.
10,000 to 99,999 99.99% to 99.999% Very strong evidence. Standard inclusion reporting range.
100,000 to 999,999 99.999% to 99.9999% Common range for a healthy CPI across 20+ markers. Practical certainty.
1,000,000+ 99.9999%+ Extremely strong evidence. Approaches the limit of the formula.

USDC's home paternity test analyzes up to 28 genetic markers, which pushes CPI values into the higher ranges of this table when there is a true biological relationship. You can read more about how the home paternity process works on the paternity testing overview page.

Edge Cases Where the Standard Formula Isn't Enough

The formula assumes the tested man is being compared against a random unrelated man in the same population. That assumption breaks down in a handful of situations, and the standard PP number can be misleading unless additional testing is done.

Closely related possible fathers

If the alleged father has a brother or father who might also be the biological father, standard STR testing may not distinguish them cleanly. Full siblings share about 50% of their DNA, and a father and son share 50%. This means a brother or a paternal grandfather can produce a CPI that looks high but doesn't actually rule out the other relative. In these cases, labs will often recommend testing both candidates directly, or adding Y-STR testing (which tracks the paternal lineage and is identical among father, son, and brother, so it can rule out an unrelated man but not distinguish among the male-line relatives).

Incest scenarios

When the possible fathers are first-degree relatives (father-son, brothers), or in cases involving incest, allele-sharing gets complicated. Specialized calculations are needed, and standard PP formulas can overstate certainty. AABB-accredited labs have protocols for these cases that require reporting the results with explicit reference to the assumed relationships between potential fathers.

Rare-allele populations

The allele frequencies used to calculate PI and CPI come from population reference databases. If the tested individuals belong to a small, isolated, or under-represented population, the reference frequencies may not accurately reflect that group. In practice, most labs use pooled U.S. population databases or population-specific databases when available. Reports from accredited labs typically disclose which population database was used.

Mutation events

Occasionally an alleged father will mismatch a child at one marker due to a spontaneous mutation during sperm formation. Mutation rates at STR markers are low (roughly 1 in 1,000 to 1 in 500 per marker per generation) but not zero. When a single mismatch appears against many matches, labs apply a mutation correction to the PI at that marker rather than reporting an exclusion. Two or more mismatches across a full panel almost always means exclusion.

CPI vs. PP: The Source of Most Reader Questions

Because the two numbers are related but answer different questions, they frequently get conflated. Here is the cleanest way to keep them straight:

  • CPI is a ratio. It says how much more likely the DNA evidence is under the "he is the father" hypothesis than under the "random unrelated man is the father" hypothesis. It has no upper bound and no built-in probability interpretation. A CPI of 500,000 is not "500,000%." It is a multiplier on the odds.
  • Probability of Paternity is a percentage between 0 and 100. It converts the CPI into a probability by combining it with the prior probability (usually 0.5). It answers "given the evidence, how likely is it that he is the father?"

If a report shows CPI = 500,000 and PP = 99.9998%, they are two views of the same underlying evidence. The CPI is the raw ratio. The PP is the ratio converted to a probability using Bayes' theorem with a neutral prior. Neither is more "correct" than the other. They are complementary.

For readers who want to dig deeper into the underlying statistics, the ISFG DNA Commission has published recommendations on the biostatistical evaluation of parentage tests that walk through the full derivation. It is technical but worth reading for anyone in a legal or genetic counseling context.

What This Means for a Reader Holding a Report

If your report shows a Probability of Paternity of 99.99% or higher, and the CPI is in the tens of thousands or higher, the evidence is what statisticians call "overwhelming" and what courts call "practical certainty." The residual difference between your number and 100% is not a real-world estimate of doubt in your specific case. It is the mathematical residue of a formula designed to be honest about the limits of population statistics.

If your report shows an exclusion (typically a PP of 0% and a note explaining that multiple markers do not match), the tested man is not the biological father. The formula generates a low or zero probability because the CPI is effectively zero.

If your report shows something in between (a PP below the standard inclusion threshold), the result may be inconclusive and additional testing might resolve it. This is uncommon in modern paternity testing because 20 or more markers usually produce a decisive result one way or the other.

USDC offers home paternity testing that analyzes up to 28 markers and returns Probability of Paternity values calculated using the standard Bayesian formula described in this article. You can browse the full range of DNA relationship tests on the home DNA tests page, or learn more about how the science works on our understanding DNA testing guide.

Frequently Asked Questions

Why doesn't my report just say 100%?

Because the formula labs use, PP = CPI / (CPI + 1) × 100, mathematically cannot reach 100. No matter how large the Combined Paternity Index gets, the denominator is always one larger than the numerator, so the result is always slightly less than 100%. This isn't a lab being cautious; it is the math itself. Additionally, autosomal STR testing cannot rule out an identical twin of the alleged father or, in extremely rare cases, an unrelated man who happens to share the same allele set by chance.

Is 99.9999% really different from 99.99%?

Both are effectively practical certainty for legal and personal purposes. The extra "nines" reflect how strong the DNA evidence is, but once you cross about 99.9%, courts and most people treat it as conclusive. The difference between 99.99% and 99.9999% comes down to how many markers were tested and how rare the shared alleles are in the reference population. It doesn't change the practical interpretation.

What is the prior probability, and can I change it?

The prior probability is what you (or a court) believed about paternity before the DNA test. Labs default to 0.5, meaning they treat the odds as 50/50 before running the test, because that is the most neutral starting point. If a court has other evidence suggesting the tested man is likely (or unlikely) to be the father, it can plug in a different prior. The DNA evidence then updates that belief. For most cases with strong DNA evidence, the choice of prior barely affects the final number.

Can the number ever be exactly 0%?

Reports often show 0% Probability of Paternity when the tested man is excluded, meaning multiple markers do not match in a way that cannot be explained by mutation. Technically, the number is a very small positive value, but it rounds to zero for reporting purposes. An exclusion at 3 or more markers on a modern STR panel is considered conclusive.

What if the alleged father has a brother who could also be the father?

This is one of the edge cases where the standard formula can be misleading. Full brothers share about 50% of their DNA, so a brother of the true father can produce a CPI that looks high but doesn't cleanly rule him out. If both brothers are possible fathers, the cleanest resolution is to test both directly. Some labs also run Y-STR testing, which is identical among all males in a paternal line but can rule out an unrelated man. If a brother, father, or son of the alleged father is a possibility, mention this to the lab before ordering the test so they can flag it and, if needed, recommend an alternative testing approach.

0 comments

Leave a comment

Please note, comments need to be approved before they are published.