Allele Frequency Calculator
Estimate healthy and mutant allele frequencies, genotype proportions, and recessive-carrier prevalence from a population disease frequency.
Population input
Live results
Expected genotype composition
Genotype breakdown
| Population group | Expression | Frequency | Approximate occurrence |
|---|
How to use the allele frequency calculator
What this calculator does
This calculator converts the prevalence of an autosomal recessive condition into estimated allele and genotype frequencies using the Hardy – Weinberg model. It treats the affected frequency as q², takes its square root to obtain the mutant allele frequency q, and then uses p = 1 – q. From those two allele frequencies it estimates unaffected non-carriers p² and heterozygous carriers 2pq. It is a population-level mathematical estimate, not a diagnosis, genetic test, individual risk assessment, or substitute for counseling.
When to use it
Use it to explore how a rare recessive disease prevalence relates to carrier frequency, to check classroom or population-genetics exercises, to compare hypothetical populations, or to prepare a transparent spreadsheet for further analysis. It is most appropriate when the Hardy – Weinberg assumptions are a reasonable teaching approximation and the prevalence applies to the same population you want to study.
How to calculate
- The calculator opens with a ready-to-use demonstration: “1 in 10,000 people.” Its results and Excel workbook are immediately available.
- Under Disease frequency given as, choose either 1 in N people or Percentage of population. The current value is converted when you switch modes.
- Replace Occurrence of the disease with a valid positive value. Results update live.
- Read the carrier estimate first, then the p, q, p², and q² values. The chart and table show the same model in complementary forms.
- Select Download Excel to export the current inputs, formulas, outputs, and breakdown. Reset clears the demonstration and results; Excel export remains disabled until a complete valid value is entered again.
Input guide
Disease frequency given as is required and accepts one of two modes. “1 in N people” expects a denominator greater than or equal to 1, such as 10,000. “Percentage of population” expects a percentage greater than 0 and no greater than 100, such as 0.01. Switching modes converts the current value rather than merely changing its label. A common mistake is entering 0.01 in the “1 in N” mode when 0.01% was intended.
Occurrence of the disease is required. It accepts ordinary decimal notation with a period as the decimal separator and optional comma thousands grouping in the “1 in N” mode. Scientific notation and decimal-comma input are rejected to prevent silent reinterpretation. A higher affected prevalence raises q and generally raises the estimated carrier share until q approaches 0.5; at extreme prevalences, carrier frequency later declines because nearly everyone carries two mutant alleles.
Output guide
Estimated carrier frequency (2pq) is the expected heterozygous share and appears as “1 in N” plus a percentage. Healthy allele frequency (p) and Mutant allele frequency (q) are allele proportions that sum exactly to 1. Unaffected, non-carrier (p²), Carrier (2pq), and Affected (q²) are genotype proportions that sum to 1. The summary pills repeat the disease prevalence, q, and carrier odds. The genotype table shows each group's label, algebraic expression, percentage, and approximate “1 in N” occurrence. The donut chart represents the same three mutually exclusive genotype groups; very small slices may be visually subtle, so the table is the precise reference.
Worked example
With the startup value of 1 affected person in 10,000, q² = 1 ÷ 10,000 = 0.0001. Therefore q = √0.0001 = 0.01, or 1.00%. Then p = 1 – 0.01 = 0.99, or 99.00%. The estimated carrier frequency is 2pq = 2 × 0.99 × 0.01 = 0.0198, or 1.98%, which is approximately 1 in 51 people. The unaffected non-carrier proportion is p² = 0.9801, or 98.01%, and the affected proportion remains q² = 0.01%.
Learn more
The National Human Genome Research Institute explains what an allele is, while MedlinePlus describes autosomal recessive inheritance. For a deeper technical discussion, review the peer-reviewed overview of Hardy – Weinberg equilibrium in genomic analysis.
Model assumptions and interpretation
The Hardy – Weinberg relationship p² + 2pq + q² = 1 describes an idealized equilibrium. It assumes a very large population, random mating, negligible migration, no selection, no mutation changing the locus materially, and no strong genetic drift. Those conditions are useful as a baseline, but they are not guaranteed in a real population. Founder effects, population stratification, assortative mating, differential survival, and ascertainment can all shift observed frequencies.
Population prevalence also needs careful definition. The denominator should refer to the same ancestry, geography, age range, diagnostic criteria, and time period as the population of interest. A prevalence estimated from diagnosed cases may undercount people with mild, late-onset, or unrecognized disease. The NHGRI founder-effect explanation shows why a variant can be unusually common in a smaller population.
For personal or reproductive decisions, population estimates are not enough. Family history, ancestry, molecular test sensitivity, variant classification, penetrance, and the partner's status can materially change risk. A genetics professional can interpret those details and explain what a screening or diagnostic result does and does not establish.