HaploSample report: Ancestry estimated by this pipeline, not self-reported. Not yours.
hu706427

Your ancestry report

Your ancestry is mostly West Asian (52%), with some European (28%) and South Asian (19%) and a little from elsewhere.

Analysed 31 August 2026

ItalianIT20%Armenian Highland and the South CaucasusAH18%LevantineLV14%Sindh and the Lower IndusSL14%Dagestan and the Nakh highlandsDG8%Southeast EuropeanEB6%South CaucasusSQ5%BedouinBD2%Sardinia and CorsicaSD2%

Tap any bar or region to see how sure we are.

Ancestry

Your mother's line is U7a4a and your father's line is R2a.

Go to Ancestry

Health & traits

Your chip covers 17 of the 24 health variants we check. We found 5. Of 23 trait variants, your chip covers 19; 7 found.

Go to Health & traits

Technical

Your file was graded B, with 638,463 markers read. Closest reference groups, your file and other calculators are here.

Go to Technical

Ancestry

Where you come from

Recent regions first, then your chromosomes, your mother's and father's lines, and thousands of years back.

ItalianIT20%Armenian Highland and the South CaucasusAH18%LevantineLV14%Sindh and the Lower IndusSL14%Dagestan and the Nakh highlandsDG8%Southeast EuropeanEB6%South CaucasusSQ5%BedouinBD2%Sardinia and CorsicaSD2%

Where your ancestors lived

Each bar is the share of your DNA that matches people living in that region today.

West Asian52.2%
Armenian Highland and the South Caucasus
18.4%
Levantine
14.3%
Dagestan and the Nakh highlands
7.6%
South Caucasus
4.8%
Bedouin
2.3%
Druze, Anatolia and Iran, North Caucasus, Arabian, Ashkenazi and Parsi: in the model, but this file cannot tell them apart from zero.
European28.2%
Italian
19.7%
Southeast European
5.5%
Sardinia and Corsica
1.7%
Basque Country, Finland and Karelia, Iberian, Northwest European, Roma, Eastern European and Volga and Ural: in the model, but this file cannot tell them apart from zero.
South Asian18.6%
Indo-Gangetic Plain14.0%
Sindh and the Lower Indus
13.7%
The Ganges Plain and Gujarat and Punjab and Kashmir: in the model, but this file cannot tell them apart from zero.
Makran and the western ranges, Himalayas and Northeast India, Chota Nagpur Plateau, Bengal and the central belt, Southern Peninsula and Andaman and Nicobar Islands: in the model, but this file cannot tell them apart from zero.
11.9% could not be placed in one region. It belongs to regions this file cannot tell apart.
How we worked this out

Each percentage is a share of the 10,292 markers we could read in your file (81% of 12,770), compared with people sampled today by the 1000 Genomes Project. The solid part of a bar is the share we are sure of. The faint part is how much higher it might be. The line is our best estimate. The range is a 95% interval from resampling your markers. If a population is missing from our reference set, your ancestry from it is counted under the nearest region we do have.

Each region and its range. West Asian 52.2% (likely 43–60%); European 28.2% (likely 21–37%); South Asian 18.6% (likely 14–24%).

South Asian is one region. These seven regional components share 96–99% of their allele frequencies, so how much South Asian ancestry you have is measured far more precisely than how it divides between the seven. Its total (18.6%) is firmer than the split inside it. The parts drawn inside it add to 14.0%, because some are too small to tell from zero and are not drawn.

European is one region. How much of your ancestry is from this region is measured from the whole panel at once. How it divides between the areas inside it is a second, harder question, fitted separately — so the total is the firmer number of the two. Its total (28.2%) is firmer than the split inside it. The parts drawn inside it add to 26.9%, because some are too small to tell from zero and are not drawn.

West Asian is one region. How much of your ancestry is from this region is measured from the whole panel at once. How it divides between the areas inside it is a second, harder question, fitted separately — so the total is the firmer number of the two. Its total (52.2%) is firmer than the split inside it. The parts drawn inside it add to 47.5%, because some are too small to tell from zero and are not drawn.

What “could not be placed” means. A region whose range reaches zero is not drawn on its own. The ancestry is still yours; this file does not have enough markers to say which region it belongs to. Here the regions that could hold it are North African (up to 6%) and Central Asian (up to 1%).

The map. Shading marks where each ancestry lives, at your own share of it. The outlines are geographic regions and the numbers are genetic, so an edge is approximate: ancestry shades into its neighbours. A faint region is one this file cannot tell from zero.

Which reference people each region is. regional reference groups, not communities — these are broad regional-linguistic samples and 6 of 7 were recruited in India

RegionReference populationsPeopleRecruited
Balochistan and the Makran coastBalochi, Brahui, Makrani307Balochistan, Pakistan
Indus and Gangetic plainsPJL1,038Punjab, Sindh, Uttar Pradesh and the Gangetic plain
Andaman IslandsOnge, Jarawa68the Andaman Islands
RegionFitted fromEstimateRange
ItalianItaly and Sardinia, fitted from Italian_North, Sardinian19.7%15–26%
Armenian Highland and the South Caucasusindividuals recruited in Armenia, Georgia and Abkhazia18.4%9–25%
LevantinePalestine, Jordan, Lebanon, Syria and northern Iraq, fitted from Palestinian, Druze, Jordanian, Assyrian14.3%8–18%
Sindh and the Lower Indusindividuals recruited in Sindh and along the lower Indus13.7%6–23%
Dagestan and the Nakh highlandsindividuals recruited in Dagestan, Chechnya and Ingushetia7.6%3–12%
Southeast EuropeanThe Balkans and the Carpathian basin, fitted from Greek, Bulgarian, Croatian, Serbian_Serb, Romanian, Hungarian and others5.5%0–10%
South Caucasusindividuals recruited in Georgia, Abkhazia and Azerbaijan4.8%2–7%
Makran and the western rangesindividuals recruited across Balochistan and Makranbelow resolution0–16%
Bedouinindividuals recruited among Bedouin communities of the Negev and Sinai2.3%0–3%
DruzeDruze individuals recruited in the Carmel and the Golanbelow resolution0–6%
Sardinia and Corsicaindividuals recruited on Sardinia and Corsica1.7%0–3%
Anatolia and Iranindividuals recruited in Turkey, Iran and northern Iraqbelow resolution0–6%
Basque CountryBasque individuals recruited in the western Pyreneesbelow resolution0–6%
North AfricanThe Maghreb — Morocco, Algeria, Tunisia and Libya, fitted from Mozabitebelow resolution0–6%
North CaucasusThe northern slope of the Caucasus, fitted from Ossetian, Lezgin, Karachai, Tabasaran, Lak, Ingushian and othersbelow resolution0–3%
ArabianThe Arabian peninsula and its desert margins, fitted from BedouinA, BedouinBbelow resolution0–5%
The Ganges Plain and Gujaratindividuals recruited in Uttar Pradesh, Bihar, Bengal and Gujaratbelow resolution0–3%
AfricanYoruba, Luhya, Gambian, Mende and Esan reference samplesbelow resolutionunder 0.01%
Indigenous AmericanPeruvian, Mexican, Colombian and Puerto Rican reference samplesbelow resolutionunder 0.01%
East AsianHan, Japanese, Dai and Kinh reference samplesbelow resolutionunder 0.01%
Oceanian20 pooled Papuan and Bougainville individuals — the whole public supplybelow resolutionunder 0.01%
Southern Peninsulaindividuals recruited in Tamil Nadu, Andhra Pradesh, Telangana, Karnataka and Keralabelow resolutionunder 0.01%
Bengal and the central beltindividuals recruited in Bengal, Madhya Pradesh, Chhattisgarh and Maharashtrabelow resolutionunder 0.01%
Himalayas and Northeast Indiaindividuals recruited in Nepal, the terai, and the northeastern hillsbelow resolutionunder 0.01%
Chota Nagpur Plateauindividuals recruited in Jharkhand, Odisha and the eastern Ghatsbelow resolutionunder 0.01%
Andaman and Nicobar IslandsOnge, Jarawa and Great Andamanese individualsbelow resolution0–1%
Siberian21 West Siberian, South Siberian and Amur populations — 498 individualsbelow resolutionunder 0.01%
Central Asian155 individuals from six Central Asian populations — the oases and the Kazakh steppebelow resolution0–1%
Finland and KareliaFinland and Karelia, fitted from Finnishbelow resolution0.0–0.5%
IberianIberia, fitted from Spanish, Basquebelow resolution0.0–0.1%
Northwest EuropeanBritain, Ireland, the Low Countries, northern Germany, fitted from CEU, GBRbelow resolutionunder 0.01%
RomaRoma individuals recruited in Barcelona, Bilbao, Granada, Madrid and Portobelow resolution0–2%
Eastern EuropeanThe East European plain and the Baltic, fitted from Russian, Ukrainian, Belarusian, Czech, Estonian, Lithuanianbelow resolutionunder 0.01%
Volga and UralThe middle Volga and the southern Urals, fitted from Bashkir, Mordovian, Chuvash, Tatar_Kazan, Udmurt, Tatar_Mishar and othersbelow resolutionunder 0.01%
Punjab and Kashmirindividuals recruited in Punjab, Haryana and the Kashmir valleybelow resolutionunder 0.01%
Ashkenaziindividuals of Ashkenazi Jewish descent, recruited in Europe, Israel and the United Statesbelow resolution0–3%
ParsiParsi individuals recruited in Gujarat and Sindhbelow resolution0–3%

Your chromosomes, painted

Each chromosome coloured by which region its stretches match. Long stretches point to recent ancestors; this sees back about nine generations.

South AsianSA 58.9%EuropeanEU 39.0%East AsianEA 1.3%AfricanAF 0.6%Indigenous AmericanAM 0.2%Not painted: 25.6% of the genome
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22

Two bars per chromosome, one for each copy. Grey is a stretch that matched no region for long enough to name.

How we worked this out

Painted in 4 cM windows against African, European, East Asian, South Asian, Indigenous American and Oceanian reference haplotypes. A label has to hold for at least two consecutive windows to count; single-window matches are what deep shared ancestry looks like, while a real ancestor leaves blocks of 13 to 74 centimorgans. Roughly a sixth of the genome (25.6%) matched no group for long enough to name and is left grey rather than guessed.

Against held-out people the painter had never seen, a named stretch is right: African 99.6%, Indigenous American 99.9%, East Asian 98.4%, European 90.2%, Oceanian 99.9%, South Asian 83.2%.

This is a recent-ancestry picture. Older mixtures are still in you, but recombination has cut them into stretches too short to name, so they show here as the ancestry around them. Shares are of what could be painted. The ancient peoples chapter measures something else, in another way, so its percentages do not line up with these.

Your mother's line and your father's line

Two threads pass down almost unchanged: one from mother to child, one from father to son. Each follows a single ancestor across thousands of years.

Maternal line
U7a4a
Your mother's mother's mother, and so on back.
mt-MRCAL1'2'3'4'5'6L2'3'4'5'6L2'3'4'6L3'4'6L3'4L3NRUU2'3'4'7'8'9U7U7aU7a4U7a4a
Paternal line
R2aR-M124
Your father's father's father, and so on back. Men on this branch are most often from South Asian today.
A1A1bBTCTCFFIJKKK2K2bPP1RR2R2a
How we worked this out

Maternal. Your file covers 2,137 mitochondrial positions. 20 of them match U7a4a, and 5 variants are not explained by it. That is normal: every family line carries a few mutations of its own. Positions the chip does not read are not counted either way. Scored against PhyloTree Build 17 over the mtDNA positions this file genotyped. Positions the array does not cover are not counted as evidence either way. Tree: van Oven & Kayser 2009, PhyloTree Build 17. U7a4a is 14 branches from the root. No ages are printed because we do not yet hold a dated tree whose licence we have cleared.

Where the branch points, and how often it is right. Naming a region from the branch is right 69% of the time, against 28% for always guessing the commonest. This was measured by holding each man out in turn. Men on clade R2 are from South Asian 93% of the time among the 40 we hold. This names a continent. Inside a continent the method is no better than guessing, so no finer region is named.

Paternal. Descent through the published haplogroup tree, entering a branch only where the marker defining it was genotyped and derived. 2,012 informative markers. Why it stops at R2a: the chip this file came from carries no marker that separates the branches below this point — it is not something missing from your file, and a test on a chip with denser Y coverage could go further. The full label R-M124 changes as the tree is revised; the marker name does not.

An empty region on a map is one nobody has excavated, or excavated and not published.

Thousands of years back

The ancient peoples you come from

Scientists have read DNA from people who lived long ago. Here your DNA is fitted as a mix of those ancient groups.

Not measured for this file

The ancestry model we have does not fit this genome. That is a result rather than an error — it means these ancient sources cannot account for this person's ancestry, and any percentages we showed would be describing a model we have already rejected.

About fifty thousand years ago

Your Neanderthal and Denisovan DNA

When early humans left Africa they met Neanderthals and Denisovans and had children together. Almost everyone outside Africa carries a little of both.

Not measured for this file

We hold no archaic call set covering Balochi, the reference population your genome most resembles, nor one for its region. We will not borrow an unrelated population's — a percentile measured against the wrong cohort looks like an answer but is not yours.

How we worked this out

This depends on how many of the SPrime archaic positions your file happens to carry. It says nothing about the quality of the rest of the file. Below a minimum count the rank among other people moves too much to print.

Health & traits

A few variants, found or not found

We only say whether your file carries each variant, and what that usually means. There is no risk score here, and nothing on this page is medical advice.

Traits

A handful of well-studied variants, whether your file carries them, and what each one usually means.

We check 23 trait positions. Your chip covers 19 of them: 7 found, 12 read and not carried, 4 it cannot read.

KITLG, hair colourKITLG rs642742
Found, one copy

You carry this variant.

A regulatory change near KITLG associated with lighter hair colour. The same gene carries an independent pigmentation change in stickleback fish, which is why the 2007 paper reporting it is about both.

establisheduntested in South Asian studies

The hair-colour association was reported in European and Asian cohorts; its effect in South Asians specifically has not been measured.

cis-Regulatory changes in Kit ligand expression and parallel evolution of pigmentation in sticklebacks and humans. Cell, 2007. European and Asian PMID 18083106
OCA2 R419Q, eye colourOCA2 rs1800407
Found, one copy

You carry this variant.

Associated with green and hazel eye colour, and reported to shift eye colour away from blue in people who otherwise carry the blue-eye haplotype at HERC2.

establisheduntested in South Asian studies

The eye-colour association comes from European cohorts. Eye colour in South Asians is far less studied and this variant's effect there has not been measured.

Allele variations in the OCA2 gene (pink-eyed-dilution locus) are associated with genetic susceptibility to melanoma. Eur J Hum Genet, 2005. European PMID 15889046
Fast-twitch muscle (ACTN3)ACTN3 rs1815739
Found, one copy

You carry one copy of the version linked to endurance rather than sprint muscle.

Two copies of this variant produce no alpha-actinin-3 in fast-twitch muscle fibres. Associated with sprint and power performance in studies of elite athletes.

moderateuntested in South Asian studies

The association was found by comparing elite athletes with controls, which is a claim about the extreme tail of a distribution and not about ordinary training. Roughly one person in five worldwide carries two copies and they are not distinguishable in everyday life.

BEB 57.6%GBR 46.7%GIH 56.8%ITU 61.3%PJL 52.1%STU 65.2%
ACTN3 genotype is associated with human elite athletic performance. American Journal of Human Genetics, 2003. Australian elite athletes and controls PMID 12879365
Caffeine metabolism, CYP1A2 -163C>ACYP1A2 rs762551
Found, two copies

You are likely a fast caffeine metaboliser: coffee wears off sooner.

CYP1A2 clears about 95% of ingested caffeine. The A allele (the *1F haplotype) is associated with higher inducibility of the enzyme and faster clearance in smokers and heavy coffee drinkers.

establisheduntested in South Asian studies

The effect is CONDITIONAL, which is the whole finding and the part most easily lost. The genotype separates fast from slow metabolisers mainly among people whose enzyme is being induced -- by smoking, or by heavy intake.

BEB 57.6%CEU 72.7%CHB 63.6%GBR 73.1%GIH 52.4%ITU 54.4%
Functional significance of a C-->A polymorphism in intron 1 of the cytochrome P450 CYP1A2 gene tested with caffeine. British Journal of Clinical Pharmacology, 1999. German cohort PMID 10233211
MTHFR C677TMTHFR rs1801133
Found, one copy

Your body processes folate slightly more slowly. For most people this makes no practical difference.

Reduces the activity of methylenetetrahydrofolate reductase. ClinVar records it under "drug response".

moderateuntested in South Asian studies

This variant is the subject of a large amount of overstated commercial testing. The major clinical and genetics bodies advise against testing it outside specific contexts, and this row exists to say the variant is present or absent, not to be acted on.

gnomAD amr 47.8%gnomAD eas 34.8%gnomAD nfe 33.7%gnomAD sas 14.9%
A candidate genetic risk factor for vascular disease: a common mutation in methylenetetrahydrofolate reductase. Nature Genetics, 1995. Canadian and Dutch PMID 7647779
Long-chain fatty acid conversion, FADS1FADS1 rs174537
Found, one copy

You carry this variant.

Associated with the efficiency of converting plant-derived short-chain omega-3 and omega-6 fatty acids into the long-chain forms the body uses. The T allele is associated with lower conversion.

establisheduntested in South Asian studies

The dietary reading of this variant is where it stops being simple. Conversion efficiency interacts with what a person actually eats, and a vegetarian diet -- which a large share of South Asia follows -- supplies only the short-chain forms this enzyme has to convert.

BEB 20.9%CEU 36.4%CHB 35.4%GBR 38.5%GIH 10.2%ITU 13.7%
Genome-wide association study of plasma polyunsaturated fatty acids in the InCHIANTI study. PLoS Genetics, 2009. Italian cohort PMID 19851445
Skin pigmentation, SLC24A5 A111TSLC24A5 rs1426654
Found, two copies

You carry this variant.

The single largest-effect common variant on skin pigmentation known. The A allele reduces melanin production and accounts for a substantial share of the pigmentation difference between European and West African populations.

establisheduntested in South Asian studies

South Asia is where this variant is most informative and least deterministic. It segregates INSIDE the region rather than between it and anywhere else, so two people from the same state can carry different genotypes.

BEB 53.5%CEU 100.0%CHB 2.9%GBR 100.0%GIH 95.2%ITU 65.2%
SLC24A5, a putative cation exchanger, affects pigmentation in zebrafish and humans. Science, 2005. Human populations and zebrafish PMID 16357253
12 read, variant not carried
HERC2 pigmentation variantHERC2 rs1667394
Not found

You do not carry this variant.

One of the pigmentation positions in the HERC2/OCA2 region reported for hair and eye colour in a genome-wide study of Europeans. It sits near, and is inherited with, the better-known blue-eye variant.

establisheduntested in South Asian studies

Established in Icelandic and Dutch cohorts. Its effect in South Asians has not been measured, and pigmentation variants in this region behave differently outside Europe.

Genetic determinants of hair, eye and skin pigmentation in Europeans. Nat Genet, 2007. Icelandic and Dutch PMID 17952075
MC1R R151C, red hair and fair skinMC1R rs1805007
Not found

You do not carry this variant.

One of the three MC1R variants most consistently reported with red hair and freckling. The effect is strongest with two copies, and carrying one copy of any of the three is common in people who are not red-haired.

establisheduntested in South Asian studies

The MC1R red-hair association was established in British and Irish cohorts. This variant is rare in South Asians and its effect there has not been measured.

Variants of the melanocyte-stimulating hormone receptor gene are associated with red hair and fair skin in humans. Nat Genet, 1995. British and Irish PMID 7581459
MC1R R160W, red hair and fair skinMC1R rs1805008
Not found

You do not carry this variant.

A second of the three MC1R variants reported with red hair and freckling, in the same 1995 study that named the first.

establisheduntested in South Asian studies

As with the other MC1R variants here, the association was established in British and Irish cohorts and has not been measured in South Asians.

Variants of the melanocyte-stimulating hormone receptor gene are associated with red hair and fair skin in humans. Nat Genet, 1995. British and Irish PMID 7581459
Earwax typeABCC11 rs17822931
Not found

You likely have wet earwax.

The clearest single-variant trait known in human genetics. One amino acid change disables the ABCC11 transporter, and two copies of the disabling allele give dry, flaky earwax and markedly reduced underarm odour; one or no copies give wet earwax.

establisheduntested in South Asian studies

Deterministic for earwax; only strongly associated for body odour, which also depends on skin bacteria, washing and clothing. The variant is readable on every consumer array we have checked.

BEB 54.7%CHB 97.1%GBR 11.0%GIH 39.8%ITU 56.4%PJL 37.0%
A SNP in the ABCC11 gene is the determinant of human earwax type. Nature Genetics, 2006. Japanese, and a 33-population global survey PMID 16444273The impact of natural selection on an ABCC11 SNP determining earwax type. Molecular Biology and Evolution, 2011. Worldwide populations, selection analysis PMID 20937735
Alcohol flushALDH2 rs671
Not found

You clear alcohol the usual way; no flush from this variant.

Reduces aldehyde dehydrogenase 2 activity; associated with facial flushing after alcohol. ClinVar records it under "drug response".

establisheduntested in South Asian studies

Carried as a deliberate negative, like the G6PD A- row. In gnomAD v4 this allele is at 0.24364 in East Asians and 0.00025 in South Asians — roughly one in four against roughly one in four thousand.

gnomAD eas 24.4%gnomAD nfe 0.00%gnomAD sas 0.03%
The alcohol flushing response: an unrecognized risk factor for esophageal cancer from alcohol consumption. PLoS Medicine, 2009. Review, East Asian populations PMID 19320537
Hair thickness and shovel-shaped incisors, EDAR V370AEDAR rs3827760
Not found

You do not carry this variant.

The derived allele is associated with thicker, straighter hair shafts, more eccrine sweat glands and shovel-shaped upper incisors. It is one of the strongest signals of recent selection in East Asian populations.

establisheduntested in South Asian studies

Effectively an East Asian variant: 94% in Han Chinese and 0-5% across the South Asian cohorts, reaching 5% only in Bengali. For most readers here the answer is 'you do not carry it', which is a fact about the allele's distribution and not about their hair.

BEB 5.2%CEU 0.00%CHB 93.7%GBR 0.00%GIH 1.5%ITU 0.00%
A scan for genetic determinants of human hair morphology: EDAR is associated with Asian hair thickness. Human Molecular Genetics, 2008. Asian populations PMID 18065779
Blue-eye variant (HERC2)HERC2 rs12913832
Not found

You carry the common version, linked to brown eyes.

The single variant explaining most of the blue-brown eye colour difference in European populations, by regulating OCA2 expression.

establisheduntested in South Asian studies

The evidence is almost entirely European and the variant is uncommon here — gnomAD v4 puts it at 0.09097 in South Asians against 0.76357 in non-Finnish Europeans. Eye colour in populations where brown predominates is not well explained by this variant, and a confident call from it would be a European model applied where it has not been tested.

gnomAD afr 12.6%gnomAD eas 0.17%gnomAD nfe 76.4%gnomAD sas 9.1%
A single SNP in an evolutionary conserved region within intron 86 of the HERC2 gene determines human blue-brown eye color. American Journal of Human Genetics, 2008. Australian and Dutch European-ancestry PMID 18252222
Lactase persistenceMCM6 rs4988235
Not found

Milk may be harder on your stomach as an adult. This is the usual pattern across much of South and East Asia.

The variant upstream of LCT most strongly associated with continued lactase production into adulthood. ClinVar records it under "association".

establisheduntested in South Asian studies

Two separate limits, and they point in opposite directions. No lactase association in the GWAS Catalog has ever had a South Asian discovery cohort, so the association itself is untested here.

BEB 5.8%CEU 73.7%GBR 72.0%GIH 14.1%ITU 5.9%PJL 26.0%
Identification of a variant associated with adult-type hypolactasia. Nature Genetics, 2002. Finnish families PMID 11788828Herders of Indian and European cattle share their predominant allele for lactase persistence. Molecular Biology and Evolution, 2012. Indian populations PMID 21836184
Freckling and hair colour, IRF4IRF4 rs12203592
Not found

You do not carry this variant.

Associated with freckling, lighter hair colour and skin sensitivity to sun. One of the few pigmentation loci that acts on patterning rather than on overall tone.

establisheduntested in South Asian studies

Common in Europe at 9-18% and essentially absent in the South Asian, African and East Asian cohorts. It is on the panel because the report is global, not because it is expected to say much to a South Asian reader.

BEB 0.00%CEU 16.2%CHB 0.00%GBR 18.1%GIH 0.00%ITU 0.98%
A genome-wide association study identifies novel alleles associated with hair color and skin pigmentation. PLoS Genetics, 2008. European cohorts PMID 18483556
Hair and eye colour, SLC24A4SLC24A4 rs12896399
Not found

You do not carry this variant.

A potassium-dependent sodium-calcium exchanger locus associated with lighter hair and eye colour, and one of the markers in the published HIrisPlex eye and hair colour models.

establisheduntested in South Asian studies

Unlike most pigmentation variants on this panel, this one is common in South Asia too -- 30-33% against 39% in Britain -- so it does separate people here. It is still one marker of the several a real prediction model uses, and this report prints the genotype rather than running that model.

BEB 31.4%CEU 56.1%CHB 29.1%GBR 39.0%GIH 33.5%ITU 32.8%
Genetic determinants of hair, eye and skin pigmentation in Europeans. Nature Genetics, 2007. Icelandic and Dutch cohorts PMID 17952075
Skin pigmentation and freckling, TYR S192YTYR rs1042602
Not found

You do not carry this variant.

A coding change in tyrosinase, the rate-limiting enzyme of melanin synthesis. The A allele is associated with lighter skin, freckling and sun sensitivity.

establisheduntested in South Asian studies

At 36% in Britain, 13% in Punjab and absent in the East Asian and African cohorts, this is largely a European-gradient variant. Pigmentation is polygenic; one tyrosinase variant shifts a tendency and does not set a tone.

BEB 2.9%CEU 39.9%CHB 0.00%GBR 35.7%GIH 11.2%ITU 2.9%
Genetic determinants of hair, eye and skin pigmentation in Europeans. Nature Genetics, 2007. Icelandic and Dutch cohorts PMID 17952075
Hair and eye colour, TYR upstreamTYR rs1393350
Not found

You do not carry this variant.

A regulatory variant upstream of tyrosinase, associated with red and blond hair, blue eyes and freckling in the same genome-wide scans that found the coding change above.

establisheduntested in South Asian studies

27% in Britain and effectively absent everywhere else we hold. It is on the panel because the report is global, and it will tell most South Asian readers only that they do not carry it.

BEB 4.1%CEU 24.2%CHB 0.00%GBR 27.5%GIH 11.2%ITU 0.49%
Genetic determinants of hair, eye and skin pigmentation in Europeans. Nature Genetics, 2007. Icelandic and Dutch cohorts PMID 17952075
4 this chip cannot read
SLC45A2 L374F, skin and hair pigmentationSLC45A2 rs16891982
Not read: ambiguous position

One of the largest single contributions to the difference in skin and hair pigmentation between European and non-European populations.

Established in European cohorts, where the lighter-pigmentation allele is common. It is rare in South Asians and its contribution to pigmentation there has not been measured.

establisheduntested in South Asian studies

Variants of the MATP/SLC45A2 gene are protective for melanoma in the French population. Hum Mutat, 2008. French PMID 18683857New common variants affecting susceptibility to basal cell carcinoma. Nat Genet, 2009. European PMID 19578363
Alcohol metabolism, ADH1B His48ArgADH1B rs1229984
Not read: not readable

The ADH1B*2 allele encodes an enzyme that converts alcohol to acetaldehyde far faster than the common form.

Reads one step of a two-step pathway. The flush most people mean is usually ALDH2 (rs671, also on this panel), which clears acetaldehyde; ADH1B controls how fast it arrives.

establisheduntested in South Asian studies

A global perspective on genetic variation at the ADH genes reveals unusual patterns of linkage disequilibrium and diversity. American Journal of Human Genetics, 2002. Global populations PMID 11774072
Skin pigmentation, OCA2 H615ROCA2 rs1800414
Not read: not on this chip

An East Asian-specific pigmentation variant.

Your chip does not include this position. Not read is different from not found.

establisheduntested in South Asian studies

A common variant in the OCA2 gene is associated with skin pigmentation in East Asians. Human Genetics, 2010. East Asian cohorts PMID 20049473
Bitter taste (PTC)TAS2R38 rs713598
Not read: ambiguous position

One of three coding variants in the TAS2R38 bitter receptor that together determine sensitivity to phenylthiocarbamide and related compounds.

One of three variants, and this file reads one of them. The phenotype depends on the combination, so a call here is partial by construction rather than by accident.

establisheduntested in South Asian studies

Positional cloning of the human quantitative trait locus underlying taste sensitivity to phenylthiocarbamide. Science, 2003. Utah families of European ancestry PMID 12595690
How we worked this out

Not found: we read the position and the variant is on neither copy. That is a measurement. Found, one copy or two copies: also a measurement. Not read: the chip does not type this position, or could not read it cleanly. Nobody looked, so it is not the same as not found. Could not be read: several variants share this position or the strand is unmeasured, so the row says what blocked it.

Most associations here were established in European or East Asian cohorts. Each row's tags say which population the association has been tested in. Catalogue v0.1.0. Nothing here is a prediction about you.

Health variants

Whether your file carries variants that studies have linked to health. Found or not found, with the studies behind each one.

We check 24 health positions. Your chip covers 17 of them: 5 found, 12 read and not carried, 7 it cannot read.

Not medical advice, and not a reason to start, stop or change any medicine. A variant we could not read may still be there, and a variant missing from this short list was never examined. If a doctor has prescribed something, keep taking it and show them this page.
CYP2C19 Clopidogrel and other CYP2C19-activated drugs2 of 585 chip-readable variants
1 found
CYP2C19*2, the commonest no-function alleleCYP2C19 rs4244285
Found, one copy

Tell a doctor about this variant before starting clopidogrel or some antidepressants. It changes how the body processes them.

A splice-site change that produces no working CYP2C19 enzyme from the affected copy. CYP2C19 converts several drugs into their active form, so carrying two copies means the enzyme activity is absent rather than reduced.

establishedreplicated in South Asian studies

This is not medical advice, and not a reason to change any medicine. If a doctor has prescribed something, keep taking it and show them this page. Metaboliser status is assigned from a person's full CYP2C19 star-allele diplotype; this file reads two of those alleles, so a result here is partial. Genotype is also not a measurement of enzyme activity, which is what a clinical test would give you.

BEB 32.6%GBR 14.3%GIH 33.0%ITU 37.3%PJL 34.4%STU 41.2%
Clinical Pharmacogenetics Implementation Consortium Guideline for CYP2C19 Genotype and Clopidogrel Therapy: 2022 Update. Clinical Pharmacology and Therapeutics, 2022. Guideline, systematic literature review PMID 35034351Prevalence of CYP2C19 Poor Metabolisers Among South Indian Psychiatric Patients. Annals of Neurosciences, 2025. South Indian patients PMID 41063924
CYP2C19*3, a second no-function alleleCYP2C19 rs4986893
Not found

You do not carry this variant.

A premature stop codon that truncates the CYP2C19 protein. Same consequence as *2 -- no working enzyme from that copy -- by a different mechanism.

establishedreplicated in South Asian studies

Not medical advice. Much rarer than *2 in South Asians -- gnomAD puts it at 0.5% against 33% -- so for most readers the *2 row is the informative one. A full diplotype needs more alleles than this file reads.

BEB 2.3%GIH 0.48%ITU 0.49%PJL 1.6%STU 1.5%gnomAD afr 0.04%
Clinical Pharmacogenetics Implementation Consortium Guideline for CYP2C19 Genotype and Clopidogrel Therapy: 2022 Update. Clinical Pharmacology and Therapeutics, 2022. Guideline, systematic literature review PMID 35034351Prevalence of CYP2C19 Poor Metabolisers Among South Indian Psychiatric Patients. Annals of Neurosciences, 2025. South Indian patients PMID 41063924
CYP2C19*17, faster metabolismCYP2C19 rs12248560
Not found

You do not carry this variant.

Increases CYP2C19 activity, the opposite direction to the *2 and *3 variants already in this report. The Clinical Pharmacogenetics Implementation Consortium publishes prescribing guidance based on the combination of these variants; this report shows what was found and does not adjust any dose.

establisheduntested in South Asian studies

The prescribing guidance behind this variant is built mainly on European and East Asian cohorts. Its frequency and effect in South Asians are less well measured.

Clinical Pharmacogenetics Implementation Consortium guidelines for cytochrome P450-2C19 (CYP2C19) genotype and clopidogrel therapy. Clin Pharmacol Ther, 2011. CPIC guideline PMID 21716271
VKORC1 Warfarin sensitivity1 of 9 chip-readable variants
1 found
VKORC1 -1639, reduced enzyme expressionVKORC1 rs9923231
Found, one copy

You carry this variant.

This promoter variant lowers how much VKORC1 enzyme the liver produces. VKORC1 is the target of one class of anticoagulant, so the amount present matters to anyone taking one.

establisheduntested in South Asian studies

Not medical advice. Anticoagulant dosing is managed by blood tests that measure the actual effect in the actual person, which is far more informative than any genotype. Never change a dose on the basis of this page. Effect evidence is largely European and East Asian; no South Asian cohort is cited here.

BEB 15.7%GBR 35.7%GIH 17.5%ITU 9.3%PJL 19.8%STU 10.8%
Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Pharmacogenetics-Guided Warfarin Dosing: 2017 Update. Clinical Pharmacology and Therapeutics, 2017. Guideline, systematic literature review PMID 28198005
ACKR1 1 variant
1 found
Duffy-null blood groupACKR1 rs2814778
Found, one copy

You carry this variant.

The C allele abolishes Duffy antigen expression on red cells. It confers resistance to Plasmodium vivax malaria and is associated with a lower baseline neutrophil count that is benign -- 'benign ethnic neutropenia'.

establisheduntested in South Asian studies

The clinically important part is the neutrophil count, because a normal result for a Duffy-null person can be read as abnormal against a reference range built on people who are not. This is close to fixed in West African populations, absent in the South Asian and East Asian cohorts, and is on the panel because the report is global.

BEB 0.00%CEU 0.00%CHB 0.00%GBR 0.00%GIH 0.00%ITU 0.00%
Disruption of a GATA motif in the Duffy gene promoter abolishes erythroid gene expression in Duffy-negative individuals. Nature Genetics, 1995. African populations PMID 7663520
IFNL4 hepatitis C treatment response1 variant
1 found
IFNL4 (IL28B), hepatitis C treatment responseIFNL4 rs12979860
Found, one copy

You carry this variant.

Genotype at this locus predicts response to interferon-based hepatitis C therapy and spontaneous clearance of the virus. The C allele is the favourable one; the T allele is associated with poorer response.

establisheduntested in South Asian studies

This is a variant whose clinical relevance has largely PASSED. Interferon-based regimens have been superseded by direct-acting antivirals, which cure across genotypes, and it is reported here as a well-established association rather than as a treatment decision anybody should now make.

BEB 19.8%CEU 27.3%CHB 6.3%GBR 30.8%GIH 23.8%ITU 23.5%
Genetic variation in IL28B predicts hepatitis C treatment-induced viral clearance. Nature, 2009. Multi-ancestry cohort PMID 19684573
CYP3A5 tacrolimus metabolism1 variant
1 found
CYP3A5*3, tacrolimus metabolismCYP3A5 rs776746
Found, two copies

You carry this variant.

The commonest reason CYP3A5 produces no working enzyme. It is the main genetic factor in tacrolimus dosing after transplantation.

establisheduntested in South Asian studies

Frequencies for this variant differ sharply between populations, and the South Asian estimate is less well measured than the European and African ones.

Sequence diversity in CYP3A promoters and characterization of the genetic basis of polymorphic CYP3A5 expression. Nat Genet, 2001. Multi-population PMID 11279519CYP3A5 genotype predicts renal CYP3A activity and blood pressure in healthy adults. J Appl Physiol (1985), 2003. Healthy adults PMID 12754175
G6PD 3 of 225 chip-readable variants; 248 more no chip can see
2 read, not found
G6PD deficiency, A- variantG6PD rs1050828
Not found

You do not carry this variant.

The commonest G6PD-deficiency variant in populations of African ancestry. ClinVar classifies it Pathogenic/Likely pathogenic.

establisheduntested in South Asian studies

Carried here as a deliberate negative. In gnomAD v4 this variant is at 0.00030 in South Asians and 0.12277 in Africans — a 400-fold difference — while the Mediterranean variant on the row above runs the other way, 0.019 against 0.0002. A G6PD panel built on the well-known African variant would miss almost every deficient South Asian and would look like a working panel while doing it. That is the same failure as reporting lactase persistence from one European SNP, in a gene where the consequence is a drug reaction.

gnomAD afr 12.3%gnomAD eas 0.00%gnomAD nfe 0.01%gnomAD sas 0.03%
Glucose-6-phosphate dehydrogenase deficiency. Lancet, 2008. Review PMID 18177777
G6PD deficiency, Mediterranean variantG6PD rs5030868
Not found

You do not carry this variant.

This variant reduces glucose-6-phosphate dehydrogenase activity and is the commonest cause of G6PD deficiency reported in Indian populations. ClinVar classifies it Pathogenic/Likely pathogenic for G6PD deficiency.

establishedreplicated in South Asian studies

G6PD deficiency is diagnosed by an enzyme activity test, not by a genotype. This tells you a variant is present. It is also not the only cause of G6PD deficiency, and the others are not read by this file. Measured as readable: the GSA probe here is [A/G] on the plus strand, which interrogates the correct alternate, and the site is present in every real 23andMe and AncestryDNA file tested.

gnomAD afr 0.02%gnomAD eas 0.00%gnomAD mid 4.2%gnomAD nfe 0.02%gnomAD sas 1.9%
Glucose-6-phosphate dehydrogenase deficiency. Lancet, 2008. Review PMID 18177777Prevalence and spectrum of mutations causing G6PD deficiency in Indian populations. Infection, Genetics and Evolution, 2020. Indian populations PMID 33069889
G6PD deficiency, Orissa variantG6PD rs78478128
Not read: not on this chip

A variant reducing glucose-6-phosphate dehydrogenase activity, classified Pathogenic/Likely pathogenic in ClinVar, and the most frequently reported cause of G6PD deficiency in Indian series.

establishedreplicated in South Asian studies

The measurement worth quoting precisely, because it is easy to misread. Among 350 molecularly characterised G6PD-deficient individuals drawn from a screen of 20,896 people across India, this variant accounted for 56.5% of deleterious alleles and the Mediterranean variant for 23.6%. That is a share of the alleles found IN DEFICIENT PEOPLE, not a frequency in the population — overall deficiency prevalence in the same screen was 1.9%, ranging 0.8 to 6.3% by region. The two numbers answer different questions and only the second says how common deficiency is.

Prevalence and spectrum of mutations causing G6PD deficiency in Indian populations. Infection, Genetics and Evolution, 2020. 20,896 individuals screened across India; 350 deficient individuals characterised PMID 33069889Glucose-6-phosphate dehydrogenase deficiency. Lancet, 2008. Review PMID 18177777
HBB Beta-thalassaemia and sickle cell6 of 223 chip-readable variants; 272 more no chip can see
1 read, not found
Beta-thalassemia, Cap+1 A>CHBB rs34305195
Not found

You do not carry this variant.

A variant in the HBB transcription start region, classified Pathogenic/Likely pathogenic in ClinVar, and reported in Indian series as a mild beta-thalassemia allele.

establishedreplicated in South Asian studies

Reported in the literature as producing a milder phenotype than the splice-site and nonsense alleles, which is a statement about published series and not a prediction about any individual.

Genetic Heterogeneity of Beta Globin Mutations among Asian-Indians and Importance in Genetic Counselling and Diagnosis. Mediterranean Journal of Hematology and Infectious Diseases, 2013. Asian-Indian PMID 23350016Spectrum of beta-thalassemia mutations and their association with allelic sequence polymorphisms at the beta-globin gene cluster in an Eastern Indian population. American Journal of Hematology, 2002. Eastern Indian PMID 12210807
Beta-thalassemia, codon 15 G>AHBB rs33986703
Not read: ambiguous position

A nonsense variant in HBB, classified Pathogenic in ClinVar, and one of the beta-thalassemia alleles reported in Indian series.

establishedreplicated in South Asian studies

Two limits. This site is absent from the base Illumina GSA manifest and present in all three real 23andMe v5 files, because 23andMe adds custom content on top of the GSA platform and does not publish its site list — so the probe alleles here are unknown and unknowable from any source we are willing to use. It is also a T/A pair, which is strand-ambiguous, and until now we discarded such sites rather than resolve them — our policy rather than a limit of the array, whose three designs state the strand unanimously here. The row declines rather than reporting an absence.

gnomAD eas 0.04%gnomAD sas 0.00%
Genetic Heterogeneity of Beta Globin Mutations among Asian-Indians and Importance in Genetic Counselling and Diagnosis. Mediterranean Journal of Hematology and Infectious Diseases, 2013. Asian-Indian PMID 23350016
Beta-thalassemia, codon 30 G>CHBB rs33960103
Not read: ambiguous position

A variant at the codon 30 splice junction of HBB, classified Pathogenic in ClinVar and reported in Indian beta-thalassemia series.

establishedreplicated in South Asian studies

Two independent limits. ClinVar records two further pathogenic alleles at this coordinate, so an array whose probe reads one of those has said nothing about ours, and the row never reports an absence here. It is also a C/G pair, which is strand-ambiguous; whether the probe designs at this position state a strand unanimously has not been measured, so the drop is attributed to our own policy rather than to the array until it is.

Spectrum of beta-thalassemia mutations and their association with allelic sequence polymorphisms at the beta-globin gene cluster in an Eastern Indian population. American Journal of Hematology, 2002. Eastern Indian PMID 12210807
Beta-thalassemia, IVS1-1 G>THBB rs33971440
Not read: not readable

A splice-donor variant in HBB, classified Pathogenic in ClinVar, and among the beta-thalassemia alleles most frequently reported in Indian series after IVS1-5.

establishedreplicated in South Asian studies

Three separate pathogenic alleles sit at this coordinate, so an array whose probe reads a different one has said nothing about this variant. The row therefore never reports an absence here.

The phenotypic and molecular diversity of hemoglobinopathies in India: A review of 15 years of experience. International Journal of Laboratory Hematology, 2019. Indian, 15-year single-centre series PMID 30489691Genetic Heterogeneity of Beta Globin Mutations among Asian-Indians and Importance in Genetic Counselling and Diagnosis. Mediterranean Journal of Hematology and Infectious Diseases, 2013. Asian-Indian PMID 23350016
Beta-thalassemia, IVS1-5 G>CHBB rs33915217
Not read: ambiguous position

A splice-site variant in HBB, classified Pathogenic in ClinVar, and the most frequently reported beta-thalassemia allele in Indian series.

establishedreplicated in South Asian studies

Whether this variant can be read from a consumer array is unknown, and that is the honest answer rather than a hedge. Three different pathogenic alleles share this rsID at chr11:5226925 — C>A, C>G and C>T are separate ClinVar records — and the Illumina GSA manifest carries four probe designs at the position, one of which does interrogate C>G. Which design a vendor actually shipped is not stated in the file, which names the marker plainly with no suffix. So the site is typed, and what was typed cannot be determined.

The phenotypic and molecular diversity of hemoglobinopathies in India: A review of 15 years of experience. International Journal of Laboratory Hematology, 2019. Indian, 15-year single-centre series PMID 30489691Genetic Heterogeneity of Beta Globin Mutations among Asian-Indians and Importance in Genetic Counselling and Diagnosis. Mediterranean Journal of Hematology and Infectious Diseases, 2013. Asian-Indian PMID 23350016Spectrum of beta-thalassemia mutations and their association with allelic sequence polymorphisms at the beta-globin gene cluster in an Eastern Indian population. American Journal of Hematology, 2002. Eastern Indian PMID 12210807
Sickle cell variantHBB rs334
Not read: not on this chip

The HBB variant that produces haemoglobin S, classified Pathogenic in ClinVar for sickle cell disease.

establishedreplicated in South Asian studies

Not typed by 23andMe v5 at all — absent from all three real files tested, while present in a 2023 AncestryDNA file. So whether this row can be answered depends on which vendor and which chip version produced the upload, and the answer has to be computed per file rather than stated per vendor. It is also a T/A pair, which is strand-ambiguous, and unlike the other rows here the three probe designs at this position disagree about which strand they read, so no declaration exists to resolve it against. That is a limit of the array rather than a policy of ours.

The phenotypic and molecular diversity of hemoglobinopathies in India: A review of 15 years of experience. International Journal of Laboratory Hematology, 2019. Indian, 15-year single-centre series PMID 30489691
NUDT15 Thiopurine sensitivity1 of 3 chip-readable variants
1 read, not found
NUDT15 R139C, reduced enzyme activityNUDT15 rs116855232
Not found

You do not carry this variant.

The R139C change reduces NUDT15 enzyme activity, so an active metabolite the enzyme normally clears persists longer. CPIC assigns metaboliser status from NUDT15 and TPMT together.

establishedreplicated in South Asian studies

Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment is more dangerous than any genotype on this page. Show this to the prescribing doctor and let them decide. This variant is roughly 24 times commoner in South Asians than in Europeans, which is why it is here and why guidance derived from European cohorts has historically under-served this population.

BEB 6.4%GIH 3.9%ITU 7.8%PJL 8.3%STU 8.3%gnomAD afr 0.09%
Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Thiopurine Dosing Based on TPMT and NUDT15 Genotypes. Clinical Pharmacology and Therapeutics, 2026. Guideline, systematic literature review PMID 41618934Optimizing mercaptopurine therapy in indian pediatric ALL: The role of TPMT and NUDT15 genotyping. Cancer Treatment and Research Communications, 2026. Indian paediatric acute lymphoblastic leukaemia PMID 41819030
TPMT Thiopurine sensitivity1 of 9 chip-readable variants
1 read, not found
TPMT*3C, thiopurine S-methyltransferase activityTPMT rs1142345
Not found

You do not carry this variant.

TPMT is the second enzyme CPIC reads alongside NUDT15. The *3C allele reduces its activity.

establisheduntested in South Asian studies

Not medical advice. TPMT*3C is one of several TPMT alleles and this file reads one of them, so a normal result here does not establish normal TPMT activity. Enzyme activity is measurable directly and that test, not this one, is what a clinic would use. No South Asian cohort study of this allele's effect is cited here -- the frequency data below is South Asian, the effect evidence is not.

BEB 2.9%GIH 2.4%ITU 2.0%PJL 1.0%STU 0.49%gnomAD afr 5.5%
Clinical Pharmacogenetics Implementation Consortium (CPIC) Guideline for Thiopurine Dosing Based on TPMT and NUDT15 Genotypes. Clinical Pharmacology and Therapeutics, 2026. Guideline, systematic literature review PMID 41618934
SLCO1B1 Statin transport1 of 1 chip-readable variants
1 read, not found
SLCO1B1 V174A, reduced hepatic transportSLCO1B1 rs4149056
Not found

You do not carry this variant.

SLCO1B1 is a liver transporter: it moves certain compounds out of the blood and into hepatocytes, where they act and are cleared. The V174A change reduces that transport, so what the transporter handles stays in circulation longer.

establisheduntested in South Asian studies

Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment on the basis of a web page is more dangerous than anything this row describes. The variant is LESS common in South Asians than in Europeans -- 4.9% against 15.9% -- so for most readers here this row will be uninformative. Effect evidence is European; no South Asian cohort is cited.

BEB 5.2%GBR 14.3%GIH 1.9%ITU 6.4%PJL 3.6%STU 4.4%
The Clinical Pharmacogenetics Implementation Consortium Guideline for SLCO1B1, ABCG2, and CYP2C9 genotypes and Statin-Associated Musculoskeletal Symptoms. Clinical Pharmacology and Therapeutics, 2022. Guideline, systematic literature review PMID 35152405
ASPA Canavan disease carrier1 variant
1 read, not found
ASPA E285A, Canavan disease carrierASPA rs28940279
Not found

You do not carry this variant.

The commonest ASPA variant in Canavan disease, a recessive condition affecting the white matter of the brain. ClinVar classifies it Pathogenic/Likely pathogenic.

establisheduntested in South Asian studies

This variant's frequency has not been measured in South Asian populations. The Canavan literature is largely Ashkenazi Jewish, where it is commonest; what carrier frequency looks like in South Asia is not something this report can tell you.

Cloning of the human aspartoacylase cDNA and a common missense mutation in Canavan disease. Nat Genet, 1993. Canavan families PMID 8252036The frequency of the C854 mutation in the aspartoacylase gene in Ashkenazi Jews in Israel. Am J Hum Genet, 1994. Ashkenazi Jewish PMID 8037206
HFE hereditary haemochromatosis2 variants
1 read, not found
HFE C282Y, hereditary haemochromatosisHFE rs1800562
Not found

You do not carry this variant.

The commonest cause of hereditary haemochromatosis, in which the body absorbs more iron than it needs. ClinVar classifies it Pathogenic.

establisheduntested in South Asian studies

Almost all HFE research is in European-ancestry cohorts, where this variant is commonest. It is rare in South Asians and its frequency there is not well measured, so an absence here says less than it would for a European genome.

A novel MHC class I-like gene is mutated in patients with hereditary haemochromatosis. Nat Genet, 1996. Families with haemochromatosis PMID 8696333
HFE H63DHFE rs1799945
Not read: ambiguous position

A second HFE change described in the same 1996 paper as C282Y. It is common in many populations and, on its own, is not generally associated with iron overload; the combination most reported is one copy of this alongside one copy of C282Y.

moderateuntested in South Asian studies

ClinVar records conflicting classifications for this variant, and it is reported far more often than iron overload occurs, so it is shown as a finding rather than a conclusion. Its frequency in South Asians is not well measured.

A novel MHC class I-like gene is mutated in patients with hereditary haemochromatosis. Nat Genet, 1996. Families with haemochromatosis PMID 8696333
ABCG2 transporter function1 variant
1 read, not found
ABCG2 Q141K, transporter functionABCG2 rs2231142
Not found

You do not carry this variant.

The T allele reduces ABCG2 transporter function. It is associated with higher serum urate and with reduced response to allopurinol, and CPIC uses it in guidance on rosuvastatin dosing.

establisheduntested in South Asian studies

Urate and gout have large dietary and renal components that this variant does not read. The allopurinol association describes response at a given dose rather than whether the drug works.

BEB 12.2%CEU 11.6%CHB 31.1%GBR 14.3%GIH 6.8%ITU 10.8%
Identification of a urate transporter, ABCG2, with a common functional polymorphism causing gout. Proceedings of the National Academy of Sciences, 2009. European and African-American cohorts PMID 19506252
CYP2C9 reduced-function allele1 variant
1 read, not found
CYP2C9*3, reduced-function alleleCYP2C9 rs1057910
Not found

You do not carry this variant.

CYP2C9*3 substantially reduces enzyme activity. CPIC guidelines use CYP2C9 genotype in dosing recommendations for warfarin and for phenytoin.

establisheduntested in South Asian studies

A genotype is not a dose. CPIC recommendations combine CYP2C9 with VKORC1 and with clinical factors, and warfarin is titrated on INR measurement whatever the genotype says. Carrying *3 is a reason to expect a lower dose requirement, not a reason to change one.

BEB 11.6%CEU 6.6%CHB 3.9%GBR 7.1%GIH 13.1%ITU 10.3%
The role of the CYP2C9-Leu359 allelic variant in the tolbutamide polymorphism. Pharmacogenetics, 1996. Human liver samples PMID 8946472
How we worked this out

Rows report variant presence against published work. The catalogue (v0.1.0) is a small selection of the variants known in each gene: ClinVar records hundreds more, and roughly half of those are deletions, duplications or repeat expansions that no genotyping chip can detect at all. Metaboliser status for CYP2C19 or TPMT needs a full star-allele diplotype and an enzyme activity test, which this file cannot give. G6PD deficiency is diagnosed by an enzyme test, not a genotype.

Whether a position can be read depends on the vendor and chip version that produced the upload, and it is measured for each file. Frequencies shown are from gnomAD v4 and the 1000 Genomes Project. The evidence for an effect often comes from European studies, and each row's tags say where it has been tested.

Technical

For the curious

The detail behind the numbers: the reference groups, your file, and the same file through other tools.

Reference groups closest to you

How close your DNA sits to each group in our reference set.

Closest areas

Indo-Gangetic PlainPuerto Rican in Puerto RicoMexican Ancestry in Los Angeles

The areas whose reference samples your DNA is nearest to. No order, and no ranking.

typical rangefar outside
Balochi
Iberian
Mexican ancestry
Colombian
Puerto Rican
7 further groups, none of them close
Makrani
Tuscan
Pathan
Peruvian
Punjabi
British
African ancestry
Inside the group's rangeClose to the groupOutsideColour is the group's region.
Groups are drawn where they are from. Several were sampled abroad; hover a row for its sampling place.
How we worked this out

Each file and every individual reference genome is projected into the first ten principal components of our panel. The chart is Mahalanobis distance in that space, so it scales each group by its own spread. The green band is where a typical member sits, the amber band is close but outside it, and a marker farther right is less like that group. Hover or tap a row for its distance. The closest areas come from the nearest three individual references, coarsened and left unranked.

On the same 289 people and 54 groups, this distance placed the correct group first 60.2% of the time against 35.3% for G25, and in the top three 74.0% against 63.0%. The G25 coordinates in this report are an export for other tools.

Present-day map

Where your file falls among living reference groups

Each dot is one person in our reference panel, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.

2,000 reference individuals across 16 regions, on the first two principal componentsEastern ChineseWest African forestDeccan and the Tamil PlainsWest African savannaIberianThe Ganges Plain and GujaratEastern and Southern AfricanMainland Southeast AsiaFinland and KareliaBengal and the central beltYou

2,000 reference individuals, the same cloud a report draws, on the panel fitted today. The first two components carry 13.9% and 5.4% of total variance. Every point is one published reference individual and names itself on hover. Clusters are labelled with the region each cohort is pooled into and coloured by the continental component that region sits in. 504 individuals from cohorts pooled into no region are not drawn.

The picture is turned a quarter clockwise and mirrored, which is why it resembles a map: Europe upper left, East Asia upper right, Africa along the bottom, South Asia between them. That is a rotation and a reflection of the plane, so no point has moved relative to any other and every distance is what it was. It is a reading aid and nothing more. The axes have no meaning of their own: they are directions of maximum variance, their signs are arbitrary, and a region's position depends on which other regions are in the panel. The first component runs vertically here and the second horizontally.

Every region carries a numbered disc at its centre and an outline around its members, and the key below repeats the number — so a region too crowded to spell out on the plot can still be found on it. The outline is the convex hull of that region's own individuals with the furthest 8% trimmed off, which keeps one stray from dragging a boundary across the chart; it encloses people, not territory. Spelled out only in the key at this size: 4 Northwest European, 14 Tai and Kadai, 13 Punjab and Kashmir, 16 Sierra Leone and the Upper Guinea coast, 8 Japanese, 7 Italian.

African

  • 2West African forest207
  • 5West African savanna113
  • 10Eastern and Southern African99
  • 16Sierra Leone and the Upper Guinea coast85

Bengal and the central belt

  • 15Bengal and the central belt86

East Asian

  • 1Eastern Chinese208
  • 8Japanese104
  • 11Mainland Southeast Asia99
  • 14Tai and Kadai93

European

  • 4Northwest European190
  • 6Iberian107
  • 7Italian107
  • 12Finland and Karelia99

Indo-Gangetic Plain

  • 9The Ganges Plain and Gujarat103
  • 13Punjab and Kashmir96

Southern Peninsula

  • 3Deccan and the Tamil Plains204

10,279 of 12,770 panel markers placed you on the map.

The groups nearest your position, closest first:

  1. Balochi
  2. Iberian
  3. Mexican ancestry

This is a position on a plot. Sitting near a group means your genome resembles that group on these axes. It does not say who your ancestors were.

How we worked this out

The axes come from the reference panel alone. Your file is projected onto them using the markers it shares with the panel, so it moves no point on the map.

Nearest is measured over more axes than the two drawn, to the middle of each group, so a group that looks close here can rank lower in the list.

Ancient source map

Where your file falls among ancient genomes

Each dot is one ancient person, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.

558 ancient individuals in 14 source groups, on the first two principal componentsChina, c. 2,000 BPChina, c. 4,500 BPTurkey, c. 8,500 BPRussia, c. 8,500 BPPapua New Guinea, c. 500 BPIran, c. 10,000 BPChina, c. 1,500 BPTaiwan, c. 1,500 BPSteppe and Siberia, c. 5,500 BPRussia, c. 32,500 BPCameroon, c. 8,000 BPSouth Africa, c. 2,500 BPYou

558 excavated individuals with at least 100,000 1240K SNPs, on the first two principal components (5.1% and 1.7% of variance). 41 individuals sit far from their group's centroid and are left out of every outline and centre; they are not drawn here, and the methods page shows them. Axes have no meaning of their own, and a group's position depends on which other groups are in the map. 4 sources pooled from several regions are not drawn: a period is not a people.

Named only in the key at this size: 6 Sweden, c. 7,500 BP, 9 Japan, c. 4,500 BP.

Before 10,000 BP

  • 1Russia, c. 32,500 BP12

10,000 to 7,000 BP

  • 2Iran, c. 10,000 BP26 +2
  • 3Turkey, c. 8,500 BP85 +9
  • 4Russia, c. 8,500 BP48 +3
  • 5Cameroon, c. 8,000 BP8
  • 6Sweden, c. 7,500 BP17 +2

7,000 to 5,000 BP

  • 7Steppe and Siberia, c. 5,500 BP15 +1

5,000 to 3,500 BP

  • 8China, c. 4,500 BP113 +3
  • 9Japan, c. 4,500 BP14 +1

After 3,500 BP

  • 10South Africa, c. 2,500 BP7
  • 11China, c. 2,000 BP136 +14
  • 12Taiwan, c. 1,500 BP21 +4
  • 13China, c. 1,500 BP24
  • 14Papua New Guinea, c. 500 BP32 +2

Counts are individuals kept; +n is individuals dropped. Thin: fewer than 5 usable individuals. An outline is drawn around each excavation group of six or more people, not around a source.

130,334 of 138,630 sites in your file are on the map.

The sources nearest your position, closest first:

  1. Iran, c. 10,000 BP
  2. Turkey, c. 8,500 BP
  3. Russia, c. 32,500 BP

This is a position on a plot. Sitting near a source means your genome resembles that group on these axes. It does not say who your ancestors were.

How we worked this out

The axes come from the ancient individuals alone. Your file is projected onto them using the sites it shares with them, so it moves no point on the map.

Nearest is measured over the first four axes, to the middle of each source. The plot shows two of them, so a source that looks close here can rank lower in the list.

On eleven test genomes the nearest source fell in the same region as the qpAdm fit. Inside a region, the order of the nearest sources does not track the qpAdm proportions.

The same ancestor, twice

Stretches inherited twice

Places where both copies of a chromosome came from the same ancestor. A long stretch points to a recent shared ancestor.

13stretches
1.34%of your genome
8.6 cMlongest
7.8 cM
25–50 generations700–1,400 years ago
5 stretches
13.7 cM
12.5–25 generations350–700 years ago
4 stretches
25.9 cM
5–12.5 generations140–350 years ago
4 stretches
none
under 5 generationswithin 140 years
0 stretches

Oldest first. Bar heights compare your own bands with each other.

Where they sit

Each bar is one chromosome, drawn to length. A mark is a stretch where both of your copies match, coloured by how far back the shared ancestor sits.

1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
1-2 cM · 700–1,400 years2-4 cM · 350–700 years4-10 cM · 140–350 years10 cM and longer · under 140 years
How we worked this out

Runs of homozygosity need no reference panel. The question is whether your two copies match, so this runs on 596,515 of your own markers, more than the 12,770 that reach the ancestry estimate. It is the one chapter in this report that our reference data does not limit. In total the stretches cover 47.4 cM of 3,545 cM.

Two different detectors were run over your file and they agree on 99% of the length. They fail in different ways, so where they agree the result is not an artefact of either one.

Lengths are in centimorgans, read off a recombination map, not in megabases. Near the middle of a chromosome a megabase can be a fifth of a centimorgan and near the end it can be two, so a generation count computed from physical distance is wrong by a factor that changes along every chromosome. Generations follow from genetic length as roughly 50 divided by the length in centimorgans, and a year figure uses 28 years to a generation.

Each stretch is shown as a range of generations. A 20 cM stretch averages about two and a half generations back, but over a very wide spread, and one stretch is a single draw from it. A range is what the evidence supports.

What this does not say. It says nothing about whether your parents are related. That question needs a comparison against the population you descend from, and we do not publish it. In a community that has married within itself for centuries the background is high with no close relatedness at all, so a fixed threshold would tell many people something untrue about their own family.

Your file

How complete your raw DNA file was. This sets how precise every other chapter can be.

B

B: some chapters are limited; each one says so.

638,463markers in the file
97.2%read successfully
23andMe, build 37recognised from the file's contents; Y chromosome present

Some checks flagged this file; the affected chapters say so where you read them. Nothing was sequenced by us. We read the export you already had.

How we worked this out

Heterozygosity is 16.6%: the share of read positions where your two copies differ. A consumer chip usually reads between 24% and 31%; well outside that range points to a file problem rather than to anything about you.

Your file is kept and used to build future reference panels, under a Creative Commons licence. We never sell or release an individual genome, and no genetic data is ever written to a log.

Who you were compared against

524 groups, 10,922 people. The reference behind every number in this report.

Every study behind the panel, and its regions

People alive today: 6,745 people, 18 studies. The reference this report measures your ancestry against.

The1000GenomesProjectConsortiumNature2015
2,320
unrecorded
1,544
PattersonReichGenetics2012
1,003
NakatsukaReichNatGenet2017
619
LazaridisKrauseNature2014
494
WangReichNature2021
207
BiaginiCalafellEJHG2019
120
LazaridisReichNature2016
83
EBC India 2023
81
YangZhangJHumGenet2018
80
MondalBertranpetitNatGenet2016
70
JeongKrauseNatEcolEvol2019
65
MallickReichNature2016
21
BergströmTyler-SmithScience2020
15
EBC Munda 2023
10
DamgaardWillerslevScience2018
10
RaghavanWillerslevNature2014
2
RasmussenWillerslevNature2014
1

Excavated individuals: 13,454 people, 309 studies. Used by the ancient peoples chapters. The region percentages do not use them.

PattersonReichNature2021
670
LazaridisReichScience2022
602
WangHofmanováNature2025
543
NarasimhanReichScience2019
476
Gnecchi-RusconeHofmanováNature2024
399
MargaryanWillerslevNature2020
394
BenekerKivisildGenomeBiol2025
332
OlaldeReichNature2018
318
GretzingerSchiffelsNature2022
297
AllentoftWillerslevNature2024
281
LazaridisReichNature2025
265
GelabertReichBioRxiv2023
246
the other 297 studies
MarótiTörökCurrBiol2022
210
MathiesonReichNature2018
205
OlaldeReichScience2019
199
PapacHaakSciAdv2021
197
AntonioPritchardeLife2024
181
JeongWarinnerCell2020
177
Maravall-LópezNoresNature2025
169
FernandesReichNature2021
165
StolarekFiglerowiczGenomeBiol
165
HuiKivisildSciAdv2024
148
ZengReichNature2025
146
RingbauerReichNature2025
145
KumarFuScience2022
139
LiuReichScience2022
133
Villalba-MoucoHaakSciAdv2021
131
GyurisSzécsényi-NagyCell2025
129
AntonioPritchardScience2019
129
DamgaardWillerslevNature2018
128
MathiesonReichNature2015
124
PenskeHaakNature2023
123
Szécsényi-NagySiklósiNatComm2025
118
OlaldeReichCell2023
115
LipsonReichNature2017
106
WangReichNature2021
100
Gnecchi-RusconeKrauseSciAdv2021
98
GhalichiHaakNature2024
97
NakatsukaReichNature2023
96
FurtwänglerHerbigBioTechniques2020
89
SkourtaniotiStockhammerNatEcolEvol2023
88
SeersholmSikoraNature2024
88
SkourtaniotiKrauseCell2020
86
RivollatHaakSciAdv2020
83
SaagThomasSciAdv2025
83
GerberSzécsényi-NagySciAdv2024
81
BarrieWillerslevNature2024
78
WangFuSciAdv2023
75
PosthKrauseNature2023
74
LiuFuNatComm2025
74
AllentoftWillerslevNature2015
73
RivollatHaakNature2023
70
DamgaardWillerslevScience2018
68
PosthKrauseSciAdv2021
65
Gnecchi-RusconeKrauseCell2022
60
BrunelPruvostPNAS2020
58
AmorimVeeramahNatComm2018
56
SirakReichNatComm2021
56
CassidyBradleyNature2025
56
RavasiniTrombettaGenomeBiol2024
54
Agranat-TamirReichCell2020
54
FernandesReichNatEcolEvol2020
53
MarcusNovembreNatComm2020
52
ScheibKivisildScience2018
51
NikitinReichNature2025
51
NingCuiNatComm2020
48
WangHaakNatComm2019
48
NakatsukaFehren-SchmitzCell2020
48
ReitsemaReichPNAS2022
47
PenskeHaakSciRep2024
46
NägeleSchroederScience2020
45
BrielleKusimbaNature2023
44
SassoKivisildPNAS2024
41
MichelKrauseNature2024
41
FlegontovSchiffelsNature2019
40
LazaridisReichNature2016
40
CassidyBradleyNature2020
38
VeeramahBurgerPNAS2018
37
FisherPruvostiScience2022
37
PosthReichCell2018
37
ImmelKrause-KyoraCommunBiol2021
37
Rodríguez-VarelaGötherströmCell2023
37
ChyleńskiMalmströmNatComm2023
37
HarneyRaiNatComm2019
36
BaiFuCurrBiol2024
35
NovakReichPLoSOne2021
33
PrendergastReichScience2019
33
MattilaJakobssonCommunBiol2023
33
OlaldeReichNature2026
32
Gnecchi-RusconeHofmanováPNAS2025
32
ChildebayevaHaakMolBiolEvol2022
32
CarlhoffKrauseNatComm2023
31
FerrazPosthNatEcolEvol2023
30
FowlerReichNature2021
30
BarqueraKrauseNature2024
30
BlöcherBurgerPNAS2023
30
PeltolaOnkamoCurrBiol2023
29
SirakReichNatEcolEvol2024
29
SaagMetspaluSciAdv2021
28
KoptekinSomelCurrBiol2023
28
LiuJeongNatComm2022
28
MootsPinhasiNatEcolEvol2023
28
ArzelieriPruvostProcBiolSci2024
27
WaldmanReichCell2022
27
MittnikKrauseScience2019
26
Rodríguez-VarelaGötherströmSciAdv2024
26
GretzingerSchiffelsNatHumBehav2024
25
FangWangCellRep2025
25
JärveVillemsCurrBiol2019
25
DuliasRichardsPNAS2022
24
ŽegaracBurgerSciRep2021
24
SchroederAllentoftPNAS2019
24
Seguin-OrlandoOrlandoCurrBiol2021
23
WangFuCell2021
23
EbenesersdóttirHelgasonScience2018
23
McCollWillerslevScience2018
23
KrzewińskaGötherströmSciAdv2018
23
HarneyReichBioRxiv2022
22
FreilichPinhasiSciRep2021
22
MaoFuCell2021
22
KrettekPosthSciAdv2025
21
HarneyReichNatComm2018
21
ZagorcPinhasiArchaeolAnthropolSci2024
20
JeongWarinnerPNAS2018
19
MittnikKrauseNatComm2018
19
GelabertPinhasiSciRep2022
19
Moreno-MayarWillerslevScience2018
18
SikoraWillerslevNature2019
18
MarciniakPerryPNAS2022
18
NakatsukaReichNatComm2020
18
KrzewińskaGötherströmCurrBiol2018
18
GambaPinhasiNatComm2014
17
PopovićBacaSciAdv2021
17
FernandesPinhasiSciRep2018
17
YuKrauseiScience2022
17
unrecorded
16
OliveiraStonekingNatEcolEvol2022
16
LazaridisStamatoyannopoulosNature2017
16
BraceBarnesNatEcolEvol2019
16
Sánchez-QuintoJakobssonPNAS2019
16
KennettReichNatComm2022
15
ZhangCuiNature2021
14
MartinianoBradleyPLosGenet2017
14
WangSchiffelsSciAdv2020
14
LiuReichCell2026
14
KumarFuMolBioEvol2021
14
LinderholmKrzewińskaSciRep2020
14
MarchiExcoffierCell2022
13
FregelBustamantePNAS2018
13
YangFuScience2020
13
HarneyPinhasiGenomeRes2021
13
KılınçGötherströmSciAdv2021
13
HaberTyler-SmithAJHG2019
13
HaberTyler-SmithAJHG2020
13
BurgerWegmannCurrBiol2020
13
RobbeetsNingNature2021
12
YuKrauseCell2020
12
CookeNakagomeSciAdv2021
12
RaghavanWillerslevScience2015
11
LamnidisSchiffelsNatComm2018
11
TaoWangCurrBiol2023
11
PosthPowellNatEcolEvol2018
11
SkoglundReichCell2017
11
LipsonReichCurrBiol2018
11
MalmströmJakobssonProcBiolSci2019
11
RymbekovaKuhlwilmBioRxiv2025
10
YakaSomelCurrBiol2021
10
SchiffelsDurbinNatComm2016
10
MartinianoBradleyNatComm2016
9
FeldmanKrauseSciAdv2019
9
SaupeScheibCurrBiol2021
9
NingCuiCurrBiol2019
9
LeeGakuhariHPGG2024
9
TieslerReichAntiquity2022
9
CoutinhoJakobssonAJBA2020
9
ZhuWeniScience2022
8
ParasayanGeiglSciAdv2024
8
FuReichNature2016
8
UnterländerBurgerNatComm2017
8
LipsonReichScience2018
8
LipsonReichNature2025
8
ArmitReichAntiquity2023
8
JonesBradleyCurrBiol2017
8
WrightLambertSciAdv2018
8
O’SullivanMaixnerSciAdv2018
8
CapodiferroAchilliCell2021
8
VanDeLoosdrechtKrauseScience2018
8
SimõesJakobssonPNAS2024
8
GelabertPinhasiCurrBiol2022
7
FeldmanKrauseNatComm2019
7
ChildebayevaHaakCommunBiol2024
7
KılınçGötherströmCurrBiol2016
7
WangPosthCurrBiol2023
7
LvWangBMCBiol2024
7
SaagTambetsCurrBiol2019
6
IngmanStockhammerPLoSOne2021
6
SaagMetspaluCurrBiol2017
6
Villalba-MoucoHaakCurrBiol2019
6
González-FortesHofreiterCurrBiol2017
6
ClementePapageorgopoulouCell2021
6
ZhangNingArchaeolAnthropolSci2024
6
GretzingerSchiffelsNatEcolEvol2024
6
BongersFehren-SchmitzPNAS2020
6
SimõesJakobssonNature2023
6
BroushakiBurgerScience2016
5
WangStockhammerPNAS2023
5
HofmanováBurgerPNAS2016
5
JeongWarinnerPNAS2016
5
LipsonReichCurrBiol2020
5
VeselkaCattelainAntiquity2024
5
BagnascoSciRep2024
5
AneliPaganiMolBiolEvol2022
5
BraceBarnesCurrBiol2022
5
Villa-IslasÁvila-ArcosScience2023
4
MartinianoDurbinCellGenom2024
4
WhiteOlaldeBiology2021
4
SpyrouKrauseNature2022
4
HaberTyler-SmithAJHG2017
4
GüntherJakobssonPLoSBiol2018
4
LipsonReichNature2020
4
KennettPerryNatComm2017
4
PilliMittnikCurrBiol2024
4
FernandesMegalithic_Unpublished
4
LindoDiRienzoSciAdv2018
4
delaFuenteMoragaPNAS2018
4
SharkoNedoluzhkoEJHG2024
4
SikoraWillerslevScience2017
4
CassidyBradleyPNAS2016
4
SchlebuschJakobssonScience2017
4
SkoglundJakobssonScience2014
3
JonesBradleyNatComm2015
3
ReichPääboNature2010
3
ImmelKrause-KyoraSciRep2020
3
LazaridisKrauseNature2014
3
HaakReichNature2015
3
LipsonPrendergastNature2022
3
González-FortesBarbujaniProcBiolSci2019
3
Guarino-VignonBonCommunBiol2023
3
SümerKrauseNature2024
3
BarqueraKrauseCurrBiol2020
3
Sandoval-VelescoSchroederAmJHumGenet2023
3
SchroederGilbertPNAS2015
3
AltınışıkSomelSciAdv2022
3
Rodríguez-VarelaGirdland-FlinkCurrBiol2017
3
ScheibKivisildHumBiol2019
2
GüntherJakobssonPNAS2015
2
RaghavanWillerslevNature2013
2
PrüferPääboNature2013
2
MorezGirdland-FlinkPLoSGenet2023
2
HajdinjakPääboNature2021
2
UllingerIngramNEA2022
2
SkoglundReichNature2016
2
RohlandReichGenomeRes2022
2
Seguin-OrlandoWillerslevScience2014
2
Seguin-OrlandoOrlandoiScience2021
2
RaghavanWillerslevScience2014
2
ScorranoSikoraCommunBiol2022
2
GreenPääboScience2010
2
Nieves-ColonMolBioEvol2020
2
DeAngelisRickardsGenes2022
2
Moreno-MayarWillerslevNature2018
2
AlvesDinaNatComm2024
2
OmrakGötherströmCurrBiol2016
2
LlorenteManicaScience2015
2
MassilaniPääboScience2020
2
SrigyanValdioseraCommunBiol2022
2
BennettGeiglNatEcolEvol2023
1
MalaspinasWillerslevCurrBiol2014
1
OlaldeLalueza-FoxMolBiolEvol2015
1
LindoFigueiroPNASNexus2022
1
MafessoniPääboPNAS2020
1
EsselMeyerNature2023
1
SlonPääboNature2018
1
SiskaManicaSciAdv2017
1
WeberPinhasiSciRep2025
1
RivollatDeguillouxPNAS2022
1
SilvaRichardsJArchSci2026
1
FoodyEdwardsAntiquity2025
1
vandenBrinkReichLevant2017
1
SedigReichAntiquity2024
1
Teschler-NicolaPinhasiCommunBiol2020
1
AgelarakisAgelarakisJournalModernHellenism2026
1
ShindeReichCell2019
1
NikitinReichPLoSOne2023
1
KellerZinkNatComm2012
1
RasmussenWillerslevNature2010
1
SchuenemannKrauseNatComm2017
1
OlaldeLalueza-FoxNature2014
1
KeyKrauseNatEcolEvol2020
1
ZallouaMatisoo-SmithSciRep2018
1
GerberSzécsényi-NagyBioRxiv2024
1
SlimakSikoraCellGenom2024
1
HigginsCaramelliNatComm2024
1
SatoIshidaGenomeBiolEvol2021
1
ZhurTrifonoviScience2024
1
FuPääboNature2015
1
SchroederWillerslevPNAS2018
1
EgfjordAllentoftPLoSOne2021
1
JensenSchroederNatComm2019
1
YangFuCurrBiol2017
1
Guarino-VignonBonFrontGenet2022
1
FuPääboNature2014
1
ArianoBradleyCurrBiol2022
1
PrüferKrauseNatEcolEvol2021
1
MargaryanGilbertAdvGenet2021
1
ScorranoMacciardiSciRep2022
1
NedoluzhkoIlinskyResSq2022
1
ValdioseraJakobssonPNAS2018
1
LarenaJakobssonPNAS2021
1

Behind the South Asian regions: 2,148 people across 11 studies, counted by the study that collected them.

NIBMG Indian population panel
557
1000 Genomes Project
472
Coorg and South Indian cohorts
424
South Indian cohort (GSE242813)
181
Allen Ancient DNA Resource
162
Estonian Biocentre, India
140
India-recruited (mixed sources)
89
Indian arrays (GSE93037)
65
Estonian Biocentre, Munda
26
Xing et al. worldwide panel
24
Khasi (northeast India)
8
South Asia 489Europe 503East Asia 504The Americas 347Africa 504Recently mixed 157
100%correct on 120 people held out of the fit
100%correct at the marker density of a typical consumer file
How we worked this out

Accuracy is the share of held-out people whose largest region came out right: 120 people from 24 groups at the full panel of 12,770 markers, and again at 10,317 markers, the density a real 23andMe export reaches. The people tested are part of the reference set they are scored against, so their own data helped define the groups. Someone uploading a file gets no such help, so the real-world rate is a little lower. Five individuals per population. A perfect score on 120 people is consistent with a true rate above about 97%. The Indigenous American component is built from four samples, three of them majority European, so a correct call there is an easier test than it sounds. It measures which region comes out largest. Whether the printed range covers the true value is a separate measurement we have not run. 10 of the held-out set are not scored (African Caribbean and African American samples, which define no component). Measured 2026-09-09 against the 12,770-marker panel this report used. A rebuilt panel re-runs it.

A few dozen people per group is what makes our ranges as wide as they are. A narrower range would need more people. West Asia is fitted from Bedouin, Palestinian and Druze samples, and Oceania from pooled Papuan and Bougainville individuals. Central Asia is not yet a region of its own.

Coordinates for other tools

Twenty-five numbers you can paste into community ancestry tools.

Not measured for this file

this map was fitted and checked on South Asian genomes, and this file is 19% South Asian. Placing it would mean extrapolating a map nobody has verified there

Other calculators

The same file run through popular community calculators. Where they agree with our result, the answer does not depend on the model.

Not measured for this file

Community calculators need more of their own markers than this file carries, so none were run.