Overview
Your ancestry report
Your ancestry is mostly East Asian (49%), with some Oceanian (42%) and a little from elsewhere.
Analysed 31 August 2026
Tap any bar or region to see how sure we are.
Health & traits
Your chip covers 6 of the 24 health variants we check. We found 2. Of 23 trait variants, your chip covers 10; 2 found.
Go to Health & traitsTechnical
Your file was graded A, with 562,605 markers read. Closest reference groups, your file and other calculators are here.
Go to TechnicalAncestry
Ancestry
Where you come from
Recent regions first, then your chromosomes, your mother's and father's lines, and thousands of years back.
Where your ancestors lived
Each bar is the share of your DNA that matches people living in that region today.
How we worked this out
Each percentage is a share of the 4,046 markers we could read in your file (32% of 12,770), compared with people sampled today by the 1000 Genomes Project. The solid part of a bar is the share we are sure of. The faint part is how much higher it might be. The line is our best estimate. The range is a 95% interval from resampling your markers. If a population is missing from our reference set, your ancestry from it is counted under the nearest region we do have.
Each region and its range. East Asian 49.5% (likely 42–56%); Oceanian 42.3% (likely 38–46%); South Asian 8.2% (likely 2–16%).
South Asian is one region. These seven regional components share 96–99% of their allele frequencies, so how much South Asian ancestry you have is measured far more precisely than how it divides between the seven. Its total (8.2%) is firmer than the split inside it. The parts drawn inside it add to 0.0%, because some are too small to tell from zero and are not drawn.
East Asian is one region. How much of your ancestry is from this region is measured from the whole panel at once. How it divides between the areas inside it is a second, harder question, fitted separately — so the total is the firmer number of the two. Its total (49.5%) is firmer than the split inside it. The parts drawn inside it add to 38.9%, because some are too small to tell from zero and are not drawn.
Oceanian is one region. How much of your ancestry is from this region is measured from the whole panel at once. How it divides between the areas inside it is a second, harder question, fitted separately — so the total is the firmer number of the two. Its total (42.3%) is firmer than the split inside it. The parts drawn inside it add to 40.0%, because some are too small to tell from zero and are not drawn.
What “could not be placed” means. A region whose range reaches zero is not drawn on its own. The ancestry is still yours; this file does not have enough markers to say which region it belongs to. Here the regions that could hold it are African (up to 2%).
The map. Shading marks where each ancestry lives, at your own share of it. The outlines are geographic regions and the numbers are genetic, so an edge is approximate: ancestry shades into its neighbours. A faint region is one this file cannot tell from zero.
| Region | Fitted from | Estimate | Range |
|---|---|---|---|
| Lesser Sunda Islands | individuals recruited in Flores, Sumba, Alor, Pantar, Lembata and Timor | 35.3% | 28–37% |
| Island Southeast Asia | individuals recruited in Sulawesi, the Philippines and northern Borneo | 22.5% | 10–26% |
| Taiwan | individuals recruited among the indigenous peoples of Taiwan | 16.4% | 10–23% |
| Nias and the Mentawai Islands | individuals recruited on the islands west of Sumatra | 4.7% | 3–5% |
| Andaman and Nicobar Islands | Onge, Jarawa and Great Andamanese individuals | below resolution | 0–8% |
| Chota Nagpur Plateau | individuals recruited in Jharkhand, Odisha and the eastern Ghats | below resolution | 0–8% |
| Hmong and Mien | individuals recruited in the uplands of southern China | below resolution | 0–6% |
| Orang Asli | Semang and Senoi individuals recruited in peninsular Malaysia | below resolution | 0–3% |
| Southeast Asian | Mainland Southeast Asia, fitted from Cambodian, Thai, Kinh_Vietnamese | below resolution | 0–4% |
| Polynesia | individuals recruited in Samoa and Tonga | below resolution | 0–2% |
| Tai and Kadai | individuals recruited in Yunnan, Thailand, Laos and southern China | below resolution | 0–13% |
| Semang | Bateq, Jehai, Kintaq, Mendriq and Maniq individuals recruited in peninsular Malaysia and southern Thailand | below resolution | 0–4% |
| Western Indonesia | individuals recruited in Sumatra, Java, Bali, Nias and the Mentawai islands | below resolution | 0–1% |
| Timor and Lembata | individuals recruited on Timor and Lembata | below resolution | 0–7% |
| Himalayas and Northeast India | individuals recruited in Nepal, the terai, and the northeastern hills | below resolution | 0–8% |
| Mainland Southeast Asia | individuals recruited in Cambodia and Vietnam | below resolution | 0.00–0.01% |
| Bengal and the central belt | individuals recruited in Bengal, Madhya Pradesh, Chhattisgarh and Maharashtra | below resolution | 0–6% |
| African | Yoruba, Luhya, Gambian, Mende and Esan reference samples | below resolution | 0–2% |
| Indigenous American | Peruvian, Mexican, Colombian and Puerto Rican reference samples | below resolution | under 0.01% |
| European | Utah, Tuscan, Finnish, British and Iberian reference samples | below resolution | under 0.01% |
| West Asian | Bedouin, Palestinian and Druze reference samples | below resolution | 0.0–0.1% |
| Indo-Gangetic Plain | individuals recruited in Punjab, Sindh, Kashmir, Rajasthan, Gujarat and along the Ganges | below resolution | under 0.01% |
| Southern Peninsula | individuals recruited in Tamil Nadu, Andhra Pradesh, Telangana, Karnataka and Kerala | below resolution | under 0.01% |
| Makran and the western ranges | individuals recruited across Balochistan and Makran | below resolution | under 0.01% |
| Siberian | 21 West Siberian, South Siberian and Amur populations — 498 individuals | below resolution | under 0.01% |
| Central Asian | 155 individuals from six Central Asian populations — the oases and the Kazakh steppe | below resolution | under 0.01% |
| North African | The Maghreb — Morocco, Algeria, Tunisia and Libya, fitted from Mozabite | below resolution | under 0.01% |
| Eastern Chinese | Eastern and central China, fitted from Han | below resolution | under 0.01% |
| Japanese | The Japanese archipelago, fitted from JPT, Japanese | below resolution | under 0.01% |
| Mlabri | Mlabri individuals recruited in Nan province, Thailand | below resolution | under 0.01% |
| Northern East Asian | Northern China, Mongolia and the Amur basin, fitted from CHB, Mongola, Buryat, Tu | below resolution | under 0.01% |
| Northwest China | individuals recruited in Gansu, Qinghai and the Hexi corridor | below resolution | under 0.01% |
| Southwest China | individuals recruited in Yunnan and Sichuan | below resolution | under 0.01% |
| Taimyr | Nganasan individuals recruited on the Taimyr peninsula | below resolution | under 0.01% |
| Tibetan Plateau | The Tibetan plateau, fitted from Tibetan | below resolution | under 0.01% |
| Burmese Uplands | Akha, Burmese, Karen and Tai Lue individuals recruited in Myanmar and northern Thailand | below resolution | under 0.01% |
| Lower Amur | Hezhen and Oroqen individuals recruited on the lower Amur | below resolution | under 0.01% |
| Island Melanesia | individuals recruited in Bougainville and New Britain | below resolution | under 0.01% |
| Papuan and Near Oceanian | New Guinea and the Bismarck archipelago, fitted from Papuan, Nasioi | below resolution | under 0.01% |
| Luzon and the Visayas | Luzon and the Visayas, fitted from Igorot, Luz, Vizaya | below resolution | under 0.01% |
| Remote Oceania | individuals recruited in Vanuatu, Micronesia, Samoa and Tonga | below resolution | under 0.01% |
| Sumba and Flores | individuals recruited on Sumba and Flores | below resolution | under 0.01% |
Your chromosomes, painted
Each chromosome coloured by which region its stretches match. Long stretches point to recent ancestors; this sees back about nine generations.
This genome has not been painted yet. Painting is computed separately from the rest of the report and is being rolled out across the demo board.
Your mother's line and your father's line
Two threads pass down almost unchanged: one from mother to child, one from father to son. Each follows a single ancestor across thousands of years.
this file's mitochondrial positions do not line up with the rCRS, so a haplogroup from it would be meaningless
How we worked this out
Paternal. Descent through the published haplogroup tree, entering a branch only where the marker defining it was genotyped and derived. 281 informative markers. Why it stops at K2b: the chip this file came from carries no marker that separates the branches below this point — it is not something missing from your file, and a test on a chip with denser Y coverage could go further. The full label K-M1221 changes as the tree is revised; the marker name does not.
An empty region on a map is one nobody has excavated, or excavated and not published.
Thousands of years back
The ancient peoples you come from
Scientists have read DNA from people who lived long ago. Here your DNA is fitted as a mix of those ancient groups.
The ancestry model we have does not fit this genome. That is a result rather than an error — it means these ancient sources cannot account for this person's ancestry, and any percentages we showed would be describing a model we have already rejected.
About fifty thousand years ago
Your Neanderthal and Denisovan DNA
When early humans left Africa they met Neanderthals and Denisovans and had children together. Almost everyone outside Africa carries a little of both.
You carry more than 60% of people worldwide.
205 callable sites is below the 500 this needs
How we worked this out
Archaic variants are counted at 12,106 positions in the SPrime call set (Browning et al. 2018, Cell 173(2), CC BY 4.0). This file carries a call at 928 of them (7.7%). A position where the call set and our reference disagree about the variant allele is dropped (0 here). Each count is divided by what its own set of positions can see, so it is a count of variants, and neither number is a share of your genome. A zero would not show that you carry no archaic ancestry.
The rank is among 2,504 reference genomes from around the world, counted on the same positions. Within Kinh, you carry more Neanderthal variants than 0% of 99.
172 variants trace to either source, counted once, of 617 positions read. That places you above 55% of the reference genomes.
Health & traits
Health & traits
A few variants, found or not found
We only say whether your file carries each variant, and what that usually means. There is no risk score here, and nothing on this page is medical advice.
Traits
A handful of well-studied variants, whether your file carries them, and what each one usually means.
We check 23 trait positions. Your chip covers 10 of them: 2 found, 8 read and not carried, 13 it cannot read.
Earwax typeABCC11 rs17822931Found, one copy
You carry one copy. With one copy, earwax is usually wet.
The clearest single-variant trait known in human genetics. One amino acid change disables the ABCC11 transporter, and two copies of the disabling allele give dry, flaky earwax and markedly reduced underarm odour; one or no copies give wet earwax.
Deterministic for earwax; only strongly associated for body odour, which also depends on skin bacteria, washing and clothing. The variant is readable on every consumer array we have checked.
Hair thickness and shovel-shaped incisors, EDAR V370AEDAR rs3827760Found, two copies
You carry this variant.
The derived allele is associated with thicker, straighter hair shafts, more eccrine sweat glands and shovel-shaped upper incisors. It is one of the strongest signals of recent selection in East Asian populations.
Effectively an East Asian variant: 94% in Han Chinese and 0-5% across the South Asian cohorts, reaching 5% only in Bengali. For most readers here the answer is 'you do not carry it', which is a fact about the allele's distribution and not about their hair.
8 read, variant not carried
HERC2 pigmentation variantHERC2 rs1667394Not found
You do not carry this variant.
One of the pigmentation positions in the HERC2/OCA2 region reported for hair and eye colour in a genome-wide study of Europeans. It sits near, and is inherited with, the better-known blue-eye variant.
Established in Icelandic and Dutch cohorts. Its effect in South Asians has not been measured, and pigmentation variants in this region behave differently outside Europe.
OCA2 R419Q, eye colourOCA2 rs1800407Not found
You do not carry this variant.
Associated with green and hazel eye colour, and reported to shift eye colour away from blue in people who otherwise carry the blue-eye haplotype at HERC2.
The eye-colour association comes from European cohorts. Eye colour in South Asians is far less studied and this variant's effect there has not been measured.
Alcohol flushALDH2 rs671Not found
You clear alcohol the usual way; no flush from this variant.
Reduces aldehyde dehydrogenase 2 activity; associated with facial flushing after alcohol. ClinVar records it under "drug response".
Carried as a deliberate negative, like the G6PD A- row. In gnomAD v4 this allele is at 0.24364 in East Asians and 0.00025 in South Asians — roughly one in four against roughly one in four thousand.
Lactase persistenceMCM6 rs4988235Not found
Milk may be harder on your stomach as an adult. This is the usual pattern across much of South and East Asia.
The variant upstream of LCT most strongly associated with continued lactase production into adulthood. ClinVar records it under "association".
Two separate limits, and they point in opposite directions. No lactase association in the GWAS Catalog has ever had a South Asian discovery cohort, so the association itself is untested here.
Freckling and hair colour, IRF4IRF4 rs12203592Not found
You do not carry this variant.
Associated with freckling, lighter hair colour and skin sensitivity to sun. One of the few pigmentation loci that acts on patterning rather than on overall tone.
Common in Europe at 9-18% and essentially absent in the South Asian, African and East Asian cohorts. It is on the panel because the report is global, not because it is expected to say much to a South Asian reader.
Skin pigmentation, SLC24A5 A111TSLC24A5 rs1426654Not found
You do not carry this variant.
The single largest-effect common variant on skin pigmentation known. The A allele reduces melanin production and accounts for a substantial share of the pigmentation difference between European and West African populations.
South Asia is where this variant is most informative and least deterministic. It segregates INSIDE the region rather than between it and anywhere else, so two people from the same state can carry different genotypes.
Skin pigmentation and freckling, TYR S192YTYR rs1042602Not found
You do not carry this variant.
A coding change in tyrosinase, the rate-limiting enzyme of melanin synthesis. The A allele is associated with lighter skin, freckling and sun sensitivity.
At 36% in Britain, 13% in Punjab and absent in the East Asian and African cohorts, this is largely a European-gradient variant. Pigmentation is polygenic; one tyrosinase variant shifts a tendency and does not set a tone.
Hair and eye colour, TYR upstreamTYR rs1393350Not found
You do not carry this variant.
A regulatory variant upstream of tyrosinase, associated with red and blond hair, blue eyes and freckling in the same genome-wide scans that found the coding change above.
27% in Britain and effectively absent everywhere else we hold. It is on the panel because the report is global, and it will tell most South Asian readers only that they do not carry it.
13 this chip cannot read
KITLG, hair colourKITLG rs642742Not read: not on this chip
A regulatory change near KITLG associated with lighter hair colour.
Your chip does not include this position. Not read is different from not found.
MC1R R151C, red hair and fair skinMC1R rs1805007Not read: not on this chip
One of the three MC1R variants most consistently reported with red hair and freckling.
Your chip does not include this position. Not read is different from not found.
MC1R R160W, red hair and fair skinMC1R rs1805008Not read: not on this chip
A second of the three MC1R variants reported with red hair and freckling, in the same 1995 study that named the first.
Your chip does not include this position. Not read is different from not found.
SLC45A2 L374F, skin and hair pigmentationSLC45A2 rs16891982Not read: ambiguous position
One of the largest single contributions to the difference in skin and hair pigmentation between European and non-European populations.
Established in European cohorts, where the lighter-pigmentation allele is common. It is rare in South Asians and its contribution to pigmentation there has not been measured.
Fast-twitch muscle (ACTN3)ACTN3 rs1815739Not read: not on this chip
Two copies of this variant produce no alpha-actinin-3 in fast-twitch muscle fibres.
Your chip does not include this position. Not read is different from not found.
Blue-eye variant (HERC2)HERC2 rs12913832Not read: not on this chip
The single variant explaining most of the blue-brown eye colour difference in European populations, by regulating OCA2 expression.
Your chip does not include this position. Not read is different from not found.
Alcohol metabolism, ADH1B His48ArgADH1B rs1229984Not read: not on this chip
The ADH1B*2 allele encodes an enzyme that converts alcohol to acetaldehyde far faster than the common form.
Your chip does not include this position. Not read is different from not found.
Caffeine metabolism, CYP1A2 -163C>ACYP1A2 rs762551Not read: not on this chip
CYP1A2 clears about 95% of ingested caffeine.
Your chip does not include this position. Not read is different from not found.
MTHFR C677TMTHFR rs1801133Not read: not on this chip
Reduces the activity of methylenetetrahydrofolate reductase.
Your chip does not include this position. Not read is different from not found.
Long-chain fatty acid conversion, FADS1FADS1 rs174537Not read: not on this chip
Associated with the efficiency of converting plant-derived short-chain omega-3 and omega-6 fatty acids into the long-chain forms the body uses.
Your chip does not include this position. Not read is different from not found.
Skin pigmentation, OCA2 H615ROCA2 rs1800414Not read: not on this chip
An East Asian-specific pigmentation variant.
Your chip does not include this position. Not read is different from not found.
Hair and eye colour, SLC24A4SLC24A4 rs12896399Not read: not on this chip
A potassium-dependent sodium-calcium exchanger locus associated with lighter hair and eye colour, and one of the markers in the published HIrisPlex eye and hair colour models.
Your chip does not include this position. Not read is different from not found.
Bitter taste (PTC)TAS2R38 rs713598Not read: ambiguous position
One of three coding variants in the TAS2R38 bitter receptor that together determine sensitivity to phenylthiocarbamide and related compounds.
One of three variants, and this file reads one of them. The phenotype depends on the combination, so a call here is partial by construction rather than by accident.
How we worked this out
Not found: we read the position and the variant is on neither copy. That is a measurement. Found, one copy or two copies: also a measurement. Not read: the chip does not type this position, or could not read it cleanly. Nobody looked, so it is not the same as not found. Could not be read: several variants share this position or the strand is unmeasured, so the row says what blocked it.
Most associations here were established in European or East Asian cohorts. Each row's tags say which population the association has been tested in. Catalogue v0.1.0. Nothing here is a prediction about you.
Health variants
Whether your file carries variants that studies have linked to health. Found or not found, with the studies behind each one.
We check 24 health positions. Your chip covers 6 of them: 2 found, 4 read and not carried, 18 it cannot read.
SLCO1B1 Statin transport1 of 1 chip-readable variants1 found
SLCO1B1 V174A, reduced hepatic transportSLCO1B1 rs4149056Found, one copy
You carry this variant.
SLCO1B1 is a liver transporter: it moves certain compounds out of the blood and into hepatocytes, where they act and are cleared. The V174A change reduces that transport, so what the transporter handles stays in circulation longer.
Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment on the basis of a web page is more dangerous than anything this row describes. The variant is LESS common in South Asians than in Europeans -- 4.9% against 15.9% -- so for most readers here this row will be uninformative. Effect evidence is European; no South Asian cohort is cited.
ABCG2 transporter function1 variant1 found
ABCG2 Q141K, transporter functionABCG2 rs2231142Found, one copy
You carry this variant.
The T allele reduces ABCG2 transporter function. It is associated with higher serum urate and with reduced response to allopurinol, and CPIC uses it in guidance on rosuvastatin dosing.
Urate and gout have large dietary and renal components that this variant does not read. The allopurinol association describes response at a given dose rather than whether the drug works.
G6PD 3 of 225 chip-readable variants; 248 more no chip can see1 read, not found
G6PD deficiency, A- variantG6PD rs1050828Not found
You do not carry this variant.
The commonest G6PD-deficiency variant in populations of African ancestry. ClinVar classifies it Pathogenic/Likely pathogenic.
Carried here as a deliberate negative. In gnomAD v4 this variant is at 0.00030 in South Asians and 0.12277 in Africans — a 400-fold difference — while the Mediterranean variant on the row above runs the other way, 0.019 against 0.0002. A G6PD panel built on the well-known African variant would miss almost every deficient South Asian and would look like a working panel while doing it. That is the same failure as reporting lactase persistence from one European SNP, in a gene where the consequence is a drug reaction.
G6PD deficiency, Mediterranean variantG6PD rs5030868Not read: not on this chip
This variant reduces glucose-6-phosphate dehydrogenase activity and is the commonest cause of G6PD deficiency reported in Indian populations. ClinVar classifies it Pathogenic/Likely pathogenic for G6PD deficiency.
G6PD deficiency is diagnosed by an enzyme activity test, not by a genotype. This tells you a variant is present. It is also not the only cause of G6PD deficiency, and the others are not read by this file. Measured as readable: the GSA probe here is [A/G] on the plus strand, which interrogates the correct alternate, and the site is present in every real 23andMe and AncestryDNA file tested.
G6PD deficiency, Orissa variantG6PD rs78478128Not read: not on this chip
A variant reducing glucose-6-phosphate dehydrogenase activity, classified Pathogenic/Likely pathogenic in ClinVar, and the most frequently reported cause of G6PD deficiency in Indian series.
The measurement worth quoting precisely, because it is easy to misread. Among 350 molecularly characterised G6PD-deficient individuals drawn from a screen of 20,896 people across India, this variant accounted for 56.5% of deleterious alleles and the Mediterranean variant for 23.6%. That is a share of the alleles found IN DEFICIENT PEOPLE, not a frequency in the population — overall deficiency prevalence in the same screen was 1.9%, ranging 0.8 to 6.3% by region. The two numbers answer different questions and only the second says how common deficiency is.
CYP2C19 Clopidogrel and other CYP2C19-activated drugs2 of 585 chip-readable variants1 read, not found
CYP2C19*2, the commonest no-function alleleCYP2C19 rs4244285Not read: not on this chip
A splice-site change that produces no working CYP2C19 enzyme from the affected copy. CYP2C19 converts several drugs into their active form, so carrying two copies means the enzyme activity is absent rather than reduced.
This is not medical advice, and not a reason to change any medicine. If a doctor has prescribed something, keep taking it and show them this page. Metaboliser status is assigned from a person's full CYP2C19 star-allele diplotype; this file reads two of those alleles, so a result here is partial. Genotype is also not a measurement of enzyme activity, which is what a clinical test would give you.
CYP2C19*3, a second no-function alleleCYP2C19 rs4986893Not found
You do not carry this variant.
A premature stop codon that truncates the CYP2C19 protein. Same consequence as *2 -- no working enzyme from that copy -- by a different mechanism.
Not medical advice. Much rarer than *2 in South Asians -- gnomAD puts it at 0.5% against 33% -- so for most readers the *2 row is the informative one. A full diplotype needs more alleles than this file reads.
CYP2C19*17, faster metabolismCYP2C19 rs12248560Not read: not on this chip
Increases CYP2C19 activity, the opposite direction to the *2 and *3 variants already in this report. The Clinical Pharmacogenetics Implementation Consortium publishes prescribing guidance based on the combination of these variants; this report shows what was found and does not adjust any dose.
The prescribing guidance behind this variant is built mainly on European and East Asian cohorts. Its frequency and effect in South Asians are less well measured.
HFE hereditary haemochromatosis2 variants1 read, not found
HFE C282Y, hereditary haemochromatosisHFE rs1800562Not found
You do not carry this variant.
The commonest cause of hereditary haemochromatosis, in which the body absorbs more iron than it needs. ClinVar classifies it Pathogenic.
Almost all HFE research is in European-ancestry cohorts, where this variant is commonest. It is rare in South Asians and its frequency there is not well measured, so an absence here says less than it would for a European genome.
HFE H63DHFE rs1799945Not read: not on this chip
A second HFE change described in the same 1996 paper as C282Y. It is common in many populations and, on its own, is not generally associated with iron overload; the combination most reported is one copy of this alongside one copy of C282Y.
ClinVar records conflicting classifications for this variant, and it is reported far more often than iron overload occurs, so it is shown as a finding rather than a conclusion. Its frequency in South Asians is not well measured.
CYP2C9 reduced-function allele1 variant1 read, not found
CYP2C9*3, reduced-function alleleCYP2C9 rs1057910Not found
You do not carry this variant.
CYP2C9*3 substantially reduces enzyme activity. CPIC guidelines use CYP2C9 genotype in dosing recommendations for warfarin and for phenytoin.
A genotype is not a dose. CPIC recommendations combine CYP2C9 with VKORC1 and with clinical factors, and warfarin is titrated on INR measurement whatever the genotype says. Carrying *3 is a reason to expect a lower dose requirement, not a reason to change one.
8 this chip cannot read
HBB Beta-thalassaemia and sickle cell6 of 223 chip-readable variants; 272 more no chip can see6 not read
Beta-thalassemia, Cap+1 A>CHBB rs34305195Not read: not on this chip
A variant in the HBB transcription start region, classified Pathogenic/Likely pathogenic in ClinVar, and reported in Indian series as a mild beta-thalassemia allele.
Reported in the literature as producing a milder phenotype than the splice-site and nonsense alleles, which is a statement about published series and not a prediction about any individual.
Beta-thalassemia, codon 15 G>AHBB rs33986703Not read: not on this chip
A nonsense variant in HBB, classified Pathogenic in ClinVar, and one of the beta-thalassemia alleles reported in Indian series.
Two limits. This site is absent from the base Illumina GSA manifest and present in all three real 23andMe v5 files, because 23andMe adds custom content on top of the GSA platform and does not publish its site list — so the probe alleles here are unknown and unknowable from any source we are willing to use. It is also a T/A pair, which is strand-ambiguous, and until now we discarded such sites rather than resolve them — our policy rather than a limit of the array, whose three designs state the strand unanimously here. The row declines rather than reporting an absence.
Beta-thalassemia, codon 30 G>CHBB rs33960103Not read: not on this chip
A variant at the codon 30 splice junction of HBB, classified Pathogenic in ClinVar and reported in Indian beta-thalassemia series.
Two independent limits. ClinVar records two further pathogenic alleles at this coordinate, so an array whose probe reads one of those has said nothing about ours, and the row never reports an absence here. It is also a C/G pair, which is strand-ambiguous; whether the probe designs at this position state a strand unanimously has not been measured, so the drop is attributed to our own policy rather than to the array until it is.
Beta-thalassemia, IVS1-1 G>THBB rs33971440Not read: not on this chip
A splice-donor variant in HBB, classified Pathogenic in ClinVar, and among the beta-thalassemia alleles most frequently reported in Indian series after IVS1-5.
Three separate pathogenic alleles sit at this coordinate, so an array whose probe reads a different one has said nothing about this variant. The row therefore never reports an absence here.
Beta-thalassemia, IVS1-5 G>CHBB rs33915217Not read: not on this chip
A splice-site variant in HBB, classified Pathogenic in ClinVar, and the most frequently reported beta-thalassemia allele in Indian series.
Whether this variant can be read from a consumer array is unknown, and that is the honest answer rather than a hedge. Three different pathogenic alleles share this rsID at chr11:5226925 — C>A, C>G and C>T are separate ClinVar records — and the Illumina GSA manifest carries four probe designs at the position, one of which does interrogate C>G. Which design a vendor actually shipped is not stated in the file, which names the marker plainly with no suffix. So the site is typed, and what was typed cannot be determined.
Sickle cell variantHBB rs334Not read: not on this chip
The HBB variant that produces haemoglobin S, classified Pathogenic in ClinVar for sickle cell disease.
Not typed by 23andMe v5 at all — absent from all three real files tested, while present in a 2023 AncestryDNA file. So whether this row can be answered depends on which vendor and which chip version produced the upload, and the answer has to be computed per file rather than stated per vendor. It is also a T/A pair, which is strand-ambiguous, and unlike the other rows here the three probe designs at this position disagree about which strand they read, so no declaration exists to resolve it against. That is a limit of the array rather than a policy of ours.
NUDT15 Thiopurine sensitivity1 of 3 chip-readable variants1 not read
NUDT15 R139C, reduced enzyme activityNUDT15 rs116855232Not read: not on this chip
The R139C change reduces NUDT15 enzyme activity, so an active metabolite the enzyme normally clears persists longer. CPIC assigns metaboliser status from NUDT15 and TPMT together.
Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment is more dangerous than any genotype on this page. Show this to the prescribing doctor and let them decide. This variant is roughly 24 times commoner in South Asians than in Europeans, which is why it is here and why guidance derived from European cohorts has historically under-served this population.
TPMT Thiopurine sensitivity1 of 9 chip-readable variants1 not read
TPMT*3C, thiopurine S-methyltransferase activityTPMT rs1142345Not read: not on this chip
TPMT is the second enzyme CPIC reads alongside NUDT15. The *3C allele reduces its activity.
Not medical advice. TPMT*3C is one of several TPMT alleles and this file reads one of them, so a normal result here does not establish normal TPMT activity. Enzyme activity is measurable directly and that test, not this one, is what a clinic would use. No South Asian cohort study of this allele's effect is cited here -- the frequency data below is South Asian, the effect evidence is not.
VKORC1 Warfarin sensitivity1 of 9 chip-readable variants1 not read
VKORC1 -1639, reduced enzyme expressionVKORC1 rs9923231Not read: not on this chip
This promoter variant lowers how much VKORC1 enzyme the liver produces. VKORC1 is the target of one class of anticoagulant, so the amount present matters to anyone taking one.
Not medical advice. Anticoagulant dosing is managed by blood tests that measure the actual effect in the actual person, which is far more informative than any genotype. Never change a dose on the basis of this page. Effect evidence is largely European and East Asian; no South Asian cohort is cited here.
ASPA Canavan disease carrier1 variant1 not read
ASPA E285A, Canavan disease carrierASPA rs28940279Not read: not on this chip
The commonest ASPA variant in Canavan disease, a recessive condition affecting the white matter of the brain. ClinVar classifies it Pathogenic/Likely pathogenic.
This variant's frequency has not been measured in South Asian populations. The Canavan literature is largely Ashkenazi Jewish, where it is commonest; what carrier frequency looks like in South Asia is not something this report can tell you.
ACKR1 1 variant1 not read
Duffy-null blood groupACKR1 rs2814778Not read: not on this chip
The C allele abolishes Duffy antigen expression on red cells. It confers resistance to Plasmodium vivax malaria and is associated with a lower baseline neutrophil count that is benign -- 'benign ethnic neutropenia'.
The clinically important part is the neutrophil count, because a normal result for a Duffy-null person can be read as abnormal against a reference range built on people who are not. This is close to fixed in West African populations, absent in the South Asian and East Asian cohorts, and is on the panel because the report is global.
IFNL4 hepatitis C treatment response1 variant1 not read
IFNL4 (IL28B), hepatitis C treatment responseIFNL4 rs12979860Not read: not on this chip
Genotype at this locus predicts response to interferon-based hepatitis C therapy and spontaneous clearance of the virus. The C allele is the favourable one; the T allele is associated with poorer response.
This is a variant whose clinical relevance has largely PASSED. Interferon-based regimens have been superseded by direct-acting antivirals, which cure across genotypes, and it is reported here as a well-established association rather than as a treatment decision anybody should now make.
CYP3A5 tacrolimus metabolism1 variant1 not read
CYP3A5*3, tacrolimus metabolismCYP3A5 rs776746Not read: not on this chip
The commonest reason CYP3A5 produces no working enzyme. It is the main genetic factor in tacrolimus dosing after transplantation.
Frequencies for this variant differ sharply between populations, and the South Asian estimate is less well measured than the European and African ones.
How we worked this out
Rows report variant presence against published work. The catalogue (v0.1.0) is a small selection of the variants known in each gene: ClinVar records hundreds more, and roughly half of those are deletions, duplications or repeat expansions that no genotyping chip can detect at all. Metaboliser status for CYP2C19 or TPMT needs a full star-allele diplotype and an enzyme activity test, which this file cannot give. G6PD deficiency is diagnosed by an enzyme test, not a genotype.
Whether a position can be read depends on the vendor and chip version that produced the upload, and it is measured for each file. Frequencies shown are from gnomAD v4 and the 1000 Genomes Project. The evidence for an effect often comes from European studies, and each row's tags say where it has been tested.
Technical
Technical
For the curious
The detail behind the numbers: the reference groups, your file, and the same file through other tools.
Reference groups closest to you
How close your DNA sits to each group in our reference set.
Closest areas
The areas whose reference samples your DNA is nearest to. No order, and no ranking.
7 further groups, none of them close
How we worked this out
Each file and every individual reference genome is projected into the first ten principal components of our panel. The chart is Mahalanobis distance in that space, so it scales each group by its own spread. The green band is where a typical member sits, the amber band is close but outside it, and a marker farther right is less like that group. Hover or tap a row for its distance. The closest areas come from the nearest three individual references, coarsened and left unranked.
On the same 289 people and 54 groups, this distance placed the correct group first 60.2% of the time against 35.3% for G25, and in the top three 74.0% against 63.0%. The G25 coordinates in this report are an export for other tools.
Present-day map
Where your file falls among living reference groups
Each dot is one person in our reference panel, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.
2,000 reference individuals, the same cloud a report draws, on the panel fitted today. The first two components carry 13.9% and 5.4% of total variance. Every point is one published reference individual and names itself on hover. Clusters are labelled with the region each cohort is pooled into and coloured by the continental component that region sits in. 504 individuals from cohorts pooled into no region are not drawn.
The picture is turned a quarter clockwise and mirrored, which is why it resembles a map: Europe upper left, East Asia upper right, Africa along the bottom, South Asia between them. That is a rotation and a reflection of the plane, so no point has moved relative to any other and every distance is what it was. It is a reading aid and nothing more. The axes have no meaning of their own: they are directions of maximum variance, their signs are arbitrary, and a region's position depends on which other regions are in the panel. The first component runs vertically here and the second horizontally.
Every region carries a numbered disc at its centre and an outline around its members, and the key below repeats the number — so a region too crowded to spell out on the plot can still be found on it. The outline is the convex hull of that region's own individuals with the furthest 8% trimmed off, which keeps one stray from dragging a boundary across the chart; it encloses people, not territory. Spelled out only in the key at this size: 4 Northwest European, 14 Tai and Kadai, 13 Punjab and Kashmir, 16 Sierra Leone and the Upper Guinea coast, 8 Japanese, 7 Italian.
African
- 2West African forest207
- 5West African savanna113
- 10Eastern and Southern African99
- 16Sierra Leone and the Upper Guinea coast85
Bengal and the central belt
- 15Bengal and the central belt86
East Asian
- 1Eastern Chinese208
- 8Japanese104
- 11Mainland Southeast Asia99
- 14Tai and Kadai93
European
- 4Northwest European190
- 6Iberian107
- 7Italian107
- 12Finland and Karelia99
Indo-Gangetic Plain
- 9The Ganges Plain and Gujarat103
- 13Punjab and Kashmir96
Southern Peninsula
- 3Deccan and the Tamil Plains204
4,046 of 12,770 panel markers placed you on the map.
The groups nearest your position, closest first:
- Kinh
- Bengali
- Dai
This is a position on a plot. Sitting near a group means your genome resembles that group on these axes. It does not say who your ancestors were.
How we worked this out
The axes come from the reference panel alone. Your file is projected onto them using the markers it shares with the panel, so it moves no point on the map.
Nearest is measured over more axes than the two drawn, to the middle of each group, so a group that looks close here can rank lower in the list.
Ancient source map
Where your file falls among ancient genomes
Each dot is one ancient person, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.
558 excavated individuals with at least 100,000 1240K SNPs, on the first two principal components (5.1% and 1.7% of variance). 41 individuals sit far from their group's centroid and are left out of every outline and centre; they are not drawn here, and the methods page shows them. Axes have no meaning of their own, and a group's position depends on which other groups are in the map. 4 sources pooled from several regions are not drawn: a period is not a people.
Named only in the key at this size: 9 Japan, c. 4,500 BP.
Before 10,000 BP
- 1Russia, c. 32,500 BP12
10,000 to 7,000 BP
- 2Iran, c. 10,000 BP26 +2
- 3Turkey, c. 8,500 BP85 +9
- 4Russia, c. 8,500 BP48 +3
- 5Cameroon, c. 8,000 BP8
- 6Sweden, c. 7,500 BP17 +2
7,000 to 5,000 BP
- 7Steppe and Siberia, c. 5,500 BP15 +1
5,000 to 3,500 BP
- 8China, c. 4,500 BP113 +3
- 9Japan, c. 4,500 BP14 +1
After 3,500 BP
- 10South Africa, c. 2,500 BP7
- 11China, c. 2,000 BP136 +14
- 12Taiwan, c. 1,500 BP21 +4
- 13China, c. 1,500 BP24
- 14Papua New Guinea, c. 500 BP32 +2
Counts are individuals kept; +n is individuals dropped. Thin: fewer than 5 usable individuals. An outline is drawn around each excavation group of six or more people, not around a source.
40,476 of 138,630 sites in your file are on the map.
The sources nearest your position, closest first:
- China, c. 1,500 BP
- Papua New Guinea, c. 500 BP
- Japan, c. 4,500 BP
This is a position on a plot. Sitting near a source means your genome resembles that group on these axes. It does not say who your ancestors were.
How we worked this out
The axes come from the ancient individuals alone. Your file is projected onto them using the sites it shares with them, so it moves no point on the map.
Nearest is measured over the first four axes, to the middle of each source. The plot shows two of them, so a source that looks close here can rank lower in the list.
On eleven test genomes the nearest source fell in the same region as the qpAdm fit. Inside a region, the order of the nearest sources does not track the qpAdm proportions.
The same ancestor, twice
Stretches inherited twice
Places where both copies of a chromosome came from the same ancestor. A long stretch points to a recent shared ancestor.
Oldest first. Bar heights compare your own bands with each other.
Where they sit
Each bar is one chromosome, drawn to length. A mark is a stretch where both of your copies match, coloured by how far back the shared ancestor sits.
How we worked this out
Runs of homozygosity need no reference panel. The question is whether your two copies match, so this runs on 540,179 of your own markers, more than the 12,770 that reach the ancestry estimate. It is the one chapter in this report that our reference data does not limit. In total the stretches cover 76.4 cM of 3,545 cM.
Two different detectors were run over your file and they agree on 96% of the length. They fail in different ways, so where they agree the result is not an artefact of either one.
Lengths are in centimorgans, read off a recombination map, not in megabases. Near the middle of a chromosome a megabase can be a fifth of a centimorgan and near the end it can be two, so a generation count computed from physical distance is wrong by a factor that changes along every chromosome. Generations follow from genetic length as roughly 50 divided by the length in centimorgans, and a year figure uses 28 years to a generation.
Each stretch is shown as a range of generations. A 20 cM stretch averages about two and a half generations back, but over a very wide spread, and one stretch is a single draw from it. A range is what the evidence supports.
What this does not say. It says nothing about whether your parents are related. That question needs a comparison against the population you descend from, and we do not publish it. In a community that has married within itself for centuries the background is high with no close relatedness at all, so a fixed threshold would tell many people something untrue about their own family.
Your file
How complete your raw DNA file was. This sets how precise every other chapter can be.
A: complete enough for every chapter.
No problems found in the checks we run. Nothing was sequenced by us. We read the export you already had.
How we worked this out
Heterozygosity is 24.8%: the share of read positions where your two copies differ. A consumer chip usually reads between 24% and 31%; well outside that range points to a file problem rather than to anything about you.
Your file is kept and used to build future reference panels, under a Creative Commons licence. We never sell or release an individual genome, and no genetic data is ever written to a log.
Who you were compared against
524 groups, 10,922 people. The reference behind every number in this report.
Every study behind the panel, and its regions
People alive today: 6,745 people, 18 studies. The reference this report measures your ancestry against.
Excavated individuals: 13,454 people, 309 studies. Used by the ancient peoples chapters. The region percentages do not use them.
the other 297 studies
Behind the South Asian regions: 2,148 people across 11 studies, counted by the study that collected them.
The full list, the sampling map and the panel drawn by genetic similarity: How it works.
How we worked this out
Accuracy is the share of held-out people whose largest region came out right: 120 people from 24 groups at the full panel of 12,770 markers, and again at 10,317 markers, the density a real 23andMe export reaches. The people tested are part of the reference set they are scored against, so their own data helped define the groups. Someone uploading a file gets no such help, so the real-world rate is a little lower. Five individuals per population. A perfect score on 120 people is consistent with a true rate above about 97%. The Indigenous American component is built from four samples, three of them majority European, so a correct call there is an easier test than it sounds. It measures which region comes out largest. Whether the printed range covers the true value is a separate measurement we have not run. 10 of the held-out set are not scored (African Caribbean and African American samples, which define no component). Measured 2026-09-09 against the 12,770-marker panel this report used. A rebuilt panel re-runs it.
A few dozen people per group is what makes our ranges as wide as they are. A narrower range would need more people. West Asia is fitted from Bedouin, Palestinian and Druze samples, and Oceania from pooled Papuan and Bougainville individuals. Central Asia is not yet a region of its own.
Coordinates for other tools
Twenty-five numbers you can paste into community ancestry tools.
this map was fitted and checked on South Asian genomes, and this file is 8% South Asian. Placing it would mean extrapolating a map nobody has verified there
Other calculators
The same file run through popular community calculators. Where they agree with our result, the answer does not depend on the model.
Community calculators need more of their own markers than this file carries, so none were run.