Overview
Your ancestry report
Your ancestry is almost entirely East Asian (86%), with a little from elsewhere.
Analysed 31 August 2026
Tap any bar or region to see how sure we are.
Health & traits
Your chip covers 0 of the 24 health variants we check. We found none. Of 23 trait variants, your chip covers 0; none found.
Go to Health & traitsTechnical
Your file was graded A, with 55,746 markers read. Closest reference groups, your file and other calculators are here.
Go to TechnicalAncestry
Ancestry
Where you come from
Recent regions first, then your chromosomes, your mother's and father's lines, and thousands of years back.
Where your ancestors lived
Each bar is the share of your DNA that matches people living in that region today.
How we worked this out
Each percentage is a share of the 3,043 markers we could read in your file (24% of 12,770), compared with people sampled today by the 1000 Genomes Project. The solid part of a bar is the share we are sure of. The faint part is how much higher it might be. The line is our best estimate. The range is a 95% interval from resampling your markers. If a population is missing from our reference set, your ancestry from it is counted under the nearest region we do have.
Each region and its range. East Asian 85.9% (likely 79–96%).
East Asian is one region. How much of your ancestry is from this region is measured from the whole panel at once. How it divides between the areas inside it is a second, harder question, fitted separately — so the total is the firmer number of the two. Its total (85.9%) is firmer than the split inside it. The parts drawn inside it add to 85.1%, because some are too small to tell from zero and are not drawn.
What “could not be placed” means. A region whose range reaches zero is not drawn on its own. The ancestry is still yours; this file does not have enough markers to say which region it belongs to. Here the regions that could hold it are South Asian (up to 7%), Siberian (up to 19%), Indigenous American (up to 6%), Oceanian (up to 3%) and North African (up to 2%).
The map. Shading marks where each ancestry lives, at your own share of it. The outlines are geographic regions and the numbers are genetic, so an edge is approximate: ancestry shades into its neighbours. A faint region is one this file cannot tell from zero.
| Region | Fitted from | Estimate | Range |
|---|---|---|---|
| Eastern Chinese | Eastern and central China, fitted from Han | 51.2% | 57% |
| Southwest China | individuals recruited in Yunnan and Sichuan | 8.9% | 9% |
| Amur and Sakhalin | The lower Amur and Sakhalin, fitted from Ulchi, Nanai, Nivh, Even, Evenk_Transbaikal | 8.7% | 11% |
| Japanese | The Japanese archipelago, fitted from JPT, Japanese | 7.5% | 2% |
| Tibetan Plateau | The Tibetan plateau, fitted from Tibetan | 7.2% | 9% |
| Taimyr | Nganasan individuals recruited on the Taimyr peninsula | 6.3% | 7% |
| Taiwan | individuals recruited among the indigenous peoples of Taiwan | 3.2% | 0% |
| Chukotka and Beringia | Chukotka, Kamchatka and the Aleutian chain, fitted from Chukchi, Eskimo_Naukan, Eskimo_ChaplinSireniki, Koryak, Aleut, Yukagir_Tundra | below resolution | under 0.01% |
| Indigenous American | Peruvian, Mexican, Colombian and Puerto Rican reference samples | below resolution | 0–6% |
| Mongolia and the Buryat steppe | individuals recruited in Mongolia, Buryatia and Kalmykia | below resolution | 0% |
| Mlabri | Mlabri individuals recruited in Nan province, Thailand | 0.9% | 2% |
| Hmong and Mien | individuals recruited in the uplands of southern China | below resolution | under 0.01% |
| Northern East Asian | Northern China, Mongolia and the Amur basin, fitted from CHB, Mongola, Buryat, Tu | below resolution | under 0.01% |
| Northwest China | individuals recruited in Gansu, Qinghai and the Hexi corridor | below resolution | under 0.01% |
| Lower Amur | Hezhen and Oroqen individuals recruited on the lower Amur | below resolution | 0% |
| African | Yoruba, Luhya, Gambian, Mende and Esan reference samples | below resolution | 0.0–0.1% |
| European | Utah, Tuscan, Finnish, British and Iberian reference samples | below resolution | under 0.01% |
| Oceanian | 20 pooled Papuan and Bougainville individuals — the whole public supply | below resolution | 0–3% |
| West Asian | Bedouin, Palestinian and Druze reference samples | below resolution | under 0.01% |
| Indo-Gangetic Plain | individuals recruited in Punjab, Sindh, Kashmir, Rajasthan, Gujarat and along the Ganges | below resolution | 0–1% |
| Southern Peninsula | individuals recruited in Tamil Nadu, Andhra Pradesh, Telangana, Karnataka and Kerala | below resolution | under 0.01% |
| Makran and the western ranges | individuals recruited across Balochistan and Makran | below resolution | 0–2% |
| Bengal and the central belt | individuals recruited in Bengal, Madhya Pradesh, Chhattisgarh and Maharashtra | below resolution | under 0.01% |
| Himalayas and Northeast India | individuals recruited in Nepal, the terai, and the northeastern hills | below resolution | 0–7% |
| Chota Nagpur Plateau | individuals recruited in Jharkhand, Odisha and the eastern Ghats | below resolution | 0.0–0.5% |
| Andaman and Nicobar Islands | Onge, Jarawa and Great Andamanese individuals | below resolution | 0.00–0.02% |
| Central Asian | 155 individuals from six Central Asian populations — the oases and the Kazakh steppe | below resolution | under 0.01% |
| North African | The Maghreb — Morocco, Algeria, Tunisia and Libya, fitted from Mozabite | below resolution | 0–2% |
| Mainland Southeast Asia | individuals recruited in Cambodia and Vietnam | below resolution | under 0.01% |
| Orang Asli | Semang and Senoi individuals recruited in peninsular Malaysia | below resolution | under 0.01% |
| Semang | Bateq, Jehai, Kintaq, Mendriq and Maniq individuals recruited in peninsular Malaysia and southern Thailand | below resolution | under 0.01% |
| Southeast Asian | Mainland Southeast Asia, fitted from Cambodian, Thai, Kinh_Vietnamese | below resolution | under 0.01% |
| Tai and Kadai | individuals recruited in Yunnan, Thailand, Laos and southern China | below resolution | under 0.01% |
| Burmese Uplands | Akha, Burmese, Karen and Tai Lue individuals recruited in Myanmar and northern Thailand | below resolution | under 0.01% |
| Island Southeast Asia | individuals recruited in Sulawesi, the Philippines and northern Borneo | below resolution | under 0.01% |
| Inuit and Aleut | individuals recruited in Chukotka, St Lawrence Island and the Aleutians | below resolution | under 0.01% |
| South Siberian and Mongolian | The Altai, the Sayan and the Mongolian steppe, fitted from Tuvinian, Tubalar, Tofalar, Altaian, Khakass, Khakass_Kachin and others | below resolution | under 0.01% |
| West Siberian | The Ob and Yenisei basins, fitted from Ket, Selkup, Nganasan, Mansi, Khanty, Tatar_Siberian | below resolution | under 0.01% |
| Ob Ugra | Khanty and Mansi individuals recruited on the lower Ob | below resolution | under 0.01% |
| Sakha | Sakha individuals recruited in the Lena basin | below resolution | under 0.01% |
Your chromosomes, painted
Each chromosome coloured by which region its stretches match. Long stretches point to recent ancestors; this sees back about nine generations.
Two bars per chromosome, one for each copy. Grey is a stretch that matched no region for long enough to name.
How we worked this out
Painted in 4 cM windows against African, European, East Asian, South Asian, Indigenous American and Oceanian reference haplotypes. A label has to hold for at least two consecutive windows to count; single-window matches are what deep shared ancestry looks like, while a real ancestor leaves blocks of 13 to 74 centimorgans. Roughly a sixth of the genome (22.9%) matched no group for long enough to name and is left grey rather than guessed.
Against held-out people the painter had never seen, a named stretch is right: African 99.6%, Indigenous American 99.9%, East Asian 98.4%, European 90.2%, Oceanian 99.9%, South Asian 83.2%.
This is a recent-ancestry picture. Older mixtures are still in you, but recombination has cut them into stretches too short to name, so they show here as the ancestry around them. Shares are of what could be painted. The ancient peoples chapter measures something else, in another way, so its percentages do not line up with these.
Your mother's line and your father's line
Two threads pass down almost unchanged: one from mother to child, one from father to son. Each follows a single ancestor across thousands of years.
This file has no mitochondrial positions, so we cannot read a mother's line from it.
the chip this file came from carries no Y markers at all. A chip with Y coverage would give a result
How we worked this out
Thousands of years back
The ancient peoples you come from
Scientists have read DNA from people who lived long ago. Here your DNA is fitted as a mix of those ancient groups.
We can say which ancient ancestries are present, but not put numbers on them. The threshold that stops us is ours: below it we cannot test the mixture model strictly enough to trust a proportion, and a proportion from a model we could not properly test is not a weaker answer, it is an untested one. There is nothing to fix at your end.
About fifty thousand years ago
Your Neanderthal and Denisovan DNA
When early humans left Africa they met Neanderthals and Denisovans and had children together. Almost everyone outside Africa carries a little of both.
122 callable sites is below the 500 this needs
How we worked this out
This depends on how many of the SPrime archaic positions your file happens to carry. It says nothing about the quality of the rest of the file. Below a minimum count the rank among other people moves too much to print.
Health & traits
Health & traits
A few variants, found or not found
We only say whether your file carries each variant, and what that usually means. There is no risk score here, and nothing on this page is medical advice.
Traits
A handful of well-studied variants, whether your file carries them, and what each one usually means.
All 23 positions in the trait catalogue are missing from this chip. A different vendor's file usually answers some of them.
23 this chip cannot read
HERC2 pigmentation variantHERC2 rs1667394Not read: not on this chip
One of the pigmentation positions in the HERC2/OCA2 region reported for hair and eye colour in a genome-wide study of Europeans.
Your chip does not include this position. Not read is different from not found.
KITLG, hair colourKITLG rs642742Not read: not on this chip
A regulatory change near KITLG associated with lighter hair colour.
Your chip does not include this position. Not read is different from not found.
MC1R R151C, red hair and fair skinMC1R rs1805007Not read: not on this chip
One of the three MC1R variants most consistently reported with red hair and freckling.
Your chip does not include this position. Not read is different from not found.
MC1R R160W, red hair and fair skinMC1R rs1805008Not read: not on this chip
A second of the three MC1R variants reported with red hair and freckling, in the same 1995 study that named the first.
Your chip does not include this position. Not read is different from not found.
OCA2 R419Q, eye colourOCA2 rs1800407Not read: not on this chip
Associated with green and hazel eye colour, and reported to shift eye colour away from blue in people who otherwise carry the blue-eye haplotype at HERC2.
Your chip does not include this position. Not read is different from not found.
SLC45A2 L374F, skin and hair pigmentationSLC45A2 rs16891982Not read: not on this chip
One of the largest single contributions to the difference in skin and hair pigmentation between European and non-European populations.
Your chip does not include this position. Not read is different from not found.
Earwax typeABCC11 rs17822931Not read: not on this chip
The clearest single-variant trait known in human genetics.
Your chip does not include this position. Not read is different from not found.
Fast-twitch muscle (ACTN3)ACTN3 rs1815739Not read: not on this chip
Two copies of this variant produce no alpha-actinin-3 in fast-twitch muscle fibres.
Your chip does not include this position. Not read is different from not found.
Alcohol flushALDH2 rs671Not read: not on this chip
Reduces aldehyde dehydrogenase 2 activity; associated with facial flushing after alcohol.
Your chip does not include this position. Not read is different from not found.
Hair thickness and shovel-shaped incisors, EDAR V370AEDAR rs3827760Not read: not on this chip
The derived allele is associated with thicker, straighter hair shafts, more eccrine sweat glands and shovel-shaped upper incisors.
Your chip does not include this position. Not read is different from not found.
Blue-eye variant (HERC2)HERC2 rs12913832Not read: not on this chip
The single variant explaining most of the blue-brown eye colour difference in European populations, by regulating OCA2 expression.
Your chip does not include this position. Not read is different from not found.
Lactase persistenceMCM6 rs4988235Not read: not on this chip
The variant upstream of LCT most strongly associated with continued lactase production into adulthood.
Your chip does not include this position. Not read is different from not found.
Alcohol metabolism, ADH1B His48ArgADH1B rs1229984Not read: not on this chip
The ADH1B*2 allele encodes an enzyme that converts alcohol to acetaldehyde far faster than the common form.
Your chip does not include this position. Not read is different from not found.
Caffeine metabolism, CYP1A2 -163C>ACYP1A2 rs762551Not read: not on this chip
CYP1A2 clears about 95% of ingested caffeine.
Your chip does not include this position. Not read is different from not found.
MTHFR C677TMTHFR rs1801133Not read: not on this chip
Reduces the activity of methylenetetrahydrofolate reductase.
Your chip does not include this position. Not read is different from not found.
Long-chain fatty acid conversion, FADS1FADS1 rs174537Not read: not on this chip
Associated with the efficiency of converting plant-derived short-chain omega-3 and omega-6 fatty acids into the long-chain forms the body uses.
Your chip does not include this position. Not read is different from not found.
Freckling and hair colour, IRF4IRF4 rs12203592Not read: not on this chip
Associated with freckling, lighter hair colour and skin sensitivity to sun.
Your chip does not include this position. Not read is different from not found.
Skin pigmentation, OCA2 H615ROCA2 rs1800414Not read: not on this chip
An East Asian-specific pigmentation variant.
Your chip does not include this position. Not read is different from not found.
Hair and eye colour, SLC24A4SLC24A4 rs12896399Not read: not on this chip
A potassium-dependent sodium-calcium exchanger locus associated with lighter hair and eye colour, and one of the markers in the published HIrisPlex eye and hair colour models.
Your chip does not include this position. Not read is different from not found.
Skin pigmentation, SLC24A5 A111TSLC24A5 rs1426654Not read: not on this chip
The single largest-effect common variant on skin pigmentation known.
Your chip does not include this position. Not read is different from not found.
Skin pigmentation and freckling, TYR S192YTYR rs1042602Not read: not on this chip
A coding change in tyrosinase, the rate-limiting enzyme of melanin synthesis.
Your chip does not include this position. Not read is different from not found.
Hair and eye colour, TYR upstreamTYR rs1393350Not read: not on this chip
A regulatory variant upstream of tyrosinase, associated with red and blond hair, blue eyes and freckling in the same genome-wide scans that found the coding change above.
Your chip does not include this position. Not read is different from not found.
Bitter taste (PTC)TAS2R38 rs713598Not read: not on this chip
One of three coding variants in the TAS2R38 bitter receptor that together determine sensitivity to phenylthiocarbamide and related compounds.
Your chip does not include this position. Not read is different from not found.
How we worked this out
Not found: we read the position and the variant is on neither copy. That is a measurement. Found, one copy or two copies: also a measurement. Not read: the chip does not type this position, or could not read it cleanly. Nobody looked, so it is not the same as not found. Could not be read: several variants share this position or the strand is unmeasured, so the row says what blocked it.
Most associations here were established in European or East Asian cohorts. Each row's tags say which population the association has been tested in. Catalogue v0.1.0. Nothing here is a prediction about you.
Health variants
Whether your file carries variants that studies have linked to health. Found or not found, with the studies behind each one.
We check 24 health positions. Your chip covers 0 of them: 0 found, 0 read and not carried, 24 it cannot read.
14 this chip cannot read
G6PD 3 of 225 chip-readable variants; 248 more no chip can see3 not read
G6PD deficiency, A- variantG6PD rs1050828Not read: not on this chip
The commonest G6PD-deficiency variant in populations of African ancestry. ClinVar classifies it Pathogenic/Likely pathogenic.
Carried here as a deliberate negative. In gnomAD v4 this variant is at 0.00030 in South Asians and 0.12277 in Africans — a 400-fold difference — while the Mediterranean variant on the row above runs the other way, 0.019 against 0.0002. A G6PD panel built on the well-known African variant would miss almost every deficient South Asian and would look like a working panel while doing it. That is the same failure as reporting lactase persistence from one European SNP, in a gene where the consequence is a drug reaction.
G6PD deficiency, Mediterranean variantG6PD rs5030868Not read: not on this chip
This variant reduces glucose-6-phosphate dehydrogenase activity and is the commonest cause of G6PD deficiency reported in Indian populations. ClinVar classifies it Pathogenic/Likely pathogenic for G6PD deficiency.
G6PD deficiency is diagnosed by an enzyme activity test, not by a genotype. This tells you a variant is present. It is also not the only cause of G6PD deficiency, and the others are not read by this file. Measured as readable: the GSA probe here is [A/G] on the plus strand, which interrogates the correct alternate, and the site is present in every real 23andMe and AncestryDNA file tested.
G6PD deficiency, Orissa variantG6PD rs78478128Not read: not on this chip
A variant reducing glucose-6-phosphate dehydrogenase activity, classified Pathogenic/Likely pathogenic in ClinVar, and the most frequently reported cause of G6PD deficiency in Indian series.
The measurement worth quoting precisely, because it is easy to misread. Among 350 molecularly characterised G6PD-deficient individuals drawn from a screen of 20,896 people across India, this variant accounted for 56.5% of deleterious alleles and the Mediterranean variant for 23.6%. That is a share of the alleles found IN DEFICIENT PEOPLE, not a frequency in the population — overall deficiency prevalence in the same screen was 1.9%, ranging 0.8 to 6.3% by region. The two numbers answer different questions and only the second says how common deficiency is.
HBB Beta-thalassaemia and sickle cell6 of 223 chip-readable variants; 272 more no chip can see6 not read
Beta-thalassemia, Cap+1 A>CHBB rs34305195Not read: not on this chip
A variant in the HBB transcription start region, classified Pathogenic/Likely pathogenic in ClinVar, and reported in Indian series as a mild beta-thalassemia allele.
Reported in the literature as producing a milder phenotype than the splice-site and nonsense alleles, which is a statement about published series and not a prediction about any individual.
Beta-thalassemia, codon 15 G>AHBB rs33986703Not read: not on this chip
A nonsense variant in HBB, classified Pathogenic in ClinVar, and one of the beta-thalassemia alleles reported in Indian series.
Two limits. This site is absent from the base Illumina GSA manifest and present in all three real 23andMe v5 files, because 23andMe adds custom content on top of the GSA platform and does not publish its site list — so the probe alleles here are unknown and unknowable from any source we are willing to use. It is also a T/A pair, which is strand-ambiguous, and until now we discarded such sites rather than resolve them — our policy rather than a limit of the array, whose three designs state the strand unanimously here. The row declines rather than reporting an absence.
Beta-thalassemia, codon 30 G>CHBB rs33960103Not read: not on this chip
A variant at the codon 30 splice junction of HBB, classified Pathogenic in ClinVar and reported in Indian beta-thalassemia series.
Two independent limits. ClinVar records two further pathogenic alleles at this coordinate, so an array whose probe reads one of those has said nothing about ours, and the row never reports an absence here. It is also a C/G pair, which is strand-ambiguous; whether the probe designs at this position state a strand unanimously has not been measured, so the drop is attributed to our own policy rather than to the array until it is.
Beta-thalassemia, IVS1-1 G>THBB rs33971440Not read: not on this chip
A splice-donor variant in HBB, classified Pathogenic in ClinVar, and among the beta-thalassemia alleles most frequently reported in Indian series after IVS1-5.
Three separate pathogenic alleles sit at this coordinate, so an array whose probe reads a different one has said nothing about this variant. The row therefore never reports an absence here.
Beta-thalassemia, IVS1-5 G>CHBB rs33915217Not read: not on this chip
A splice-site variant in HBB, classified Pathogenic in ClinVar, and the most frequently reported beta-thalassemia allele in Indian series.
Whether this variant can be read from a consumer array is unknown, and that is the honest answer rather than a hedge. Three different pathogenic alleles share this rsID at chr11:5226925 — C>A, C>G and C>T are separate ClinVar records — and the Illumina GSA manifest carries four probe designs at the position, one of which does interrogate C>G. Which design a vendor actually shipped is not stated in the file, which names the marker plainly with no suffix. So the site is typed, and what was typed cannot be determined.
Sickle cell variantHBB rs334Not read: not on this chip
The HBB variant that produces haemoglobin S, classified Pathogenic in ClinVar for sickle cell disease.
Not typed by 23andMe v5 at all — absent from all three real files tested, while present in a 2023 AncestryDNA file. So whether this row can be answered depends on which vendor and which chip version produced the upload, and the answer has to be computed per file rather than stated per vendor. It is also a T/A pair, which is strand-ambiguous, and unlike the other rows here the three probe designs at this position disagree about which strand they read, so no declaration exists to resolve it against. That is a limit of the array rather than a policy of ours.
CYP2C19 Clopidogrel and other CYP2C19-activated drugs2 of 585 chip-readable variants3 not read
CYP2C19*2, the commonest no-function alleleCYP2C19 rs4244285Not read: not on this chip
A splice-site change that produces no working CYP2C19 enzyme from the affected copy. CYP2C19 converts several drugs into their active form, so carrying two copies means the enzyme activity is absent rather than reduced.
This is not medical advice, and not a reason to change any medicine. If a doctor has prescribed something, keep taking it and show them this page. Metaboliser status is assigned from a person's full CYP2C19 star-allele diplotype; this file reads two of those alleles, so a result here is partial. Genotype is also not a measurement of enzyme activity, which is what a clinical test would give you.
CYP2C19*3, a second no-function alleleCYP2C19 rs4986893Not read: not on this chip
A premature stop codon that truncates the CYP2C19 protein. Same consequence as *2 -- no working enzyme from that copy -- by a different mechanism.
Not medical advice. Much rarer than *2 in South Asians -- gnomAD puts it at 0.5% against 33% -- so for most readers the *2 row is the informative one. A full diplotype needs more alleles than this file reads.
CYP2C19*17, faster metabolismCYP2C19 rs12248560Not read: not on this chip
Increases CYP2C19 activity, the opposite direction to the *2 and *3 variants already in this report. The Clinical Pharmacogenetics Implementation Consortium publishes prescribing guidance based on the combination of these variants; this report shows what was found and does not adjust any dose.
The prescribing guidance behind this variant is built mainly on European and East Asian cohorts. Its frequency and effect in South Asians are less well measured.
NUDT15 Thiopurine sensitivity1 of 3 chip-readable variants1 not read
NUDT15 R139C, reduced enzyme activityNUDT15 rs116855232Not read: not on this chip
The R139C change reduces NUDT15 enzyme activity, so an active metabolite the enzyme normally clears persists longer. CPIC assigns metaboliser status from NUDT15 and TPMT together.
Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment is more dangerous than any genotype on this page. Show this to the prescribing doctor and let them decide. This variant is roughly 24 times commoner in South Asians than in Europeans, which is why it is here and why guidance derived from European cohorts has historically under-served this population.
TPMT Thiopurine sensitivity1 of 9 chip-readable variants1 not read
TPMT*3C, thiopurine S-methyltransferase activityTPMT rs1142345Not read: not on this chip
TPMT is the second enzyme CPIC reads alongside NUDT15. The *3C allele reduces its activity.
Not medical advice. TPMT*3C is one of several TPMT alleles and this file reads one of them, so a normal result here does not establish normal TPMT activity. Enzyme activity is measurable directly and that test, not this one, is what a clinic would use. No South Asian cohort study of this allele's effect is cited here -- the frequency data below is South Asian, the effect evidence is not.
SLCO1B1 Statin transport1 of 1 chip-readable variants1 not read
SLCO1B1 V174A, reduced hepatic transportSLCO1B1 rs4149056Not read: not on this chip
SLCO1B1 is a liver transporter: it moves certain compounds out of the blood and into hepatocytes, where they act and are cleared. The V174A change reduces that transport, so what the transporter handles stays in circulation longer.
Not medical advice, and not a reason to stop or change any medicine. Stopping a prescribed treatment on the basis of a web page is more dangerous than anything this row describes. The variant is LESS common in South Asians than in Europeans -- 4.9% against 15.9% -- so for most readers here this row will be uninformative. Effect evidence is European; no South Asian cohort is cited.
VKORC1 Warfarin sensitivity1 of 9 chip-readable variants1 not read
VKORC1 -1639, reduced enzyme expressionVKORC1 rs9923231Not read: not on this chip
This promoter variant lowers how much VKORC1 enzyme the liver produces. VKORC1 is the target of one class of anticoagulant, so the amount present matters to anyone taking one.
Not medical advice. Anticoagulant dosing is managed by blood tests that measure the actual effect in the actual person, which is far more informative than any genotype. Never change a dose on the basis of this page. Effect evidence is largely European and East Asian; no South Asian cohort is cited here.
ASPA Canavan disease carrier1 variant1 not read
ASPA E285A, Canavan disease carrierASPA rs28940279Not read: not on this chip
The commonest ASPA variant in Canavan disease, a recessive condition affecting the white matter of the brain. ClinVar classifies it Pathogenic/Likely pathogenic.
This variant's frequency has not been measured in South Asian populations. The Canavan literature is largely Ashkenazi Jewish, where it is commonest; what carrier frequency looks like in South Asia is not something this report can tell you.
ACKR1 1 variant1 not read
Duffy-null blood groupACKR1 rs2814778Not read: not on this chip
The C allele abolishes Duffy antigen expression on red cells. It confers resistance to Plasmodium vivax malaria and is associated with a lower baseline neutrophil count that is benign -- 'benign ethnic neutropenia'.
The clinically important part is the neutrophil count, because a normal result for a Duffy-null person can be read as abnormal against a reference range built on people who are not. This is close to fixed in West African populations, absent in the South Asian and East Asian cohorts, and is on the panel because the report is global.
HFE hereditary haemochromatosis2 variants2 not read
HFE C282Y, hereditary haemochromatosisHFE rs1800562Not read: not on this chip
The commonest cause of hereditary haemochromatosis, in which the body absorbs more iron than it needs. ClinVar classifies it Pathogenic.
Almost all HFE research is in European-ancestry cohorts, where this variant is commonest. It is rare in South Asians and its frequency there is not well measured, so an absence here says less than it would for a European genome.
HFE H63DHFE rs1799945Not read: not on this chip
A second HFE change described in the same 1996 paper as C282Y. It is common in many populations and, on its own, is not generally associated with iron overload; the combination most reported is one copy of this alongside one copy of C282Y.
ClinVar records conflicting classifications for this variant, and it is reported far more often than iron overload occurs, so it is shown as a finding rather than a conclusion. Its frequency in South Asians is not well measured.
ABCG2 transporter function1 variant1 not read
ABCG2 Q141K, transporter functionABCG2 rs2231142Not read: not on this chip
The T allele reduces ABCG2 transporter function. It is associated with higher serum urate and with reduced response to allopurinol, and CPIC uses it in guidance on rosuvastatin dosing.
Urate and gout have large dietary and renal components that this variant does not read. The allopurinol association describes response at a given dose rather than whether the drug works.
CYP2C9 reduced-function allele1 variant1 not read
CYP2C9*3, reduced-function alleleCYP2C9 rs1057910Not read: not on this chip
CYP2C9*3 substantially reduces enzyme activity. CPIC guidelines use CYP2C9 genotype in dosing recommendations for warfarin and for phenytoin.
A genotype is not a dose. CPIC recommendations combine CYP2C9 with VKORC1 and with clinical factors, and warfarin is titrated on INR measurement whatever the genotype says. Carrying *3 is a reason to expect a lower dose requirement, not a reason to change one.
IFNL4 hepatitis C treatment response1 variant1 not read
IFNL4 (IL28B), hepatitis C treatment responseIFNL4 rs12979860Not read: not on this chip
Genotype at this locus predicts response to interferon-based hepatitis C therapy and spontaneous clearance of the virus. The C allele is the favourable one; the T allele is associated with poorer response.
This is a variant whose clinical relevance has largely PASSED. Interferon-based regimens have been superseded by direct-acting antivirals, which cure across genotypes, and it is reported here as a well-established association rather than as a treatment decision anybody should now make.
CYP3A5 tacrolimus metabolism1 variant1 not read
CYP3A5*3, tacrolimus metabolismCYP3A5 rs776746Not read: not on this chip
The commonest reason CYP3A5 produces no working enzyme. It is the main genetic factor in tacrolimus dosing after transplantation.
Frequencies for this variant differ sharply between populations, and the South Asian estimate is less well measured than the European and African ones.
How we worked this out
Rows report variant presence against published work. The catalogue (v0.1.0) is a small selection of the variants known in each gene: ClinVar records hundreds more, and roughly half of those are deletions, duplications or repeat expansions that no genotyping chip can detect at all. Metaboliser status for CYP2C19 or TPMT needs a full star-allele diplotype and an enzyme activity test, which this file cannot give. G6PD deficiency is diagnosed by an enzyme test, not a genotype.
Whether a position can be read depends on the vendor and chip version that produced the upload, and it is measured for each file. Frequencies shown are from gnomAD v4 and the 1000 Genomes Project. The evidence for an effect often comes from European studies, and each row's tags say where it has been tested.
Technical
Technical
For the curious
The detail behind the numbers: the reference groups, your file, and the same file through other tools.
Reference groups closest to you
How close your DNA sits to each group in our reference set.
Closest areas
The areas whose reference samples your DNA is nearest to. No order, and no ranking.
7 further groups, none of them close
How we worked this out
Each file and every individual reference genome is projected into the first ten principal components of our panel. The chart is Mahalanobis distance in that space, so it scales each group by its own spread. The green band is where a typical member sits, the amber band is close but outside it, and a marker farther right is less like that group. Hover or tap a row for its distance. The closest areas come from the nearest three individual references, coarsened and left unranked.
On the same 289 people and 54 groups, this distance placed the correct group first 60.2% of the time against 35.3% for G25, and in the top three 74.0% against 63.0%. The G25 coordinates in this report are an export for other tools.
Present-day map
Where your file falls among living reference groups
Each dot is one person in our reference panel, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.
2,000 reference individuals, the same cloud a report draws, on the panel fitted today. The first two components carry 13.9% and 5.4% of total variance. Every point is one published reference individual and names itself on hover. Clusters are labelled with the region each cohort is pooled into and coloured by the continental component that region sits in. 504 individuals from cohorts pooled into no region are not drawn.
The picture is turned a quarter clockwise and mirrored, which is why it resembles a map: Europe upper left, East Asia upper right, Africa along the bottom, South Asia between them. That is a rotation and a reflection of the plane, so no point has moved relative to any other and every distance is what it was. It is a reading aid and nothing more. The axes have no meaning of their own: they are directions of maximum variance, their signs are arbitrary, and a region's position depends on which other regions are in the panel. The first component runs vertically here and the second horizontally.
Every region carries a numbered disc at its centre and an outline around its members, and the key below repeats the number — so a region too crowded to spell out on the plot can still be found on it. The outline is the convex hull of that region's own individuals with the furthest 8% trimmed off, which keeps one stray from dragging a boundary across the chart; it encloses people, not territory. Spelled out only in the key at this size: 4 Northwest European, 14 Tai and Kadai, 13 Punjab and Kashmir, 11 Mainland Southeast Asia, 16 Sierra Leone and the Upper Guinea coast, 8 Japanese, 7 Italian.
African
- 2West African forest207
- 5West African savanna113
- 10Eastern and Southern African99
- 16Sierra Leone and the Upper Guinea coast85
Bengal and the central belt
- 15Bengal and the central belt86
East Asian
- 1Eastern Chinese208
- 8Japanese104
- 11Mainland Southeast Asia99
- 14Tai and Kadai93
European
- 4Northwest European190
- 6Iberian107
- 7Italian107
- 12Finland and Karelia99
Indo-Gangetic Plain
- 9The Ganges Plain and Gujarat103
- 13Punjab and Kashmir96
Southern Peninsula
- 3Deccan and the Tamil Plains204
3,043 of 12,770 panel markers placed you on the map.
The groups nearest your position, closest first:
- Han Chinese
- Southern Han Chinese
- Japanese
This is a position on a plot. Sitting near a group means your genome resembles that group on these axes. It does not say who your ancestors were.
How we worked this out
The axes come from the reference panel alone. Your file is projected onto them using the markers it shares with the panel, so it moves no point on the map.
Nearest is measured over more axes than the two drawn, to the middle of each group, so a group that looks close here can rank lower in the list.
Ancient source map
Where your file falls among ancient genomes
Each dot is one ancient person, placed by their genome on the first two axes of a principal component analysis. The ringed point marked You is your file, placed the same way.
558 excavated individuals with at least 100,000 1240K SNPs, on the first two principal components (5.1% and 1.7% of variance). 41 individuals sit far from their group's centroid and are left out of every outline and centre; they are not drawn here, and the methods page shows them. Axes have no meaning of their own, and a group's position depends on which other groups are in the map. 4 sources pooled from several regions are not drawn: a period is not a people.
Named only in the key at this size: 9 Japan, c. 4,500 BP.
Before 10,000 BP
- 1Russia, c. 32,500 BP12
10,000 to 7,000 BP
- 2Iran, c. 10,000 BP26 +2
- 3Turkey, c. 8,500 BP85 +9
- 4Russia, c. 8,500 BP48 +3
- 5Cameroon, c. 8,000 BP8
- 6Sweden, c. 7,500 BP17 +2
7,000 to 5,000 BP
- 7Steppe and Siberia, c. 5,500 BP15 +1
5,000 to 3,500 BP
- 8China, c. 4,500 BP113 +3
- 9Japan, c. 4,500 BP14 +1
After 3,500 BP
- 10South Africa, c. 2,500 BP7
- 11China, c. 2,000 BP136 +14
- 12Taiwan, c. 1,500 BP21 +4
- 13China, c. 1,500 BP24
- 14Papua New Guinea, c. 500 BP32 +2
Counts are individuals kept; +n is individuals dropped. Thin: fewer than 5 usable individuals. An outline is drawn around each excavation group of six or more people, not around a source.
8,853 of 138,630 sites in your file are on the map.
The sources nearest your position, closest first:
- China, c. 4,500 BP
- China, c. 2,000 BP
- Japan, c. 4,500 BP
This is a position on a plot. Sitting near a source means your genome resembles that group on these axes. It does not say who your ancestors were.
How we worked this out
The axes come from the ancient individuals alone. Your file is projected onto them using the sites it shares with them, so it moves no point on the map.
Nearest is measured over the first four axes, to the middle of each source. The plot shows two of them, so a source that looks close here can rank lower in the list.
On eleven test genomes the nearest source fell in the same region as the qpAdm fit. Inside a region, the order of the nearest sources does not track the qpAdm proportions.
The same ancestor, twice
Stretches inherited twice
Places where both copies of a chromosome came from the same ancestor. A long stretch points to a recent shared ancestor.
This file carries 55,688 autosomal markers we could place, and runs of homozygosity need at least 200,000 to measure rather than to guess at. Below that the short runs vanish and the detector starts inventing others, so the honest answer is that we cannot tell.
Your file
How complete your raw DNA file was. This sets how precise every other chapter can be.
A: complete enough for every chapter.
The checks above are the ones that flagged; the rest passed. Nothing was sequenced by us. We read the export you already had.
How we worked this out
Heterozygosity is 28.9%: the share of read positions where your two copies differ. A consumer chip usually reads between 24% and 31%; well outside that range points to a file problem rather than to anything about you.
Your file is kept and used to build future reference panels, under a Creative Commons licence. We never sell or release an individual genome, and no genetic data is ever written to a log.
Who you were compared against
524 groups, 10,922 people. The reference behind every number in this report.
Every study behind the panel, and its regions
People alive today: 6,745 people, 18 studies. The reference this report measures your ancestry against.
Excavated individuals: 13,454 people, 309 studies. Used by the ancient peoples chapters. The region percentages do not use them.
the other 297 studies
Behind the South Asian regions: 2,148 people across 11 studies, counted by the study that collected them.
The full list, the sampling map and the panel drawn by genetic similarity: How it works.
How we worked this out
Accuracy is the share of held-out people whose largest region came out right: 120 people from 24 groups at the full panel of 12,770 markers, and again at 10,317 markers, the density a real 23andMe export reaches. The people tested are part of the reference set they are scored against, so their own data helped define the groups. Someone uploading a file gets no such help, so the real-world rate is a little lower. Five individuals per population. A perfect score on 120 people is consistent with a true rate above about 97%. The Indigenous American component is built from four samples, three of them majority European, so a correct call there is an easier test than it sounds. It measures which region comes out largest. Whether the printed range covers the true value is a separate measurement we have not run. 10 of the held-out set are not scored (African Caribbean and African American samples, which define no component). Measured 2026-09-09 against the 12,770-marker panel this report used. A rebuilt panel re-runs it.
A few dozen people per group is what makes our ranges as wide as they are. A narrower range would need more people. West Asia is fitted from Bedouin, Palestinian and Druze samples, and Oceania from pooled Papuan and Bougainville individuals. Central Asia is not yet a region of its own.
Coordinates for other tools
Twenty-five numbers you can paste into community ancestry tools.
this map was fitted and checked on South Asian genomes, and this file is 0% South Asian. Placing it would mean extrapolating a map nobody has verified there
Other calculators
The same file run through popular community calculators. Where they agree with our result, the answer does not depend on the model.
Community calculators need more of their own markers than this file carries, so none were run.