Risposte alle domande frequenti e un glossario dei termini chiave della genealogia genetica.
Domande frequenti
FAQ
My Y haplogroup is O2a - what does that mean?
O2a is a node on the paternal haplogroup tree: it means that along your purely paternal line you can trace back to the ancient male ancestor who carried the O2a defining mutation. It shows that you belong to the O2a paternal subclade, but it does not mean all your ancestors came from a single population. To get more specific, look at the downstream branches of O2a and the defining mutation that belongs to you.
Is O2a the same as O-M122?
Most likely they are two names for the same subclade: O2a is a hierarchical (ISOGG-style) name and O-M122 is a mutation name (M122 is the defining mutation of that subclade). To decide whether two names are the same subclade, check whether they point to the same parent-node chain and share the same defining mutations, rather than just comparing the names.
Why is my haplogroup so different from my friend's?
There are many downstream branches under one paternal haplogroup, so you and your friend may share a large upstream clade (both in O, for example) but fall into different downstream subclades, which makes the names look very different. One of you may also sit more upstream while the other sits more downstream (different resolution depth). This is normal; check whether your lineage paths meet at some node to judge how closely related you are.
Why does my subclade have so few (or so many) samples?
The further downstream in the tree, the finer the split, and a single fine branch usually has very few samples, whereas clades near the root or large popular clades have many. Few samples does not mean you are a rare or unusual population - it just means that few people from that fine branch have been collected or tested so far. As more samples arrive, the branch keeps being refined.
What is the earliest ancestor haplogroup I can trace back to?
Follow your lineage path up to the root: the most upstream node on that chain is the earliest paternal haplogroup you can trace back to (usually one of the big letter clades such as O, C, N, D or R). The closer to the root, the earlier the TMRCA and the broader the scope.
Where does the data come from, and how often is it updated?
The paternal tree and mutation data combine public studies, ancient DNA samples and user-submitted samples, and are continuously integrated and updated; after new samples arrive, branches are refined, so haplogroups and the tree evolve over time. For the precise data sources and update policy, refer to the site's documentation.
How do I look up my haplogroup path on the website?
On a haplogroup page you can see that node's defining mutations (equivalent mutations), the lineage path up to the root and the downstream subtree; the Samples page can filter by haplogroup, surname, ethnic group or country to look at the population distribution of a subclade. For programmatic queries, use the site's public read-only API (/api/v1/agent/*).
Is my sample in the same subclade as a particular ancient DNA sample?
If you share a subclade with that ancient sample - both matching the defining mutation of the same node, or of the same node further upstream - then you belong to the same paternal subclade and can compare its age and region. Ancient samples are often used to calibrate the tree's dates and distributions, and are an important reference for studying the history of a subclade.
Can a haplogroup tell me my surname or ethnic group?
It can only give a tendency, not a definite answer. One surname or ethnic group usually contains several paternal haplogroups, and one haplogroup spans many surnames and ethnic groups. You may observe that a haplogroup is more frequent among a certain surname or region, but that is a probabilistic association, not a conclusion.
Why does the haplogroup tree keep changing?
Because new samples, new tested sites and new research keep arriving: discovering a new defining mutation adds or splits a node, and naming and TMRCA may also be revised. This is normal and healthy evolution - it shows the data is being continuously improved.
How is my private data protected?
Personal test data and ancestry information are sensitive personal information. User data is managed under the site's privacy policy and is not exposed together with the public knowledge base; public tree/sample distributions use aggregated information. For the exact data-use and privacy terms, refer to the site's privacy statement.
Glossario
Terminologia
Haplogroup
A haplogroup is a lineage of people defined by a shared set of genetic markers (mutations). Paternal (Y-DNA) haplogroups are defined by mutations on the Y chromosome (Y-SNPs); men in the same paternal haplogroup are considered to descend from one ancient male ancestor.
- Each haplogroup is identified by one or more defining mutations (see 'Equivalent mutation'); for example O2a is defined by markers such as O-M122.
- Haplogroups form a tree: the root node is the earliest ancestral haplogroup, and the further downstream (child nodes) you go, the more recent and more finely resolved the branches are.
- Names are usually a letter plus digits (e.g. O2a) or a letter plus a representative mutation name (e.g. O-M122).
- A paternal haplogroup reflects only the single line father's father's father... It is not the full ancestry of an individual.
Paternal (Y) haplogroup vs maternal (mtDNA) haplogroup
A person's ancestry has two single lines:
- Paternal haplogroup (Y haplogroup): passed along the Y chromosome from father to son, carried only by men; it reflects the paternal-line ancestor.
- Maternal haplogroup (mtDNA haplogroup): passed through mitochondria from mother to children, carried by both men and women, but only women pass it on; it reflects the maternal-line ancestor.
The two are independent: your Y haplogroup and your mtDNA haplogroup come from different distant ancestors. The ancestor tree / paternal tree on this site is based on Y haplogroups. Beyond these two single lines, the contribution of all other ancestors is reflected by autosomal DNA (see 'Autosomal ancestry composition vs Y haplogroup').
SNP (single nucleotide polymorphism)
An SNP (single nucleotide polymorphism) is a difference of a single base at one position in the genome, for example a C that has become a T at a particular position. In Y-chromosome research, SNPs are the main basis for defining paternal haplogroups.
- An SNP is usually referred to by its coordinate or its mutation name (e.g. M122, F15365).
- A 'defining mutation' is the SNP used to delimit a haplogroup; the same haplogroup may have more than one equivalent mutation.
- Mutation records on this site include coordinates (with coordinate sets for several reference genomes such as hg19 and CP086569), the reference base (ref) and the alternate base (alt); see 'Mutation names and reference genomes'.
Y-STR (short tandem repeat)
A Y-STR (short tandem repeat) is a short sequence on the Y chromosome that is repeated over and over (for example 'AGAT' repeated a number of times). Different men can have different repeat counts at the same locus, so Y-STRs are widely used for paternal kinship comparison and genealogical identification (a Y-STR haplotype).
- Unlike Y-SNPs: STRs change quickly (high mutation rate) and suit paternal relationships within the last few hundred to few thousand years; SNPs change slowly and suit defining haplogroups on a large scale.
- Men in the same Y haplogroup usually have similar but not identical Y-STR values; identical STRs do not prove the same haplogroup (SNPs are needed as well).
Mutation names and reference genomes (hg19 / CP086569)
The same mutation (SNP) has different coordinates under different reference genomes, so coordinates must always be labelled with their reference set, otherwise they get mixed up. Mutation records on this site usually store several coordinate sets at once. The common reference genomes are:
- hg19 (GRCh37): historically the most widely used human reference genome.
- hg38 (GRCh38): a newer version, whose coordinates are a different set of numbers from hg19.
- CP086569 (e.g. CP086569.1 / CP086569.2): a newer T2T-style reference genome, where '1'/'2' are different versions.
Always check which reference a coordinate belongs to; comparing coordinates from different references directly will misalign them. Mutation names (e.g. M122, F15365) do not change with the reference genome and are a more stable identifier.
ISOGG nomenclature and commercial -MF names
The same lineage often carries two parallel naming systems:
- ISOGG nomenclature: the international haplogroup nomenclature system (e.g. O2a2b1a1a1a1), built from letters and digits in a hierarchy.
- Commercial name (-MF): a name given by a commercial testing company from a mutation name, often prefixed with -MF (e.g. O-MF12345), which is easy to remember and cross-reference.
When both refer to the same lineage they are equivalent. Nodes from different sources on this site may use one or the other, so two names can look different while being the same lineage (see 'Why do I get a different haplogroup from different companies or databases?'). About 13% of the tree nodes on this site carry an -MF name.
Equivalent mutation
Equivalent mutations are several mutation names that represent the same event at the same position on the phylogenetic tree. They are equivalents of one another, are counted only once, but appear side by side under multiple names (usually joined by a slash, e.g. FGC24753/YP1711).
- Effect: as soon as any one of them is detected, the sample can be assigned to the haplogroup that carries that mutation.
- In the tree and in records, equivalent mutations are usually stored merged (the merged mutation table on this site has roughly 380,000 rows, about 35% of which contain slash-joined merged names).
- Different from a 'parallel mutation': an equivalent mutation is an alias of the same ancestral event, while a parallel mutation is a different event that happens to fall at a similar position (see 'Parallel mutation').
Parallel mutation
Parallel mutations are mutations that occur independently, multiple times, in different lineages; they can give different haplogroups a marker at a similar coordinate (for example one coordinate/base combination shared by several haplogroups).
- Consequence: a single SNP coordinate may map to several haplogroups. In measurements on this site, one SNP coordinate can map to more than ten different haplogroups.
- For this reason, resolving an SNP to a haplogroup must return all matches, never just one, otherwise the call will be wrong.
- Different from an 'equivalent mutation': equivalents are aliases of one event, parallels are coincidence/convergence of different events.
Subclade (downstream branch / child node)
A subclade is a node in the tree together with all of its downstream (deeper, more recent) haplogroups. Position is usually described as:
- Upstream: closer to the root, earlier, broader (e.g. O).
- Downstream: farther from the root, more recent, more finely split (e.g. O2a2b1a1).
Samples inside one subclade share that subclade's defining mutations. The further downstream you go, the more branches and the fewer samples there are; so 'my subclade has very few samples' is common and normal.
TMRCA (node age)
TMRCA (Time to Most Recent Common Ancestor) is the age of the most recent common paternal ancestor of everyone within a haplogroup (subclade), i.e. the 'age' of that node.
- Basis of estimation: the number of defining mutations, mutation rates and a molecular clock model, calibrated with samples of known age.
- The result is an interval estimate (e.g. 'about 5,000-8,000 years ago') and carries error; different algorithms or parameters give different values.
- Nodes closer to the root are older; the further downstream, the more recent.
- It is often cross-checked against archaeological and ancient DNA evidence. See 'How is TMRCA (node age) estimated, and how large is the error?'.
Lineage path (upward lineage)
A lineage path is the sequence of haplogroups running from a given haplogroup up through its parent nodes all the way to the root (e.g. J-ZS1805 -> ... -> J -> ... -> root). It shows the evolutionary path and hierarchical relationships of that subclade.
- This site offers a lineage query that traces a node up to the root.
- The opposite is the 'subtree' (all downstream nodes of a node).
- To decide who is earlier or who contains whom, rely on the parent-node chain (the lineage relationship), not on a field named 'level': that field may not equal the true depth.
Mutation under observation
'Under observation' means that the evidence for that mutation (or the branch it belongs to) is not yet sufficient: it has not been formally placed at a level, or samples are still being collected to verify it.
- Mutation records on this site flag this with the under_observation field (several thousand records are under observation).
- Queries and displays can usually exclude under-observation mutations to show only confirmed ones, or include them to follow the latest progress.
- Different from the status field: in the mutation table on this site the status field is always 0, so do NOT filter with status=1 or you will get an empty set; to exclude unconfirmed branches use under_observation=0.
Samples and reference samples
- User sample: an individual's Y-DNA test result, which can be annotated with its haplogroup/subclade, surname, region and so on.
- Reference sample / ancient DNA sample: a sample from a published study or from ancient human remains, used to calibrate tree structure, TMRCA and distributions.
The Samples page on this site brings user samples and ancient samples together and can filter by haplogroup, surname, ethnic group or country, so you can examine the population distribution and history of a subclade.
Autosomal ancestry composition vs Y haplogroup
The two describe different facts and must never be conflated:
- Autosomal ancestry (e.g. '60% northern Han, 30% southern Han'): the mixed proportions contributed by all of your ancestors; a whole-genome, multi-line ancestry portrait.
- Y haplogroup: only your purely paternal line; a single-line, point-like fact.
So a person can easily have autosomal ancestry that is mostly from one population while their Y haplogroup comes from another direction. The ancestor tree on this site is based on paternal haplogroups and complements autosomal ancestry composition.
Metodologia
How to read a Y haplogroup tree
The basics of reading a Y haplogroup tree:
1. The root is at the top (or on one side) and is the earliest ancestor; the further downstream, the more recent.
2. Each node is a haplogroup, identified by one or more defining mutations.
3. Adjacent nodes are linked by a father-to-son inheritance relationship; going up from a node to the root gives its lineage path.
4. The subtree of a node is all of its downstream haplogroups; the subtree of a large clade can be enormous (tens of thousands of nodes) and normally has to be browsed level by level.
5. Assignment is decided by defining mutations; names (ISOGG or commercial -MF) are only labels and may differ between sources.
How haplogroups are named
There are two common naming systems for paternal haplogroups:
- Hierarchical naming (ISOGG style): letters plus digits, such as O -> O2 -> O2a -> O2a2b1a1; the deeper the level, the longer the code, which indicates a more downstream position.
- Mutation-name naming: named after a representative mutation, such as O-M122 or O-F15365, which gives the defining mutation directly.
In practice the two systems coexist and can be mapped onto each other (two names for the same lineage). Commercial companies also use internal names (the commercial names seen on this site are marked with -MF). When comparing, go by 'defining mutation + parent-node chain', not by name alone.
Why one SNP can map to several haplogroups
There are two main reasons:
1. Parallel mutation: the same genomic coordinate/base change arose independently in different lineages, so it is shared by several haplogroups.
2. Equivalence/merging: different names are in fact the same mutation, or merged storage means that one name, when expanded, covers several downstream nodes.
So it is dangerous to equate 'one SNP' with 'one haplogroup'. The correct approach: when resolving an SNP, return all matching haplogroups (on this site one SNP can match more than ten haplogroups at most), and disambiguate using the parent-node chain and additional markers.
How a DNA sample is assigned to a haplogroup
The general workflow (simplified):
1. Sequencing yields the mutation sites (SNPs) and Y-STR values on that man's Y chromosome.
2. The detected mutations are compared level by level against the defining mutations in the tree, from upstream to downstream, to see which level's defining mutations the sample matches.
3. The deeper the defining mutation matched, the more downstream a haplogroup/subclade can be determined.
4. If the match involves parallel mutations or the evidence is insufficient, the sample is temporarily labelled 'under observation' or kept at a safer, more upstream level.
5. Adding more samples from the same subclade allows finer resolution. So haplogroup assignment is a process that keeps refining as data and samples accumulate.
How TMRCA (node age) is estimated and how large the error is
How TMRCA (node age) is estimated:
- Count the mutations accumulated in that subclade and invert them through a molecular clock model ('mutation rate x time').
- Calibrate with samples of known age (e.g. radiocarbon-dated ancient DNA) to narrow the error.
- The result is usually an interval rather than a single value, and the error can reach hundreds to thousands of years, depending on mutation rate, sample size and model.
So the same subclade may get different ages under different algorithms or databases, which is normal. Read the interval and the order of magnitude rather than insisting on an exact year.
Why different companies or databases give a different haplogroup
Common reasons:
1. Different naming systems: ISOGG hierarchical name vs mutation name vs commercial internal name (-MF) can make the same subclade look different.
2. Different tree versions or resolution depth: new samples keep being added and the tree keeps being refined, so you now sit more downstream with a longer name.
3. Different sets of tested sites: the more probes/sites tested, the deeper the subclade you can land on.
4. Parallel mutations or insufficient evidence can lead to different upstream/downstream calls.
Conclusion: if haplogroups from different sources have a consistent parent-node chain and overlapping defining mutations, they are in fact the same; the differences come mostly from naming and versions, not from you belonging to a different population.
The role of equivalent and parallel mutations in the tree
- Equivalent mutation: several aliases of the same ancestral event, equivalent to one another. A node in the tree often lists several equivalent names (joined by slashes); detecting any one of them assigns the sample to that node. Merged storage keeps them consistent.
- Parallel mutation: similar mutations occurring independently in different lineages, which makes different nodes appear to share a coordinate. It requires resolution to return all matches and disambiguation via the parent-node chain or more markers.
Both affect how many haplogroups one mutation maps to, and both are core issues that SNP resolution has to handle.
Why the tree is expanded level by level (large clades expand only a few levels by default)
A mature paternal tree can reach tens of thousands of nodes, and some large clades (for example one branch of O) have more than ten thousand descendants. Returning every downstream node at once would produce more data than can be read or transferred.
Common handling:
- Large clades expand only a few levels relative to the root by default (e.g. 6 levels); deeper child nodes are collapsed and can be drilled into level by level on demand.
- Small clades (few descendants), or an explicit user request, expand fully.
That way you can see the backbone structure and still go into fine branches when needed. Subtree queries on this site use the strategy 'limit levels by default + boundary nodes can be drilled into further'.