Introduction
This study presents a detailed analysis of Steppe admixture in Indo-Aryan populations, with a particular focus on the Jatt Sikh population in Punjab. Utilizing both personal genetic data and the Allen Ancient DNA Resource, we model the genome at two depths: a proximal three-source model — the field-standard frame for South Asian genomes (Narasimhan et al., 2019) — and a distal four-source model that resolves the Steppe source into its own ancestral streams. In the proximal frame the genome resolves as 42.8% Steppe from the Andronovo (Steppe_MLBA) horizon, 33.2% Iranian farmer-related, and 24.0% Ancient Ancestral South Indian (AASI). The distal model opens the Steppe end into its EHG forager core and a 13.4% European farmer (EEF) stream, matching the roughly one-third EEF that Steppe_MLBA populations carry. A Bactria-Margiana Archaeological Complex (BMAC) contribution cannot be resolved as a separate stream.
Key Findings
The genome resolves into the same ancestry streams at two model depths. On the left, the distal four-source model resolves the deep ancestral streams themselves: the EHG forager core of the Steppe, the European farmer (EEF) ancestry the Steppe absorbed on its way east, the Iranian/CHG-related axis, and AASI. On the right, the proximal three-source model — the field-standard frame of published South Asian work (Narasimhan et al., 2019) — recombines the Steppe-side streams into their Bronze Age carrier, the Andronovo (Steppe_MLBA) horizon, at 42.8%:
| Model | Source | Role | Weight (%) | Z |
|---|---|---|---|---|
| Proximal 3-way — Andronovo (right panel) CORE12 · p = 0.060 (best of its rotation) | Andronovo (Steppe_MLBA) | Steppe | 42.8 ± 3.1 | 13.8 |
| Iran_GanjDareh_N | Iranian farmer-related | 33.2 ± 3.5 | 9.4 | |
| ONG (Onge) | AASI proxy | 24.0 ± 1.8 | 13.6 | |
| Distal 4-way (left panel) deepened right set · p = 0.191 | Russia_Samara_Eneolithic | EHG (Steppe forager core) | 24.6 ± 1.9 | 12.8 |
| Stuttgart_LBK | European farmer (EEF) | 13.4 ± 4.0 | 3.3 | |
| Iran_GanjDareh_N | Iranian-related | 35.7 ± 6.5 | 5.5 | |
| ONG (Onge) | AASI proxy | 26.3 ± 3.1 | 8.5 |
- Only the Andronovo horizon fits. With a Neolithic farmer (Iran_GanjDareh_N) in the model, the Steppe source itself must supply the genome’s European-farmer-related ancestry. Andronovo (Steppe_MLBA) does, and fits at p = 0.060. The Steppe ancestry came through the Sintashta–Andronovo horizon, which had already absorbed European farmer ancestry (Allentoft et al., 2015).
- The four-source model splits the 42.8% open: EHG 24.6% + European farmer (EEF) 13.4% + Iranian-related 35.7% + AASI 26.3% (p = 0.191). The EEF stream only becomes visible at this depth. And the numbers agree: Steppe_MLBA populations are roughly one-third European farmer (Narasimhan et al., 2019), so 42.8% Andronovo predicts about 13% EEF — the model measures 13.4%.
- The Steppe share depends on the farmer used; the AASI share does not. With a Bronze Age farmer the Steppe reads 33–38%; with the Neolithic farmer, 42.8%. The AASI share barely moves — 24–29% in every model.
- Three lines of evidence converge on Sintashta–Andronovo. The autosomal fits point to Steppe_MLBA, the class published work identifies for South Asia (Narasimhan et al., 2019). The paternal lineage, R1a-Z93 > L657, is the Asian branch of the R1a expansion: its European sister branch (Z282) went with Corded Ware into Europe (Underhill et al., 2015), and the South Asian subclades date to the Bronze Age (Silva et al., 2017). And the horizon’s southward expansion through Central Asia (c. 2100–1500 BCE) is the only Steppe movement that reaches the Punjab in the required window. Which branch within the horizon remains open.
- Three sources are enough. Adding an explicit European farmer term as a fourth source fails in every configuration — that ancestry is already inside the Steppe source, which formed as a mixture of Yamnaya-related and European farmer ancestry (Haak et al., 2015; Allentoft et al., 2015).
- No separate BMAC (Oxus) stream. Bronze Age Oxus genomes (Gonur, Sapalli-tepe, Ulug-depe) behave like the other farmer sources; swapping them into the farmer slot changes nothing (detailed under Investigating BMAC Influence). Narasimhan et al. (2019) reached the same conclusion for South Asia generally.
- The Onge stand in for AASI. No ancient AASI genome has been sequenced, so the lineage is reconstructed, not sampled — and clustering accordingly never recovers a discrete AASI component. The Onge are the closest sampled relatives and carry no West Eurasian ancestry (Reich et al., 2009), which is what makes the farmer and AASI streams separable; the AASI figure is measured through that proxy.
Methodology
The sections above present the findings; this one documents how they were produced, in enough detail to re-run them.
Data
The target is a Jatt Sikh genome genotyped independently on two consumer panels, 23andMe v5 and AncestryDNA, each with over 500,000 SNPs. The two panels serve as mutual checks. qpAdm compares the target against ancient sources over the intersection of SNPs present in both the query genotype and the Allen Ancient DNA Resource 1240k panel. Because the Bronze Age ancient samples are the limiting factor, the informative overlap is nearly identical for the two chips, and their ancestry estimates converge to the same values with slightly different standard errors. The current models were fitted on the merged target against AADR version 66; raw outputs are in the Supplementary Data.
Tools Used
The analysis rests on the following tools and datasets:
- ADMIXTOOLS 2 (R package; Maier et al., 2023) — qpAdm model fitting, f-statistics, and cached f2 blocks. docs
- Allen Ancient DNA Resource — ancient and present-day genotypes. Fitted on v54.1.p1 1240k, re-fitted on v66. AADR
- 23andMe v5 chip — target genotypes, over 500,000 SNPs.
- AncestryDNA — second target genotype set, over 500,000 SNPs, used to cross-check the first.
- Big Y-700 — Y-DNA haplogroup: R1a-Z93 > R-L657 > R-FTF40903.
- ADAMIXTURE 1.7.5 — unsupervised ancestry clustering, K = 5 to 12, on AADR v66. repo
- f4-statistics — tests of what each source population itself carries.
- IllustrativeDNA G25 — independent admixture check. results
Model fitting
Admixture models were fitted with qpAdm in ADMIXTOOLS 2 (Maier et al., 2023), run in genotype mode with allsnps = TRUE: each f-statistic uses the SNPs available for its own populations, the appropriate regime when a single array-genotyped target meets ancient capture data of heterogeneous coverage. The regime is stamped in every raw model receipt in the Supplementary Data. qpAdm models the target as a mixture of a small number of source populations (the "left" set) and evaluates the model against a set of reference populations (the "right" set) chosen to be differentially related to the sources. The right set does the discriminating work, so its composition matters more than its size.
The reference set used here contains twelve populations, one anchor per ancestral stream the sources draw from: deep African outgroups (Mbuti, Mota), Upper Palaeolithic Eurasians (Ust-Ishim, Tianyuan, Kostenki), Papuan for the deep eastern lineage, Karitiana for Ancient North Eurasian ancestry, an East Asian Neolithic anchor (Mongolia_N), Caucasus and Western hunter-gatherers (Kotias, Iron Gates), and two farmer anchors (Levant PPNB, Anatolian Neolithic). This follows the design guidance of Harney et al. (2021): no population directly ancestral to a source, and the AASI proxy (Onge) kept out of the reference set it is judged against. The set was validated before use: it is stable on the accepted model and correctly rejects deliberately wrong sources, including an East-Asian-shifted steppe population and an EEF-bearing one.
Source selection
The two headline models in Key Findings use the distal design of published South Asian work. The proximal three-source model pairs Andronovo (Steppe_MLBA; Russia_KrasnoyarskKrai_MLBA_Andronovo) with Iran_GanjDareh_N (Neolithic Zagros) and the Onge, on the twelve-population reference set above. Iran_GanjDareh_N is the standard distal farmer choice: it predates every other source in the model and carries no Anatolian or Steppe admixture of its own. The distal four-source model replaces Andronovo with its ancestral streams — Russia_Samara_Eneolithic for EHG and Stuttgart_LBK for the European farmer (EEF) stream — alongside Iran_GanjDareh_N and the Onge. Because Turkey_N and Kotias in the twelve-population set are ancestral to two of these sources, that model is evaluated on an adjusted reference set with those two removed and Natufian and Afontova Gora added.
For the supporting time-window analysis, each Bronze Age window supplies its own Steppe source: five Yamnaya and Afanasievo populations for the Early Bronze Age, three Corded Ware populations for the Middle, and Sintashta and Srubnaya for the Late, with Andronovo (Steppe_MLBA) fitted additionally (p = 0.280). The figure shows one representative per window; the full grid of 70 models (7 farmer candidates × 10 Steppe sources, distributed 5/3/2 across the windows) is in the Supplementary Data.
The Iranian farmer-related source is held fixed across all three windows at Iran_ShahTepe_BA (Gorgan plain, 3235-3150 BCE), so that the only variable in the comparison is the Steppe era. Three criteria led to this choice. First, chronology: the farmer source must predate the earliest window, and every Bactria-Margiana era candidate (2300-1000 BCE) is younger than the Early Bronze Age window itself. Second, feasibility: the one other old-enough candidate, Chalcolithic Anau (Turkmenistan_C), fails every Early Bronze Age combination (p = 0.011-0.037), while ShahTepe is feasible in all ten window-by-Steppe combinations. Third, data quality: ShahTepe is the best-covered candidate in the set (7 individuals at 2.14x median coverage; the Gonur group has 34 individuals at 0.01x).
The deep South Asian source is the Onge (ONG). As discussed under Key Findings, the Onge stand in for the unsampled AASI lineage; they carry no detectable West Eurasian ancestry, which is what allows the farmer and AASI streams to be separated.
Model evaluation
A qpAdm p-value tests whether the proposed model is consistent with the data given the reference set. A model is retained when p exceeds 0.05 and every ancestry coefficient is positive and resolved; a high p-value does not prove the model, and p-values do not rank passing models against one another (Harney et al., 2021). Models were additionally stress-tested by adding fourth sources (European farmer, Bactria-Margiana populations) to check whether three streams suffice, and cross-checked against the model-free ADAMIXTURE clustering shown below, which requires no source assumptions.
Finally, because the target is a consumer genotyping array co-analysed with capture and shotgun ancient data, the models were checked for sequencing-platform bias using the Compatibility SNP panel of Fournier, Fulton, and Reich (2026), which restricts analysis to positions with minimal technology-specific bias. Re-fitting the three Bronze Age window models on the 246,529 SNPs shared with the panel (59% of the full set), under matched extraction settings for both arms, moves no ancestry estimate by more than 2.0 percentage points, within one standard error in every case, while standard errors widen by about 10% as expected from the reduced SNP count. The reported proportions are therefore not artifacts of mixing genotyping platforms.
Results
The three subsections below work through the Bronze Age windows in order. Each follows the same shape: the historical setting, the fitted models for that window, the four-source stress test, and what the window adds to the overall picture.
Early Bronze Age (3300-2600 BCE): The Yamnaya Influence
The Yamnaya culture emerged on the Pontic-Caspian steppe in the Early Bronze Age. Known for their wheeled vehicles, advanced metallurgy, and distinctive kurgan burial practices, the Yamnaya people reshaped Eurasian genetics and culture. Recent genomic research shows that Yamnaya populations drew approximately four fifths of their ancestry from the Caucasus-Lower Volga (CLV) cline and the remainder from Ukrainian Neolithic hunter-gatherers (UNHG) (Lazaridis et al., 2025). The CLV cline represents a gradient of admixture between Caucasus hunter-gatherer (CHG) ancestry and steppe populations, while the UNHG component reflects interactions with local European groups. From this steppe base the Yamnaya spread widely across Eurasia (Haak et al., 2015).
That deep structure is what the basal model reads in this genome. Fitted from the streams that existed before the Bronze Age began, the genome resolves into the Steppe’s EHG forager core (24.6%), a European farmer (EEF) stream (13.4%), Iranian-related ancestry (35.7%) — the axis that also carries the CHG side of the CLV cline — and AASI (26.3%):
Five Early Bronze Age Steppe sources were also tested against the fixed farmer source (Iran_ShahTepe_BA) and the Onge. All five models fit, with closely similar proportions (AADR v66):
| Steppe source | p | Steppe (%) | Iranian farmer-related (%) | AASI (%) |
|---|---|---|---|---|
| Russia_Samara_EBA_Yamnaya (representative) | 0.414 | 35.7 ± 3.8 | 39.2 ± 3.8 | 25.2 ± 1.5 |
| Russia_Khakassia_Afanasievo | 0.633 | 37.8 | 36.9 | 25.3 |
| Russia_Altai_Afanasievo | 0.510 | 38.1 | 36.5 | 25.3 |
| Russia_Kalmykia_EBA_Yamnaya | 0.376 | 36.6 | 37.9 | 25.5 |
| Russia_Orenburg_EBA_Yamnaya | 0.528 | 36.2 | 38.5 | 25.3 |
A four-source model adding an explicit European farmer term (Austria_N_LBK, n = 103) was also tested for this window. With ShahTepe as the farmer source the added coefficient is 2.1 ± 3.4 percent (Z = 0.61), indistinguishable from zero. The behaviour of this term across model variants is informative in both directions. With European-farmer-free farmer sources (Ganj Dareh, Anau, Fergana) alongside Yamnaya, the term stays at 0–4 percent with |Z| ≤ 1.3: no separate European farmer stream is required. With classic Steppe_MLBA sources such as Sintashta or Alakul in the Steppe slot, the term goes negative (−4 to −12 percent depending on the farmer): those populations carry more European farmer ancestry than the target's Steppe stream did. Only one configuration produces a positive term (6–7 percent, with the AASI-rich Indus-Periphery-like farmer Iran_ShahriSokhta_BA2, whose own composition redistributes the coefficients). Read together, the tests bound the European farmer content of the incoming Steppe stream: present, but below the level of the sampled classic Sintashta–Andronovo populations, pointing to a vehicle at the horizon's earlier, less-admixed edge.
| Ancestry component | Source | Estimate (%) | Z |
|---|---|---|---|
| Steppe | Russia_Samara_EBA_Yamnaya | 37.5 ± 3.9 | 9.6 |
| Iranian farmer-related | Iran_ShahTepe_BA | 33.6 ± 4.9 | 6.8 |
| AASI | ONG (Onge) | 26.9 ± 1.8 | 15.2 |
| European farmer (EEF) | Austria_N_LBK | 2.1 ± 3.4 | 0.61 |
Standard errors are shown for the representative models; the full grid output is in the Supplementary Data.
Observations for this window:
- The Steppe share is 36-38% under every Early Bronze Age source, including Afanasievo. This uniformity reflects the coherence of the early Steppe gene pool rather than identifying any of these populations as the migrating group; the transmitting culture is constrained separately (see Key Findings). The scale of the component supports a substantial demographic contribution from Steppe-related populations of the kind proposed by Anthony (2007).
- The Iranian farmer-related component is the largest in this window (36-39%). It reflects genetic connections between South Asia and the Iranian plateau region that predate the Indo-Aryan migrations (Broushaki et al., 2016).
- AASI ancestry is stable near 25% in every model (Basu et al., 2016).
Middle Bronze Age (2900-2350 BCE): Corded Ware
The Early Bronze Age sources above are chronological stand-ins; the Middle Bronze Age supplies the first admixed European Steppe populations. The Corded Ware culture represents a fusion of Steppe ancestry with European Neolithic farmer populations (Haak et al., 2015). Our qpAdm analysis for this window tests three Corded Ware populations (Poland, Czechia, and Esperstedt in Germany) as the Steppe source.
The Corded Ware populations themselves are admixed. Before modelling the target, we resolved Czech Corded Ware as a two-source mixture: Yamnaya-related Steppe ancestry plus European farmer-related ancestry represented by Globular Amphora (Allentoft et al., 2015). This matters for everything that follows: from the Middle Bronze Age onward, the Steppe source itself carries European farmer ancestry folded inside it.
| Ancestry component | Source | Estimate (%) |
|---|---|---|
| Yamnaya-related Steppe | Russia_Samara_EBA_Yamnaya | 71.1 ± 1.5 |
| European farmer-related (EEF) | Ukraine_EBA_GlobularAmphora | 28.9 ± 1.5 |
This two-source structure is well established in the ancient-DNA literature. Corded Ware communities formed on the North European Plain around 2900 BCE as Yamnaya-related migrants mixed with local Neolithic farming populations, and genomic studies consistently recover roughly a quarter to a third European farmer ancestry in them (Haak et al., 2015; Papac et al., 2021). Our two-source fit above reproduces that published range from an independent direction, which is a useful check on the reference setup before the target is modelled.
With the decomposition in hand, the target was modelled with each of three Corded Ware populations in the Steppe slot. All three fit, with closely similar proportions; the figure shows the Poland_CordedWare model.
| Steppe source | p | Steppe (%) | Iranian farmer-related (%) | AASI (%) |
|---|---|---|---|---|
| Poland_CordedWare (shown in figure) | 0.628 | 36.4 ± 3.8 | 35.9 ± 4.0 | 27.7 ± 1.5 |
| Germany_Esperstedt_CordedWare | 0.665 | 36.0 | 35.5 | 28.5 |
| Czechia_EBA_CordedWare | 0.320 | 35.8 | 36.2 | 28.0 |
In the four-source test for this window, the explicit European farmer term goes negative (−4.2 ± 3.7 percent) and the model is infeasible. This is the expected consequence of the Corded Ware decomposition above: the Steppe source now carries European farmer ancestry of its own, so adding a separate term double-counts it.
Observations for this window:
- The proportions are statistically unchanged from the Early Bronze Age window (Steppe 36%, farmer 36%, AASI 28%), even though the Steppe source is now an admixed European population. The models continue to see one Steppe signal.
- The European farmer ancestry inside Corded Ware does not inflate any component. With the farmer slot held at ShahTepe across all windows, the accounting stays clean; earlier versions of this analysis, which used an AASI-rich farmer source, showed apparent shifts between windows that were artifacts of the source composition, not of the target's history.
Late Bronze Age (1900-1200 BCE): Andronovo Culture
This window corresponds to the Sintashta–Andronovo horizon itself: the period in which Steppe_MLBA populations moved south through Central Asia, appearing in the Turan corridor between roughly 1800 and 1500 BCE; in South Asia itself, directly dated Steppe ancestry first appears in the Swat valley by 1200–800 BCE (Narasimhan et al., 2019). Unlike the earlier windows, the Steppe sources tested here are contemporaries of the migration, not stand-ins from an earlier era.
The Steppe source here is Andronovo (Steppe_MLBA) — the classic branches of Kazakhstan and the Yenisei, eastern kin of Sintashta (Chelyabinsk, the fortified-settlement culture at the horizon's root) and Srubnaya (Samara, its western sibling). All are, like Corded Ware before them, fusions of Yamnaya-related and European farmer ancestry (Allentoft et al., 2015). The figure shows the headline model: Andronovo with the distal farmer Iran_GanjDareh_N and the Onge.
| Source | Role | Weight (%) | Z |
|---|---|---|---|
| Andronovo (Steppe_MLBA) | Steppe | 42.8 ± 3.1 | 13.8 |
| Iran_GanjDareh_N | Iranian farmer-related | 33.2 ± 3.5 | 9.4 |
| ONG (Onge) | AASI proxy | 24.0 ± 1.8 | 13.6 |
The four-source test is decisive in this window. Adding an explicit European farmer term to the Bronze Age farmer models drives it to −10.1 ± 4.0 percent: Late Bronze Age Steppe populations such as Srubnaya and Sintashta carry substantial European farmer ancestry of their own (Wang et al., 2019), more than the target’s Steppe stream requires, so forcing an additional term overcorrects. The headline model above reads the same fact from the other side: with the Neolithic farmer (Iran_GanjDareh_N), the Steppe source itself must supply the European farmer ancestry, and Andronovo — carrying it in the right amount — is the source that fits.
Observations for this window:
- The headline model resolves this window: Steppe 42.8%, Iranian farmer-related 33.2%, AASI 24.0% (p = 0.060). The chronological alignment is the meaningful part: the Andronovo horizon (1900–1200 BCE) is contemporary with the migration window itself (Narasimhan et al., 2019), coinciding with the period proposed for Indo-European language spread into South Asia (Anthony, 2007).
- AASI ancestry stays at 24–29% across every configuration, continuous with the indigenous substrate documented from the Indus period onward (Shinde et al., 2019).
Investigating BMAC Influence
The Bactria-Margiana Archaeological Complex (BMAC), or Oxus civilization, sat directly on the route the Steppe stream travelled into South Asia, and its urban centres interacted with both Steppe and South Asian populations. Whether it contributed ancestry to Indo-Aryan groups is therefore a natural question, and one this analysis tests directly rather than inherits.
The test mirrors the European farmer design used in the Results: each Bronze Age window model (the supporting time-window analysis) is extended with a BMAC population as an explicit fourth source, against the same twelve-population reference set. Three BMAC representatives span the available samples: Gonur, the capital (Turkmenistan_BA1-1, 34 individuals at 0.01x coverage), Dzharkutan, a late steppe-admixed site (Uzbekistan_BA1-1, 9 individuals at 2.11x), and the Bustan/Sapalli group (Uzbekistan_BA, 31 individuals). If BMAC contributed a distinct stream, its coefficient should resolve positive; nine models (three windows, three sources) put that to the test on AADR v66.
| Window | BMAC source | p | BMAC estimate (%) | Z | Feasible |
|---|---|---|---|---|---|
| Early Bronze Age | Gonur | 0.321 | 6.4 ± 43.0 | 0.15 | yes |
| Early Bronze Age | Dzharkutan | 0.334 | 14.6 ± 42.9 | 0.34 | yes |
| Early Bronze Age | Sapalli | 0.404 | 30.6 ± 39.5 | 0.77 | yes |
| Middle Bronze Age | Gonur | 0.573 | 18.2 ± 34.5 | 0.53 | yes |
| Middle Bronze Age | Dzharkutan | 0.711 | 37.6 ± 39.0 | 0.96 | yes |
| Middle Bronze Age | Sapalli | 0.526 | 3.9 ± 38.5 | 0.10 | yes |
| Late Bronze Age | Gonur | 0.340 | −37.1 ± 39.1 | −0.95 | no |
| Late Bronze Age | Dzharkutan | 0.199 | 6.1 ± 72.8 | 0.08 | yes |
| Late Bronze Age | Sapalli | 0.454 | −72.3 ± 67.1 | −1.08 | no |
No model resolves a BMAC stream: every coefficient is statistically indistinguishable from zero (|Z| ≤ 1.1). The standard errors tell the more precise story. They run from ±34 to ±73 percentage points, ten to twenty times the ±3.4 of the European farmer term tested the same way, and when a BMAC source enters, the Iranian farmer coefficient's error inflates in step (±42 in the figure) while the Steppe and AASI estimates do not move at all. The model is telling us that BMAC ancestry and Iranian farmer-related ancestry are close to interchangeable from this genome's point of view; the two coefficients trade freely against each other, and the data cannot apportion ancestry between them.
The same holds for the headline model. Extending Andronovo + Iran_GanjDareh_N + Onge with each BMAC source in turn, the Steppe and AASI estimates stay put while the farmer and BMAC coefficients trade against each other: the farmer term collapses toward zero in every case (Z = 0.1–0.9), and with the steppe-admixed Dzharkutan the swap is nearly total (farmer 2.2 ± 15.1, BMAC 40.0 ± 19.0). BMAC substitutes for the Iranian farmer-related stream; it never resolves as a stream of its own.
The interchangeability is confirmed from the other direction: substituted directly into the farmer slot instead of added as a fourth source, BMAC populations fit well (twenty source combinations, p = 0.77–0.92, with proportions shifting only a few points). BMAC ancestry is, from South Asia's vantage, largely the same Central Asian farmer-related ancestry the model already carries. The contrast with the European farmer test is the internal control: the same design resolved that term crisply, so the failure to resolve a BMAC term reflects genuine genetic redundancy, not a weak instrument.
The conclusion matches the ancient-DNA record. Narasimhan et al. (2019) found that the main BMAC population contributed little ancestry to South Asian gene pools, and that gene flow ran detectably the other way: BMAC individuals carry 2–5% Andamanese hunter-gatherer (AHG)-related ancestry from South Asia, while a reciprocal BMAC signal in South Asians is undetectable. The Steppe stream crossed the Oxus world; genetically, it did not stay.
Model-Free Clustering (ADAMIXTURE)
The qpAdm models in this analysis all specify their sources. As a model-free cross-check, ADAMIXTURE clustering was run on AADR v66 at K = 5 to 12, letting the components emerge from the reference panel with no source assumptions. ADAMIXTURE represents each genome as a mixture of K components learned from the data; the components are statistical constructs of the panel, not ancient populations, and reading them well means watching how they behave as K changes rather than fixating on any single K.
The sweep tells a story in itself. At K = 5 the Iranian farmer and Steppe streams fuse into a single component carrying 48% of the genome. K = 8 splits them apart, and briefly resolves a fused EEF-plus-Steppe component (22.8%) whose shape is exactly the Corded Ware fusion documented in the Results: the clustering rediscovers, on its own, the population structure the supervised models specify. By K = 10 and 12 the streams under discussion are cleanly separated, at a cost: a formal prediction criterion prefers K = 5, and the growing unassigned slivers at high K are the over-splitting it warns about. Four panels spanning the sweep (K = 5, 8, 10, 12) are shown for that reason; no single K is privileged.
| Ancestry stream (rolled up) | K=5 | K=8 | K=10 | K=12 |
|---|---|---|---|---|
| Western Steppe Herder | — | 28.8 | 25.7 | 29.3 |
| Iranian-Neolithic / BMAC farmer | — | 20.2 | 17.2 | 19.1 |
| European farmer (EEF) affinity | 27.3 | 2.8 | 22.5 | 23.5 |
| Iranian + Steppe composite | 48.0 | — | 6.0 | — |
| EEF + Steppe composite | — | 22.8 | — | — |
| Deep South Eurasian composite (Papuan, Tianyuan, Ami anchors) | 11.1 | 10.6 | 10.7 | 10.7 |
| Deep African + Andamanese composite | 5.3 | 4.2 | 4.2 | 4.2 |
| ANE / Native American composite | 8.3 | 7.7 | 7.8 | 7.8 |
| unassigned | — | 2.7 | 6.0 | 5.4 |
First, the Steppe component stabilises at 26–29% once K separates it from the farmer streams, bracketing the Bronze Age farmer qpAdm estimates (33.5–36.4%) from below; on that like-for-like comparison the two methods agree that roughly a third of the genome traces to the Steppe. The headline distal frame reads higher (42.8%) because its Steppe term also carries the EEF share that clustering counts separately. Second, the EEF-related component is large (22–27%) at every K, once the K = 8 EEF-plus-Steppe composite is counted with it: this is the total Anatolian-farmer affinity discussed under Key Findings, carried inside the farmer and Steppe streams rather than arriving as its own migration. Third, and most diagnostic, no discrete AASI cluster ever forms at any K. The deep ancestry surfaces only as composites anchored by Papuan, Tianyuan, and Andamanese references, which is exactly what an unsampled lineage from a near-simultaneous three-way split of the deep eastern lineages into AASI, Andamanese, and East Asian branches (the eastern trifurcation) should look like when forced through a reference panel that does not contain it.
The trifurcation reading was tested directly rather than left as interpretation. In f4 tests of the form f4(Sohi, Mbuti; Andamanese, X), where Sohi is the label for the target genome, the genome is statistically symmetric between the Andamanese branch and the East Asian branches (Ami, Dai, and the 40,000-year-old Tianyuan genome; |Z| ≤ 0.5 in every comparison), with only the pre-split Ust-Ishim genome showing the weak asymmetry its earlier divergence predicts (Z = 2.1). The deep component is at the resolution floor of current references: no available population can subdivide it further.
Population context comes from the same toolkit. Compared against 87 unrelated Punjabi genomes from the 1000 Genomes Project with f4 tests of the form f4(Mbuti, Sintashta; Sohi, Punjabi), the target shares more drift with Sintashta than every one of the 87, significantly so for 77 of them (median Z = −5.2), while no Punjabi sample is significantly more Steppe-shifted than the target. A Jatt genome sitting at the high-Steppe end of the Punjabi distribution is what the community's agro-pastoralist history predicts, and it is the same placement the qpAdm proportions give in absolute terms.
Additional Genetic Insights
qpAdm measures autosomal mixture only. Two independent data types check it: the paternal lineage from Big Y-700, and a G25 fit computed with entirely different machinery.
Y-DNA Haplogroup Analysis via Big Y-700

The Big Y-700 test resolved the paternal lineage to R-FTF40903, a subclade of R1a-Z93 through L657. The lineage's placement can be inspected on both public phylogenies: the YFull tree and the FTDNA Discover haplotree. Its chronology, from published work:
- R1a-M417 expansion, c. 3500 BCE: the parent lineage expands with the early Steppe world and splits into a European branch (Z282, later carried by Corded Ware into Europe) and an Asian branch (Z93) (Underhill et al., 2015).
- Z93, c. 2900–2600 BCE: the Asian branch forms; ancient carriers appear in Sintashta and Andronovo contexts (Silva et al., 2017).
- L657, c. 2200 BCE: the largest South Asian subclade forms within Z93 > Z94, its date squarely inside the migration window (Silva et al., 2017).
- South Asian prevalence by the late 2nd millennium BCE: multiple closely related founder clades expand across the subcontinent, consistent with arrival through the 1800–1500 BCE corridor window the ancient-DNA record identifies.
One limit should be stated plainly: no sampled ancient individual has yet been confirmed to carry L657 itself; the ancient carriers sit on neighbouring Z93 branches, so the paternal evidence corroborates the horizon without pinning the transmitting population. The paternal line is in any case a co-witness rather than independent proof: a single lineage cannot measure ancestry proportions. Its value is that it names the same source world as the autosomal fits, through an entirely different inheritance system: Z93's sister branch travelled west with Corded Ware while Z93 itself went east and south with the Sintashta–Andronovo horizon, the two directions of one dispersal described under Key Findings.
IllustrativeDNA Results

The third cross-check uses IllustrativeDNA's G25 fit, a least-distance model in a 25-dimensional PCA space. Its two-way model gives:
- Andronovo Culture: 34.1%
- Indus Valley Civilization: 65.9%
This two-way model maps directly onto the Late Bronze Age qpAdm results, using an entirely different method and reference frame:
- The Andronovo component (34.1%) matches the Steppe estimate of the Late Bronze Age window under its Bronze Age farmer configuration (Andronovo 33.5 ± 3.7%, p = 0.280).
- The Indus Valley Civilization component (65.9%) matches the combined Iranian farmer-related and AASI ancestry of that configuration (39.5% + 27.0% = 66.5%), which is exactly what an IVC reference should absorb: the Indus population was itself a mixture of those two streams.
Two methods with different assumptions converging within a percentage point is meaningful corroboration. It also illustrates the reference-frame point made throughout this analysis: G25's "Indus Valley" component is not a fourth ancestry stream but a repackaging of two of the three streams the qpAdm models resolve separately.
A caution on where this cross-method agreement does and does not extend. G25 fits are constrained least-distance solutions in a 25-dimensional PCA space, and that machinery happily produces clean-looking percentages even from mutually collinear sources. Its deep "Neolithic ingredient" models (Zagros, EHG, CHG, EEF, and similar) have no valid qpAdm equivalent: run through f-statistics, those source sets are degenerate, returning negative or greater-than-100% weights with every coefficient statistically empty. We verified this directly on the present target. Proximal Bronze Age models built from real, genetically distinguishable populations, like the two-way fit above, are the level at which the two methods can check each other, and there they agree.
Conclusion
Modelled against AADR v66, with the Onge standing for the unsampled AASI lineage, the Jatt Sikh genome resolves into three ancestry streams: 42.8% Steppe, 33.2% Iranian farmer-related, and 24.0% AASI. The principal findings:
- The Steppe came through the Andronovo (Steppe_MLBA) horizon. In the headline model (42.8 ± 3.1% Steppe, p = 0.060) the farmer end is Neolithic, so the Steppe source itself must supply the genome’s European farmer ancestry — and only the Andronovo horizon does. The R1a-Z93 > L657 paternal lineage and corridor chronology point to the same horizon. The specific branch lies beyond the resolution of current samples.
- The four-source basal model splits that Steppe open. Fitted from the streams that existed before the Bronze Age began, the genome resolves as EHG 24.6% + European farmer (EEF) 13.4% + Iranian-related 35.7% + AASI 26.3% (p = 0.191). The arithmetic closes: Steppe_MLBA populations are roughly one-third European farmer, so 42.8% Andronovo predicts about 13% EEF — the basal model measures 13.4%.
- Three sources are sufficient. Explicit fourth terms fail in informative ways: the European farmer term is unresolvable or negative depending on what the other sources already carry, and no BMAC term can be resolved at all — BMAC ancestry substitutes for the Iranian farmer stream rather than adding to it.
- The indigenous substrate persists. AASI ancestry stands near a quarter of the genome in every model tested. The migrations added to South Asia’s genetic foundation; they did not replace it.
These results sit squarely within the current ancient-DNA picture. Steppe ancestry in South Asia runs highest in the northwest and declines southeastward, with genome-wide surveys of the Indian cline reporting Steppe proportions from near zero to ~45% (Kerdoncuff et al., 2025); a Steppe share in the low forties is what the cline’s northwestern end carries. The earliest Steppe ancestry directly dated in South Asian remains appears in the Swat valley by 1200–800 BCE, consistent with arrival in the window modelled here (Narasimhan et al., 2019). On the Steppe side, the record is deep and consistent: Yamnaya formation is now traced to the Caucasus–lower Volga cline (Lazaridis et al., 2025), the steppe's later population sequence is documented across 137 ancient genomes (Damgaard et al., 2018), and the Early Bronze Age expansions eastward are known to have left little genetic trace south of the steppe (de Barros Damgaard et al., 2018), the same asymmetry reproduced here: paired with the Neolithic farmer, Early Bronze Age Steppe sources fail outright, while the Middle-to-Late Bronze Age populations — the migration’s contemporaries — fit. The reference record is still growing, with over ten thousand newly reported West Eurasian ancient genomes in the latest release (Akbari et al., 2026).
What the genome records, it records plainly: an Iranian farmer-related and AASI foundation of the kind that built the Indus world, joined in the second millennium BCE by a Steppe stream whose paternal lineage still runs unbroken from the Bronze Age steppe to the Punjab. The clearest open question is branch-level attribution within the Sintashta–Andronovo horizon, and it will be settled by denser ancient sampling of the Central Asian corridor, not by further modelling of the samples that exist.
Supplementary Data
qpAdm Model Outputs
Complete outputs for the 146 qpAdm models behind the current figures, fitted on AADR v66 with the twelve-population reference set described under Methodology, plus the unsupervised clustering sweep:
- Farmer × period grid — 70 models: 7 farmer candidates × 10 Steppe sources, distributed 5/3/2 across the three windows. The source-selection evidence for the figures.
- Final window models — 12 models: the window three- and four-source fits, with all standard errors.
- BMAC four-source tests — 12 models: three windows × three BMAC sources plus baselines.
- Late Bronze Age farmer grid — 30 models across ten farmer candidates.
- European farmer succession tests — 20 models: the EEF term across farmer and Steppe regimes.
- ADAMIXTURE K-sweep — component loadings and stream roll-ups at K = 5, 8, 10, 12.
Raw per-model outputs (AADR v66)
One receipt file per published model, in the figures' exact regime: full weights with standard errors, the rank-drop table behind each p-value, and the pop-drop table of all nested submodels. Each file states its own dataset, target, and complete left and right population lists.
Headline models (shown in Key Findings):
- Proximal 3-way — Andronovo (Steppe_MLBA) + Iran_GanjDareh_N + Onge (42.8% Steppe, p = 0.060)
- Basal 4-way — EHG + EEF + Iranian-related + AASI, deepened reference set (p = 0.191)
Three-source window models:
- EBA — Samara Yamnaya (representative)
- EBA — Khakassia Afanasievo
- EBA — Altai Afanasievo
- EBA — Kalmykia Yamnaya
- EBA — Orenburg Yamnaya
- MBA — Poland Corded Ware (shown in figure)
- MBA — Czechia Corded Ware
- MBA — Esperstedt Corded Ware
- LBA — Samara Srubnaya
- LBA — Chelyabinsk Sintashta
Four-source European farmer tests:
Four-source BMAC tests:
- EBA + Gonur
- EBA + Dzharkutan
- EBA + Sapalli
- MBA + Gonur
- MBA + Dzharkutan
- MBA + Sapalli
- LBA + Gonur
- LBA + Dzharkutan
- LBA + Sapalli
Archived v54 raw outputs
The original 2024 analysis ran on AADR v54.1.p1 with per-chip targets and different source populations, including the AASI-rich farmer source whose composition effects are discussed in the Results. Its raw outputs are preserved unmodified:
- 23andMe - Russia Samara EBA Yamnaya
- AncestryDNA - Russia Samara EBA Yamnaya
- 23andMe - Czech Corded Ware
- AncestryDNA - Czech Corded Ware
- 23andMe - Russia Srubnaya Alakul
- AncestryDNA - Russia Srubnaya Alakul
- 23andMe - 3-way qpAdm Sintashta
- AncestryDNA - 3-way qpAdm Sintashta
- 23andMe - Russia MBA Poltavka
- 23andMe - Turkmenistan Gonur BA 1
- 23andMe - Mongolia EIA Slab Grave 1
- 23andMe - Kazakhstan Kumsay EBA
- 23andMe - BMAC
- AncestryDNA - BMAC
- 23andMe - 4-way qpAdm Gonur BA1
- AncestryDNA - 4-way qpAdm Gonur BA1
- 23andMe - 4-way qpAdm Gonur BA2
- AncestryDNA - 4-way qpAdm Gonur BA2
- 23andMe - 4-way qpAdm Geoksyur
- AncestryDNA - 4-way qpAdm Geoksyur
References
- Allen Ancient DNA Resource (AADR), version 66. Harvard Dataverse.
- Akbari, Ali, et al. "Ancient DNA reveals pervasive directional selection across West Eurasia." Nature 654 (2026): 419-428. Link to Article
- Allentoft, Morten E., et al. "Population genomics of Bronze Age Eurasia." Nature 522.7555 (2015): 167-172. Link to Article
- Anthony, David W. The Horse, the Wheel, and Language: How Bronze-Age Riders from the Eurasian Steppes Shaped the Modern World. Princeton University Press, 2007. Link to Book
- Basu, Analabha, et al. "Genomic reconstruction of the history of extant populations of India reveals five distinct ancestral components and a complex structure." Proceedings of the National Academy of Sciences 113.6 (2016): 1594-1599. Link to Article
- Broushaki, Farnaz, et al. "Early Neolithic genomes from the eastern Fertile Crescent." Science 353.6298 (2016): 499-503. Link to Article
- Damgaard, Peter de Barros, et al. "137 ancient human genomes from across the Eurasian steppes." Nature 557 (2018): 369-374. Link to Article
- de Barros Damgaard, Peter, et al. "The first horse herders and the impact of early Bronze Age steppe expansions into Asia." Science 360.6396 (2018): eaar7711. Link to Article
- Fournier, Romain, Alice Fulton, and David Reich. "A SNP panel for coanalysis of capture and shotgun ancient DNA data." Genome Research 36 (2026). Link to Article
- Haak, Wolfgang, et al. "Massive migration from the steppe was a source for Indo-European languages in Europe." Nature 522.7555 (2015): 207-211. Link to Article
- Harney, Éadaoin, et al. "Assessing the performance of qpAdm: a statistical tool for studying population admixture." Genetics 217.4 (2021): iyaa045. Link to Article
- Kerdoncuff, Élise, et al. "50,000 years of evolutionary history of India: Impact on health and disease variation." Cell 188 (2025). Link to Article
- Lazaridis, Iosif, et al. "The genetic origin of the Indo-Europeans." Nature 639 (2025): 132-142. Link to Article
- Maier, Robert, et al. "On the limits of fitting complex models of population history to f-statistics." eLife 12 (2023): e85492. Link to Article
- Narasimhan, Vagheesh M., et al. "The formation of human populations in South and Central Asia." Science 365.6457 (2019): eaat7487. Link to Article
- Papac, Luka, et al. "Dynamic changes in genomic and social structures in third-millennium BCE central Europe." Science Advances 7.35 (2021): eabi6941. Link to Article
- Reich, David, et al. "Reconstructing Indian population history." Nature 461 (2009): 489-494. Link to Article
- Shinde, Vasant, et al. "An ancient Harappan genome lacks ancestry from Steppe pastoralists or Iranian farmers." Cell 179.3 (2019): 729-735. Link to Article
- Silva, Marina, et al. "A genetic chronology for the Indian Subcontinent points to heavily sex-biased dispersals." BMC Evolutionary Biology 17 (2017): 88. Link to Article
- Underhill, Peter A., et al. "The phylogenetic and geographic structure of Y-chromosome haplogroup R1a." European Journal of Human Genetics 23 (2015): 124-131. Link to Article
- Wang, Chuan-Chao, et al. "Ancient human genome-wide data from a 3000-year interval in the Caucasus corresponds with eco-geographic regions." Nature Communications 10 (2019): 590. Link to Article