Swaveda

Genes Aren't Grammar: What Ancient DNA Can and Cannot Say About Vedic Culture

Ancient DNA has revolutionized what we know about who moved into South Asia and when — but it cannot tell us who composed the Rigveda, what language they spoke, or what rituals they practiced.

Vikram Joshi for SwavedaAugust 1, 2026

Want to correct something or add a source? Sign in below to contribute.

A genome doesn't recite verses. It doesn't perform a fire ritual. It doesn't decline a noun. That sounds obvious, but it's the exact point that gets lost every time a new ancient-DNA paper on South Asia goes viral and both sides of the "Aryan migration" argument mine it for vindication.

The past decade has produced genuinely landmark science. A 2019 study led by geneticist Vagheesh Narasimhan and colleagues, including David Reich, analyzed ancient genomes from Central and South Asia and found that populations related to Steppe pastoralists — herders from the grasslands north of the Black and Caspian Seas — mixed into South Asian gene pools sometime after the decline of the Indus Valley Civilization, with genetic signals consistent with movement in the second millennium BCE. A separate 2019 paper by archaeologist Vasant Shinde and colleagues sequenced DNA from a skeleton at Rakhigarhi, one of the largest Indus Valley sites, and found no detectable Steppe ancestry or Iranian farmer-related ancestry in that individual — pointing instead to a lineage already present in South Asia before those later admixture events. An earlier 2009 paper by Reich and coauthors in Nature had already established that most present-day Indian populations descend from mixtures of two ancestral groups, commonly labeled Ancestral North Indian and Ancestral South Indian.

These are real, peer-reviewed, replicated findings. What they are not is a verdict on the Rigveda (the oldest Vedic text, composed in early Sanskrit), on who first spoke Sanskrit, or on the origins of Vedic ritual practice.

Ancestry Is Not a Language Tag

DNA records descent — who mated with whom, and when. Language is learned, not inherited. A community can absorb new ancestry over generations while keeping its language, or shift its language entirely while its gene pool barely changes. Ireland speaks English today with almost no English ancestry replacement; large parts of Latin America speak Spanish with minimal Iberian ancestry. Genetic continuity and linguistic continuity are separate variables, and conflating them is a category error, not a discovery.

The Narasimhan et al. paper itself is explicit about this distinction, noting that the genetic data can identify population movements and mixture events but cannot independently establish which languages those populations spoke. That caveat tends to disappear the moment a headline turns "Steppe ancestry entered South Asia around the second millennium BCE" into "Aryans invaded India" or, in the opposite direction, "no Steppe ancestry at Rakhigarhi" into "the Rigveda was composed entirely within the Indus Valley Civilization." Neither leap is supported by the underlying paper.

What Rakhigarhi Does and Doesn't Settle

The Rakhigarhi genome is a single individual from one site at one point in time. It tells us that this person's ancestry, at that moment, lacked the Steppe and Iranian-farmer-related components later found widely across South Asia. It does not tell us what language that person spoke, whether Indus Valley cities practiced anything resembling Vedic fire ritual, or whether the broader Indus Valley population was linguistically uniform at all — the Indus script itself remains undeciphered, so we have no direct textual evidence of the language spoken at Rakhigarhi or any other Harappan site.

Meanwhile, the dating of the Rigveda rests on internal linguistic evidence — meter, vocabulary layers, references to rivers and geography — evaluated by philologists working with the text itself, not on ancient DNA. Genetic data and textual dating are produced by entirely different methods answering entirely different questions. A genome can't confirm or refute a philological argument about verb forms, and a verse can't confirm or refute an admixture date estimated from genome-wide statistics.

Two Kinds of Overreach

The confusion runs in both directions. One version treats any detection of Steppe-related ancestry as proof of a violent "Aryan invasion" wiping out an indigenous civilization — a framing the ancient-DNA papers themselves don't support, since they describe gradual admixture over generations, not a single conquest event. The other version treats the absence of Steppe ancestry at a single Harappan site as proof that Vedic culture is purely indigenous to the Indus Valley, collapsing archaeology, genetics, and philology into one triumphant story. Both moves use the same trick: borrowing the authority of a hard science to settle a question that science wasn't designed to answer.

What the genetic record actually supports is narrower and more useful: population movement and mixture across South Asia and Central Asia over several millennia, documented with increasing precision as more ancient genomes are sequenced. What language those populations spoke, what they believed, and what they wrote down are questions for linguists, epigraphists, and textual historians — working, as they always have, from grammar, meter, and inscription, not from a genome.

Contribute to this article

Spotted an error, want to add a source, or have a correction? Sign in to send a contribution. Submissions are evaluated for factual accuracy before they change the article. Contrarian views are preserved publicly as reader notes.