What does the journey look like from the DNA of a human who lived 5,000 years ago to the story of the migration of an entire ancient population?
It begins in the laboratory. DNA is most commonly extracted from the petrous bone or teeth. It is then sequenced, and the resulting DNA fragments are aligned to the human reference genome.
The result is a massive data matrix consisting of individuals and their genetic variants. Because these are extremely high-dimensional data, we use visualization methods that allow us to identify population structure and genetic similarities.
But visualization alone does not show population history. For that, we need models of admixture and migration. Longer stretches of DNA inherited from common ancestors are also important.
Ultimately, we combine genetic evidence with archaeological and chronological information. Keep in mind, however: we are reconstructing the history represented by the samples we obtained, not the history of an entire ancient population.
Today, we can read the past from DNA. But how easy is it to read it correctly?
It is surprisingly difficult. Certain methods used to study ancient human DNA can produce results that seem convincing at first glance, yet are actually incorrect.
This is well illustrated by one of our simulations published in the journal Genetics. In the simulated scenario, a later population inherited genetic information from an older one. However, a commonly used analysis workflow incorrectly modelled the older population as a mixture involving ancestry represented by that later population. It produced a plausible-looking result, but the ancestry flow underlying it went in the opposite direction.
What is even more misleading is that two other commonly used methods appeared to support this incorrect scenario. Agreement between three different approaches therefore created a false sense of security. If a model passes a statistical test, it does not automatically mean it is correct; it merely means it has not been rejected.
How do you then verify that your methods really work?
We cannot compare the result with what actually happened thousands of years ago. We create a realistic environment in which we can test how genetic methods behave.
In our simulations, for instance, small populations live in a given region, exchange migrants every generation, and can be separated by barriers of varying permeability. Occasionally, long-distance migration events can also occur between them.
The advantage is that for genetic data generated this way, we know the exact history behind their creation. Thus, we can test whether the methods in use can actually reconstruct it.
Your newer preprint shows that misleading impressions can also arise at the visualization stage. How did you arrive at that conclusion?
Using spatial simulations, we found that sparse sampling of individuals combined with the loss of rare genetic variants can distort principal component plots, the most common visualization method. They can give rise to, for example, triangles, Y-shaped clusters, or artificial outliers.
An observer might interpret such a triangle as evidence for the mixture of three distinct ancestral populations. In reality, however, it may arise simply as a consequence of how the data were sampled and subsequently compressed into two or three dimensions.
That is why we are trying to develop better ways of visualization. Even so, a convincing picture is only the beginning of an argument, not its conclusion.
One of the questions you investigated is the origin of the people of the Yamnaya culture. What did you find?
Ancient-DNA research had already shown that about five thousand years ago, people genetically related to Yamnaya pastoralists expanded from the Pontic–Caspian steppe both westward into Europe and eastward across Eurasia. This migration profoundly transformed the genetic landscape of Europe and was very likely associated with the spread of at least some Indo-European languages.
We also showed that about four-fifths of Yamnaya ancestry can be traced to populations living between the Caucasus and the Lower Volga. These expanded westward and mixed with local hunter-gatherers in the Dnipro and Don regions.
How do these genetic findings relate to the history and spread of early Indo-European languages?
While genetic data cannot directly prove which language specific individuals spoke, they offer important clues. We detected ancestry related to these Caucasus–Lower Volga populations in Bronze Age Anatolia, a region where Hittite—the earliest documented Indo-European language—was spoken. This creates a plausible genetic connection between the steppe and the Anatolian branch of Indo-European languages that had previously been difficult to explain.
What first drew you to archaeogenetics, and what interests you about it today?
I was always interested in prehistory and the effort to reconstruct the history and relationships of languages through the history of human populations.
Over time, however, I became more skeptical toward some of the standard methods used in our field. Today, I am attracted to the applied mathematics behind method development. The details of human migrations are of course still fascinating, but I am increasingly interested in the question of how to ensure that the tools we use to study the past actually work.
What has surprised you most in your research into human history?
Far more than any specific research finding, I have been genuinely surprised by certain research practices. I did not expect to see so many papers in top journals in our field relying on questionable assumptions or on methods that have never been properly validated.
That is why, as part of our project, we developed best-practice approaches and general principles for testing simple admixture models. We want to contribute to making the methods used in archaeogenetics more reliable.
You received the GACR President’s Award for this project. What does the award mean to you?
It is the first national award of my career, so I was both surprised and delighted to receive it. I also view it as recognition for the entire field of archaeogenetics in the Czech Republic.
As for our project, what I value most is our perseverance and attention to detail—both statistical and archaeological. I am fortunate to have people on my team who see it similarly and for whom quality is more important than speed.
You had an offer to stay at Harvard, yet you chose to return to Ostrava. What do the city and the University of Ostrava mean to you?
I am happy in Ostrava and would like to stay here. At this point, I have a research group here, my children are going to school here, and I’m rather reluctant to uproot everything and move somewhere else. There are certainly more vibrant research environments elsewhere now. But, to be honest, these days I often get more insights from talking to frontier AI models than from talking to another researcher.
What would you like to build here in the coming years?
If we are talking about a larger dream, I would like to build a framework that helps us manage the exponential growth of data in archaeogenetics.
We have more and more data, but we do not yet have enough methods, people, or infrastructure to make this growth sustainable long-term. And the problem is not just in the genomes themselves; large databases contain errors or lack archaeological metadata. Therefore, I think one of the major tasks of archaeogenetics will be to learn not only how to process this massive volume of genetic and archaeological data, but above all how to understand it correctly.



