Systematic benchmark of ancient DNA read mapping

Adrien Oliva, Raymond Tobler, Alan Cooper, Bastien Llamas, Yassine Souilmi

    Research output: Contribution to journalArticle


    The current standard practice for assembling individual genomes involves mapping millions of short DNA sequences (also known as DNA 'reads') against a pre-constructed reference genome. Mapping vast amounts of short reads in a timely manner is a computationally challenging task that inevitably produces artefacts, including biases against alleles not found in the reference genome. This reference bias and other mapping artefacts are expected to be exacerbated in ancient DNA (aDNA) studies, which rely on the analysis of low quantities of damaged and very short DNA fragments (~30-80 bp). Nevertheless, the current gold-standard mapping strategies for aDNA studies have effectively remained unchanged for nearly a decade, during which time new software has emerged. In this study, we used simulated aDNA reads from three different human populations to benchmark the performance of 30 distinct mapping strategies implemented across four different read mapping software - BWA-aln, BWA-mem, NovoAlign and Bowtie2 - and quantified the impact of reference bias in downstream population genetic analyses. We show that specific NovoAlign, BWA-aln and BWA-mem parameterizations achieve high mapping precision with low levels of reference bias, particularly after filtering out reads with low mapping qualities. However, unbiased NovoAlign results required the use of an IUPAC reference genome. While relevant only to aDNA projects where reference population data are available, the benefit of using an IUPAC reference demonstrates the value of incorporating population genetic information into the aDNA mapping process, echoing recent results based on graph genome representations.
    Original languageEnglish
    Pages (from-to)1-12
    JournalBriefings in Bioinformatics
    Issue number5
    Publication statusPublished - 2021


    Dive into the research topics of 'Systematic benchmark of ancient DNA read mapping'. Together they form a unique fingerprint.

    Cite this