The massive global datasets created by Next-Generation Sequencing (NGS), such as the American SRA and European ENA, contain around 100 petabytes of DNA/RNA data, a volume equivalent to all text on the internet. Traditionally, searching this “mountain of data” has been slow, expensive, and resource-intensive, often requiring researchers to download entire datasets.
Computer scientists at ETH Zurich have solved this problem by developing a tool called MetaGraph, which functions as a “Google for DNA.” Published in the journal Nature, the method allows researchers to enter a sequence as a full-text query and find where it has appeared in the raw data within seconds or minutes.
MetaGraph achieves this efficiency through a complex mathematical process that compresses the data by a factor of 300 and stores it in compressed form using linked graphs. This index-based system makes searching highly precise, efficient, and cost-effective.
The tool’s scalability means it requires less additional computing power as the data volume grows, making it an ideal catalyst for accelerating genetic research. MetaGraph can help rapidly identify little-researched pathogens, track pandemic spread, and identify antibiotic resistance genes or therapeutic bacteriophages in the databases. The open-source tool is already available and currently indexes nearly half of the world’s sequence data.
Image credit: Nature (2025). DOI: 10.1038/s41586-025-09603-w (Medical Xpress)
By Andres Eberhard, ETH Zurich





