Explorer›Biotechnology›Biology
Research PaperResearchia:202610.08017

Taxonomic Classification with Complete Tag Arrays

Travis Gagie

Abstract

Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensi...

Submitted: October 8, 2026Subjects: Biology; Biotechnology

Description / Details

Taxonomic classifiers such as Kraken assign each kk-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more kk-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44,GB and can be built in minutes on a desktop computer; it classifies a read in 66,μμs with one thread and reaches 93.8% genus-level accuracy, close to what Cliffy reports for its 9,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04,GB, at 75,μμs per read. On the same machine and reads, it is more accurate than Kraken2 (79.3%) and Tagger (81.7 to 92.8%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.


Source: arXiv:2610.10500v1 - http://arxiv.org/abs/2610.10500v1 PDF: https://arxiv.org/pdf/2610.10500v1 Original Link: http://arxiv.org/abs/2610.10500v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Oct 8, 2026
Topic:
Biotechnology
Area:
Biology
Comments:
0
Bookmark