Taxonomic Classification with Complete Tag Arrays
Abstract
Taxonomic classifiers such as Kraken assign each $k$-mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more $k$-mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken~2, but its index is large and expensi...
Description / Details
Taxonomic classifiers such as Kraken assign each -mer of a reference database to the lowest common ancestor (LCA) of the genomes containing it, but this works less well as databases grow, because more and more -mers are shared across species. Cliffy (Ahmed, Boucher and Langmead, 2025) instead uses variable-length exact matches found with an r-index, and can list approximately the genera containing each match; on 16S rRNA it is more accurate than Kraken2, but its index is large and expensive to build. We present KATKA, which finds the maximal exact matches (MEMs) of at least a given length in each read with Boyer--Moore--Li on a run-length compressed suffix array, counts the occurrences of each MEM in each genus exactly with a complete, run-length compressed tag array, and gives each genus credit in proportion to those counts. On the SILVA 16S rRNA database, KATKA's default index takes 1.44,GB and can be built in minutes on a desktop computer; it classifies a read in 66,s with one thread and reaches 93.8% genus-level accuracy, close to what Cliffy reports for its 9,GB index. Grammar-compressing the runs of the tag array shrinks the index to 1.04,GB, at 75,s per read. On the same machine and reads, it is more accurate than Kraken2 (79.3%) and Tagger (81.7 to 92.8%, depending on how mates that disagree are scored). Indexing minimizer digests instead of the sequences makes the index three times smaller and classification 1.7 times faster, at a cost of 1.3 points of accuracy. KATKA is available at https://github.com/TravisGagie/KATKA.
Source: arXiv:2610.10500v1 - http://arxiv.org/abs/2610.10500v1 PDF: https://arxiv.org/pdf/2610.10500v1 Original Link: http://arxiv.org/abs/2610.10500v1
Please sign in to join the discussion.
No comments yet. Be the first to share your thoughts!
Oct 8, 2026
Biotechnology
Biology
0