Explorerβ€ΊBiotechnologyβ€ΊBiology
Research PaperResearchia:202609.30017

CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts

Guang Yang

Abstract

Genome foundation models are growing into sparse mixture-of-experts (MoE) networks whose expert weights no longer fit on the machines that hold the sequences, yet sending a private genome to rented accelerators exposes it: we show that a single server hosting one expert recovers the input nucleotides with 99.8% top-1 accuracy. We present CipherGenome, a protocol that keeps the embedding, attention and router of a 15.1B-parameter MoE genome model on a trusted thin client and outsources every expe...

Submitted: September 30, 2026Subjects: Biology; Biotechnology

Description / Details

Genome foundation models are growing into sparse mixture-of-experts (MoE) networks whose expert weights no longer fit on the machines that hold the sequences, yet sending a private genome to rented accelerators exposes it: we show that a single server hosting one expert recovers the input nucleotides with 99.8% top-1 accuracy. We present CipherGenome, a protocol that keeps the embedding, attention and router of a 15.1B-parameter MoE genome model on a trusted thin client and outsources every expert projection, 95.8% of the parameters, to untrusted and possibly colluding GPU servers under module-LWE encryption. The design exploits three structural facts: expert layers are linear between two SwiGLU gates, expert weights are public, and GPU integer tensor cores can evaluate a ciphertext-weight product exactly modulo 2482^{48} in a single GEMM. The client evaluates the nonlinearity exactly and re-encrypts with fresh secrets, so no polynomial approximation or bootstrapping is ever needed. On 72 windows from 12 bacterial genomes, encryption adds 2.54Γ—10βˆ’42.54 \times 10^{-4} nats per token of KL divergence (95% CI upper bound 3.95Γ—10βˆ’43.95 \times 10^{-4}), below a pre-registered non-inferiority margin and indistinguishable from bf16 inference, while the same inversion attack falls to chance level. A reusable public hint cuts end-to-end latency by 3.54 times, wire compression reduces traffic 6.8 times, per-layer padding reduces routing leakage from 54.9% to 8.9% accuracy, and HE-compatible int4 experts remain non-inferior to their plaintext counterparts. Per expert and token, the server-side cost is more than six orders of magnitude below a CKKS baseline.


Source: arXiv:2609.35883v1 - http://arxiv.org/abs/2609.35883v1 PDF: https://arxiv.org/pdf/2609.35883v1 Original Link: http://arxiv.org/abs/2609.35883v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 30, 2026
Topic:
Biotechnology
Area:
Biology
Comments:
0
Bookmark
CipherGenome: Homomorphic Inference for Genomic Mixture-of-Experts | Researchia