Explorerโ€บComputer Visionโ€บComputer Vision
Research PaperResearchia:202610.03002

Sphere Encoder 2

Kaiyu Yue

Abstract

Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to a...

Submitted: October 3, 2026Subjects: Computer Vision; Computer Vision

Description / Details

Sphere Encoder is an autoencoder that generates images by decoding random points from a high-dimensional latent sphere. We identify two limitations of the original formulation that reduce its generation quality. First, random points concentrate near the equator relative to the pole on an encoded latent, but the training rotation never reaches this region, leaving a gap that limits one-step generation. Second, training for generation with pixel-wise reconstruction loss encourages the decoder to average over plausible images, producing blurry images that lack high-frequency details. We present Sphere Encoder 2 to address both limitations, substantially improving image generation quality while maintaining the speed and simplicity of a autoencoder. Models are released at \href{https://github.com/kaiyuyue/sphere2}{github.com/kaiyuyue/sphere2}.


Source: arXiv:2610.02208v1 - http://arxiv.org/abs/2610.02208v1 PDF: https://arxiv.org/pdf/2610.02208v1 Original Link: http://arxiv.org/abs/2610.02208v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Oct 3, 2026
Topic:
Computer Vision
Area:
Computer Vision
Comments:
0
Bookmark