Explorerβ€ΊRoboticsβ€ΊRobotics
Research PaperResearchia:202609.23067

Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation

John Church

Abstract

We present an improvement on previous spacecraft pose estimation architectures that results in the lowest published mean rotation errors we know of on the SPEED+ lightbox and sunlamp test sets for a known, non-cooperative spacecraft. By using a previously established heatmap-based pose estimation architecture and adapting a large self-supervised ViT foundation model (DINOv3) in place of the smaller convolutional and ViT encoders of previous work, we show that pose estimation accuracy improves fr...

Submitted: September 23, 2026Subjects: Robotics; Robotics

Description / Details

We present an improvement on previous spacecraft pose estimation architectures that results in the lowest published mean rotation errors we know of on the SPEED+ lightbox and sunlamp test sets for a known, non-cooperative spacecraft. By using a previously established heatmap-based pose estimation architecture and adapting a large self-supervised ViT foundation model (DINOv3) in place of the smaller convolutional and ViT encoders of previous work, we show that pose estimation accuracy improves from 300M to 840M parameters with no saturation yet observed. We also evaluate our 840M model on a Jetson Orin NX 16GB, measuring single-pass network inference at 133.8 ms per crop with a board draw of 32.0 W. These measurements demonstrate embedded inference feasibility on a processor family with orbital flight heritage. Our resulting model outperforms previous models across lightbox and sunlamp domains while training only on synthetic data. Our best model, using DINOv3 840M adapted with LoRA as the encoder (rank 64, three-seed ensemble with four-rotation test-time augmentation), results in 1.56∘1.56^\circ mean rotation error on sunlamp and 1.17∘1.17^\circ on lightbox, compared to the previous best mean rotation errors we know of on these test sets, 2.66∘2.66^\circ and 1.75∘1.75^\circ by EagerNet.


Source: arXiv:2609.26561v1 - http://arxiv.org/abs/2609.26561v1 PDF: https://arxiv.org/pdf/2609.26561v1 Original Link: http://arxiv.org/abs/2609.26561v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 23, 2026
Topic:
Robotics
Area:
Robotics
Comments:
0
Bookmark
Vision Foundation Models with Synthetic-Only Training for Monocular Spacecraft Pose Estimation | Researchia