Explorerโ€บArtificial Intelligenceโ€บAI
Research PaperResearchia:202609.22058

Rare Event Estimation via Iterative Unalignment

Hanming Yang

Abstract

As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Mont...

Submitted: September 22, 2026Subjects: AI; Artificial Intelligence

Description / Details

As agents are deployed with increased autonomy, even extremely rare events along their stochastic output trajectories can occur and prove catastrophic. Safe deployment therefore does not depend on whether these events can occur, but on how often they might. We study the problem of estimating the probability of rare events that arise from stochastic variation in the agent's own actions. Estimating this type of risk requires searching over the combinatorially vast space of trajectories. Naive Monte Carlo is computationally prohibitive in this regime, and constructing effective importance sampling (IS) proposals requires coordinated changes to a context-dependent chain of conditional distributions. We develop a new IS method that perturbs the original model's weights to construct the proposal. The proposal is itself a differentiably parameterized language model, enabling gradient-based search over weight space. We formulate an objective that combines a differentiable surrogate for event amplification and an adaptive regularization scheme that dynamically balances amplification against estimator stability. We evaluate our approach on โˆผ\sim120M and โˆผ\sim2.6B models across three event families spanning 300+ rare events as rare as 10โˆ’910^{-9}, with reference probabilities computed with <10%<10\% relative standard error. In our most verifiable settings, we observe that our IS estimator achieves over 800ร—800\times compute-weighted efficiency gains over naive Monte Carlo for events with probabilities lower than 10โˆ’710^{-7}. Our implementation is available at https://github.com/namkoong-lab/iterative-unalignment.


Source: arXiv:2609.24969v1 - http://arxiv.org/abs/2609.24969v1 PDF: https://arxiv.org/pdf/2609.24969v1 Original Link: http://arxiv.org/abs/2609.24969v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 22, 2026
Topic:
Artificial Intelligence
Area:
AI
Comments:
0
Bookmark
Rare Event Estimation via Iterative Unalignment | Researchia