Next-token functional estimation
Milind Nakul, Vidya Muthukumar, Ashwin Pananjady
Source abstract
Suppose we observe the first points of a sequence of random variables having length , and wish to estimate a functional of the unobserved final point and the empirical measure of the observed training points. Such next-token functionals include the probability that the next token is novel (also known as the surprise probability), the tail probability of the minimum distance between the next token and training points, and the test error of a classifier trained on the observed points. All of these quantities are classically estimated by the leave-one-out method, which is inconsistent under temporal dependence. We propose a leave-a-window-out estimator, which deletes a window of length after each index before forming the empirical measure and reduces to leave-one-out at . Under natural assumptions, we show that the error of our estimator decays at a parametric rate for any stationary -mixing process that also admits a Marton coupling. Our results thus cover several natural functionals on a large class of stochastic processes. We complement these upper bounds with a sharp minimax lower bound for estimating the surprise probability on mixing Markov chains. Simulations on Markov chains, moving-average processes, and autoregressive processes show that our estimator succeeds in many scenarios where leave-one-out and add-constant baselines fail.
Evidence graph
No public relationships recorded yet.
Integrity note: This page is a factual metadata record created by deterministic ingestion. It is not a claim that the work moves a mathematical frontier or has been independently verified.