AI Insight
This paper introduces SmoothFBO, the first algorithm designed to handle functional bilevel optimization in non-stationary online settings where the environment changes over time. The algorithm uses a time-smoothed stochastic hypergradient estimator with a window parameter to reduce variance and achieve stable updates with sublinear regret guarantees. The method generalizes classical parametric bilevel optimization and demonstrates superior performance compared to existing methods in hyperparameter optimization and model-based reinforcement learning tasks.
Why it matters
This advancement enables more robust machine learning in real-world scenarios where data distributions and objectives evolve over time, such as adaptive hyperparameter tuning and dynamic reinforcement learning environments. The theoretical guarantees combined with practical scalability make this approach applicable to a wide range of hierarchical learning problems that require continuous adaptation.
Understand the Science
⚠️ Preprint – Noch nicht peer-reviewed
Dieser Artikel wurde noch nicht von unabhängigen Experten begutachtet. Die Ergebnisse sind vorläufig und sollten mit Vorsicht interpretiert werden.
-cross
Abstract: Functional bilevel optimization (FBO) provides a powerful framework for hierarchical learning in function spaces, yet current methods are limited to static offline settings and perform suboptimally in online, non-stationary scenarios. We propose SmoothFBO, the first algorithm for non-stationary FBO with both theoretical guarantees and practical scalability. SmoothFBO introduces a time-smoothed stochastic hypergradient estimator that reduces variance through a window parameter, enabling stable outer-loop updates with sublinear regret. Importantly, the classical parametric bilevel case is a special reduction of our framework, making SmoothFBO a natural extension to online, non-stationary settings. Empirically, SmoothFBO consistently outperforms existing FBO methods in non-stationary hyperparameter optimization and model-based reinforcement learning, demonstrating its practical effectiveness. Together, these results establish SmoothFBO as a general, theoretically grounded, and practically viable foundation for bilevel optimization in online, non-stationary scenarios.