We study local stationary solutions of finite-horizon discrete-time Pontryagin systems near a steady extremal. Suppose that the stationarity equation for the control is regular, the reduced state--costate map is hyperbolic, and the endpoint conditions satisfy a scaled transversality condition with respect to the stable and unstable subspaces. Then the linearized boundary-value problem admits an inverse whose Green estimate is uniform in the horizon. The Green kernel separates interior decay from the two reflections induced by the endpoint conditions. For $x_0=x_{\rm in}$ and $p_T=r_x(x_T,y)$, a contraction argument in a weighted norm proves existence and uniqueness in a neighborhood independent of $T$, together with uniform Lipschitz estimates and a pointwise quadratic remainder. We also derive an explicit admissible data radius and an a posteriori criterion for existence and local uniqueness near an approximate trajectory. For these graph boundary conditions, a one-sided Green estimate shows that a perturbation of the terminal reward changes the initial control and the gradient with respect to the initial state of the stationary objective by $O(e^{-\alpha_{\rm ter} T})$ for every $\alpha_{\rm ter}$ below the dichotomy rate. For linear-quadratic systems with invertible $A$, stabilizable $(A,B)$, $Q\succ0$, $R\succ0$, and a nonpositive terminal Hessian, a symplectic graph condition verifies the assumptions, and the finite-horizon Riccati matrix and initial feedback gain converge at rate $O(e^{-2\gamma T})$. Numerical experiments verify the certificates and the predicted decay rates.
This paper studies the existence of relaxed equilibria for finite-horizon continuous-time time-inconsistent mean field games. We work directly on the product space of relaxed feedback policies and population flows. The policy component is endowed with the stable topology of Young measures, while the population component is restricted to a compact convex set of Wasserstein-continuous flows with uniform moment and time-regularity bounds. For every policy-flow pair, we establish uniform Sobolev and H\"older estimates for the associated auxiliary value function and prove its stability under Young-measure convergence of policies and uniform Wasserstein convergence of population flows. We also establish continuity of the induced population-flow map. The latter requires a duality argument for the Fokker-Planck equations because Young-measure convergence yields only weak-$*$ convergence of the controlled drifts. We then construct a set-valued best-response/consistency map with nonempty compact convex values and apply the Kakutani-Fan-Glicksberg fixed-point theorem. The resulting fixed point satisfies both the equilibrium response condition of the intra-personal game and the mean field consistency condition, , thereby establishing the existence of a relaxed equilibrium
We study infinite-horizon time-inconsistent Markov decision processes with a countably infinite state space and unbounded reward functions. The reward is allowed to depend explicitly on the initial time and initial state, thereby accommodating general sources of time inconsistency. We seek relaxed feedback equilibria, and our approach is based on entropy regularization and weighted functional analytic methods. With entropy regularization, we characterize a regular relaxed equilibrium through a fixed-point operator. By introducing two weight functions with distinct roles, one controlling the growth of rewards and values and the other defining the ambient weighted space, we construct a compact invariant set under a product topology and apply the Schauder-Tychonoff fixed-point theorem to establish existence of regularized equilibria. Importantly, the invariant set can be chosen uniformly for small entropy weight $\lambda\in(0,1]$. We then let $\lambda\to0+$ and show, through compactness, concentration of Gibbs policies, and uniform-integrability arguments, that a subsequential limit is a relaxed equilibrium of the original unregularized problem. We further study a policy iteration algorithm (PIA) for the entropy-regularized equilibrium problem. Under a weighted-discounting structure and sufficiently strong discounting, we establish exponential convergence and uniqueness of the regularized equilibrium in a suitable weighted Banach space. Combining the policy-iteration error with a quantitative soft-max approximation bound, we show that the iterated policies constitute weighted $\varepsilon$-equilibria for the original unregularized problem and derive an explicit regret estimate. A numerical example illustrating the convergence of PIA under strong discounting and a counterexample demonstrating its failure under weak discounting are also provided.
We study contraction properties of non-stationary continuous-time mean-field games (MFGs) under discounting and entropy regularization. The state of the representative agent evolves according to a controlled continuous-time Markov chain, and both the state and action spaces are finite. In contrast to the undiscounted case, we show that, under a sufficiently large discount rate, finite-horizon MFGs admit a horizon-independent contraction condition, which also coincides with the corresponding infinite-horizon non-stationary contraction condition. As a byproduct, we obtain an explicit convergence rate between finite- and infinite-horizon mean-field equilibria. For each finite horizon, we further derive a refined contraction criterion from the spectral radius of a positive operator that majorizes the propagation of policy errors, and show that its large-horizon limit agrees with the horizon-independent contraction factor. Finally, we provide an explicit error bound between discounted and undiscounted finite-horizon regularized equilibria.
In this paper, we propose and study a class of differential stochastic variational inequalities (DSVIs), in which an ordinary differential equation (ODE) is coupled with history-dependent stochastic variational inequalities (SVI). This framework models closed-loop stochastic systems with time-varying random equilibria and includes optimization-constrained ODEs as special cases. Under appropriate technical conditions, we establish uniqueness, measurability, and Lipschitz continuity with respect to the state of the second-stage response, and consequently the existence and uniqueness of the induced state trajectory. Moreover, we construct a sample average approximation (SAA) based on independent sample paths and prove uniform convergence of the approximate trajectories. For transfer between related stochastic environments, we derive a local $1/2$-H\"older estimate for parametric variational inequalities with moving feasible sets and a quantitative trajectory-stability bound in terms of the initial-state difference and the Wasserstein distance between exogenous path laws. Numerical experiments illustrate the SAA convergence and transfer-stability results. We further apply the framework to an elderly-health monitoring system. Similarity-weighted reuse of precomputed responses achieves an accuracy close to the full-recomputation benchmark of 0.97, while reducing the online batch runtime from 86 seconds to less than one second. Perturbation and delayed-update experiments additionally characterize robustness to sensor noise and the trade-off between response freshness, predictive accuracy, and computational cost. These results provide theoretical and computational support for efficient transfer learning in history-dependent DSVI systems.
Xiaojun Chen, Jian Guo, Xin Guo et al.· 0 citations
We revisit the problem of computing mean-field equilibria (MFEs) in discrete-time, monotone, finite-horizon mean-field games (MFGs). We show that, when the transition kernel is independent of the state-measure term and the reward function satisfies the usual weak monotonicity condition and is Lipschitz continuous, anchored proximal gradient descent methods can be used to compute a monotone MFE. We also establish last-iterate convergence results for these methods. Our approach relies on formulating the computation problem as an optimization problem over the space of occupation measures. Using this formulation, we show that the problem is equivalent to a class of constrained Lipschitz monotone inclusion problems. We then apply iterative methods for this monotone inclusion formulation to derive a tractable algorithm. The resulting algorithm achieves a convergence rate of \(O(1/\sqrt{T})\) after \(T\) iterations, without requiring any regularization. This rate holds even in the absence of a uniqueness assumption for the corresponding MFE.
We propose a Jacobi-like relative value iteration (RVI) algorithm and a Gauss-Seidel-like implementation for the ergodic risk-sensitive control (ERSC) problem of a controlled discrete time Markov chain (DTMC) on a finite state space. Under the assumption that the DTMC is irreducible and recurrent under every stationary Markov policy, we prove that the iterates of the proposed RVI algorithms converge at a geometric rate. The main challenge stems from the multiplicative structure of the ERSC cost criterion and the associated Bellman-like operators, which prevents us from adapting the analogous global contraction and bi-Lipschitz continuity properties that underlie the proof of convergence in the average cost setting. We overcome this by establishing local contraction properties for the risk-sensitive Bellman-like operators and a local bi-Lipschitz continuity property for their fixed points, and use these properties to show the iterates converge geometrically. We conclude by implementing our proposed RVI algorithms on two examples: service effort control for a single-server queue of finite capacity, and maximizing the exit rate from a finite domain (on a graph).