Reinforcement learning (RL) trains representations on data selected by the agent's policy, which then uses the resulting returns to guide its next choices. We show that this loop can sustain a lower-return policy even when representation fitting is globally optimal on those data. In a self-confirming superposition trap...
This work presents an exposition of the OSQ problem by summarizing its various formulations in the current literature and categorizing existing solutions into three different types, and summarizes the empirical methods proposed by existing works to verify the efficiency of OSQ mitigation approaches.
Dai Shi, Andi Han, Lequan Lin et al.· IEEE Transactions on Pattern...· 0 citations
Continual knowledge graph embedding updates entity and relation representations as a graph grows. Existing methods primarily address catastrophic forgetting, but entity admission also changes the candidate universe of every compatible query. A historical answer can therefore lose rank even when its score and its orderi...
An input may activate few hidden units even when different inputs collectively use an entire network. We study the statistical complexity of this input-dependent sparsity in the one-hidden-layer ReLU model of Awasthi et al. (COLT 2024). For width $s$, at most $k$ active units per input, and effective weight and bias bo...
Xiao-Yu Li, Zhizhou Sha, Jiao-Jiao Jiang et al.· 0 citations
We determine exactly what a kurtosis bound buys for one-sided tail control. For the class $\mathcal{C}(\kappa)$ of real random variables with mean $0$, variance $1$, and fourth moment at most $\kappa$, the skewness left free, we compute the worst-case tail probability $V_1(t,\kappa)=\sup_{X\in\mathcal{C}(\kappa)}\mathb...
Xiaoyu Li, Andi Han, Jiaojiao Jiang et al.· 0 citations
Worst-case multiclass bounds do not become smaller when the best classifier is already nearly correct: what is missing is an optimistic rate, a guarantee whose fluctuation scales with the oracle risk itself. For a class of Natarajan dimension $d_N$ and Daniely-Shalev-Shwartz dimension $d_{DS}$, the optimal excess risk...
Xiao-Yu Li, Andi Han, Jiao-Jiao Jiang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.