SequenceO1: End-to-End Ultra-Long (100K) Sequence Modeling in Recommendation with Low-Rank Caching
Modern short-video recommenders must exploit ultra-long user histories—which can reach hundreds of thousands or even millions of interactions per user—but are constrained by strict latency and training-throughput budgets. As a result, production systems typically truncate histories or rely on two-stage retrieve-then-ra...