This work introduces Video2Skill, a benchmark that covers robot tabletop manipulation and human kitchen activity and tests three core capabilities: locating manipulation events in time, grouping events of the same transformation, and deciding when to reuse an existing skill or create a new one.
Jian-Shu Zhang, Ce Zhang, Xi-Yuan Yang et al.· 0 citations
LENS is a training-free keyframe sampling framework that dynamically decides when to zoom in for fine-grained details and when to zoom out for broader context based on the text query, enabling the model to reason across multiple granularities while capturing both high-fidelity details and long-range context.
Ce Zhang, Jinxi He, Katia Sycara et al.· arXiv.org· 2 citations
Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily focus on what to retain in a bounded memory, yet access that memory using the same fixed-cost procedure for every query, despite substantial variation in the evidence re...
Ce Zhang, Jing Bi, Jinxi He et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.