Streaming video understanding requires Vision Language Models (VLLMs) to process growing video streams and answer user questions under tight latency constraints. Existing methods improve efficiency through token pruning and memory-bank schemes, but mainly reduce visual tokens after visual encoding. Consequently, downst...
Jing-Chi Jiang, Yi-Ran Ling, Ruo-Nan Li et al.· 0 citations
This work proposes H2Table (Hierarchical Hypergraph-Enhanced Table Reasoning), a novel framework that represents complex tables as hierarchical nested hypergraphs, and designs a tailored hypergraph encoder to facilitate message passing between hyperedges and nodes within complex tables.
Jia Ling, Yang-Fan Wang, Chen Tang et al.· 0 citations
FAIRY is developed to execute and evaluate agentic agronomic operations on full-season spatiotemporal workflows that span ridge preparation, planting, irrigation, fertilization, pest and disease treatment, harvest, grain handling, drying, and storage, and an evaluation suite that combines agentic success, full-path spa...
Ao Qu, Panagiotis Michelakis, Lin-Yuan Han et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.