This work systematically study MoE designs for vision encoder scaling and finds that fine-grained MoE topologies yield substantial gains over both dense and standard MoE counterparts, and proposes an auxiliary-loss-free balancing variant for better expert utilization, and designs a specialized MoE kernel to mitigate inference latency overhead.
Bonan Zhang, Shiyu Dong, Quan Hung Tran et al.· 0 citations
A structured meta-rubric framework that captures the grading criteria at authoring time, and fixed mechanical rules compile it into a flat checklist of binary, machine-gradable checks that an LLM judge scores reliably at evaluation time is instantiated.
Xilun Chen, Zhaleh Feizollahi, Ross Goodwin et al.· 0 citations
This work introduces S-EMBER (Streaming Egocentric Memory Benchmark for Episodic Retrieval), a large-scale benchmark comprising 3,141 videos totaling 388 hours of organic activity captured via Ray-Ban Meta smart glasses that establishes a hardware-authentic foundation for developing grounded, reliable episodic memory in the next generation of wearable AI agents.
Xiaodong Wang, Xuanyi Zhao, Pedro Rodriguez et al.· 0 citations