Unmanned aerial vehicle (UAV) multimodal perception integrates visible (RGB), infrared (IR), synthetic aperture radar (SAR), and depth sensors for scene understanding under diverse conditions. However, differences in optics, resolution, and mounting often limit practical systems to global or image-center alignment. Aft...
Jing-Pu Yang, Deming Tang, Yi-Lin Sun et al.· 1 citation
UAV video question answering requires separating camera motion from changes in the scene, but RGB-only multimodal models receive no explicit, stable reference for that separation. We present EgoSIS, a pose-free adapter that converts RGB-derived bidirectional flow into motion-canonical visual evidence in three stages. F...
Jing-Pu Yang, Feng-Xian Ji, Ming-Xuan Cui et al.· 0 citations
This work introduces FinCUABuildBench, a benchmark for evaluating financial CUA task construction, and introduces FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks.
Jing-Pu Yang, Feng-Xian Ji, Jinri Guo et al.· 0 citations
J-Access is proposed, an inference-time audit that uses the Jacobian lens to map intermediate representations into vocabulary space and measures how often target concepts remain accessible along the model's output pathway, positioning J-Access as a model-level diagnostic for assessing residual susceptibility in unlearn...
Zirui Song, Hua-Xing Liu, Xiang Wang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.