Skip to content

Author

Jing-Hao Wang

We have 5 of 9 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

DeepShare: Assurance-Driven Deep Learning Job Scheduling for Multi-Tenant Clusters

DeepShare is a scheduler that uses a continuous tenant-assurance signal to coordinate these decisions at runtime to achieve a more advantageous utilization-QoS trade-off than optimizing quotas, scheduling, and resource sharing independently.

Jing-Hao Wang, Yi-Hang Zhou, Xiao Zhou et al. · 0 citations
Preprint Sep 2026

Latency-Aware Orchestration for Multi-Agent LLM Workflows on Heterogeneous GPUs

Concurrent multi-agent workflows expose future dependencies and serving-state requirements while running on heterogeneous GPU pools with time-varying load, model residency, and resource availability. The logical workflow defines the required computation, whereas its physical scheduling units, model-lifecycle actions, r...

Jing-Hao Wang, Yi-Feng Zhang, Xiao Zhou et al. · 0 citations
Jul 2026

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

This work presents SpecBox, a runtime built around speculative sandbox preallocation tailored for dynamic LLM agent execution pipelines, and implements keyword matching and streaming semantic embedding to enable intent-driven sandbox prewarming, which identifies pending tool execution demands mid-LLM token generation a...

Yi-Hui Zhang, Tian-Yu Wo, Jing-Hao Wang et al. · 1 citation
Preprint Aug 2026

HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization

A hierarchical search-space planning framework for GPU kernel optimization that delivers stronger overall implementation validity, sample efficiency, and optimization performance than existing training-free methods, while remaining competitive with the training-based CUDA-L1 without additional model training is propose...

Jing-Hao Wang, Qiqi Gu, Chenpeng Wu et al. · 0 citations
Preprint Aug 2026

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters

ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.

Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.