CT-PPO Experimental Artifacts for Multi-Tenant ISAC Sensing-Session Consolidation
Experimental artifacts supporting the manuscript “Sense Once, Serve Many: Common-Trace Factorized Constrained PPO for Online Sensing-Session Consolidation in Multi-Tenant ISAC Networks.” This deposit contains the final artifacts used for the learned-policy comparisons, component ablation, arrival-load robustness evaluation, and heuristic-reference analyses reported in the manuscript. The archive includes 20 learned-policy training/evaluation bundles covering four methods across training seeds 0–4: Common-Trace Factorized Constrained PPO (CT-PPO), Joint-Credit PPO (JC-PPO), Factorized-JC, and CT-Reward. Each learned-policy run uses a budget of 1,000,000 physical interaction slots. Five additional arrival-load robustness bundles contain frozen-checkpoint CT-PPO and JC-PPO evaluations at the low and high tested arrival rates for seeds 0–4; the nominal-load evaluations are contained in the corresponding learned-policy bundles. The deposit also includes report_heuristic.json, containing the evaluation results for the four deterministic heuristic baselines and Random Valid used in the manuscript. Together, these artifacts support the primary CT-PPO/JC-PPO comparison, the four-way component ablation, regime-wise and consolidation analyses, arrival-load stress evaluation, and heuristic-reference comparisons reported in the manuscript. The associated implementation and training/evaluation scripts are available in the linked GitHub software repository.