Skip to content

Author

Shengbo Lu

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#reinforcement learning Open access Sep 2026

An Intelligent Construction Method for Petrochemical Datasets Based on RLHF and Data De-Identification

To mitigate high expert-annotation costs, domain-preference misalignment, and the inherent trade-off between sensitive-data protection and training utility in petrochemical dataset construction, an iterative framework combining human-feedback-aligned reinforcement learning (RLHF) with post hoc data de-identification is proposed. Direct scoring and pairwise preference feedback are generated using two high-capability language models. A reward model is subsequently trained via a joint Bradley–Terry and mean-squared-error loss, followed by three rounds of closed-loop proximal policy optimization (PPO) constrained by a fixed supervised fine-tuning reference model. Retained high-quality samples are then processed through a four-stage post-RLHF de-identification pipeline. Experimental results demonstrate that the PPO-V3 model achieves a reward score increase of 1.183 over the baseline alongside a 96.2% pairwise win rate, while the sensitivity-aware adaptive differential privacy with context-aware token-level injection (SA-ADP-CTI) post-RLHF de-identification method attains a composite score of 0.9636. The reliability of both the automated feedback and privacy-preservation mechanisms is further validated through blind expert review and manual spot checks.

Yimin Liu, Qike Ji, Shengbo Lu et al. · 0 citations