Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

Residual Toxicity of Iodinated Disinfection Byproducts in Boiled Drinking Water.

Disinfection byproducts (DBPs) represent a persistent challenge to supply safe drinking water. While boiling drinking water has been confirmed to reduce human exposure risks to DBPs, the remaining risk, particularly waters with iodide, has been underestimated. Herein, we examined how volatilization and thermally accelerated transformations could shift the profiles of the priority DBPs in boiled water. We comprehensively evaluated the toxicity of samples before and after boiling, the results of which suggested that 40-70% of cytotoxicity and 59-89% of genotoxicity remain after boiling. The non-targeted analysis revealed that the phenolic iodinated DBPs (I-DBPs) accumulated and contributed significantly to the remaining total organic iodine. Specifically, the concentrations of 3,5-diiodosalicylic acid and 2,4,6-triiodophenol increased by 124.8 and 139.0% during boiling, respectively, which may be attributed to thermally accelerated reactions from large-molecule-weight precursors and low volatilization of these I-DBPs. In addition, continuous boiling, the addition of vitamin C, and baking soda were investigated to be effective in mitigating toxicity and phenolic I-DBPs resistant to boiling, with the degradation rates of I-DBPs ranging from 1.1% to 33.8%. Overall, we highlighted the occurrence and transformation of thermally stable DBPs during boiling, providing important insights for mitigating the residual risks in boiled drinking water.

Wenyuan Yang, Shengkun Dong, Chao Fang et al. · 0 citations
Aug 2026

Zipf-like Statistical Regularities in Molecular Sequence Representations for Chemical Language Models

Large language model (LLM)-based approaches increasingly use molecular strings such as SMILES and SELFIES for molecular generation and property prediction. However, the statistical properties of molecular token distributions have not been systematically characterized. Here, we analyzed rank–frequency distributions of tokens derived by byte-pair encoding (BPE) across large molecular databases. BPE-derived tokens showed approximate Zipf-like rank–frequency scaling across the examined representations and chemical spaces, with fitted slopes moderately steeper than the canonical value of −1, paralleling a statistical pattern widely observed in natural language. Moreover, when BERT models were pretrained using BPE vocabularies with different rank–frequency slopes, the closeness of these slopes to the ideal Zipf value of −1 strongly correlated with performance on molecular property prediction tasks (Pearson r = 0.91, p < 0.001) and remained associated after adjustment for vocabulary size (partial r = 0.86, p < 0.001). Together, these findings show that molecular BPE vocabularies exhibit an approximate Zipf-like rank–frequency regularity and that slope closeness provides an empirical diagnostic for comparing vocabulary sizes within the examined SMILES/SELFIES BPE framework.

Anyu Liu, Chao Fang, Yuntao Li et al. · 0 citations