Skip to content

TripPattern: A Pattern-based Text Watermarking Method for Large Language Models

Sep 2026 · 0 citations · 48 references
Computer Science

TL;DR

TripPattern is a watermarking framework that formulates text watermarking as a pattern-based matching task using three vocabulary partitions, which maintains LLM generation quality while achieving robust watermark detectability.

Abstract

Text watermarking techniques have gained significant attention for identifying machine-generated text and mitigating risks from large language models (LLMs). Existing methods typically divide an LLM's vocabulary into green and red tokens, but encouraging generation toward green tokens can reduce text quality and naturalness. To address this, we propose TripPattern, a watermarking framework that formulates text watermarking as a pattern-based matching task using three vocabulary partitions. TripPattern divides the vocabulary into one neutral group and two pattern groups. During generation, the model alternates token selection between the two pattern groups to embed detectable patterns, while neutral tokens are selected independently to improve flexibility and preserve naturalness. For detection, TripPattern uses pattern-based statistical tests that provide interpretable p-values by measuring how often adjacent tokens alternate between the pattern groups. Theoretical analysis and empirical evaluations on four multilingual datasets show that TripPattern maintains LLM generation quality while achieving robust watermark detectability.

View source

Similar papers

#artificial intelligence Preprint Aug 2026

OpenStamp: A Watermark for Open-Source Language Models

This work introduces OpenStamp, a watermarking technique that encodes the watermarking logic directly into the model weights by modifying only the final projection, or unembedding, layer, and shows that OpenStamp achieves superior detection performance, with minimal degradation in model capabilities compared to prior m...

Miroojin Bakshi, Saksham Rastogi, Danish Pruthi · 0 citations
#machine learning Preprint Sep 2026

WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading

WeaveMark, a robust and scalable multi-bit LLM watermarking scheme based on coded payload spreading, improves payload capacity through multi-bit-per-token spreading (weaving), improving extraction accuracy through soft-decision error-correcting codes, and preserving text quality through unbiased multilayer reweighting.

Gang-Hyun Park, Ju-Hyeong Lee, Heeyoul Kwak et al. · 0 citations
Preprint Aug 2026

Optimal Watermark Localization in Mixed-Source Large Language Model Texts

Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied gl...

José H. Blanchet, T. Cai, Xiang Li et al. · 0 citations
#machine learning Preprint Sep 2026

TANGO: Watermarking Masked Diffusion Language Models in Token Pairs

TANGO is presented, a watermark for masked-diffusion language models that keys each new token to a nearby token that is already unmasked, and TANGO biases the new token toward a color determined by the key and the nearby token's color.

Kasra Arabi, Nir Weinberger, Micah Goldblum et al. · 0 citations
#natural language process... Preprint Sep 2026

CertMark: Distortion-Free Multi-Bit Watermarking with Certified Decoding

Leading multi-bit watermarking methods for language models encode messages by biasing the model's next-token probabilities, creating a trade-off between message recovery and text quality. Their decoders typically return the highest-scoring candidate from accumulated token-level evidence, without a certified abstention...

Pawel Batorski, P. Spurek, Paul Swoboda · 0 citations

An Experimental Study on Attacks and Vulnerabilities of Text Watermarks

Results show that the copy–paste attack degrades watermark detectability when using very sparingly and only short sentences are used, while deletion becomes effective only at extreme levels on short texts, and underscore the need for future watermark designs to prioritise resilience to semantic and cross-lingual transf...

Jakob Elias Albrecht, Andreas Jakoby, Benno Maria et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.