Home
Scholarly Works
Blind Watermarking for Tabular Datasets in Machine...
Journal article

Blind Watermarking for Tabular Datasets in Machine Learning: A Primary Key Free Method

Abstract

Watermarking is widely used to protect the ownership of tabular datasets. Comparing to non-blind watermarking that requires the original dataset to detect watermark, blind watermarking is more secure because it can accurately detect watermark without using the original dataset, which restricts public access to the original dataset. Existing blind watermarking methods rely on either a primary key or a virtual primary key to watermark a tabular dataset. However, these watermarks can be easily removed by an attacker with little to no impact on the dataset’s machine learning utility, because a primary key can be significantly modified without affecting the machine learning utility, and a virtual primary key is fragile to slight modifications on the dataset. Can we design a blind watermarking method without relying on a primary key or virtual primary key? In this article, we tackle this challenging task by a novel primary key-free method that embeds a sinusoidal signal as the watermark into a discrete-time signal constructed from the tabular dataset. We theoretically analyzed the robustness of our watermark against six challenging attacks, and empirically validated the outstanding performance of our method through comprehensive experiments on two real-world datasets.

Authors

Che X; Akbari M; Li S; Yue D; Zhang Y; Chu L

Journal

ACM Transactions on Knowledge Discovery from Data, Vol. 20, No. 6, pp. 1–32

Publisher

Association for Computing Machinery (ACM)

Publication Date

July 31, 2026

DOI

10.1145/3820040

ISSN

1556-4681