Signing an AI model: model watermarking, animated

← Back to Research Lab

Machine Learning Security

A stolen model runs inside someone else’s product. How does its owner show it is theirs? This film compares three kinds of model watermark, where the signature is trained into the model itself: black-box (answers to a set of secret inputs), white-box (a sparse pattern in the weights), and architecture-level (a non-intrusive auxiliary head). It ends on the weakness they share, and on tying the mark to the training record. A companion to the Model Heist Detector.

Scientific Reference: The author’s SecurePoL paper (Ural & Yoshigoe, IEEE Access 2025) and dissertation, which train one design from each of the three watermark families under one regime and compare them; trigger sets follow Adi et al. (2018). The examples in the film are illustrations, not a shared benchmark.
🧠 What did you just learn?

The Threat Model Dictates the Defense. If a thief steals your weights and deploys them publicly, you can download the weights and run a statistical test (White-box). But if they hide the model behind an API, you must prove ownership using only queries and responses (Black-box).

Sparse Parameter Perturbations (White-box). The verifier needs access to model weights. Robustness depends on the mark, threshold, and attack budget; a statistical test on the weights does not make false positives impossible.

Feature-Based Triggers (Black-box). Secret input-label associations provide an ownership signal through model queries. Adi et al. evaluate utility and robustness empirically. A matching label alone is not conclusive evidence of theft.

Not covered here: watermarking generated text. Marking what a language model writes (for example a keyed green list of tokens, Kirchenbauer et al., 2023) is a separate technique with a separate purpose. This page is about marking the model itself.

Non-Intrusive Auxiliary Head (SecurePoL). A separate classifier uses shared features for verification and may be pruned by an attacker. SecurePoL combines this signal with a separate trajectory check; it is not an unremovable watermark.

Conditional comparison. Capacity, utility impact, and robustness depend on access, detector calibration, and the attack tested. The film does not establish one universal ranking or bound across these mechanisms.