Signing an AI model: model watermarking, animated
A stolen model runs inside someone else’s product. How does its owner show it is theirs? This film compares three kinds of model watermark, where the signature is trained into the model itself: black-box (answers to a set of secret inputs), white-box (a sparse pattern in the weights), and architecture-level (a non-intrusive auxiliary head). It ends on the weakness they share, and on tying the mark to the training record. A companion to the Model Heist Detector.
🧠 What did you just learn?
The Threat Model Dictates the Defense. If a thief steals your weights and deploys them publicly, you can download the weights and run a statistical test (White-box). But if they hide the model behind an API, you must prove ownership using only queries and responses (Black-box).
Sparse Parameter Perturbations (White-box). The verifier needs access to model weights. Robustness depends on the mark, threshold, and attack budget; a statistical test on the weights does not make false positives impossible.
Feature-Based Triggers (Black-box). Secret input-label associations provide an ownership signal through model queries. Adi et al. evaluate utility and robustness empirically. A matching label alone is not conclusive evidence of theft.
Not covered here: watermarking generated text. Marking what a language model writes (for example a keyed green list of tokens, Kirchenbauer et al., 2023) is a separate technique with a separate purpose. This page is about marking the model itself.
Non-Intrusive Auxiliary Head (SecurePoL). A separate classifier uses shared features for verification and may be pruned by an attacker. SecurePoL combines this signal with a separate trajectory check; it is not an unremovable watermark.
Conditional comparison. Capacity, utility impact, and robustness depend on access, detector calibration, and the attack tested. The film does not establish one universal ranking or bound across these mechanisms.
