Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization

1Sydney AI Centre, The University of Sydney, 2University of Melbourne,
3City University of Hong Kong, 4Wuhan University,
5Mohamed bin Zayed University of Artificial Intelligence
TC-UAP teaser

Protecting personal videos against unauthorized video customization. Once posted online, a user’s videos may be collected and exploited by either tuning-based or reference-based customization. Our TC-UAP proactively protects the videos by applying imperceptible temporally consistent perturbations, making both customization methods fail to generate usable results.

Abstract

Recent diffusion-based video generation models have enabled high-quality personalized video customization through both tuning-based pipelines, which fine-tune a video diffusion model, and reference-based pipelines such as image-to-video generation. However, these capabilities raise serious concerns about personal privacy, identity ownership and intellectual property protection. Existing anti-customization works focus on protecting images, while protection for videos against both reference- and tuning-based customization remains largely underexplored. Protecting videos in this setting raises three challenges: (i) Image-level perturbations, optimized frame by frame, cannot survive temporal compression by 3D video VAE. (ii) A video-level perturbation optimized on a single video is vulnerable to temporal editing and fails to protect unseen videos. (iii) Temporally inconsistent perturbations are not robust to temporal attacks. To address these challenges, we propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization. TC-UAP optimizes an identity-level multi-frame UAP over sliding windows from multiple videos, accounting for local temporal dependencies induced by temporal compression in video VAE and enabling a single perturbation to protect unseen videos of varying lengths. Moreover, we introduce intrinsic temporal modeling and an extrinsic surrogate temporal-attack loss, which make the perturbation temporally consistent and robust to unseen temporal attacks. Empirically, quantitative and qualitative results show that TC-UAP achieves the strongest identity protection compared with existing methods under both reference- and tuning-based video customization, and remains robust under multiple unseen temporal attacks.

More Results

Each row shows the customization results when the videos are protected by different methods.

Text prompt: “The p3r5on is licking an ice-cream cone, smiling between licks while facing the camera.”

cleanPhotoGuardMistIDProtectorOurs
cleanPhotoGuardMistIDProtectorOurs
cleanPhotoGuardMistIDProtectorOurs
cleanPhotoGuardMistIDProtectorOurs
cleanPhotoGuardMistIDProtectorOurs
cleanPhotoGuardMistIDProtectorOurs

BibTeX

@article{huang2026delving,
  title={Delving into the Temporal Challenges of Unified Video Protection Against Image-to-Video and Fine-Tuning-based Customization},
  author={Yuxin Huang and Ziming Hong and Mingming Gong and Wanyu Wang and Jing Zhang and Tongliang Liu},
  year={2026},
  journal={arXiv preprint arXiv:2607.13336}
}