Towards Solving Adversarial Examples with DP-guided Diffusion Models

Drs. Mathias Lécuyer in the Department of Computer Science and Geoff Pleiss in the Department of Statistics have been awarded the DSI Postdoctoral Matching Fund for their project titled "Towards Solving Adversarial Examples with DP-guided Diffusion Models".

Summary

AI models show a worrying susceptibility to adversarial attacks, in which an attacker applies imperceptible changes to the input to arbitrarily influence a target model. In particular, such attacks can jailbreak aligned foundation models. We propose a novel technique combining denoising diffusion models and Differential Privacy to design AI models that are provably robust against adversarial attacks. Our approach will improve on existing Randomized Smoothing defences, and enable new capabilities such as joint robustness against multiple threat models, and sound adaptivity to the difficulty of each input at prediction time.

Background

AI robustness should be considered under adversarial threat models, in which we can provably assure our desired properties in worst-case scenarios. Otherwise, even seemingly well behaving models are likely to create safety failures due to interactions in competitive environments or adversarial use by bad actors. Ensuring provable model properties is challenging, as it contrasts with traditional AI learning theory and performance metrics which characterize models’ behavior in expectation. The most promising approach to date to enforce robustness guarantees in large AI models is Randomized Smoothing (RS), which prevents adversarial attacks by averaging predictions over noisy versions of the input at test time. RS also serves as a building block to enforce fairness properties, and create robust watermarks or unlearnable examples.

Challenge

RS suffers from two severe limitations. First, RS is inflexible, and its guarantees apply to one specific threat-model for which the defence is designed. Second, RS induces a trade-off between robustness and accuracy, and state-of-the-art models only defend against small attacks at the cost of degraded accuracy. Theoretical analysis shows that this limitation is fundamental to RS with symmetric noise distributions, for which RS exhibits a curse of dimensionality in the input size. Previous work has sought to use input-adaptivity to bypass this curse of dimensionality, though analyzing adaptive techniques is challenging and only yielded modest improvements.

Solution

This proposal aims to develop a new rigorous approach to RS that is fully adaptive, by combining denoising diffusion models and Differential Privacy. Intuitively, a denoising diffusion model maps a noise sample to an image by denoising the input over many steps. At each step, we will nudge the process towards our target input. Since each nudge is small and inherently noisy, small details of the original input do not matter too much, providing robustness

to adversarial changes. Our approach will drastically improve existing RS defences’ accuracy/robustness trade-off, and enable new capabilities such as joint robustness against multiple threat models, and adaptivity to the difficulty of each input at prediction time.

Musqueam First Nation land acknowledegement

We honour xwməθkwəy̓ əm (Musqueam) on whose ancestral, unceded territory UBC Vancouver is situated. UBC Science is committed to building meaningful relationships with Indigenous peoples so we can advance Reconciliation and ensure traditional ways of knowing enrich our teaching and research.

Learn more: Musqueam First Nation

Data Science Institute

EOS Main Building
6339 Stores Road, Room 113C
dsi.admin@science.ubc.ca

Faculty of Science

Office of the Dean, Earth Sciences Building
2178–2207 Main Mall
Vancouver, BC Canada
V6T 1Z4
UBC Crest The official logo of the University of British Columbia. Urgent Message An exclamation mark in a speech bubble. Arrow An arrow indicating direction. Arrow in Circle An arrow indicating direction. Bluesky The logo for the Bluesky social media service. A bookmark An ribbon to indicate a special marker. Calendar A calendar. Caret An arrowhead indicating direction. Time A clock. Chats Two speech clouds. External link An arrow pointing up and to the right. Facebook The logo for the Facebook social media service. A Facemask The medical facemask. Information The letter 'i' in a circle. Instagram The logo for the Instagram social media service. Linkedin The logo for the LinkedIn social media service. Lock, closed A closed padlock. Lock, open An open padlock. Location Pin A map location pin. Mail An envelope. Mask A protective face mask. Menu Three horizontal lines indicating a menu. Minus A minus sign. Money A money bill. Telephone An antique telephone. Plus A plus symbol indicating more or the ability to add. RSS Curved lines indicating information transfer. Search A magnifying glass. Arrow indicating share action A directional arrow. Spotify The logo for the Spotify music streaming service. Twitter The logo for the Twitter social media service. Youtube The logo for the YouTube video sharing service.