Sanitizing hidden activations for improving adversarial robustness of convolutional neural networks

Mu, Tianshi; Lin, Kequan; Zhang, Huabing; Wang, Jian

doi:10.3233/JIFS-210371

Sanitizing hidden activations for improving adversarial robustness of convolutional neural networks

Article type: Research Article

Authors: Mu, Tianshi^{; *} | Lin, Kequan | Zhang, Huabing | Wang, Jian

Affiliations: Digital Grid Research Institute, China Southern Power Grid, Guangzhou, China

Correspondence: [*] Corresponding author. Tianshi Mu, Digital Grid Research Institute, China Southern Power Grid, Guangzhou 510700, China. E-mail: [email protected].

Abstract: Deep learning is gaining significant traction in a wide range of areas. Whereas, recent studies have demonstrated that deep learning exhibits the fatal weakness on adversarial examples. Due to the black-box nature and un-transparency problem of deep learning, it is difficult to explain the reason for the existence of adversarial examples and also hard to defend against them. This study focuses on improving the adversarial robustness of convolutional neural networks. We first explore how adversarial examples behave inside the network through visualization. We find that adversarial examples produce perturbations in hidden activations, which forms an amplification effect to fool the network. Motivated by this observation, we propose an approach, termed as sanitizing hidden activations, to help the network correctly recognize adversarial examples by eliminating or reducing the perturbations in hidden activations. To demonstrate the effectiveness of our approach, we conduct experiments on three widely used datasets: MNIST, CIFAR-10 and ImageNet, and also compare with state-of-the-art defense techniques. The experimental results show that our sanitizing approach is more generalized to defend against different kinds of attacks and can effectively improve the adversarial robustness of convolutional neural networks.

Keywords: Adversarial examples, sanitizing hidden activations, adversarial robustness, convolutional neural networks

DOI: 10.3233/JIFS-210371

Journal: Journal of Intelligent & Fuzzy Systems, vol. 41, no. 2, pp. 3993-4003, 2021

Published: 15 September 2021

Price: EUR 27.50

North America

IOS Press, Inc.
6751 Tepper Drive
Clifton, VA 20124
USA

Tel: +1 703 830 6300
Fax: +1 703 830 2300
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

Europe

IOS Press
Nieuwe Hemweg 6B
1013 BG Amsterdam
The Netherlands

Tel: +31 20 688 3355
Fax: +31 20 687 0091
[email protected]

For editorial issues, permissions, book requests, submissions and proceedings, contact the Amsterdam office [email protected]

Asia

Inspirees International (China Office)
Ciyunsi Beili 207(CapitaLand), Bld 1, 7-901
100025, Beijing
China

Free service line: 400 661 8717
Fax: +86 10 8446 7947
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

如果您在出版方面需要帮助或有任何建, 件至: [email protected]

Share this:

North America

Europe

Asia