Monge blunts Bayes: Hardness Results for Adversarial Training

Cranko, Zac; Menon, Aditya Krishna; Nock, Richard; Ong, Cheng Soon; Shi, Zhan; Walder, Christian

Computer Science > Machine Learning

arXiv:1806.02977 (cs)

[Submitted on 8 Jun 2018 (v1), last revised 7 May 2019 (this version, v4)]

Title:Monge blunts Bayes: Hardness Results for Adversarial Training

Authors:Zac Cranko, Aditya Krishna Menon, Richard Nock, Cheng Soon Ong, Zhan Shi, Christian Walder

View PDF

Abstract:The last few years have seen a staggering number of empirical studies of the robustness of neural networks in a model of adversarial perturbations of their inputs. Most rely on an adversary which carries out local modifications within prescribed balls. None however has so far questioned the broader picture: how to frame a resource-bounded adversary so that it can be severely detrimental to learning, a non-trivial problem which entails at a minimum the choice of loss and classifiers.
We suggest a formal answer for losses that satisfy the minimal statistical requirement of being proper. We pin down a simple sufficient property for any given class of adversaries to be detrimental to learning, involving a central measure of "harmfulness" which generalizes the well-known class of integral probability metrics. A key feature of our result is that it holds for all proper losses, and for a popular subset of these, the optimisation of this central measure appears to be independent of the loss. When classifiers are Lipschitz -- a now popular approach in adversarial training --, this optimisation resorts to optimal transport to make a low-budget compression of class marginals. Toy experiments reveal a finding recently separately observed: training against a sufficiently budgeted adversary of this kind improves generalization.

Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
ACM classes:	I.2.6
Cite as:	arXiv:1806.02977 [cs.LG]
	(or arXiv:1806.02977v4 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1806.02977

Submission history

From: Richard Nock [view email]
[v1] Fri, 8 Jun 2018 06:00:08 UTC (2,563 KB)
[v2] Wed, 12 Sep 2018 22:16:39 UTC (2,538 KB)
[v3] Tue, 5 Feb 2019 21:49:08 UTC (3,865 KB)
[v4] Tue, 7 May 2019 23:59:29 UTC (3,865 KB)

Computer Science > Machine Learning

Title:Monge blunts Bayes: Hardness Results for Adversarial Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Monge blunts Bayes: Hardness Results for Adversarial Training

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators