Blackbox Attacks on Reinforcement Learning Agents Using Approximated Temporal Information

Zhao, Yiren; Shumailov, Ilia; Cui, Han; Gao, Xitong; Mullins, Robert; Anderson, Ross

Computer Science > Machine Learning

arXiv:1909.02918 (cs)

[Submitted on 6 Sep 2019 (v1), last revised 21 Nov 2019 (this version, v2)]

Title:Blackbox Attacks on Reinforcement Learning Agents Using Approximated Temporal Information

Authors:Yiren Zhao, Ilia Shumailov, Han Cui, Xitong Gao, Robert Mullins, Ross Anderson

View PDF

Abstract:Recent research on reinforcement learning (RL) has suggested that trained agents are vulnerable to maliciously crafted adversarial samples. In this work, we show how such samples can be generalised from White-box and Grey-box attacks to a strong Black-box case, where the attacker has no knowledge of the agents, their training parameters and their training methods. We use sequence-to-sequence models to predict a single action or a sequence of future actions that a trained agent will make. First, we show our approximation model, based on time-series information from the agent, consistently predicts RL agents' future actions with high accuracy in a Black-box setup on a wide range of games and RL algorithms. Second, we find that although adversarial samples are transferable from the target model to our RL agents, they often outperform random Gaussian noise only marginally. This highlights a serious methodological deficiency in previous work on such agents; random jamming should have been taken as the baseline for evaluation. Third, we propose a novel use for adversarial samplesin Black-box attacks of RL agents: they can be used to trigger a trained agent to misbehave after a specific time delay. This appears to be a genuinely new type of attack. It potentially enables an attacker to use devices controlled by RL agents as time bombs.

Subjects:	Machine Learning (cs.LG); Cryptography and Security (cs.CR); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (stat.ML)
Cite as:	arXiv:1909.02918 [cs.LG]
	(or arXiv:1909.02918v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1909.02918

Submission history

From: Ilia Shumailov [view email]
[v1] Fri, 6 Sep 2019 14:06:21 UTC (5,424 KB)
[v2] Thu, 21 Nov 2019 19:07:45 UTC (4,524 KB)

Computer Science > Machine Learning

Title:Blackbox Attacks on Reinforcement Learning Agents Using Approximated Temporal Information

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Blackbox Attacks on Reinforcement Learning Agents Using Approximated Temporal Information

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators