Multi-Armed Bandits for Correlated Markovian Environments with Smoothed Reward Feedback

Fiez, Tanner; Sekar, Shreyas; Ratliff, Lillian J.

Computer Science > Machine Learning

arXiv:1803.04008 (cs)

[Submitted on 11 Mar 2018 (v1), last revised 1 Mar 2019 (this version, v2)]

Title:Multi-Armed Bandits for Correlated Markovian Environments with Smoothed Reward Feedback

Authors:Tanner Fiez, Shreyas Sekar, Lillian J. Ratliff

View PDF

Abstract:We study a multi-armed bandit problem in a dynamic environment where arm rewards evolve in a correlated fashion according to a Markov chain. Different than much of the work on related problems, in our formulation a learning algorithm does not have access to either a priori information or observations of the state of the Markov chain and only observes smoothed reward feedback following time intervals we refer to as epochs. We demonstrate that existing methods such as UCB and $\varepsilon$-greedy can suffer linear regret in such an environment. Employing mixing-time bounds on Markov chains, we develop algorithms called EpochUCB and EpochGreedy that draw inspiration from the aforementioned methods, yet which admit sublinear regret guarantees for the problem formulation. Our proposed algorithms proceed in epochs in which an arm is played repeatedly for a number of iterations that grows linearly as a function of the number of times an arm has been played in the past. We analyze these algorithms under two types of smoothed reward feedback at the end of each epoch: a reward that is the discount-average of the discounted rewards within an epoch, and a reward that is the time-average of the rewards within an epoch.

Comments:	Significant revision of prior version including deeper discussion of related work, gap-independent regret bounds, and regret bounds for discounted rewards
Subjects:	Machine Learning (cs.LG)
Cite as:	arXiv:1803.04008 [cs.LG]
	(or arXiv:1803.04008v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1803.04008

Submission history

From: Tanner Fiez [view email]
[v1] Sun, 11 Mar 2018 18:44:50 UTC (3,774 KB)
[v2] Fri, 1 Mar 2019 22:03:13 UTC (5,457 KB)

Computer Science > Machine Learning

Title:Multi-Armed Bandits for Correlated Markovian Environments with Smoothed Reward Feedback

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Multi-Armed Bandits for Correlated Markovian Environments with Smoothed Reward Feedback

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators