
DJ
Dylan J. Foster, Alexander Rakhlin
· 1 min read
ResearcharXiv cs.LG
Foundations of Reinforcement Learning and Interactive Decision Making
arXiv:2312.16730v2 Announce Type: replace
Abstract: Interactive decision making is the problem of learning to act well in an unknown environment, using the data that one's own actions generate to continuously improve, and arises in situations ranging from online platforms and robotics to medical treatments. This monograph gives a statistical perspective on algorithm design and complexity for interactive decision making, building from multi-armed bandits through contextual and structured bandits to reinforcement learning with function approximation within a single, unified framework. Special attention is paid to function approximation and flexible models such as neural networks, and to the connection between supervised learning and decision making: the reader will learn how to turn any supervised learning method into a decision making algorithm, how to analyze the result, and how to determine whether a given problem can be solved with few interactions.
A unifying theme is that for interactive problems, unlike supervised learning, the question of how hard a problem is cannot be separated from the question of how to solve it. We develop this theme through two recent ideas, which complement classical approaches such as optimism and posterior sampling: the Estimation-to-Decisions principle, which reduces decision making to supervised estimation combined with an exploration rule, and the Decision-Estimation Coefficient, a complexity measure that both determines the exploration rule and lower bounds the regret of any algorithm. For contextual bandits, the resulting theory is essentially complete and its algorithms are widely deployed; for reinforcement learning, it represents the current frontier.
Original source
This story was published by arXiv cs.LG and written by Dylan J. Foster, Alexander Rakhlin. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on arxiv.org


