深度学习强化学习入门：Richard S. Sutton经典教程（2017版）

需积分: 11 14 浏览量更新于2024-07-19 2 收藏 10.85MB PDF 举报

《强化学习：一个介绍》是Richard S. Sutton和Andrew G. Barto合著的一本经典教材，于2017年发布第二版。这本书涵盖了强化学习的最新理论和发展，特别强调了DeepMind团队的创新成果，对于想要深入理解强化学习以及将其作为机器学习全面学习资料的人来说，是一本不可或缺的参考资料。在第一部分的"Introduction"章节中，作者首先定义了强化学习（Reinforcement Learning），这是一种通过与环境互动来学习如何做出决策的机器学习方法，目标是最大化长期累积奖励。学习者在这种过程中，没有显式的指导，而是通过尝试与错误来改进策略。接着，作者列举了多个强化学习的实际例子，如机器人控制、游戏策略（如围棋）、自动驾驶等，以帮助读者直观感受这一概念在实际问题中的应用。这部分内容强调了强化学习在解决复杂决策问题上的潜力。章节进一步阐述了强化学习的核心元素，包括状态、动作、环境、奖励函数、策略和价值函数等。学习者的目标是在给定状态下选择最优动作以获得最大的未来奖励。同时，书中讨论了不同类型的算法，如值迭代、策略梯度、深度Q网络等，以及它们在实践中的优缺点。然而，作者也明确了本书的局限性和研究范围，它主要关注基于模型的学习方法，对无模型或部分模型的强化学习以及现代深度强化学习的某些高级技术可能着墨较少。尽管如此，它为理解基础原理提供了坚实的基础。在"An Extended Example: Tic-Tac-Toe"部分，作者通过一个简单的棋盘游戏展示了强化学习的具体实施过程，让读者通过实例掌握强化学习的计算和决策逻辑。这一章节有助于读者建立实践操作的概念框架。最后，作者概述了强化学习的历史背景，包括早期的工作，如马尔可夫决策过程（Markov Decision Processes，MDPs）和Q-learning算法的提出，以及强化学习近年来在学术界和工业界的兴起，特别是AlphaGo的胜利等重大突破。《强化学习：一个介绍》以其系统性、实用性，为读者提供了一个全面且深入的强化学习入门指南，对于那些希望在这个领域深入研究或应用的人来说，无论是初学者还是专业人士，都能从中受益良多。

xvi Summary of Notation

|S| number of elements in set S

t discrete time step

T, T (t) ﬁnal time step of an episode, or of the episode including time step t

action at time t

state at time t, typically due, stochastically, to S

t−1

and A

t−1

reward at time t, typically due, stochastically, to S

t−1

and A

t−1

π policy, decision-making rule

π(s) action taken in state s under deterministic policy π

π(a|s) probability of taking action a in state s under stochastic policy π

π(a|s, θ) probability of taking action a in state s given parameter θ

return (cumulative discounted reward) following time t (Section 3.3)

t:h

ﬂat return (uncorrected, undiscounted) from t + 1 to h (Section 5.8)

λs

λ-return, corrected by estimated state values (Section 12.1)

λa

λ-return, corrected by estimated action values (Section 12.1)

λs

t:h

truncated, corrected λ-return, with state values (Section 12.3)

λa

t:h

truncated, corrected λ-return, with action values (Section 12.3)

p(s

, r|s, a) probability of transition to state s

with reward r, from state s and action a

p(s

|s, a) probability of transition to state s

, from state s taking action a

r(s, a, s

) expected immediate reward on transition from s to s

under action a

(s) value of state s under policy π (expected return)

∗

(s) value of state s under the optimal policy

(s, a) value of taking action a in state s under policy π

∗

(s, a) value of taking action a in state s under the optimal policy

V, V

array estimates of state-value function v

or v

∗

Q, Q

array estimates of action-value function q

or q

∗

temporal-diﬀerence error at t (a random variable) (Section 6.1)

w, w

d-vector of weights underlying an approximate value function

, w

t,i

ith component of learnable weight vector

d dimensionality—the number of components of w

alternate dimensionality—the number of components of θ

m number of 1s in a sparse binary feature vector

ˆv(s,w) approximate value of state s given weight vector w

(s) alternate notation for ˆv(s,w)

ˆq(s, a, w) approximate value of state–action pair s, a given weight vector w

x(s) vector of features visible when in state s

x(s, a) vector of features visible when in state s taking action a

(s), x

(s, a) ith component of vector x(s) or x(s, a)

shorthand for x(S

) or x(S

, A

)

x inner product of vectors, w

; e.g., ˆv(s,w)

= w

x(s)

µ(s) onpolicy distribution over states (Section 9.2)

µ |S|-vector of the µ(s)

kxk

µ-weighted norm of any vector x(s), i.e.,

µ(s)x(s)

(Section 11.4)

v, v

secondary d-vector of weights, used to learn w (Chapter 11)

d-vector of eligibility traces at time t (Chapter 12)

Chapter 1

Introduction

The idea that we learn by interacting with our environment is probably the ﬁrst to occur to us when

we think about the nature of learning. When an infant plays, waves its arms, or looks about, it has no

explicit teacher, but it does have a direct sensorimotor connection to its environment. Exercising this

connection produces a wealth of information about cause and eﬀect, about the consequences of actions,

and about what to do in order to achieve goals. Throughout our lives, such interactions are undoubtedly

a major source of knowledge about our environment and ourselves. Whether we are learning to drive a

car or to hold a conversation, we are acutely aware of how our environment responds to what we do, and

we seek to inﬂuence what happens through our behavior. Learning from interaction is a foundational

idea underlying nearly all theories of learning and intelligence.

In this book we explore a computational approach to learning from interaction. Rather than directly

theorizing about how people or animals learn, we explore idealized learning situations and evaluate the

eﬀectiveness of various learning methods. That is, we adopt the perspective of an artiﬁcial intelligence

researcher or engineer. We explore designs for machines that are eﬀective in solving learning problems of

scientiﬁc or economic interest, evaluating the designs through mathematical analysis or computational

experiments. The approach we explore, called reinforcement learning, is much more focused on goal-

directed learning from interaction than are other approaches to machine learning.

1.1 Reinforcement Learning

Reinforcement learning is learning what to do—how to map situations to actions—so as to maximize

a numerical reward signal. The learner is not told which actions to take, but instead must discover

which actions yield the most reward by trying them. In the most interesting and challenging cases,

actions may aﬀect not only the immediate reward but also the next situation and, through that, all

subsequent rewards. These two characteristics—trial-and-error search and delayed reward—are the two

most important distinguishing features of reinforcement learning.

Reinforcement learning, like many topics whose names end with “ing,” such as machine learning

and mountaineering, is simultaneously a problem, a class of solution methods that work well on the

problem, and the ﬁeld that studies this problems and its solution methods. It is convenient to use a

single name for all three things, but at the same time essential to keep the three conceptually separate.

In particular, the distinction between problems and solution methods is very important in reinforcement

learning; failing to make this distinction is the source of a many confusions.

We formalize the problem of reinforcement learning using ideas from dynamical systems theory,

speciﬁcally, as the optimal control of incompletely-known Markov decision processes. The details of this

2 CHAPTER 1. INTRODUCTION

formalization must wait until Chapter 3, but the basic idea is simply to capture the most important

aspects of the real problem facing a learning agent interacting over time with its environment to achieve

a goal. A learning agent must be able to sense the state of its environment to some extent and must be

able to take actions that aﬀect the state. The agent also must have a goal or goals relating to the state of

the environment. Markov decision processes are intended to include just these three aspects—sensation,

action, and goal—in their simplest possible forms without trivializing any of them. Any method that

is well suited to solving such problems we consider to be a reinforcement learning method.

Reinforcement learning is diﬀerent from supervised learning, the kind of learning studied in most

current research in the ﬁeld of machine learning. Supervised learning is learning from a training set

of labeled examples provided by a knowledgable external supervisor. Each example is a description of

a situation together with a speciﬁcation—the label—of the correct action the system should take to

that situation, which is often to identify a category to which the situation belongs. The object of this

kind of learning is for the system to extrapolate, or generalize, its responses so that it acts correctly

in situations not present in the training set. This is an important kind of learning, but alone it is

not adequate for learning from interaction. In interactive problems it is often impractical to obtain

examples of desired behavior that are both correct and representative of all the situations in which the

agent has to act. In uncharted territory—where one would expect learning to be most beneﬁcial—an

agent must be able to learn from its own experience.

Reinforcement learning is also diﬀerent from what machine learning researchers call unsupervised

learning, which is typically about ﬁnding structure hidden in collections of unlabeled data. The terms

supervised learning and unsupervised learning would seem to exhaustively classify machine learning

paradigms, but they do not. Although one might be tempted to think of reinforcement learning as a

kind of unsupervised learning because it does not rely on examples of correct behavior, reinforcement

learning is trying to maximize a reward signal instead of trying to ﬁnd hidden structure. Uncovering

structure in an agent’s experience can certainly be useful in reinforcement learning, but by itself does

not address the reinforcement learning problem of maximizing a reward signal. We therefore consider

reinforcement learning to be a third machine learning paradigm, alongside supervised learning and

unsupervised learning and perhaps other paradigms as well.

One of the challenges that arise in reinforcement learning, and not in other kinds of learning, is the

trade-oﬀ between exploration and exploitation. To obtain a lot of reward, a reinforcement learning

agent must prefer actions that it has tried in the past and found to be eﬀective in producing reward.

But to discover such actions, it has to try actions that it has not selected before. The agent has to

exploit what it has already experienced in order to obtain reward, but it also has to explore in order to

make better action selections in the future. The dilemma is that neither exploration nor exploitation

can be pursued exclusively without failing at the task. The agent must try a variety of actions and

progressively favor those that appear to be best. On a stochastic task, each action must be tried many

times to gain a reliable estimate of its expected reward. The exploration–exploitation dilemma has been

intensively studied by mathematicians for many decades, yet remains unresolved. For now, we simply

note that the entire issue of balancing exploration and exploitation does not even arise in supervised

and unsupervised learning, at least in their purest forms.

Another key feature of reinforcement learning is that it explicitly considers the whole problem of a

goal-directed agent interacting with an uncertain environment. This is in contrast to many approaches

that consider subproblems without addressing how they might ﬁt into a larger picture. For example, we

have mentioned that much of machine learning research is concerned with supervised learning without

explicitly specifying how such an ability would ﬁnally be useful. Other researchers have developed

theories of planning with general goals, but without considering planning’s role in real-time decision

making, or the question of where the predictive models necessary for planning would come from. Al-

though these approaches have yielded many useful results, their focus on isolated subproblems is a

signiﬁcant limitation.

剩余444页未读，继续阅读

Angel__c

粉丝: 1
资源: 6

深度学习强化学习入门：Richard S. Sutton经典教程（2017版）

Reinforcement Learning: An Introduction second edition

reinforcement-learning-an-introduction-chinese:《强化学习

Reinforcement Learning: An Introduction最新版习题解答（第一版本）

reinforcement learning an introduction 第2版 答案

reinforcement learning an introduction 答案

reinforcement learning: an introduction.pdf

reinforcement learning : an introduction

reinforcement learning: an introduction

reforcement learning an introduction电子书

强化学习入门资料algorithms for reinforcement learning

最新资源

reinforcement learning an introduction 第2版答案