强化学习入门：第二版完全草稿

5星 · 超过95%的资源需积分: 50 84 浏览量更新于2024-07-18 2 收藏 13.44MB PDF 举报

"Reinforcement Learning：An Introduction 第二版，一本深入浅出介绍强化学习的教材，适合配合David的相关课程一起学习。书中包含大量详细解析，对于理解课程内容非常有帮助。" 本书是Richard S. Sutton和Andrew G. Barto合著的《强化学习：入门》第二版的完整草稿，截至2017年11月5日已经完成，可能只待添加一个额外的案例研究到第16章。尽管如此，参考文献仍需全面校对，索引也尚未添加。作者鼓励读者发现任何错误或疏漏时向他们反馈，以便在最终版本印刷前进行修正。强化学习是一种机器学习方法，通过与环境的交互来优化策略，以最大化长期奖励。该领域涵盖了一系列元素，包括环境、状态、动作、奖励函数、策略和价值函数等。其目标是在未知环境中寻找最优行为序列，以获得最大的累积奖励。第一章介绍了强化学习的基本概念，首先定义了强化学习的概念，并通过多个例子（如机器人控制、游戏策略等）来说明其应用。接着，它列举了强化学习的核心元素，包括马尔科夫决策过程（Markov Decision Process, MDP）、动态规划（Dynamic Programming）、蒙特卡洛方法（Monte Carlo Methods）、Temporal Difference Learning（TD学习）以及Q-learning等。同时，书中讨论了强化学习的局限性及其适用范围，并通过国际象棋游戏Tic-Tac-Toe的扩展示例来进一步阐述强化学习的工作原理。早期的历史部分提到了A. Harry Klopf的工作，他被认为是强化学习领域的先驱之一。书中各章节末尾的“文献回顾与历史注解”部分，旨在记录相关领域的关键进展和重要贡献，可能会有遗漏，作者欢迎读者提供补充。此书是深入理解强化学习的重要参考资料，不仅适合初学者，也适用于有一定背景知识的学习者。它结合理论与实践，通过详尽的解释和实例，帮助读者掌握这一复杂而强大的机器学习分支。

xvi Summary of Notation

|S| number of elements in set S

t discrete time step

T, T (t) ﬁnal time step of an episode, or of the episode including time step t

action at time t

state at time t, typically due, stochastically, to S

t−1

and A

t−1

reward at time t, typically due, stochastically, to S

t−1

and A

t−1

π policy, decision-making rule

π(s) action taken in state s under deterministic policy π

π(a|s) probability of taking action a in state s under stochastic policy π

π(a|s, θ) probability of taking action a in state s given parameter θ

return (cumulative discounted reward) following time t (Section 3.3)

t:h

ﬂat return (uncorrected, undiscounted) from t + 1 to h (Section 5.8)

λs

λ-return, corrected by estimated state values (Section 12.1)

λa

λ-return, corrected by estimated action values (Section 12.1)

λs

t:h

truncated, corrected λ-return, with state values (Section 12.3)

λa

t:h

truncated, corrected λ-return, with action values (Section 12.3)

p(s

, r|s, a) probability of transition to state s

with reward r, from state s and action a

p(s

|s, a) probability of transition to state s

, from state s taking action a

r(s, a, s

) expected immediate reward on transition from s to s

under action a

(s) value of state s under policy π (expected return)

∗

(s) value of state s under the optimal policy

(s, a) value of taking action a in state s under policy π

∗

(s, a) value of taking action a in state s under the optimal policy

V, V

array estimates of state-value function v

or v

∗

Q, Q

array estimates of action-value function q

or q

∗

temporal-diﬀerence error at t (a random variable) (Section 6.1)

w, w

d-vector of weights underlying an approximate value function

, w

t,i

ith component of learnable weight vector

d dimensionality—the number of components of w

alternate dimensionality—the number of components of θ

m number of 1s in a sparse binary feature vector

ˆv(s,w) approximate value of state s given weight vector w

(s) alternate notation for ˆv(s,w)

ˆq(s, a, w) approximate value of state–action pair s, a given weight vector w

x(s) vector of features visible when in state s

x(s, a) vector of features visible when in state s taking action a

(s), x

(s, a) ith component of vector x(s) or x(s, a)

shorthand for x(S

) or x(S

, A

)

x inner product of vectors, w

; e.g., ˆv(s,w)

= w

x(s)

µ(s) onpolicy distribution over states (Section 9.2)

µ |S|-vector of the µ(s)

kxk

µ-weighted norm of any vector x(s), i.e.,

µ(s)x(s)

(Section 11.4)

v, v

secondary d-vector of weights, used to learn w (Chapter 11)

d-vector of eligibility traces at time t (Chapter 12)

离散

确定

随机

随机pi政策下在s状态采取a动作的可能性

给定参量fi下在s状态中采取a动作的可能性

平均

Chapter 1

Introduction

The idea that we learn by interacting with our environment is probably the ﬁrst to occur to us when

we think about the nature of learning. When an infant plays, waves its arms, or looks about, it has no

explicit teacher, but it does have a direct sensorimotor connection to its environment. Exercising this

connection produces a wealth of information about cause and eﬀect, about the consequences of actions,

and about what to do in order to achieve goals. Throughout our lives, such interactions are undoubtedly

a major source of knowledge about our environment and ourselves. Whether we are learning to drive a

car or to hold a conversation, we are acutely aware of how our environment responds to what we do, and

we seek to inﬂuence what happens through our behavior. Learning from interaction is a foundational

idea underlying nearly all theories of learning and intelligence.

In this book we explore a computational approach to learning from interaction. Rather than directly

theorizing about how people or animals learn, we explore idealized learning situations and evaluate the

eﬀectiveness of various learning methods. That is, we adopt the perspective of an artiﬁcial intelligence

researcher or engineer. We explore designs for machines that are eﬀective in solving learning problems of

scientiﬁc or economic interest, evaluating the designs through mathematical analysis or computational

experiments. The approach we explore, called reinforcement learning, is much more focused on goal-

directed learning from interaction than are other approaches to machine learning.

1.1 Reinforcement Learning

Reinforcement learning is learning what to do—how to map situations to actions—so as to maximize

a numerical reward signal. The learner is not told which actions to take, but instead must discover

which actions yield the most reward by trying them. In the most interesting and challenging cases,

actions may aﬀect not only the immediate reward but also the next situation and, through that, all

subsequent rewards. These two characteristics—trial-and-error search and delayed reward—are the two

most important distinguishing features of reinforcement learning.

Reinforcement learning, like many topics whose names end with “ing,” such as machine learning

and mountaineering, is simultaneously a problem, a class of solution methods that work well on the

problem, and the ﬁeld that studies this problems and its solution methods. It is convenient to use a

single name for all three things, but at the same time essential to keep the three conceptually separate.

In particular, the distinction between problems and solution methods is very important in reinforcement

learning; failing to make this distinction is the source of a many confusions.

We formalize the problem of reinforcement learning using ideas from dynamical systems theory,

speciﬁcally, as the optimal control of incompletely-known Markov decision processes. The details of this

采用

视角、观点

2 CHAPTER 1. INTRODUCTION

formalization must wait until Chapter 3, but the basic idea is simply to capture the most important

aspects of the real problem facing a learning agent interacting over time with its environment to achieve

a goal. A learning agent must be able to sense the state of its environment to some extent and must be

able to take actions that aﬀect the state. The agent also must have a goal or goals relating to the state of

the environment. Markov decision processes are intended to include just these three aspects—sensation,

action, and goal—in their simplest possible forms without trivializing any of them. Any method that

is well suited to solving such problems we consider to be a reinforcement learning method.

Reinforcement learning is diﬀerent from supervised learning, the kind of learning studied in most

current research in the ﬁeld of machine learning. Supervised learning is learning from a training set

of labeled examples provided by a knowledgable external supervisor. Each example is a description of

a situation together with a speciﬁcation—the label—of the correct action the system should take to

that situation, which is often to identify a category to which the situation belongs. The object of this

kind of learning is for the system to extrapolate, or generalize, its responses so that it acts correctly

in situations not present in the training set. This is an important kind of learning, but alone it is

not adequate for learning from interaction. In interactive problems it is often impractical to obtain

examples of desired behavior that are both correct and representative of all the situations in which the

agent has to act. In uncharted territory—where one would expect learning to be most beneﬁcial—an

agent must be able to learn from its own experience.

Reinforcement learning is also diﬀerent from what machine learning researchers call unsupervised

learning, which is typically about ﬁnding structure hidden in collections of unlabeled data. The terms

supervised learning and unsupervised learning would seem to exhaustively classify machine learning

paradigms, but they do not. Although one might be tempted to think of reinforcement learning as a

kind of unsupervised learning because it does not rely on examples of correct behavior, reinforcement

learning is trying to maximize a reward signal instead of trying to ﬁnd hidden structure. Uncovering

structure in an agent’s experience can certainly be useful in reinforcement learning, but by itself does

not address the reinforcement learning problem of maximizing a reward signal. We therefore consider

reinforcement learning to be a third machine learning paradigm, alongside supervised learning and

unsupervised learning and perhaps other paradigms as well.

One of the challenges that arise in reinforcement learning, and not in other kinds of learning, is the

trade-oﬀ between exploration and exploitation. To obtain a lot of reward, a reinforcement learning

agent must prefer actions that it has tried in the past and found to be eﬀective in producing reward.

But to discover such actions, it has to try actions that it has not selected before. The agent has to

exploit what it has already experienced in order to obtain reward, but it also has to explore in order to

make better action selections in the future. The dilemma is that neither exploration nor exploitation

can be pursued exclusively without failing at the task. The agent must try a variety of actions and

progressively favor those that appear to be best. On a stochastic task, each action must be tried many

times to gain a reliable estimate of its expected reward. The exploration–exploitation dilemma has been

intensively studied by mathematicians for many decades, yet remains unresolved. For now, we simply

note that the entire issue of balancing exploration and exploitation does not even arise in supervised

and unsupervised learning, at least in their purest forms.

Another key feature of reinforcement learning is that it explicitly considers the whole problem of a

goal-directed agent interacting with an uncertain environment. This is in contrast to many approaches

that consider subproblems without addressing how they might ﬁt into a larger picture. For example, we

have mentioned that much of machine learning research is concerned with supervised learning without

explicitly specifying how such an ability would ﬁnally be useful. Other researchers have developed

theories of planning with general goals, but without considering planning’s role in real-time decision

making, or the question of where the predictive models necessary for planning would come from. Al-

though these approaches have yielded many useful results, their focus on isolated subproblems is a

signiﬁcant limitation.

互动

important

困境

专门追求

随机

孤立

剩余444页未读，继续阅读

MK唔识扮嘢

粉丝: 0
资源: 2

强化学习入门：第二版完全草稿

Reinforcement Learning An Introduction second edition

Reinforcement Learning_An Introduction多版本合集

强化学习导论 第二版 英文版 2017最新版 Reinforcement Learning An Introduction

Reinforcement Learning: An Introduction 第二版

Reinforcement Learning: An Introduction

reinforcement learning：an introduction代码

强化学习：简介，第二版（草稿）Reinforcement Learning: An Introduction, Second Edition (Draft)

Reinforcement Learning:An Introduction （2020）第二版，原版

Reinforcement Learning: An Introduction 2nd solutions （第二版 答案）

Reinforcement Learning: An Introduction 2018年 第二版和之前2015中文翻译版

最新资源

强化学习导论第二版英文版 2017最新版 Reinforcement Learning An Introduction

Reinforcement Learning: An Introduction 2nd solutions （第二版答案）

Reinforcement Learning: An Introduction 2018年第二版和之前2015中文翻译版