加强学习：入门与进阶

需积分: 0 107 浏览量更新于2024-07-18 收藏 12.15MB PDF 举报

"Reinforcement Learning: An Introduction"是Richard S. Sutton和Andrew G. Barto合著的关于强化学习的权威书籍，第二版正在更新中。本书深入探讨了强化学习的算法，它是一种机器学习方法，通过与环境的交互来学习最优行为策略，以最大化长期奖励。强化学习在人工智能领域具有广泛的应用，例如游戏策略、机器人控制、自然语言处理等。第一部分介绍中，作者定义了强化学习的基本概念，并通过例子进行解释。他们强调了强化学习的核心在于智能体通过试错学习，不断调整其行为以获得更多的奖励。书中提到了四个基本元素：环境（Environment）、智能体（Agent）、动作（Action）、以及奖励（Reward）。在第1.2节中，作者提供了几个强化学习的例子，这些例子有助于读者理解强化学习的工作原理。通过这些例子，读者可以了解如何应用强化学习解决实际问题。第1.3节详细阐述了强化学习的组成要素，包括状态（State）、策略（Policy）、值函数（Value Function）和动态规划方法（Dynamic Programming）。策略是智能体选择行动的规则，值函数衡量了执行特定策略时预期的累计奖励，而动态规划则提供了一种优化策略的方法。第1.4节讨论了强化学习的局限性和适用范围，如离散与连续状态空间的处理、延迟奖励问题、以及探索与利用的平衡等挑战。第1.5节通过一个扩展示例——井字游戏（Tic-Tac-Toe），进一步展示了强化学习的概念如何在实际问题中应用。这个例子让读者看到智能体如何通过学习逐渐提升游戏策略。第1.6节对前面的内容进行了总结，并为后续章节的学习奠定了基础。最后，作者在每章末尾的“Bibliographical and Historical Remarks”中回顾了相关领域的文献和历史，这对于研究者来说是非常有价值的参考资料。他们鼓励读者提供遗漏的重要引用，以便在最终版本印刷前进行修正。 “Reinforcement Learning: An Introduction”是一部详尽的强化学习教程，不仅适合初学者入门，也对专业人士有着很高的参考价值。通过阅读此书，读者将能够深入理解强化学习的理论和实践，并掌握设计和实现强化学习算法的技能。

展开

xvi Summary of Notation

|S| number of elements in set S

t discrete time step

T, T (t) ﬁnal time step of an episode, or of the episode including time step t

action at time t

state at time t, typically due, stochastically, to S

t−1

and A

t−1

reward at time t, typically due, stochastically, to S

t−1

and A

t−1

π policy, decision-making rule

π(s) action taken in state s under deterministic policy π

π(a|s) probability of taking action a in state s under stochastic policy π

π(a|s, θ) probability of taking action a in state s given parameter θ

return (cumulative discounted reward) following time t (Section 3.3)

t:h

ﬂat return (uncorrected, undiscounted) from t + 1 to h (Section 5.8)

λs

λ-return, corrected by estimated state values (Section 12.1)

λa

λ-return, corrected by estimated action values (Section 12.1)

λs

t:h

truncated, corrected λ-return, with state values (Section 12.3)

λa

t:h

truncated, corrected λ-return, with action values (Section 12.3)

p(s

, r|s, a) probability of transition to state s

with reward r, from state s and action a

p(s

|s, a) probability of transition to state s

, from state s taking action a

r(s, a, s

) expected immediate reward on transition from s to s

under action a

(s) value of state s under policy π (expected return)

∗

(s) value of state s under the optimal policy

(s, a) value of taking action a in state s under policy π

∗

(s, a) value of taking action a in state s under the optimal policy

V, V

array estimates of state-value function v

or v

∗

Q, Q

array estimates of action-value function q

or q

∗

temporal-diﬀerence error at t (a random variable) (Section 6.1)

w, w

d-vector of weights underlying an approximate value function

, w

t,i

ith component of learnable weight vector

d dimensionality—the number of components of w

alternate dimensionality—the number of components of θ

m number of 1s in a sparse binary feature vector

ˆv(s,w) approximate value of state s given weight vector w

(s) alternate notation for ˆv(s,w)

ˆq(s, a, w) approximate value of state–action pair s, a given weight vector w

x(s) vector of features visible when in state s

x(s, a) vector of features visible when in state s taking action a

(s), x

(s, a) ith component of vector x(s) or x(s, a)

shorthand for x(S

) or x(S

, A

)

x inner product of vectors, w

; e.g., ˆv(s,w)

= w

x(s)

µ(s) onpolicy distribution over states (Section 9.2)

µ |S|-vector of the µ(s)

kxk

µ-weighted norm of any vector x(s), i.e.,

µ(s)x(s)

(Section 11.4)

v, v

secondary d-vector of weights, used to learn w (Chapter 11)

d-vector of eligibility traces at time t (Chapter 12)

Chapter 1

Introduction

The idea that we learn by interacting with our environment is probably the ﬁrst to occur to us when

we think about the nature of learning. When an infant plays, waves its arms, or looks about, it has no

explicit teacher, but it does have a direct sensorimotor connection to its environment. Exercising this

connection produces a wealth of information about cause and eﬀect, about the consequences of actions,

and about what to do in order to achieve goals. Throughout our lives, such interactions are undoubtedly

a major source of knowledge about our environment and ourselves. Whether we are learning to drive a

car or to hold a conversation, we are acutely aware of how our environment responds to what we do, and

we seek to inﬂuence what happens through our behavior. Learning from interaction is a foundational

idea underlying nearly all theories of learning and intelligence.

In this book we explore a computational approach to learning from interaction. Rather than directly

theorizing about how people or animals learn, we explore idealized learning situations and evaluate the

eﬀectiveness of various learning methods. That is, we adopt the perspective of an artiﬁcial intelligence

researcher or engineer. We explore designs for machines that are eﬀective in solving learning problems of

scientiﬁc or economic interest, evaluating the designs through mathematical analysis or computational

experiments. The approach we explore, called reinforcement learning, is much more focused on goal-

directed learning from interaction than are other approaches to machine learning.

1.1 Reinforcement Learning

Reinforcement learning is learning what to do—how to map situations to actions—so as to maximize

a numerical reward signal. The learner is not told which actions to take, but instead must discover

which actions yield the most reward by trying them. In the most interesting and challenging cases,

actions may aﬀect not only the immediate reward but also the next situation and, through that, all

subsequent rewards. These two characteristics—trial-and-error search and delayed reward—are the two

most important distinguishing features of reinforcement learning.

Reinforcement learning, like many topics whose names end with “ing,” such as machine learning

and mountaineering, is simultaneously a problem, a class of solution methods that work well on the

problem, and the ﬁeld that studies this problems and its solution methods. It is convenient to use a

single name for all three things, but at the same time essential to keep the three conceptually separate.

In particular, the distinction between problems and solution methods is very important in reinforcement

learning; failing to make this distinction is the source of a many confusions.

We formalize the problem of reinforcement learning using ideas from dynamical systems theory,

speciﬁcally, as the optimal control of incompletely-known Markov decision processes. The details of this

2 CHAPTER 1. INTRODUCTION

formalization must wait until Chapter 3, but the basic idea is simply to capture the most important

aspects of the real problem facing a learning agent interacting over time with its environment to achieve

a goal. A learning agent must be able to sense the state of its environment to some extent and must be

able to take actions that aﬀect the state. The agent also must have a goal or goals relating to the state of

the environment. Markov decision processes are intended to include just these three aspects—sensation,

action, and goal—in their simplest possible forms without trivializing any of them. Any method that

is well suited to solving such problems we consider to be a reinforcement learning method.

Reinforcement learning is diﬀerent from supervised learning, the kind of learning studied in most

current research in the ﬁeld of machine learning. Supervised learning is learning from a training set

of labeled examples provided by a knowledgable external supervisor. Each example is a description of

a situation together with a speciﬁcation—the label—of the correct action the system should take to

that situation, which is often to identify a category to which the situation belongs. The object of this

kind of learning is for the system to extrapolate, or generalize, its responses so that it acts correctly

in situations not present in the training set. This is an important kind of learning, but alone it is

not adequate for learning from interaction. In interactive problems it is often impractical to obtain

examples of desired behavior that are both correct and representative of all the situations in which the

agent has to act. In uncharted territory—where one would expect learning to be most beneﬁcial—an

agent must be able to learn from its own experience.

Reinforcement learning is also diﬀerent from what machine learning researchers call unsupervised

learning, which is typically about ﬁnding structure hidden in collections of unlabeled data. The terms

supervised learning and unsupervised learning would seem to exhaustively classify machine learning

paradigms, but they do not. Although one might be tempted to think of reinforcement learning as a

kind of unsupervised learning because it does not rely on examples of correct behavior, reinforcement

learning is trying to maximize a reward signal instead of trying to ﬁnd hidden structure. Uncovering

structure in an agent’s experience can certainly be useful in reinforcement learning, but by itself does

not address the reinforcement learning problem of maximizing a reward signal. We therefore consider

reinforcement learning to be a third machine learning paradigm, alongside supervised learning and

unsupervised learning and perhaps other paradigms as well.

One of the challenges that arise in reinforcement learning, and not in other kinds of learning, is the

trade-oﬀ between exploration and exploitation. To obtain a lot of reward, a reinforcement learning

agent must prefer actions that it has tried in the past and found to be eﬀective in producing reward.

But to discover such actions, it has to try actions that it has not selected before. The agent has to

exploit what it has already experienced in order to obtain reward, but it also has to explore in order to

make better action selections in the future. The dilemma is that neither exploration nor exploitation

can be pursued exclusively without failing at the task. The agent must try a variety of actions and

progressively favor those that appear to be best. On a stochastic task, each action must be tried many

times to gain a reliable estimate of its expected reward. The exploration–exploitation dilemma has been

intensively studied by mathematicians for many decades, yet remains unresolved. For now, we simply

note that the entire issue of balancing exploration and exploitation does not even arise in supervised

and unsupervised learning, at least in their purest forms.

Another key feature of reinforcement learning is that it explicitly considers the whole problem of a

goal-directed agent interacting with an uncertain environment. This is in contrast to many approaches

that consider subproblems without addressing how they might ﬁt into a larger picture. For example, we

have mentioned that much of machine learning research is concerned with supervised learning without

explicitly specifying how such an ability would ﬁnally be useful. Other researchers have developed

theories of planning with general goals, but without considering planning’s role in real-time decision

making, or the question of where the predictive models necessary for planning would come from. Al-

though these approaches have yielded many useful results, their focus on isolated subproblems is a

signiﬁcant limitation.

剩余444页未读，继续阅读

身份认证购VIP最低享 7 折!

30元优惠券

宣爷

粉丝: 0

加强学习：入门与进阶

《Reinforcement Learning: An Introduction》第二版正式封面发布

2018年强化学习经典教材：《Reinforcement Learning: An Introduction》第二版

强化学习入门经典：Reinforcement Learning_An Introduction

Reinforcement Learning: An Introduction

reinforcement learning: an introduction

Reinforcement Learning：An Introduction

Reinforcement learning: An introduction

reinforcement learning：an introduction代码

Reinforcement Learning：An Introduction PDF文档+源代码

Reinforcement Learning：An Introduction.pdf

最新资源