Hadoop生态系统与大数据处理实战

5星 · 超过95%的资源需积分: 9 64 浏览量更新于2024-07-23 收藏 8.46MB PDF 举报

"Hadoop权威指南(第三版英文)" 是一本由Tom White编写的关于Hadoop技术的详尽指南。本书涵盖了Hadoop的核心组件，包括Hadoop分布式文件系统(HDFS)和MapReduce，以及Hadoop生态系统中的其他关键工具，如Sqoop、Pig、Hive和HBase等。在Hadoop的MapReduce部分，作者介绍了如何使用MapReduce进行分布式计算。通过一个天气数据集的例子，读者可以理解数据格式和如何使用Unix工具进行分析。MapReduce的基本原理被详细阐述，包括Map和Reduce阶段，以及如何编写Java MapReduce程序。此外，还讨论了如何扩展MapReduce以适应大规模数据，并介绍了数据流的走向以及Combiner函数的角色。此外，书中还介绍了使用Hadoop Streaming执行MapReduce任务，支持使用Ruby和Python等脚本语言。在Hadoop Distributed File System (HDFS)章节，Tom White深入解析了HDFS的设计理念，包括其概念、架构和操作流程。他解释了HDFS如何为大数据存储提供高容错性和可伸缩性，讨论了数据块、NameNode和DataNode的角色，以及如何处理数据完整性、故障恢复和数据压缩。书中的其他部分涉及了如何构建和管理Hadoop集群，无论是本地部署还是在云端运行。 Sqoop的使用使得从关系型数据库导入数据到HDFS变得简单，而Pig查询语言则提供了处理大规模数据的高级抽象。Hadoop的数据仓库系统Hive被介绍为用于数据分析的工具，它允许用户使用SQL-like语法来查询和处理数据集。对于结构化和半结构化数据的处理，HBase作为NoSQL数据库被详细讲解，而ZooKeeper作为分布式协调服务，对于构建可靠的分布式系统至关重要。这本书是Hadoop学习者的宝贵资源，涵盖了从基础到高级的各种主题，旨在帮助读者理解和应用Hadoop及其生态系统中的工具，以解决大数据的存储、处理和分析问题。

Foreword

Hadoop got its start in Nutch. A few of us were attempting to build an open source

web search engine and having trouble managing computations running on even a

handful of computers. Once Google published its GFS and MapReduce papers, the

route became clear. They’d devised systems to solve precisely the problems we were

having with Nutch. So we started, two of us, half-time, to try to re-create these systems

as a part of Nutch.

We managed to get Nutch limping along on 20 machines, but it soon became clear that

to handle the Web’s massive scale, we’d need to run it on thousands of machines and,

moreover, that the job was bigger than two half-time developers could handle.

Around that time, Yahoo! got interested, and quickly put together a team that I joined.

We split off the distributed computing part of Nutch, naming it Hadoop. With the help

of Yahoo!, Hadoop soon grew into a technology that could truly scale to the Web.

In 2006, Tom White started contributing to Hadoop. I already knew Tom through an

excellent article he’d written about Nutch, so I knew he could present complex ideas

in clear prose. I soon learned that he could also develop software that was as pleasant

to read as his prose.

From the beginning, Tom’s contributions to Hadoop showed his concern for users and

for the project. Unlike most open source contributors, Tom is not primarily interested

in tweaking the system to better meet his own needs, but rather in making it easier for

anyone to use.

Initially, Tom specialized in making Hadoop run well on Amazon’s EC2 and S3 serv-

ices. Then he moved on to tackle a wide variety of problems, including improving the

MapReduce APIs, enhancing the website, and devising an object serialization frame-

work. In all cases, Tom presented his ideas precisely. In short order, Tom earned the

role of Hadoop committer and soon thereafter became a member of the Hadoop Project

Management Committee.

Tom is now a respected senior member of the Hadoop developer community. Though

he’s an expert in many technical corners of the project, his specialty is making Hadoop

easier to use and understand.

xiii

Preface

Martin Gardner, the mathematics and science writer, once said in an interview:

Beyond calculus, I am lost. That was the secret of my column’s success. It took me so

long to understand what I was writing about that I knew how to write in a way most

readers would understand.

In many ways, this is how I feel about Hadoop. Its inner workings are complex, resting

as they do on a mixture of distributed systems theory, practical engineering, and com-

mon sense. And to the uninitiated, Hadoop can appear alien.

But it doesn’t need to be like this. Stripped to its core, the tools that Hadoop provides

for building distributed systems—for data storage, data analysis, and coordination—

are simple. If there’s a common theme, it is about raising the level of abstraction—to

create building blocks for programmers who just happen to have lots of data to store,

or lots of data to analyze, or lots of machines to coordinate, and who don’t have the

time, the skill, or the inclination to become distributed systems experts to build the

infrastructure to handle it.

With such a simple and generally applicable feature set, it seemed obvious to me when

I started using it that Hadoop deserved to be widely used. However, at the time (in

early 2006), setting up, configuring, and writing programs to use Hadoop was an art.

Things have certainly improved since then: there is more documentation, there are

more examples, and there are thriving mailing lists to go to when you have questions.

And yet the biggest hurdle for newcomers is understanding what this technology is

capable of, where it excels, and how to use it. That is why I wrote this book.

The Apache Hadoop community has come a long way. Over the course of three years,

the Hadoop project has blossomed and spun off half a dozen subprojects. In this time,

the software has made great leaps in performance, reliability, scalability, and manage-

ability. To gain even wider adoption, however, I believe we need to make Hadoop even

easier to use. This will involve writing more tools; integrating with more systems; and

1. “The science of fun,” Alex Bellos, The Guardian, May 31, 2008, http://www.guardian.co.uk/science/

2008/may/31/maths.science.

writing new, improved APIs. I’m looking forward to being a part of this, and I hope

this book will encourage and enable others to do so, too.

Administrative Notes

During discussion of a particular Java class in the text, I often omit its package name,

to reduce clutter. If you need to know which package a class is in, you can easily look

it up in Hadoop’s Java API documentation for the relevant subproject, linked to from

the Apache Hadoop home page at http://hadoop.apache.org/. Or if you’re using an IDE,

it can help using its auto-complete mechanism.

Similarly, although it deviates from usual style guidelines, program listings that import

multiple classes from the same package may use the asterisk wildcard character to save

space (for example: import org.apache.hadoop.io.*).

The sample programs in this book are available for download from the website that

accompanies this book: http://www.hadoopbook.com/. You will also find instructions

there for obtaining the datasets that are used in examples throughout the book, as well

as further notes for running the programs in the book, and links to updates, additional

resources, and my blog.

What’s in This Book?

The rest of this book is organized as follows. Chapter 1 emphasizes the need for Hadoop

and sketches the history of the project. Chapter 2 provides an introduction to

MapReduce. Chapter 3 looks at Hadoop filesystems, and in particular HDFS, in depth.

Chapter 4 covers the fundamentals of I/O in Hadoop: data integrity, compression,

serialization, and file-based data structures.

The next four chapters cover MapReduce in depth. Chapter 5 goes through the practical

steps needed to develop a MapReduce application. Chapter 6 looks at how MapReduce

is implemented in Hadoop, from the point of view of a user. Chapter 7 is about the

MapReduce programming model, and the various data formats that MapReduce can

work with. Chapter 8 is on advanced MapReduce topics, including sorting and joining

data.

Chapters 9 and 10 are for Hadoop administrators, and describe how to set up and

maintain a Hadoop cluster running HDFS and MapReduce.

Later chapters are dedicated to projects that build on Hadoop or are related to it.

Chapters 11 and 12 present Pig and Hive, which are analytics platforms built on HDFS

and MapReduce, whereas Chapters 13, 14, and 15 cover HBase, ZooKeeper, and

Sqoop, respectively.

Finally, Chapter 16 is a collection of case studies contributed by members of the Apache

Hadoop community.

xvi | Preface

剩余646页未读，继续阅读

yangfan168

粉丝: 1
资源: 6

Hadoop生态系统与大数据处理实战

Hadoop权威指南[第三版:英文版]

Hadoop权威指南第三版(英文版)

hadoop 权威指南（第三版）英文版

hadoop权威指南 第三版 英文版

Hadoop权威指南第三版英文版详解

Hadoop权威指南第三版英文原版

Hadoop权威指南第三版英文原版详解

Hadoop权威指南第三版英文版：入门到精通

Hadoop权威指南第三版英文版：深入探索大数据处理

Hadoop权威指南第三版

最新资源

hadoop权威指南第三版英文版