0% found this document useful (0 votes)

19 views31 pages

Lecture 1 - Map Reduce

notes

Uploaded by

poornank05

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPT, PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

19 views31 pages

Lecture 1 - Map Reduce

notes

Uploaded by

poornank05

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PPT, PDF, TXT or read online on Scribd

You are on page 1/ 31

MapReduce

Single-node architecture

CPU
Machine Learning, Statistics

Memory

“Classical” Data Mining

Disk
Commodity Clusters
 Web data sets can be very large
 Tens to hundreds of terabytes
 Cannot mine on a single server (why?)
 Standard architecture emerging:
 Cluster of commodity Linux nodes
 Gigabit ethernet interconnect
 How to organize computations on this
architecture?
 Mask issues such as hardware failure
Cluster Architecture
2-10 Gbps backbone between racks
1 Gbps between Switch
any pair of nodes
in a rack
Switch Switch

CPU CPU CPU CPU

Mem … Mem Mem … Mem

Disk Disk Disk Disk

Each rack contains 16-64 nodes

Stable storage
 First order problem: if nodes can fail,
how can we store data persistently?
 Answer: Distributed File System
 Provides global file namespace
 Google GFS; Hadoop HDFS; Kosmix KFS
 Typical usage pattern
 Huge files (100s of GB to TB)
 Data is rarely updated in place
 Reads and appends are common
Distributed File System
Distributed File System
 Chunk Servers
 File is split into contiguous chunks
 Typically each chunk is 16-64MB
 Each chunk replicated (usually 2x or 3x)
 Try to keep replicas in different racks
 Master node
 a.k.a. Name Nodes in HDFS
 Stores metadata
 Might be replicated
 Client library for file access
 Talks to master to find chunk servers
 Connects directly to chunkservers to access data
Warm up: Word Count
 We have a large file of words, one
word to a line
 Count the number of times each
distinct word appears in the file
 Sample application: analyze web
server logs to find popular URLs
Word Count (2)
 Case 1: Entire file fits in memory
 Case 2: File too large for mem, but all
<word, count> pairs fit in mem
 Case 3: File on disk, too many
distinct words to fit in memory
 sort datafile | uniq –c
Word Count (3)
 To make it slightly harder, suppose
we have a large corpus of documents
 Count the number of times each
distinct word occurs in the corpus
 words(docs/*) | sort | uniq -c
 where words takes a file and outputs the
words in it, one to a line
 The above captures the essence of
MapReduce
 Great thing is it is naturally parallelizable
MapReduce: The Map Step
Input Intermediate
key-value pairs key-value pairs

k v
map
k v
k v
map
k v
k v

… …

k v k v
MapReduce: The Reduce Step
Output
Intermediate Key-value groups key-value pairs
key-value pairs
reduce
k v k v v v k v
reduce
k v k v v k v
group

k v

… … …

k v k v k v
MapReduce
 Input: a set of key/value pairs
 User supplies two functions:
 map(k,v)  list(k1,v1)
 reduce(k1, list(v1))  v2
 (k1,v1) is an intermediate key/value
pair
 Output is the set of (k1,v2) pairs
Map Reduce for word counting
Word Count using MapReduce
map(key, value):
// key: document name; value: text of document
for each word w in value:
emit(w, 1)

reduce(key, values):
// key: a word; value: an iterator over counts
result = 0
for each count v in values:
result += v
emit(result)
Distributed Execution Overview
User
Program

fork fork fork

assign Master
assign
map reduce
Input Data Worker
write Output
local Worker File 0
Split 0 read
write
Split 1 Worker
Split 2 Output
Worker File 1
Worker remote
read,
sort
Data flow
 Input, final output are stored on a
distributed file system
 Scheduler tries to schedule map tasks
“close” to physical storage location of
input data
 Intermediate results are stored on
local FS of map and reduce workers
 Output is often input to another map
reduce task
Coordination
 Master data structures
 Task status: (idle, in-progress, completed)
 Idle tasks get scheduled as workers
become available
 When a map task completes, it sends the
master the location and sizes of its R
intermediate files, one for each reducer
 Master pushes this info to reducers
 Master pings workers periodically to
detect failures
Failures
 Map worker failure
 Map tasks completed or in-progress at
worker are reset to idle
 Reduce workers are notified when task is
rescheduled on another worker
 Reduce worker failure
 Only in-progress tasks are reset to idle
 Master failure
 MapReduce task is aborted and client is
notified
How many Map and Reduce jobs?
 M map tasks, R reduce tasks
 Rule of thumb:
 Make M and R much larger than the
number of nodes in cluster
 One DFS chunk per map is common
 Improves dynamic load balancing and
speeds recovery from worker failure
 Usually R is smaller than M, because
output is spread across R files
Combiners
 Often a map task will produce many
pairs of the form (k,v1), (k,v2), … for
the same key k
 E.g., popular words in Word Count
 Can save network time by pre-
aggregating at mapper
 combine(k1, list(v1))  v2
 Usually same as reduce function
 Works only if reduce function is
commutative and associative
Partition Function
 Inputs to map tasks are created by
contiguous splits of input file
 For reduce, we need to ensure that
records with the same intermediate
key end up at the same worker
 System uses a default partition
function e.g., hash(key) mod R
 Sometimes useful to override
 E.g., hash(hostname(URL)) mod R
ensures URLs from a host end up in the
same output file
Map function
Reduce function
Program
Exercise 1: Host size
 Suppose we have a large web corpus
 Let’s look at the metadata file
 Lines of the form (URL, size, date, …)
 For each host, find the total number
of bytes
 i.e., the sum of the page sizes for all
URLs from that host
Exercise 2: Distributed Grep
 Find all occurrences of the given
pattern in a very large set of files
Exercise 3: Graph reversal
 Given a directed graph as an
adjacency list:
src1: dest11, dest12, …
src2: dest21, dest22, …

 Construct the graph in which all the

links are reversed
Exercise 4: Frequent Pairs
 Given a large set of market baskets,
find all frequent pairs
 Remember definitions from Association
Rules lectures
Implementations
 Google
 Not available outside Google
 Hadoop
 An open-source implementation in Java
 Uses HDFS for stable storage
 Download: http://lucene.apache.org/hadoop/
 Aster Data
 Cluster-optimized SQL Database that
also implements MapReduce
 Made available free of charge for this
class
Reading
 Jeffrey Dean and Sanjay Ghemawat,
MapReduce: Simplified Data Processing
on Large Clusters
http://labs.google.com/papers/mapreduce.html

 Sanjay Ghemawat, Howard Gobioff, and Shun-

Tak Leung, The Google File System
http://labs.google.com/papers/gfs.html

Laporan Magang
No ratings yet
Laporan Magang
49 pages
Parallel & Distributed Computing
100% (1)
Parallel & Distributed Computing
52 pages
Introduction To Map Reduce
No ratings yet
Introduction To Map Reduce
50 pages
Introduction To MapReduce
No ratings yet
Introduction To MapReduce
17 pages
User Guide For COFEE v112
100% (2)
User Guide For COFEE v112
46 pages
Blessing Project
No ratings yet
Blessing Project
38 pages
TE040 Iprocurement Test Script On Oracle Iprocurement
100% (1)
TE040 Iprocurement Test Script On Oracle Iprocurement
17 pages
Sap Bw4hana Content Add On en
No ratings yet
Sap Bw4hana Content Add On en
2,602 pages
Information Systems Vs Information Technology
No ratings yet
Information Systems Vs Information Technology
3 pages
Unit V Big Data Analytics
No ratings yet
Unit V Big Data Analytics
47 pages
AAAI2011 Tutorial Slides
No ratings yet
AAAI2011 Tutorial Slides
213 pages
Map Reduce
No ratings yet
Map Reduce
69 pages
Primary Key, Candidate Key, Alternate Key, Foreign Key, Composite Key
100% (1)
Primary Key, Candidate Key, Alternate Key, Foreign Key, Composite Key
7 pages
03 Firstmrjob Invertedindexconstruction 141206231216 Conversion Gate01 PDF
No ratings yet
03 Firstmrjob Invertedindexconstruction 141206231216 Conversion Gate01 PDF
54 pages
Parlab Parallel Boot Camp: Cloud Computing With Mapreduce and Hadoop
No ratings yet
Parlab Parallel Boot Camp: Cloud Computing With Mapreduce and Hadoop
55 pages
Parlab Parallel Boot Camp: Cloud Computing With Mapreduce and Hadoop
No ratings yet
Parlab Parallel Boot Camp: Cloud Computing With Mapreduce and Hadoop
53 pages
Distributed and Cloud Computing
No ratings yet
Distributed and Cloud Computing
58 pages
Week 02
No ratings yet
Week 02
115 pages
Module2 C MapReduceParadigm
No ratings yet
Module2 C MapReduceParadigm
74 pages
CAIM: Cerca I Anàlisi D'informació Massiva: FIB, Grau en Enginyeria Informàtica
No ratings yet
CAIM: Cerca I Anàlisi D'informació Massiva: FIB, Grau en Enginyeria Informàtica
65 pages
Chapter Five Hadoop Mapreduce & HDFS
No ratings yet
Chapter Five Hadoop Mapreduce & HDFS
44 pages
Parlab Parallel Boot Camp Cloud Computing With Mapreduce and Hadoop
No ratings yet
Parlab Parallel Boot Camp Cloud Computing With Mapreduce and Hadoop
49 pages
Chapter 9 - Processing Big Data With Mapreduce
No ratings yet
Chapter 9 - Processing Big Data With Mapreduce
157 pages
TM2 ch02 Mapreduce
No ratings yet
TM2 ch02 Mapreduce
51 pages
Map Reduce
No ratings yet
Map Reduce
42 pages
Big Data Computing
No ratings yet
Big Data Computing
36 pages
Da Unit 5 Data Analytics
No ratings yet
Da Unit 5 Data Analytics
43 pages
Ch02a Mapreduce
No ratings yet
Ch02a Mapreduce
53 pages
Take A Close Look At: Ma Ed
No ratings yet
Take A Close Look At: Ma Ed
42 pages
Introduction To: Ma Ed
No ratings yet
Introduction To: Ma Ed
42 pages
DEVELOPMENT OF AN INTERACTIVE ANDROIDdsds
No ratings yet
DEVELOPMENT OF AN INTERACTIVE ANDROIDdsds
89 pages
Map Reduce
No ratings yet
Map Reduce
44 pages
Map Reduce Notes and Learning
No ratings yet
Map Reduce Notes and Learning
48 pages
Mapreduce and Hadoop Distributed File System
No ratings yet
Mapreduce and Hadoop Distributed File System
45 pages
Map Reduce
No ratings yet
Map Reduce
25 pages
Aleksandar Lazarevic Intrusion Detection A Survey
No ratings yet
Aleksandar Lazarevic Intrusion Detection A Survey
61 pages
CS 425 / ECE 428 Distributed Systems Fall 2016: Lecture 4: Mapreduce and Hadoop
No ratings yet
CS 425 / ECE 428 Distributed Systems Fall 2016: Lecture 4: Mapreduce and Hadoop
24 pages
Map Reduce: Simplified Processing On Large Clusters
No ratings yet
Map Reduce: Simplified Processing On Large Clusters
29 pages
Lookup Cache
100% (1)
Lookup Cache
2 pages
Lecture - 3
No ratings yet
Lecture - 3
25 pages
02 Hadoop
No ratings yet
02 Hadoop
117 pages
BDA Module 3
No ratings yet
BDA Module 3
66 pages
Mapreduce Model Principles
No ratings yet
Mapreduce Model Principles
65 pages
Bda 2
No ratings yet
Bda 2
35 pages
09b - MapReduce
No ratings yet
09b - MapReduce
44 pages
Ir MR 1
No ratings yet
Ir MR 1
34 pages
Chapter 4
No ratings yet
Chapter 4
53 pages
User Guide
No ratings yet
User Guide
97 pages
Lecture 4: Mapreduce and Hadoop: Indranil Gupta (Indy)
No ratings yet
Lecture 4: Mapreduce and Hadoop: Indranil Gupta (Indy)
37 pages
Chapter 4
No ratings yet
Chapter 4
71 pages
Introduction To Hadoop
No ratings yet
Introduction To Hadoop
37 pages
Unit 5 Lecture 5
No ratings yet
Unit 5 Lecture 5
21 pages
Big Data Analytics Module 3: Mapreduce Paradigm: Faculty Name: Ms. Varsha Sanap Dr. Vivek Singh
No ratings yet
Big Data Analytics Module 3: Mapreduce Paradigm: Faculty Name: Ms. Varsha Sanap Dr. Vivek Singh
36 pages
Map Reduce Architecture: Adapted From Lectures by
No ratings yet
Map Reduce Architecture: Adapted From Lectures by
37 pages
1s07 Map Reduce Presentation 2019
No ratings yet
1s07 Map Reduce Presentation 2019
43 pages
Topic 4 Sourcing and Collecting Data
No ratings yet
Topic 4 Sourcing and Collecting Data
32 pages
Unit 5 Big Data
No ratings yet
Unit 5 Big Data
48 pages
M4 06 MapReduce
No ratings yet
M4 06 MapReduce
28 pages
Problem-Solving Using Mapreduce/Hadoop
No ratings yet
Problem-Solving Using Mapreduce/Hadoop
22 pages
Unit IV Notes
No ratings yet
Unit IV Notes
25 pages
CS403 MIDTERM SOLVED MCQS by JUNAID
No ratings yet
CS403 MIDTERM SOLVED MCQS by JUNAID
28 pages
Chapter 6
No ratings yet
Chapter 6
57 pages
Map Reduce
No ratings yet
Map Reduce
28 pages
MapReduce Introduction
No ratings yet
MapReduce Introduction
34 pages
Map Reduced B Seminar
No ratings yet
Map Reduced B Seminar
17 pages
CS 425 / ECE 428 Distributed Systems Fall 2014: Lecture 3: Mapreduce and Hadoop
No ratings yet
CS 425 / ECE 428 Distributed Systems Fall 2014: Lecture 3: Mapreduce and Hadoop
24 pages
Parallel Programming, Mapreduce Model: Unit Ii
No ratings yet
Parallel Programming, Mapreduce Model: Unit Ii
47 pages
L06 Map Reduce
No ratings yet
L06 Map Reduce
37 pages
Paper Map Reduce
No ratings yet
Paper Map Reduce
16 pages
Narvar Connect - Magento 2.x Community Extension
No ratings yet
Narvar Connect - Magento 2.x Community Extension
10 pages
UNIT III Notes
No ratings yet
UNIT III Notes
24 pages
Labanswers
No ratings yet
Labanswers
11 pages
Learn Scrum With Jira Software - Atlassian
No ratings yet
Learn Scrum With Jira Software - Atlassian
12 pages
Whitney
No ratings yet
Whitney
19 pages
Ai in Vulnerability Management
No ratings yet
Ai in Vulnerability Management
16 pages
NTT DATA AI-DX Agent Powered by Microsoft Copilot Studio Fact Sheet
No ratings yet
NTT DATA AI-DX Agent Powered by Microsoft Copilot Studio Fact Sheet
5 pages
Hathor CompanyPresentation v0.4 ForDataCustodian
No ratings yet
Hathor CompanyPresentation v0.4 ForDataCustodian
9 pages
SF Git Cheatsheet
No ratings yet
SF Git Cheatsheet
2 pages
Sab Sop
No ratings yet
Sab Sop
6 pages
Degree/ Certificate Institution Percentage/Cgpa Year: Acmegrade Internship
No ratings yet
Degree/ Certificate Institution Percentage/Cgpa Year: Acmegrade Internship
1 page
Sai Kumar CV
No ratings yet
Sai Kumar CV
3 pages
Programs
No ratings yet
Programs
3 pages
RFUMSV00 (TH) Wrong Non-Deductible Amount With Deferred Tax
No ratings yet
RFUMSV00 (TH) Wrong Non-Deductible Amount With Deferred Tax
2 pages
National Computer Center (NCC) E-Serbisyo (Philippine Government Service Portal) The National Computer Center (NCC) Was
No ratings yet
National Computer Center (NCC) E-Serbisyo (Philippine Government Service Portal) The National Computer Center (NCC) Was
2 pages
9.4 Analyzing Queries: 9.4.1 What Is Dynamic Query Analyzer
No ratings yet
9.4 Analyzing Queries: 9.4.1 What Is Dynamic Query Analyzer
2 pages
Teaching Cloud Computing Using Project-Based Learning: Linh B. Ngo
No ratings yet
Teaching Cloud Computing Using Project-Based Learning: Linh B. Ngo
1 page
BIT 2.1 May - August 2025 Teaching Timetable
No ratings yet
BIT 2.1 May - August 2025 Teaching Timetable
1 page
Advanced Supply Chain Management USA - Streamline Operations
No ratings yet
Advanced Supply Chain Management USA - Streamline Operations
3 pages
The Tech Interview Playbook: From DSA to System Design
From Everand
The Tech Interview Playbook: From DSA to System Design
Chinmoy Mukherjee
No ratings yet
DRBD-Cookbook: How to create your own cluster solution, without SAN or NAS!
From Everand
DRBD-Cookbook: How to create your own cluster solution, without SAN or NAS!
Joerg Christian Seubert
No ratings yet

Lecture 1 - Map Reduce

Uploaded by

Lecture 1 - Map Reduce

Uploaded by

MapReduce

“Classical” Data Mining

CPU CPU CPU CPU

Mem … Mem Mem … Mem

Disk Disk Disk Disk

Each rack contains 16-64 nodes

fork fork fork

 Construct the graph in which all the

 Sanjay Ghemawat, Howard Gobioff, and Shun-

You might also like