Big Data Technologies
MapReduce Paradigm, Execution Framework and Job Lifecycle
PGCP-BDA
MapReduce
MapReduce processes partitioned input with parallel map functions, groups intermediate key-value pairs by key through shuffle and sort and applies reducers.
mapper
A mapper transforms input records into zero or more intermediate key-value pairs without sharing mutable state with other map tasks.
reducer
A reducer receives one key and its grouped values after shuffle and sort and emits final or intermediate output records.
input split
An input split is a logical range assigned to one map task; an InputFormat determines splits and a RecordReader converts each split into input records.
record reader
A MapReduce component that converts an input split into key-value records supplied to the mapper.
shuffle and sort
Shuffle transfers mapper partitions to reducers, merges spill files and sorts/group records by key before reduce processing.
MapReduce job lifecycle
A MapReduce job validates input, creates splits, runs map tasks, shuffles grouped records, runs reducers and commits output.
YARN application
A YARN application has an ApplicationMaster that negotiates containers from the ResourceManager and coordinates work performed through NodeManagers.
task attempt
A task attempt is one execution of a logical map or reduce task; failed or speculative attempts may coexist, but output commit selects a successful result.
data locality
Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.
Hadoop Ecosystem
What is Hadoop?
Hadoop is an open-source framework for storing and processing Big Data in a distributed manner.
Core Components:
+--------------------------------------------------+
| HADOOP ECOSYSTEM |
| |
| HDFS → Distributed Storage |
| YARN → Resource Management |
| MapReduce → Distributed Processing |
+--------------------------------------------------+
YARN — Yet Another Resource Negotiator
- Manages cluster resources (CPU, memory) across jobs
- ResourceManager — cluster-level master; allocates resources
- NodeManager — node-level; executes tasks, manages containers
MapReduce
MapReduce is Hadoop's processing model — splits computation into two phases:
Input Data → MAP → (Intermediate key-value pairs) → SHUFFLE & SORT → REDUCE → Output
Word Count Example:
Input: "Hello World Hello"
MAP phase:
"Hello" → 1
"World" → 1
"Hello" → 1
SHUFFLE (group by key):
"Hello" → [1, 1]
"World" → [1]
REDUCE phase:
"Hello" → 2
"World" → 1
MapReduce Code Concept:
# Mapper
def map(text_line):
for word in text_line.split():
emit(word, 1)
# Reducer
def reduce(word, counts):
emit(word, sum(counts))
Hadoop Ecosystem Components
| Component | Function |
|---|---|
| HDFS | Distributed file storage |
| YARN | Resource management and scheduling |
| MapReduce | Batch processing framework |
| Hive | SQL-like interface for HDFS data |
| HBase | NoSQL database on HDFS |
| Pig | Scripting language for data transformation |
| Sqoop | Import/export between RDBMS and HDFS |
| Flume | Ingestion of streaming log data into HDFS |
| Oozie | Workflow scheduler for Hadoop jobs |
| ZooKeeper | Distributed coordination service |
MapReduce Execution
A MapReduce job divides input into logical splits. Each map task reads records through an InputFormat and RecordReader then emits intermediate key-value pairs. The framework partitions map output by reducer, sorts it by key and makes it available through shuffle. Each reducer processes one key with its grouped values and writes final records through an OutputFormat.
Map output first accumulates in memory. Thresholds trigger sorted spills to local files. Spill files are merged before reducers fetch their partitions. A combiner may reduce transferable records but must be valid when invoked zero, one or many times. The partitioner must send equal keys to the same reducer. Skewed keys create a slow reducer even when total data volume appears balanced.
Job Lifecycle and Failure
The client prepares configuration and submits job resources. A resource manager allocates containers while an application coordinator requests and monitors tasks. Scheduling prefers data-local work when possible. Failed attempts may be retried on another worker because input and intermediate data can be read again. Completed task output is committed through controlled protocols so retries do not create partial final results. Speculative execution duplicates an unusually slow attempt and accepts the first successful result; it is safe only when work and output commit are retryable.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.