Big Data Technologies

MapReduce Paradigm, Execution Framework and Job Lifecycle

PGCP-BDA

MapReduce

MapReduce processes partitioned input with parallel map functions, groups intermediate key-value pairs by key through shuffle and sort and applies reducers.

mapper

A mapper transforms input records into zero or more intermediate key-value pairs without sharing mutable state with other map tasks.

reducer

A reducer receives one key and its grouped values after shuffle and sort and emits final or intermediate output records.

input split

An input split is a logical range assigned to one map task; an InputFormat determines splits and a RecordReader converts each split into input records.

record reader

A MapReduce component that converts an input split into key-value records supplied to the mapper.

shuffle and sort

Shuffle transfers mapper partitions to reducers, merges spill files and sorts/group records by key before reduce processing.

MapReduce job lifecycle

A MapReduce job validates input, creates splits, runs map tasks, shuffles grouped records, runs reducers and commits output.

YARN application

A YARN application has an ApplicationMaster that negotiates containers from the ResourceManager and coordinates work performed through NodeManagers.

task attempt

A task attempt is one execution of a logical map or reduce task; failed or speculative attempts may coexist, but output commit selects a successful result.

data locality

Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.

Hadoop Ecosystem

What is Hadoop?

Hadoop is an open-source framework for storing and processing Big Data in a distributed manner.

Core Components:

+--------------------------------------------------+
|                HADOOP ECOSYSTEM                  |
|                                                  |
|  HDFS     →   Distributed Storage               |
|  YARN     →   Resource Management               |
|  MapReduce →  Distributed Processing            |
+--------------------------------------------------+

YARN — Yet Another Resource Negotiator

  • Manages cluster resources (CPU, memory) across jobs
  • ResourceManager — cluster-level master; allocates resources
  • NodeManager — node-level; executes tasks, manages containers

MapReduce

MapReduce is Hadoop's processing model — splits computation into two phases:

Input Data → MAP → (Intermediate key-value pairs) → SHUFFLE & SORT → REDUCE → Output

Word Count Example:
Input: "Hello World Hello"

MAP phase:
  "Hello" → 1
  "World" → 1
  "Hello" → 1

SHUFFLE (group by key):
  "Hello" → [1, 1]
  "World" → [1]

REDUCE phase:
  "Hello" → 2
  "World" → 1

MapReduce Code Concept:

# Mapper
def map(text_line):
    for word in text_line.split():
        emit(word, 1)

# Reducer
def reduce(word, counts):
    emit(word, sum(counts))

Hadoop Ecosystem Components

ComponentFunction
HDFSDistributed file storage
YARNResource management and scheduling
MapReduceBatch processing framework
HiveSQL-like interface for HDFS data
HBaseNoSQL database on HDFS
PigScripting language for data transformation
SqoopImport/export between RDBMS and HDFS
FlumeIngestion of streaming log data into HDFS
OozieWorkflow scheduler for Hadoop jobs
ZooKeeperDistributed coordination service

MapReduce Execution

A MapReduce job divides input into logical splits. Each map task reads records through an InputFormat and RecordReader then emits intermediate key-value pairs. The framework partitions map output by reducer, sorts it by key and makes it available through shuffle. Each reducer processes one key with its grouped values and writes final records through an OutputFormat.

Map output first accumulates in memory. Thresholds trigger sorted spills to local files. Spill files are merged before reducers fetch their partitions. A combiner may reduce transferable records but must be valid when invoked zero, one or many times. The partitioner must send equal keys to the same reducer. Skewed keys create a slow reducer even when total data volume appears balanced.

Job Lifecycle and Failure

The client prepares configuration and submits job resources. A resource manager allocates containers while an application coordinator requests and monitors tasks. Scheduling prefers data-local work when possible. Failed attempts may be retried on another worker because input and intermediate data can be read again. Completed task output is committed through controlled protocols so retries do not create partial final results. Speculative execution duplicates an unusually slow attempt and accepts the first successful result; it is safe only when work and output commit are retryable.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.