Big Data Technologies

Advanced MapReduce, Streaming, Compression, Scheduling and Hadoop ETL

PGCP-BDA

Hadoop Streaming

Hadoop Streaming lets external programs act as mappers and reducers by exchanging line-oriented key-value data through standard input and output.

compression codec

A compression codec encodes and decodes data; its ratio, CPU cost and splittability affect storage, transfer and mapper parallelism.

splittable compression

A splittable compressed format provides synchronization boundaries from which independent tasks can start decoding without reading the entire preceding.

speculative execution

Speculative execution launches a duplicate attempt for an unusually slow task and accepts the first successful output.

scheduler

The component that assigns runnable work to available resources while respecting capacity, locality, priorities and dependencies.

spill and merge

When a mapper buffer fills, sorted partitions spill to local files that are later merged before reducers fetch them.

job tuning

Job tuning adjusts split size, task count, memory, serialization, compression.

workflow orchestration

The scheduling and coordination of dependent data tasks, including retries, state tracking, alerts and recovery.

Hadoop ETL

An ETL pipeline that extracts source data, transforms it with Hadoop ecosystem tools and loads validated results into a target store.

failure recovery

Distributed failure recovery reconstructs state from replicas, lineage, logs or checkpoints and must prevent retried work from corrupting committed output.

Streaming, Compression and Tuning

Hadoop Streaming launches external mapper and reducer programs and exchanges records through standard input and output. Each process must write only protocol data to standard output and send diagnostics to standard error. Distributed cache mechanisms can ship supporting files. Exit status tells the framework whether an attempt succeeded.

Compression choice affects CPU cost, storage and parallelism. Gzip is common but an ordinary gzip stream is not splittable. Bzip2 can be split at higher CPU cost. Block-oriented container formats can combine compression with parallel reads. Compressing map output often saves shuffle time because intermediate data crosses both disk and network.

Tuning begins with measurement. Split size controls mapper count. Reducer count balances parallelism against scheduling and small-file overhead. Memory limits must cover sorting, user code and framework buffers. Skew may require a better partitioner or a two-stage aggregation. The scheduler shares cluster resources through queues, capacity or fairness policies while locality reduces data transfer.

Hadoop ETL Workflows

A Hadoop ETL flow extracts immutable source data, validates and normalizes it then writes curated output. Each stage should declare inputs, outputs, schema and partition rules. Orchestration represents dependencies and records state, retries and alerts. Rerunning a stage should not append duplicates or expose a partially written directory. Temporary output followed by an atomic commit supports safe publication. Row counts, rejected records and reconciliation totals provide correctness evidence beyond a successful process exit.

Scheduling and Recovery Details

FIFO scheduling is simple but lets a large job delay shorter work. Capacity scheduling divides cluster resources among queues and can lend unused capacity. Fair scheduling attempts to give active applications comparable shares over time. Priorities influence order but cannot create resources. Data locality preferences may wait briefly for a node containing an input block before accepting rack-local or remote execution.

Speculation addresses stragglers rather than failed tasks. It can waste resources when all tasks are slow for the same reason and can overload an already constrained system. External side effects are dangerous because two attempts may both perform them. Hadoop output committers isolate attempt output and publish only the accepted result.

ETL recovery needs deterministic boundaries. A source snapshot or high-water mark defines what belongs to a run. Checkpoints record completed stages. Reprocessing should use the same transformation version or record that a new version produced the result. Late records may require reopening a partition or issuing a correction. Scheduling success means all dependencies completed with validated outputs, not merely that commands exited with zero.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.