Big Data Technologies

Hadoop Evolution, Ecosystem, Architecture and Operating Modes

PGCP-BDA

Hadoop

Apache Hadoop is a framework for distributed storage and batch processing built around HDFS, YARN, MapReduce and supporting libraries.

Hadoop ecosystem

The Hadoop ecosystem combines storage, resource management, processing, SQL, ingestion, coordination and workflow tools around distributed datasets.

HDFS MapReduce and YARN

HDFS stores distributed files, MapReduce defines batch data processing and YARN allocates cluster resources and runs application containers.

Hadoop operating modes

Standalone mode uses one JVM and local storage, pseudo-distributed mode runs daemons on one host and fully distributed mode spreads daemons across a.

commodity cluster

A commodity cluster uses replaceable general-purpose servers and software replication or recomputation instead of relying only on specialized.

Hadoop configuration

Hadoop configuration is assembled from site XML files and runtime properties whose effective values must be consistent across clients and relevant daemons.

data locality

Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.

batch processing

Batch processing consumes a bounded dataset, produces a completed result and optimizes throughput rather than per-record response latency.

Hadoop limitation

Hadoop MapReduce is inefficient for low-latency queries, iterative algorithms, many small files and workloads requiring frequent in-place updates.

ecosystem tools

Interoperating platform components that provide ingestion, storage, processing, querying, coordination, security and workflow services.

Batch vs Stream Processing

Batch Processing

Batch processing processes data in discrete chunks at scheduled intervals.

12:00 AM daily → Collect all transactions → Process → Load to DWH

Tools: Apache Hadoop MapReduce, Apache Spark (batch mode), AWS EMR

Stream Processing

Stream processing processes data continuously as it arrives.

Event happens → Process immediately → Real-time result
(Latency: milliseconds to seconds)

Tools: Apache Kafka, Apache Flink, Apache Spark Streaming, AWS Kinesis

Batch vs Stream Comparison

FeatureBatchStream
WhenScheduledContinuous
LatencyHoursMilliseconds
DataBounded (finite set)Unbounded (infinite)
ExamplesMonth-end reportFraud detection
ToolsHadoop, Spark batchKafka, Flink, Spark Streaming

Evolution of Data Engineering

Data Engineering Timeline

@1970   File I/O
        Programs read/write flat files; no database
        Storage: KB-MB scale

@1980   RDBMS (E.F. Codd's rules)
        SQL databases emerge; structured data management
Storage: MB-GB scale; CURD (Create, Update, Read, Delete)

@1990   Data Warehouse (DWH)
        Separate system for analytics; OLAP
        Massive
Parallel Processing (MPP) begins

@1991-1995  Internet & Dotcom era
        Web generates new types of data; e-commerce
emerges

@1998   Database consolidation
        NoSQL concepts emerge; distributed databases

@2000   MPP at scale; 4+2 era
        Global internet; MB-GB/day per organization

@2003   Data burst
        Google publishes GFS paper; MapReduce paper
        Facebook,
YouTube launched (2004-2005)
        Social media generates PB/day

@2008    Facebook Post JSON format; social data explosion

@2010   Big Data era
        Hadoop ecosystem matures
        100s of TB to PB scale
        Cloud computing emerges

Key Historical Papers

YearPaperImpact
2003Google File System (GFS)Foundation for distributed storage
2004Google MapReduceFoundation for distributed processing
2006Bigtable (Google)Foundation for NoSQL wide-column stores
2007Amazon DynamoFoundation for NoSQL key-value stores
2010+Hadoop, Hive, Spark, KafkaOpen-source big data ecosystem

ACID Properties

What is ACID?

ACID defines properties that ensure database transactions are processed reliably:

LetterPropertyDefinition
AAtomicityTransaction is all-or-nothing; if any part fails, the entire transaction is rolled back
CConsistencyTransaction brings database from one valid state to another; all rules must be satisfied
IIsolationConcurrent transactions execute as if they are sequential; one transaction's changes not visible to others until complete
DDurabilityOnce committed, transaction persists even if system crashes; changes written to disk

ACID in Practice

Atomicity Example:

Bank Transfer: Alice → Bob (Rs. 1000)

Step 1: Debit Alice   -1000
Step 2: Credit Bob    +1000

If Step 1 passes but Step 2 FAILS:
  Without Atomicity: Alice loses 1000; Bob gets nothing (data inconsistency!)
  With Atomicity:    ROLLBACK → both steps undone → no money lost

Isolation Example:

Transaction T1: Read balance = 5000; Update balance = 4000
Transaction T2: Read balance (at same time)

Without Isolation: T2 might read 5000 (stale) or 4500 (dirty read - inconsistent)
With Isolation:    T2 waits; reads either 5000 (before T1) or 4000 (after T1)

Isolation Levels

LevelDirty ReadNon-Repeatable ReadPhantom Read
READ UNCOMMITTEDPossiblePossiblePossible
READ COMMITTEDNot PossiblePossiblePossible
REPEATABLE READNot PossibleNot PossiblePossible
SERIALIZABLENot PossibleNot PossibleNot Possible

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.