Big Data Technologies
Hadoop Evolution, Ecosystem, Architecture and Operating Modes
PGCP-BDA
Hadoop
Apache Hadoop is a framework for distributed storage and batch processing built around HDFS, YARN, MapReduce and supporting libraries.
Hadoop ecosystem
The Hadoop ecosystem combines storage, resource management, processing, SQL, ingestion, coordination and workflow tools around distributed datasets.
HDFS MapReduce and YARN
HDFS stores distributed files, MapReduce defines batch data processing and YARN allocates cluster resources and runs application containers.
Hadoop operating modes
Standalone mode uses one JVM and local storage, pseudo-distributed mode runs daemons on one host and fully distributed mode spreads daemons across a.
commodity cluster
A commodity cluster uses replaceable general-purpose servers and software replication or recomputation instead of relying only on specialized.
Hadoop configuration
Hadoop configuration is assembled from site XML files and runtime properties whose effective values must be consistent across clients and relevant daemons.
data locality
Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.
batch processing
Batch processing consumes a bounded dataset, produces a completed result and optimizes throughput rather than per-record response latency.
Hadoop limitation
Hadoop MapReduce is inefficient for low-latency queries, iterative algorithms, many small files and workloads requiring frequent in-place updates.
ecosystem tools
Interoperating platform components that provide ingestion, storage, processing, querying, coordination, security and workflow services.
Batch vs Stream Processing
Batch Processing
Batch processing processes data in discrete chunks at scheduled intervals.
12:00 AM daily → Collect all transactions → Process → Load to DWH
Tools: Apache Hadoop MapReduce, Apache Spark (batch mode), AWS EMR
Stream Processing
Stream processing processes data continuously as it arrives.
Event happens → Process immediately → Real-time result
(Latency: milliseconds to seconds)
Tools: Apache Kafka, Apache Flink, Apache Spark Streaming, AWS Kinesis
Batch vs Stream Comparison
| Feature | Batch | Stream |
|---|---|---|
| When | Scheduled | Continuous |
| Latency | Hours | Milliseconds |
| Data | Bounded (finite set) | Unbounded (infinite) |
| Examples | Month-end report | Fraud detection |
| Tools | Hadoop, Spark batch | Kafka, Flink, Spark Streaming |
Evolution of Data Engineering
Data Engineering Timeline
@1970 File I/O
Programs read/write flat files; no database
Storage: KB-MB scale
@1980 RDBMS (E.F. Codd's rules)
SQL databases emerge; structured data management
Storage: MB-GB scale; CURD (Create, Update, Read, Delete)
@1990 Data Warehouse (DWH)
Separate system for analytics; OLAP
Massive
Parallel Processing (MPP) begins
@1991-1995 Internet & Dotcom era
Web generates new types of data; e-commerce
emerges
@1998 Database consolidation
NoSQL concepts emerge; distributed databases
@2000 MPP at scale; 4+2 era
Global internet; MB-GB/day per organization
@2003 Data burst
Google publishes GFS paper; MapReduce paper
Facebook,
YouTube launched (2004-2005)
Social media generates PB/day
@2008 Facebook Post JSON format; social data explosion
@2010 Big Data era
Hadoop ecosystem matures
100s of TB to PB scale
Cloud computing emerges
Key Historical Papers
| Year | Paper | Impact |
|---|---|---|
| 2003 | Google File System (GFS) | Foundation for distributed storage |
| 2004 | Google MapReduce | Foundation for distributed processing |
| 2006 | Bigtable (Google) | Foundation for NoSQL wide-column stores |
| 2007 | Amazon Dynamo | Foundation for NoSQL key-value stores |
| 2010+ | Hadoop, Hive, Spark, Kafka | Open-source big data ecosystem |
ACID Properties
What is ACID?
ACID defines properties that ensure database transactions are processed reliably:
| Letter | Property | Definition |
|---|---|---|
| A | Atomicity | Transaction is all-or-nothing; if any part fails, the entire transaction is rolled back |
| C | Consistency | Transaction brings database from one valid state to another; all rules must be satisfied |
| I | Isolation | Concurrent transactions execute as if they are sequential; one transaction's changes not visible to others until complete |
| D | Durability | Once committed, transaction persists even if system crashes; changes written to disk |
ACID in Practice
Atomicity Example:
Bank Transfer: Alice → Bob (Rs. 1000)
Step 1: Debit Alice -1000
Step 2: Credit Bob +1000
If Step 1 passes but Step 2 FAILS:
Without Atomicity: Alice loses 1000; Bob gets nothing (data inconsistency!)
With Atomicity: ROLLBACK → both steps undone → no money lost
Isolation Example:
Transaction T1: Read balance = 5000; Update balance = 4000
Transaction T2: Read balance (at same time)
Without Isolation: T2 might read 5000 (stale) or 4500 (dirty read - inconsistent)
With Isolation: T2 waits; reads either 5000 (before T1) or 4000 (after T1)
Isolation Levels
| Level | Dirty Read | Non-Repeatable Read | Phantom Read |
|---|---|---|---|
| READ UNCOMMITTED | Possible | Possible | Possible |
| READ COMMITTED | Not Possible | Possible | Possible |
| REPEATABLE READ | Not Possible | Not Possible | Possible |
| SERIALIZABLE | Not Possible | Not Possible | Not Possible |
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.