Big Data and Data Engineering
Batch vs Stream Processing; Distributed Storage; Cloud Computing Fundamentals
C-CAT
Batch vs Stream Processing
Batch Processing
Batch processing processes data in discrete chunks at scheduled intervals.
12:00 AM daily → Collect all transactions → Process → Load to DWH
Characteristics:
| Feature | Description |
|---|---|
| Trigger | Scheduled (daily, hourly) |
| Latency | High (hours to days) |
| Throughput | High (large volumes at once) |
| Complexity | Low |
| Use case | End-of-day reports, monthly billing |
Tools: Apache Hadoop MapReduce, Apache Spark (batch mode), AWS EMR
Stream Processing
Stream processing processes data continuously as it arrives.
Event happens → Process immediately → Real-time result
(Latency: milliseconds to seconds)
Characteristics:
| Feature | Description |
|---|---|
| Trigger | Continuous; event-driven |
| Latency | Low (milliseconds to seconds) |
| Throughput | Lower (real-time) |
| Complexity | Higher |
| Use case | Fraud detection, live dashboard, real-time analytics |
Tools: Apache Kafka, Apache Flink, Apache Spark Streaming, AWS Kinesis
Batch vs Stream Comparison
| Feature | Batch | Stream |
|---|---|---|
| When | Scheduled | Continuous |
| Latency | Hours | Milliseconds |
| Data | Bounded (finite set) | Unbounded (infinite) |
| Examples | Month-end report | Fraud detection |
| Tools | Hadoop, Spark batch | Kafka, Flink, Spark Streaming |
Distributed Storage
Why Distributed Storage?
- Single machine cannot handle PB-scale data
- Traditional storage: 1 machine; limited capacity
- Distributed storage: Spread data across hundreds or thousands of machines
Single Server: Distributed Storage:
+------+ +----+ +----+ +----+ +----+
| | |Node| |Node| |Node| |Node|
| 10TB | | 1 | | 2 | | 3 | | 4 |
| | | 4TB| | 4TB| | 4TB| | 4TB|
+------+ +----+ +----+ +----+ +----+
Total: 16TB (easily expandable)
HDFS — Hadoop Distributed File System
HDFS is the distributed storage system designed for Big Data (part of Hadoop ecosystem).
Key concepts:
| Concept | Description |
|---|---|
| NameNode | Master server; stores metadata (file locations, block info) |
| DataNode | Worker servers; store actual data blocks |
| Block size | Default 128MB per block (old: 64MB) |
| Replication | Default 3 copies of each block (fault tolerance) |
File: sales.csv (500MB)
HDFS splits it into blocks:
Block 1 (0-128MB): Stored on DataNode 1, 4, 7 (replication factor 3)
Block 2 (128-256MB): Stored on DataNode 2, 5, 8
Block 3 (256-384MB): Stored on DataNode 3, 6, 9
Block 4 (384-500MB): Stored on DataNode 1, 5, 7
Cloud Storage Services
| Service | Provider | Description |
|---|---|---|
| Amazon S3 | AWS | Object storage; infinite scale; 99.999999999% durability |
| Azure Data Lake Storage | Azure | Enterprise analytics storage |
| Google Cloud Storage | GCP | Object storage with analytics integration |
| Google Cloud Bigtable | GCP | Wide-column NoSQL; millions of ops/sec |
Cloud Computing Fundamentals
What is Cloud Computing?
Cloud computing is the delivery of computing services (servers, storage, databases, networking, software) over the internet on a pay-as-you-go basis.
Key Concepts
| Concept | Definition |
|---|---|
| Virtualization | Creating virtual versions of hardware, OS, storage or network (multiple VMs on one physical server) |
| Scaling / Elasticity | Automatically increase/decrease resources based on demand |
| Horizontal Scaling | Add more machines (scale out) |
| Vertical Scaling | Add more power to existing machine (scale up) |
Cloud Service Models
+-------------------------------------------------------------+
| IaaS |
| Infrastructure as a Service |
| You manage: OS, runtime, apps, data |
| Provider manages: servers, storage, networking |
| Example: AWS EC2, Azure VMs, Google Compute Engine |
+-------------------------------------------------------------+
+-------------------------------------------+
| PaaS |
| Platform as a Service |
| You manage: app, data |
| Provider manages: OS + infrastructure |
| Example: Heroku, Google App Engine |
+-------------------------------------------+
+------------------------+
| SaaS |
| Software as a Service |
| You just USE it |
| Example: Gmail, |
| Salesforce, Office365 |
+------------------------+
Cloud Providers
| Provider | Key Services |
|---|---|
| AWS (Amazon) | EC2, S3, RDS, Redshift, EMR, Kinesis, Glue |
| Azure (Microsoft) | Virtual Machines, Blob Storage, Azure SQL, Synapse |
| GCP (Google) | GCE, GCS, BigQuery, Dataflow, Pub/Sub, Dataproc |
Virtualization
Physical Server (64 cores, 512GB RAM)
↓ Hypervisor (VMware, KVM, Hyper-V)
├── VM 1: 8 cores, 64GB (Ubuntu, web server)
├── VM 2: 16 cores, 128GB (Windows, databases)
├── VM 3: 8 cores, 64GB (CentOS, analytics)
└── VM 4: 4 cores, 32GB (Test environment)
Continue learning
Related notes
Definition of AI; Need of AI
Artificial Intelligence
Introduction to Data Engineering; Big Data — The 5 V's; Types of Data
Big Data and Data Engineering
Introduction to C Programming; C Program Structure; Data Types and Variables
C Programming
What Is a Computer?; Machine Cycle: Fetch–Decode–Execute; CPU Organization
Computer Architecture
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.