Big Data and Data Engineering

Batch vs Stream Processing; Distributed Storage; Cloud Computing Fundamentals

C-CAT

Batch vs Stream Processing

Batch Processing

Batch processing processes data in discrete chunks at scheduled intervals.

12:00 AM daily → Collect all transactions → Process → Load to DWH

Characteristics:

FeatureDescription
TriggerScheduled (daily, hourly)
LatencyHigh (hours to days)
ThroughputHigh (large volumes at once)
ComplexityLow
Use caseEnd-of-day reports, monthly billing

Tools: Apache Hadoop MapReduce, Apache Spark (batch mode), AWS EMR

Stream Processing

Stream processing processes data continuously as it arrives.

Event happens → Process immediately → Real-time result
(Latency: milliseconds to seconds)

Characteristics:

FeatureDescription
TriggerContinuous; event-driven
LatencyLow (milliseconds to seconds)
ThroughputLower (real-time)
ComplexityHigher
Use caseFraud detection, live dashboard, real-time analytics

Tools: Apache Kafka, Apache Flink, Apache Spark Streaming, AWS Kinesis

Batch vs Stream Comparison

FeatureBatchStream
WhenScheduledContinuous
LatencyHoursMilliseconds
DataBounded (finite set)Unbounded (infinite)
ExamplesMonth-end reportFraud detection
ToolsHadoop, Spark batchKafka, Flink, Spark Streaming

Distributed Storage

Why Distributed Storage?

  • Single machine cannot handle PB-scale data
  • Traditional storage: 1 machine; limited capacity
  • Distributed storage: Spread data across hundreds or thousands of machines
Single Server:          Distributed Storage:
+------+               +----+ +----+ +----+ +----+
|      |               |Node| |Node| |Node| |Node|
| 10TB |               | 1  | | 2  | | 3  | | 4  |
|      |               | 4TB| | 4TB| | 4TB| | 4TB|
+------+               +----+ +----+ +----+ +----+
                         Total: 16TB (easily expandable)

HDFS — Hadoop Distributed File System

HDFS is the distributed storage system designed for Big Data (part of Hadoop ecosystem).

Key concepts:

ConceptDescription
NameNodeMaster server; stores metadata (file locations, block info)
DataNodeWorker servers; store actual data blocks
Block sizeDefault 128MB per block (old: 64MB)
ReplicationDefault 3 copies of each block (fault tolerance)
File: sales.csv (500MB)

HDFS splits it into blocks:
Block 1 (0-128MB):   Stored on DataNode 1, 4, 7 (replication factor 3)
Block 2 (128-256MB): Stored on DataNode 2, 5, 8
Block 3 (256-384MB): Stored on DataNode 3, 6, 9
Block 4 (384-500MB): Stored on DataNode 1, 5, 7

Cloud Storage Services

ServiceProviderDescription
Amazon S3AWSObject storage; infinite scale; 99.999999999% durability
Azure Data Lake StorageAzureEnterprise analytics storage
Google Cloud StorageGCPObject storage with analytics integration
Google Cloud BigtableGCPWide-column NoSQL; millions of ops/sec

Cloud Computing Fundamentals

What is Cloud Computing?

Cloud computing is the delivery of computing services (servers, storage, databases, networking, software) over the internet on a pay-as-you-go basis.

Key Concepts

ConceptDefinition
VirtualizationCreating virtual versions of hardware, OS, storage or network (multiple VMs on one physical server)
Scaling / ElasticityAutomatically increase/decrease resources based on demand
Horizontal ScalingAdd more machines (scale out)
Vertical ScalingAdd more power to existing machine (scale up)

Cloud Service Models

+-------------------------------------------------------------+
|   IaaS                                                       |
|   Infrastructure as a Service                               |
|   You manage: OS, runtime, apps, data                       |
|   Provider manages: servers, storage, networking           |
|   Example: AWS EC2, Azure VMs, Google Compute Engine       |
+-------------------------------------------------------------+
              +-------------------------------------------+
              |   PaaS                                     |
              |   Platform as a Service                   |
              |   You manage: app, data                   |
              |   Provider manages: OS + infrastructure   |
              |   Example: Heroku, Google App Engine      |
              +-------------------------------------------+
                          +------------------------+
                          |   SaaS                  |
                          |   Software as a Service |
                          |   You just USE it       |
                          |   Example: Gmail,       |
                          |   Salesforce, Office365 |
                          +------------------------+

Cloud Providers

ProviderKey Services
AWS (Amazon)EC2, S3, RDS, Redshift, EMR, Kinesis, Glue
Azure (Microsoft)Virtual Machines, Blob Storage, Azure SQL, Synapse
GCP (Google)GCE, GCS, BigQuery, Dataflow, Pub/Sub, Dataproc

Virtualization

Physical Server (64 cores, 512GB RAM)
    ↓ Hypervisor (VMware, KVM, Hyper-V)
    ├── VM 1: 8 cores, 64GB (Ubuntu, web server)
    ├── VM 2: 16 cores, 128GB (Windows, databases)
    ├── VM 3: 8 cores, 64GB (CentOS, analytics)
    └── VM 4: 4 cores, 32GB (Test environment)

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.