Big Data and Data Engineering
Apache Kafka; Data Lake vs Data Warehouse; Big Data in the Cloud
C-CAT
Apache Kafka
What is Kafka?
Apache Kafka is a distributed event streaming platform for:
- High-throughput, low-latency message streaming
- Real-time data pipelines
- Decoupling of data producers and consumers
Kafka Core Concepts
PRODUCER (writes data) KAFKA CLUSTER CONSUMER (reads data)
+-----------+ +----------+ +----------+
| App | produces | Topic | reads | Analytics|
| Database |──────────────| (logs) |────────| ML model |
| IoT device| | Partition| | Dashboard|
| | | 1, 2, 3 | | |
+-----------+ +----------+ +----------+
Brokers (servers)
| Term | Description |
|---|---|
| Topic | Category/channel for messages |
| Partition | Topics split into partitions (parallel processing) |
| Offset | Sequential ID of messages within a partition |
| Producer | Application that sends messages to Kafka |
| Consumer | Application that reads messages from Kafka |
| Consumer Group | Group of consumers sharing topic partitions |
| Broker | Kafka server node |
| ZooKeeper/KRaft | Cluster metadata management |
Kafka Use Cases
| Use Case | Description |
|---|---|
| Log Aggregation | Collect logs from multiple servers |
| Event Streaming | Real-time event processing |
| Data Integration | Connect different systems |
| Metrics Collection | Application monitoring |
| Stream Processing | Process events with Kafka Streams / Flink |
| Website Activity | Track page views, searches |
Kafka vs Traditional Messaging
| Feature | Traditional MQ (RabbitMQ) | Kafka |
|---|---|---|
| Message retention | Deleted after consumption | Configurable retention (days) |
| Throughput | Lower | Very high (millions/sec) |
| Replay | Not possible | Yes (offset-based) |
| Order | Queue-based | Per-partition ordering |
| Scale | Limited | Massive |
| Use case | Task queues | Event streaming |
Data Lake vs Data Warehouse
Data Lake
A Data Lake stores all data in its raw, native format — structured, semi-structured and unstructured.
"Store everything; process when needed"
| Feature | Description |
|---|---|
| Storage | Raw data (any format) |
| Schema | Schema-on-read (define when querying) |
| Users | Data scientists, ML engineers |
| Technologies | HDFS, Amazon S3, Azure Data Lake |
| Cost | Very low (commodity storage) |
Risk: Data Lake can become a Data Swamp without proper governance (unusable raw data).
Data Warehouse
| Feature | Description |
|---|---|
| Storage | Processed, cleaned, structured data |
| Schema | Schema-on-write (defined at load time) |
| Users | Business analysts, BI tools |
| Technologies | Snowflake, BigQuery, Redshift |
| Cost | Higher (compute + storage) |
Lakehouse
Lakehouse = Data Lake + Data Warehouse (modern approach):
- Store raw data in Data Lake
- Apply DWH features (ACID, schema enforcement) on top
- Technologies: Delta Lake (Databricks), Apache Iceberg, Apache Hudi
Big Data in the Cloud
AWS Big Data Services
| Service | Category | Description |
|---|---|---|
| Amazon S3 | Storage | Object store; data lake foundation |
| Amazon Redshift | DWH | Cloud data warehouse; columnar |
| Amazon EMR | Processing | Managed Hadoop/Spark on AWS |
| Amazon Kinesis | Streaming | Real-time data streaming |
| AWS Glue | ETL | Serverless data integration |
| Amazon Athena | Query | Query S3 data with SQL (serverless) |
| Amazon QuickSight | BI | Cloud BI and visualization |
Google Cloud Big Data
| Service | Description |
|---|---|
| BigQuery | Serverless, massive-scale SQL DWH |
| Cloud Dataflow | Managed stream + batch processing (Apache Beam) |
| Cloud Pub/Sub | Managed Kafka-like messaging |
| Cloud Dataproc | Managed Hadoop/Spark |
| Looker | BI and data visualization |
Azure Big Data
| Service | Description |
|---|---|
| Azure Data Lake Storage | Scalable data lake |
| Azure Synapse Analytics | Integrated data warehouse + spark |
| Azure Data Factory | ETL/data integration |
| Azure Event Hubs | Kafka-compatible event streaming |
| Power BI | BI visualization |
Continue learning
Related notes
Definition of AI; Need of AI
Artificial Intelligence
Introduction to Data Engineering; Big Data — The 5 V's; Types of Data
Big Data and Data Engineering
Introduction to C Programming; C Program Structure; Data Types and Variables
C Programming
What Is a Computer?; Machine Cycle: Fetch–Decode–Execute; CPU Organization
Computer Architecture
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.