Big Data and Data Engineering

Apache Kafka; Data Lake vs Data Warehouse; Big Data in the Cloud

C-CAT

Apache Kafka

What is Kafka?

Apache Kafka is a distributed event streaming platform for:

  • High-throughput, low-latency message streaming
  • Real-time data pipelines
  • Decoupling of data producers and consumers

Kafka Core Concepts

PRODUCER (writes data)     KAFKA CLUSTER      CONSUMER (reads data)
+-----------+              +----------+        +----------+
| App       |   produces   | Topic    |  reads | Analytics|
| Database  |──────────────| (logs)   |────────| ML model |
| IoT device|              | Partition|        | Dashboard|
|           |              | 1, 2, 3  |        |          |
+-----------+              +----------+        +----------+
                            Brokers (servers)
TermDescription
TopicCategory/channel for messages
PartitionTopics split into partitions (parallel processing)
OffsetSequential ID of messages within a partition
ProducerApplication that sends messages to Kafka
ConsumerApplication that reads messages from Kafka
Consumer GroupGroup of consumers sharing topic partitions
BrokerKafka server node
ZooKeeper/KRaftCluster metadata management

Kafka Use Cases

Use CaseDescription
Log AggregationCollect logs from multiple servers
Event StreamingReal-time event processing
Data IntegrationConnect different systems
Metrics CollectionApplication monitoring
Stream ProcessingProcess events with Kafka Streams / Flink
Website ActivityTrack page views, searches

Kafka vs Traditional Messaging

FeatureTraditional MQ (RabbitMQ)Kafka
Message retentionDeleted after consumptionConfigurable retention (days)
ThroughputLowerVery high (millions/sec)
ReplayNot possibleYes (offset-based)
OrderQueue-basedPer-partition ordering
ScaleLimitedMassive
Use caseTask queuesEvent streaming

Data Lake vs Data Warehouse

Data Lake

A Data Lake stores all data in its raw, native format — structured, semi-structured and unstructured.

"Store everything; process when needed"
FeatureDescription
StorageRaw data (any format)
SchemaSchema-on-read (define when querying)
UsersData scientists, ML engineers
TechnologiesHDFS, Amazon S3, Azure Data Lake
CostVery low (commodity storage)

Risk: Data Lake can become a Data Swamp without proper governance (unusable raw data).

Data Warehouse

FeatureDescription
StorageProcessed, cleaned, structured data
SchemaSchema-on-write (defined at load time)
UsersBusiness analysts, BI tools
TechnologiesSnowflake, BigQuery, Redshift
CostHigher (compute + storage)

Lakehouse

Lakehouse = Data Lake + Data Warehouse (modern approach):

  • Store raw data in Data Lake
  • Apply DWH features (ACID, schema enforcement) on top
  • Technologies: Delta Lake (Databricks), Apache Iceberg, Apache Hudi

Big Data in the Cloud

AWS Big Data Services

ServiceCategoryDescription
Amazon S3StorageObject store; data lake foundation
Amazon RedshiftDWHCloud data warehouse; columnar
Amazon EMRProcessingManaged Hadoop/Spark on AWS
Amazon KinesisStreamingReal-time data streaming
AWS GlueETLServerless data integration
Amazon AthenaQueryQuery S3 data with SQL (serverless)
Amazon QuickSightBICloud BI and visualization

Google Cloud Big Data

ServiceDescription
BigQueryServerless, massive-scale SQL DWH
Cloud DataflowManaged stream + batch processing (Apache Beam)
Cloud Pub/SubManaged Kafka-like messaging
Cloud DataprocManaged Hadoop/Spark
LookerBI and data visualization

Azure Big Data

ServiceDescription
Azure Data Lake StorageScalable data lake
Azure Synapse AnalyticsIntegrated data warehouse + spark
Azure Data FactoryETL/data integration
Azure Event HubsKafka-compatible event streaming
Power BIBI visualization

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.