Big Data and Data Engineering
Introduction to Data Engineering; Big Data — The 5 V's; Types of Data
C-CAT
Introduction to Data Engineering
What is Data Engineering?
Data engineering is the development, implementation and maintenance of systems and processes that:
- Take in raw data
- Produce high-quality, consistent information
Support downstream use cases such as analysis and machine learning
Data Engineer manages:
- Data pipelines
- Storage systems
- Data transformations
- Data quality
| Topic | Sub-topics |
|---|---|
| Big Data Fundamentals | Evolution, 5 Vs (Volume, Velocity, Variety, Veracity, Value) |
| Databases | RDBMS (ACID, SQL), NoSQL (BASE, CAP theorem) |
| Data Warehouse | OLAP vs OLTP, cleansing, transformation, modeling, DWH vs Data mart |
| Data Engineering Lifecycle | Source → Ingestion → Storage → Transformation → Serving |
| Ingestion | ETL vs ELT |
| Storage | Distributed storage, cloud storage; Batch vs Stream processing |
| Cloud Computing | Virtualization, Scaling, Elasticity, Service models, Vendors |
| Big Data Technologies | Hadoop, Hive, Spark, Kafka |
| Applications | Job profiles, use cases |
Role of a Data Engineer
Raw Data Sources End Users
(Databases, Logs, APIs, IoT) → (Analysts, Data Scientists, BI Apps)
| ^
+------ DATA ENGINEER -------+
(builds pipelines)
Data Engineer Tasks:
1. Data Ingestion → Move data from sources to storage
2. Data Storage → Design efficient storage (DWH, Data Lake, NoSQL)
3. Data Transformation → Clean, enrich, aggregate data
4. Pipeline Building → Automate data flow
5. Quality Assurance → Ensure accuracy and consistency
Big Data — The 5 V's
What is Big Data?
Big Data refers to datasets that are so large, fast-moving and diverse that traditional database tools cannot manage them effectively.
Traditional systems (single machines, RDBMS) cannot handle:
- 100s of Terabytes and above
- Millions of records per second
- Dozens of different data formats
The 5 V's of Big Data
| V | Definition | Challenge | Example |
|---|---|---|---|
| Volume | How much data? | Storage, processing at scale | Facebook: 4 petabytes/day |
| Velocity | How fast data arrives? | Real-time processing, low latency | Twitter: 500M tweets/day (6000/sec) |
| Variety | How many types of data? | Handling structured, semi, unstructured | Text, images, video, JSON, logs |
| Veracity | How trustworthy/accurate? | Data quality, cleaning, validation | Missing values, incorrect data |
| Value | Business value derived | Extracting meaningful insights | Revenue prediction, fraud detection |
Volume Breakdown
KB → MB → GB → TB → PB → EB → ZB
1 KB = 1,024 Bytes
1 MB = 1,024 KB
1 GB = 1,024 MB
1 TB = 1,024 GB (1 hour of 4K video ≈ 7 GB; 1 TB = 140 hours)
1 PB = 1,024 TB (Google processes ~20+ PB/day)
1 EB = 1,024 PB
1 ZB = 1,024 EB (Total internet traffic expected to hit ZB scale)
Velocity Examples
| Source | Data Rate |
|---|---|
| 6,000 tweets/second | |
| YouTube | 500 hours of video uploaded/minute |
| Stock exchange | Millions of transactions/second |
| IoT sensors | Billions of sensor readings/day |
| E-commerce (sale day) | Thousands of orders/second |
Types of Data
Structured Data
- Data saved in rows, columns and tables
- Fixed schema (pre-defined structure)
Easily managed by RDBMS (SQL)
Examples:
- Bank transaction records (ID, Date, Amount, Account)
- Employee records (EmpID, Name, Salary, Department)
- Student grades table
ID | Name | Age | Salary
----|---------|-----|-------
001 | Alice | 28 | 55000
002 | Bob | 35 | 72000
Semi-Structured Data
- Flexible schema — structure varies per record
- Has some organization (tags, keys) but not fixed table format
- Serialization formats: JSON, XML, CSV, YAML
Examples:
{
"post": "...",
"location": "Mumbai",
"likes": 240,
"tags": ["travel", "tech"],
"image": "photo.jpg",
"feeling": "happy"
}
- Facebook post, Twitter tweet, MongoDB document, API response
Unstructured Data
- No fixed format — no predefined schema
- Cannot be stored in traditional rows/columns
- Examples: images (JPG/PNG), videos, audio, ML model files
Images → JPG, PNG, BMP
Video → MP4, AVI
Audio → MP3, WAV
Text → Free-form articles, emails
PDFs → Documents
Emails → Mixed content
Comparison
| Feature | Structured | Semi-Structured | Unstructured |
|---|---|---|---|
| Schema | Fixed table | Flexible (JSON/XML) | None |
| Storage | RDBMS (SQL) | NoSQL, Files | File systems, Object store |
| Searchability | Easy (SQL) | Moderate (JSONPath) | Hard (AI/ML needed) |
| Examples | Bank records | JSON API data | Images, videos |
| % of total data | ~20% | ~20% | ~80% |
Continue learning
Related notes
Definition of AI; Need of AI
Artificial Intelligence
Introduction to C Programming; C Program Structure; Data Types and Variables
C Programming
What Is a Computer?; Machine Cycle: Fetch–Decode–Execute; CPU Organization
Computer Architecture
What is a Computer?; Computer System Components; Booting Process
Computer Fundamentals
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.