Big Data Technologies
Big Data Characteristics, Adoption, Sources and Data Curation
PGCP-BDA
big data
Datasets whose scale, speed or complexity require distributed storage, parallel processing or specialized analytical methods.
volume velocity variety veracity value
The five Vs describe data quantity, arrival speed, forms, trustworthiness and the useful outcome obtained from analysis.
structured semi-structured unstructured data
Structured data follows a fixed schema, semi-structured data carries flexible tags or keys and unstructured data lacks a predefined tabular organization.
big data source
An operational system, sensor, application, log, transaction stream or external feed that produces data for a big-data platform.
distributed system
A system whose networked computers coordinate work and appear as one service while tolerating partial failures.
scale up and scale out
Scaling up adds resources to one machine; scaling out adds machines and distributes storage or computation among them.
data locality
Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.
data curation
The controlled selection, cleaning, documentation, organization and preservation of data so it remains trustworthy and reusable.
big data adoption
The organizational process of selecting valuable use cases, building suitable platforms and integrating data-driven work into operations.
governance
Policies, roles and controls that define data ownership, quality, security, permitted use, retention and accountability.
Big Data — The 5 V's
What is Big Data?
Big Data refers to datasets that are so large, fast-moving and diverse that traditional database tools cannot manage them effectively.
Traditional systems (single machines, RDBMS) cannot handle:
- 100s of Terabytes and above
- Millions of records per second
- Dozens of different data formats
The 5 V's of Big Data
| V | Definition | Challenge | Example |
|---|---|---|---|
| Volume | How much data? | Storage, processing at scale | Facebook: 4 petabytes/day |
| Velocity | How fast data arrives? | Real-time processing, low latency | Twitter: 500M tweets/day (6000/sec) |
| Variety | How many types of data? | Handling structured, semi, unstructured | Text, images, video, JSON, logs |
| Veracity | How trustworthy/accurate? | Data quality, cleaning, validation | Missing values, incorrect data |
| Value | Business value derived | Extracting meaningful insights | Revenue prediction, fraud detection |
Volume Breakdown
KB → MB → GB → TB → PB → EB → ZB
1 KB = 1,024 Bytes
1 MB = 1,024 KB
1 GB = 1,024 MB
1 TB = 1,024 GB (1 hour of 4K video ≈ 7 GB; 1 TB = 140 hours)
1 PB = 1,024 TB (Google processes ~20+ PB/day)
1 EB = 1,024 PB
1 ZB = 1,024 EB (Total internet traffic expected to hit ZB scale)
Velocity Examples
| Source | Data Rate |
|---|---|
| 6,000 tweets/second | |
| YouTube | 500 hours of video uploaded/minute |
| Stock exchange | Millions of transactions/second |
| IoT sensors | Billions of sensor readings/day |
| E-commerce (sale day) | Thousands of orders/second |
Introduction to Data Engineering
What is Data Engineering?
Data engineering is the development, implementation and maintenance of systems and processes that:
- Take in raw data
- Produce high-quality, consistent information
Support downstream use cases such as analysis and machine learning
Data Engineer manages:
- Data pipelines
- Storage systems
- Data transformations
- Data quality
| Topic | Sub-topics |
|---|---|
| Big Data Fundamentals | Evolution, 5 Vs (Volume, Velocity, Variety, Veracity, Value) |
| Databases | RDBMS (ACID, SQL), NoSQL (BASE, CAP theorem) |
| Data Warehouse | OLAP vs OLTP, cleansing, transformation, modeling, DWH vs Data mart |
| Data Engineering Lifecycle | Source → Ingestion → Storage → Transformation → Serving |
| Ingestion | ETL vs ELT |
| Storage | Distributed storage, cloud storage; Batch vs Stream processing |
| Cloud Computing | Virtualization, Scaling, Elasticity, Service models, Vendors |
| Big Data Technologies | Hadoop, Hive, Spark, Kafka |
| Applications | Job profiles, use cases |
Role of a Data Engineer
Raw Data Sources End Users
(Databases, Logs, APIs, IoT) → (Analysts, Data Scientists, BI Apps)
| ^
+------ DATA ENGINEER -------+
(builds pipelines)
Data Engineer Tasks:
1. Data Ingestion → Move data from sources to storage
2. Data Storage → Design efficient storage (DWH, Data Lake, NoSQL)
3. Data Transformation → Clean, enrich, aggregate data
4. Pipeline Building → Automate data flow
5. Quality Assurance → Ensure accuracy and consistency
Types of Data
Structured Data
- Data saved in rows, columns and tables
- Fixed schema (pre-defined structure)
Easily managed by RDBMS (SQL)
Examples:
- Bank transaction records (ID, Date, Amount, Account)
- Employee records (EmpID, Name, Salary, Department)
- Student grades table
ID | Name | Age | Salary
----|---------|-----|-------
001 | Alice | 28 | 55000
002 | Bob | 35 | 72000
Semi-Structured Data
- Flexible schema — structure varies per record
- Has some organization (tags, keys) but not fixed table format
- Serialization formats: JSON, XML, CSV, YAML
Examples:
{
"post": "...",
"location": "Mumbai",
"likes": 240,
"tags": ["travel", "tech"],
"image": "photo.jpg",
"feeling": "happy"
}
- Facebook post, Twitter tweet, MongoDB document, API response
Unstructured Data
- No fixed format — no predefined schema
- Cannot be stored in traditional rows/columns
- Examples: images (JPG/PNG), videos, audio, ML model files
Images → JPG, PNG, BMP
Video → MP4, AVI
Audio → MP3, WAV
Text → Free-form articles, emails
PDFs → Documents
Emails → Mixed content
Comparison
| Feature | Structured | Semi-Structured | Unstructured |
|---|---|---|---|
| Schema | Fixed table | Flexible (JSON/XML) | None |
| Storage | RDBMS (SQL) | NoSQL, Files | File systems, Object store |
| Searchability | Easy (SQL) | Moderate (JSONPath) | Hard (AI/ML needed) |
| Examples | Bank records | JSON API data | Images, videos |
| % of total data | ~20% | ~20% | ~80% |
RDBMS — Relational Database Management System
What is RDBMS?
RDBMS (Relational Database Management System) is software that:
- Stores data in tables (rows and columns)
- Based on E.F. Codd's relational model (1970)
- Every enterprise application needs to manage data through RDBMS
Key Concepts:
| Term | Definition |
|---|---|
| Table / Relation | Organized collection of rows (tuples) and columns (attributes) |
| Row / Tuple / Record | Single entry in a table |
| Column / Attribute / Field | A property/characteristic of the entity |
| Primary Key | Unique identifier for each row; NOT NULL |
| Foreign Key | References primary key of another table; enforces referential integrity |
| Index | Data structure for fast data retrieval |
| View | Virtual table based on a SELECT query |
| Schema | Logical structure/blueprint of the database |
Popular RDBMS Products
| Product | Vendor | Type |
|---|---|---|
| Oracle Database | Oracle | Enterprise (paid) |
| SQL Server | Microsoft | Enterprise (paid) |
| MySQL | Oracle (open source) | Very popular, web apps |
| PostgreSQL | Open source | Advanced, enterprise grade |
| SQLite | Open source | Embedded, mobile apps |
| MariaDB | Community fork of MySQL | MySQL-compatible open source |
| Db2 | IBM | Enterprise mainframe |
E.F. Codd's 12 Rules for RDBMS
Key rules (important ones):
- Information Rule — All information in tables
Guaranteed Access Rule — Every value accessible by table+column+primary key 3. NULL Values — NULL supported for missing information 4. Active Online Catalog — Database schema stored in same database 5. Data Independence — Logical and physical independence
CAP Theorem
What is CAP Theorem?
CAP Theorem (Brewer's Theorem, 2000): A distributed system can guarantee at most 2 out of 3 properties simultaneously:
| Property | Description |
|---|---|
| C — Consistency | Every read receives the most recent write or an error |
| A — Availability | Every request receives a response (not necessarily latest) |
| P — Partition Tolerance | System continues operating despite network partitions |
C (Consistency)
/\
/ \
/ \
/ \
CA ------+--------+ CP
/ \
/ \
/ AP \
+________________+
A (Availability) P (Partition Tolerance)
Network partitions ALWAYS happen in distributed systems → you must choose between C and A.
Database Classification by CAP
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.