Big Data and Data Engineering
Evolution of Data Engineering; RDBMS — Relational Database Management System; ACID Properties
C-CAT
Evolution of Data Engineering
Data Engineering Timeline
@1970 File I/O
Programs read/write flat files; no database
Storage: KB-MB scale
@1980 RDBMS (E.F. Codd's rules)
SQL databases emerge; structured data management
Storage: MB-GB scale; CURD (Create, Update, Read, Delete)
@1990 Data Warehouse (DWH)
Separate system for analytics; OLAP
Massive
Parallel Processing (MPP) begins
@1991-1995 Internet & Dotcom era
Web generates new types of data; e-commerce
emerges
@1998 Database consolidation
NoSQL concepts emerge; distributed databases
@2000 MPP at scale; 4+2 era
Global internet; MB-GB/day per organization
@2003 Data burst
Google publishes GFS paper; MapReduce paper
Facebook,
YouTube launched (2004-2005)
Social media generates PB/day
@2008 Facebook Post JSON format; social data explosion
@2010 Big Data era
Hadoop ecosystem matures
100s of TB to PB scale
Cloud computing emerges
Key Historical Papers
| Year | Paper | Impact |
|---|---|---|
| 2003 | Google File System (GFS) | Foundation for distributed storage |
| 2004 | Google MapReduce | Foundation for distributed processing |
| 2006 | Bigtable (Google) | Foundation for NoSQL wide-column stores |
| 2007 | Amazon Dynamo | Foundation for NoSQL key-value stores |
| 2010+ | Hadoop, Hive, Spark, Kafka | Open-source big data ecosystem |
RDBMS — Relational Database Management System
What is RDBMS?
RDBMS (Relational Database Management System) is software that:
- Stores data in tables (rows and columns)
- Based on E.F. Codd's relational model (1970)
- Every enterprise application needs to manage data through RDBMS
Key Concepts:
| Term | Definition |
|---|---|
| Table / Relation | Organized collection of rows (tuples) and columns (attributes) |
| Row / Tuple / Record | Single entry in a table |
| Column / Attribute / Field | A property/characteristic of the entity |
| Primary Key | Unique identifier for each row; NOT NULL |
| Foreign Key | References primary key of another table; enforces referential integrity |
| Index | Data structure for fast data retrieval |
| View | Virtual table based on a SELECT query |
| Schema | Logical structure/blueprint of the database |
Popular RDBMS Products
| Product | Vendor | Type |
|---|---|---|
| Oracle Database | Oracle | Enterprise (paid) |
| SQL Server | Microsoft | Enterprise (paid) |
| MySQL | Oracle (open source) | Very popular, web apps |
| PostgreSQL | Open source | Advanced, enterprise grade |
| SQLite | Open source | Embedded, mobile apps |
| MariaDB | Community fork of MySQL | MySQL-compatible open source |
| Db2 | IBM | Enterprise mainframe |
E.F. Codd's 12 Rules for RDBMS
Key rules (important ones):
- Information Rule — All information in tables
Guaranteed Access Rule — Every value accessible by table+column+primary key 3. NULL Values — NULL supported for missing information 4. Active Online Catalog — Database schema stored in same database 5. Data Independence — Logical and physical independence
ACID Properties
What is ACID?
ACID defines properties that ensure database transactions are processed reliably:
| Letter | Property | Definition |
|---|---|---|
| A | Atomicity | Transaction is all-or-nothing; if any part fails, the entire transaction is rolled back |
| C | Consistency | Transaction brings database from one valid state to another; all rules must be satisfied |
| I | Isolation | Concurrent transactions execute as if they are sequential; one transaction's changes not visible to others until complete |
| D | Durability | Once committed, transaction persists even if system crashes; changes written to disk |
ACID in Practice
Atomicity Example:
Bank Transfer: Alice → Bob (Rs. 1000)
Step 1: Debit Alice -1000
Step 2: Credit Bob +1000
If Step 1 passes but Step 2 FAILS:
Without Atomicity: Alice loses 1000; Bob gets nothing (data inconsistency!)
With Atomicity: ROLLBACK → both steps undone → no money lost
Isolation Example:
Transaction T1: Read balance = 5000; Update balance = 4000
Transaction T2: Read balance (at same time)
Without Isolation: T2 might read 5000 (stale) or 4500 (dirty read - inconsistent)
With Isolation: T2 waits; reads either 5000 (before T1) or 4000 (after T1)
Isolation Levels
| Level | Dirty Read | Non-Repeatable Read | Phantom Read |
|---|---|---|---|
| READ UNCOMMITTED | Possible | Possible | Possible |
| READ COMMITTED | Not Possible | Possible | Possible |
| REPEATABLE READ | Not Possible | Not Possible | Possible |
| SERIALIZABLE | Not Possible | Not Possible | Not Possible |
Continue learning
Related notes
Definition of AI; Need of AI
Artificial Intelligence
Introduction to Data Engineering; Big Data — The 5 V's; Types of Data
Big Data and Data Engineering
Introduction to C Programming; C Program Structure; Data Types and Variables
C Programming
What Is a Computer?; Machine Cycle: Fetch–Decode–Execute; CPU Organization
Computer Architecture
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.