Big Data and Data Engineering

Evolution of Data Engineering; RDBMS — Relational Database Management System; ACID Properties

C-CAT

Evolution of Data Engineering

Data Engineering Timeline

@1970   File I/O
        Programs read/write flat files; no database
        Storage: KB-MB scale

@1980   RDBMS (E.F. Codd's rules)
        SQL databases emerge; structured data management
Storage: MB-GB scale; CURD (Create, Update, Read, Delete)

@1990   Data Warehouse (DWH)
        Separate system for analytics; OLAP
        Massive
Parallel Processing (MPP) begins

@1991-1995  Internet & Dotcom era
        Web generates new types of data; e-commerce
emerges

@1998   Database consolidation
        NoSQL concepts emerge; distributed databases

@2000   MPP at scale; 4+2 era
        Global internet; MB-GB/day per organization

@2003   Data burst
        Google publishes GFS paper; MapReduce paper
        Facebook,
YouTube launched (2004-2005)
        Social media generates PB/day

@2008    Facebook Post JSON format; social data explosion

@2010   Big Data era
        Hadoop ecosystem matures
        100s of TB to PB scale
        Cloud computing emerges

Key Historical Papers

YearPaperImpact
2003Google File System (GFS)Foundation for distributed storage
2004Google MapReduceFoundation for distributed processing
2006Bigtable (Google)Foundation for NoSQL wide-column stores
2007Amazon DynamoFoundation for NoSQL key-value stores
2010+Hadoop, Hive, Spark, KafkaOpen-source big data ecosystem

RDBMS — Relational Database Management System

What is RDBMS?

RDBMS (Relational Database Management System) is software that:

  • Stores data in tables (rows and columns)
  • Based on E.F. Codd's relational model (1970)
  • Every enterprise application needs to manage data through RDBMS

Key Concepts:

TermDefinition
Table / RelationOrganized collection of rows (tuples) and columns (attributes)
Row / Tuple / RecordSingle entry in a table
Column / Attribute / FieldA property/characteristic of the entity
Primary KeyUnique identifier for each row; NOT NULL
Foreign KeyReferences primary key of another table; enforces referential integrity
IndexData structure for fast data retrieval
ViewVirtual table based on a SELECT query
SchemaLogical structure/blueprint of the database

Popular RDBMS Products

ProductVendorType
Oracle DatabaseOracleEnterprise (paid)
SQL ServerMicrosoftEnterprise (paid)
MySQLOracle (open source)Very popular, web apps
PostgreSQLOpen sourceAdvanced, enterprise grade
SQLiteOpen sourceEmbedded, mobile apps
MariaDBCommunity fork of MySQLMySQL-compatible open source
Db2IBMEnterprise mainframe

E.F. Codd's 12 Rules for RDBMS

Key rules (important ones):

  1. Information Rule — All information in tables

Guaranteed Access Rule — Every value accessible by table+column+primary key 3. NULL Values — NULL supported for missing information 4. Active Online Catalog — Database schema stored in same database 5. Data Independence — Logical and physical independence

ACID Properties

What is ACID?

ACID defines properties that ensure database transactions are processed reliably:

LetterPropertyDefinition
AAtomicityTransaction is all-or-nothing; if any part fails, the entire transaction is rolled back
CConsistencyTransaction brings database from one valid state to another; all rules must be satisfied
IIsolationConcurrent transactions execute as if they are sequential; one transaction's changes not visible to others until complete
DDurabilityOnce committed, transaction persists even if system crashes; changes written to disk

ACID in Practice

Atomicity Example:

Bank Transfer: Alice → Bob (Rs. 1000)

Step 1: Debit Alice   -1000
Step 2: Credit Bob    +1000

If Step 1 passes but Step 2 FAILS:
  Without Atomicity: Alice loses 1000; Bob gets nothing (data inconsistency!)
  With Atomicity:    ROLLBACK → both steps undone → no money lost

Isolation Example:

Transaction T1: Read balance = 5000; Update balance = 4000
Transaction T2: Read balance (at same time)

Without Isolation: T2 might read 5000 (stale) or 4500 (dirty read - inconsistent)
With Isolation:    T2 waits; reads either 5000 (before T1) or 4000 (after T1)

Isolation Levels

LevelDirty ReadNon-Repeatable ReadPhantom Read
READ UNCOMMITTEDPossiblePossiblePossible
READ COMMITTEDNot PossiblePossiblePossible
REPEATABLE READNot PossibleNot PossiblePossible
SERIALIZABLENot PossibleNot PossibleNot Possible

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.