Big Data and Data Engineering

Introduction to Data Engineering; Big Data — The 5 V's; Types of Data

C-CAT

Introduction to Data Engineering

What is Data Engineering?

Data engineering is the development, implementation and maintenance of systems and processes that:

  • Take in raw data
  • Produce high-quality, consistent information

Support downstream use cases such as analysis and machine learning

Data Engineer manages:

  • Data pipelines
  • Storage systems
  • Data transformations
  • Data quality
TopicSub-topics
Big Data FundamentalsEvolution, 5 Vs (Volume, Velocity, Variety, Veracity, Value)
DatabasesRDBMS (ACID, SQL), NoSQL (BASE, CAP theorem)
Data WarehouseOLAP vs OLTP, cleansing, transformation, modeling, DWH vs Data mart
Data Engineering LifecycleSource → Ingestion → Storage → Transformation → Serving
IngestionETL vs ELT
StorageDistributed storage, cloud storage; Batch vs Stream processing
Cloud ComputingVirtualization, Scaling, Elasticity, Service models, Vendors
Big Data TechnologiesHadoop, Hive, Spark, Kafka
ApplicationsJob profiles, use cases

Role of a Data Engineer

Raw Data Sources                    End Users
(Databases, Logs, APIs, IoT)   →  (Analysts, Data Scientists, BI Apps)
       |                                    ^
       +------ DATA ENGINEER -------+
              (builds pipelines)

Data Engineer Tasks:
1. Data Ingestion    → Move data from sources to storage
2. Data Storage      → Design efficient storage (DWH, Data Lake, NoSQL)
3. Data Transformation → Clean, enrich, aggregate data
4. Pipeline Building → Automate data flow
5. Quality Assurance → Ensure accuracy and consistency

Big Data — The 5 V's

What is Big Data?

Big Data refers to datasets that are so large, fast-moving and diverse that traditional database tools cannot manage them effectively.

Traditional systems (single machines, RDBMS) cannot handle:

  • 100s of Terabytes and above
  • Millions of records per second
  • Dozens of different data formats

The 5 V's of Big Data

VDefinitionChallengeExample
VolumeHow much data?Storage, processing at scaleFacebook: 4 petabytes/day
VelocityHow fast data arrives?Real-time processing, low latencyTwitter: 500M tweets/day (6000/sec)
VarietyHow many types of data?Handling structured, semi, unstructuredText, images, video, JSON, logs
VeracityHow trustworthy/accurate?Data quality, cleaning, validationMissing values, incorrect data
ValueBusiness value derivedExtracting meaningful insightsRevenue prediction, fraud detection

Volume Breakdown

KB → MB → GB → TB → PB → EB → ZB
1 KB = 1,024 Bytes
1 MB = 1,024 KB
1 GB = 1,024 MB
1 TB = 1,024 GB   (1 hour of 4K video ≈ 7 GB; 1 TB = 140 hours)
1 PB = 1,024 TB   (Google processes ~20+ PB/day)
1 EB = 1,024 PB
1 ZB = 1,024 EB   (Total internet traffic expected to hit ZB scale)

Velocity Examples

SourceData Rate
Twitter6,000 tweets/second
YouTube500 hours of video uploaded/minute
Stock exchangeMillions of transactions/second
IoT sensorsBillions of sensor readings/day
E-commerce (sale day)Thousands of orders/second

Types of Data

Structured Data

  • Data saved in rows, columns and tables
  • Fixed schema (pre-defined structure)

Easily managed by RDBMS (SQL)

Examples:

  • Bank transaction records (ID, Date, Amount, Account)
  • Employee records (EmpID, Name, Salary, Department)
  • Student grades table
ID  | Name    | Age | Salary
----|---------|-----|-------
001 | Alice   | 28  | 55000
002 | Bob     | 35  | 72000

Semi-Structured Data

  • Flexible schema — structure varies per record
  • Has some organization (tags, keys) but not fixed table format
  • Serialization formats: JSON, XML, CSV, YAML

Examples:

{
  "post": "...",
  "location": "Mumbai",
  "likes": 240,
  "tags": ["travel", "tech"],
  "image": "photo.jpg",
  "feeling": "happy"
}
  • Facebook post, Twitter tweet, MongoDB document, API response

Unstructured Data

  • No fixed format — no predefined schema
  • Cannot be stored in traditional rows/columns
  • Examples: images (JPG/PNG), videos, audio, ML model files
Images  → JPG, PNG, BMP
Video   → MP4, AVI
Audio   → MP3, WAV
Text    → Free-form articles, emails
PDFs    → Documents
Emails  → Mixed content

Comparison

FeatureStructuredSemi-StructuredUnstructured
SchemaFixed tableFlexible (JSON/XML)None
StorageRDBMS (SQL)NoSQL, FilesFile systems, Object store
SearchabilityEasy (SQL)Moderate (JSONPath)Hard (AI/ML needed)
ExamplesBank recordsJSON API dataImages, videos
% of total data~20%~20%~80%

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.