Big Data Technologies

Big Data Characteristics, Adoption, Sources and Data Curation

PGCP-BDA

big data

Datasets whose scale, speed or complexity require distributed storage, parallel processing or specialized analytical methods.

volume velocity variety veracity value

The five Vs describe data quantity, arrival speed, forms, trustworthiness and the useful outcome obtained from analysis.

structured semi-structured unstructured data

Structured data follows a fixed schema, semi-structured data carries flexible tags or keys and unstructured data lacks a predefined tabular organization.

big data source

An operational system, sensor, application, log, transaction stream or external feed that produces data for a big-data platform.

distributed system

A system whose networked computers coordinate work and appear as one service while tolerating partial failures.

scale up and scale out

Scaling up adds resources to one machine; scaling out adds machines and distributes storage or computation among them.

data locality

Data locality schedules computation on or near nodes holding required blocks, reducing network transfer when moving code is cheaper than moving data.

data curation

The controlled selection, cleaning, documentation, organization and preservation of data so it remains trustworthy and reusable.

big data adoption

The organizational process of selecting valuable use cases, building suitable platforms and integrating data-driven work into operations.

governance

Policies, roles and controls that define data ownership, quality, security, permitted use, retention and accountability.

Big Data — The 5 V's

What is Big Data?

Big Data refers to datasets that are so large, fast-moving and diverse that traditional database tools cannot manage them effectively.

Traditional systems (single machines, RDBMS) cannot handle:

  • 100s of Terabytes and above
  • Millions of records per second
  • Dozens of different data formats

The 5 V's of Big Data

VDefinitionChallengeExample
VolumeHow much data?Storage, processing at scaleFacebook: 4 petabytes/day
VelocityHow fast data arrives?Real-time processing, low latencyTwitter: 500M tweets/day (6000/sec)
VarietyHow many types of data?Handling structured, semi, unstructuredText, images, video, JSON, logs
VeracityHow trustworthy/accurate?Data quality, cleaning, validationMissing values, incorrect data
ValueBusiness value derivedExtracting meaningful insightsRevenue prediction, fraud detection

Volume Breakdown

KB → MB → GB → TB → PB → EB → ZB
1 KB = 1,024 Bytes
1 MB = 1,024 KB
1 GB = 1,024 MB
1 TB = 1,024 GB   (1 hour of 4K video ≈ 7 GB; 1 TB = 140 hours)
1 PB = 1,024 TB   (Google processes ~20+ PB/day)
1 EB = 1,024 PB
1 ZB = 1,024 EB   (Total internet traffic expected to hit ZB scale)

Velocity Examples

SourceData Rate
Twitter6,000 tweets/second
YouTube500 hours of video uploaded/minute
Stock exchangeMillions of transactions/second
IoT sensorsBillions of sensor readings/day
E-commerce (sale day)Thousands of orders/second

Introduction to Data Engineering

What is Data Engineering?

Data engineering is the development, implementation and maintenance of systems and processes that:

  • Take in raw data
  • Produce high-quality, consistent information

Support downstream use cases such as analysis and machine learning

Data Engineer manages:

  • Data pipelines
  • Storage systems
  • Data transformations
  • Data quality
TopicSub-topics
Big Data FundamentalsEvolution, 5 Vs (Volume, Velocity, Variety, Veracity, Value)
DatabasesRDBMS (ACID, SQL), NoSQL (BASE, CAP theorem)
Data WarehouseOLAP vs OLTP, cleansing, transformation, modeling, DWH vs Data mart
Data Engineering LifecycleSource → Ingestion → Storage → Transformation → Serving
IngestionETL vs ELT
StorageDistributed storage, cloud storage; Batch vs Stream processing
Cloud ComputingVirtualization, Scaling, Elasticity, Service models, Vendors
Big Data TechnologiesHadoop, Hive, Spark, Kafka
ApplicationsJob profiles, use cases

Role of a Data Engineer

Raw Data Sources                    End Users
(Databases, Logs, APIs, IoT)   →  (Analysts, Data Scientists, BI Apps)
       |                                    ^
       +------ DATA ENGINEER -------+
              (builds pipelines)

Data Engineer Tasks:
1. Data Ingestion    → Move data from sources to storage
2. Data Storage      → Design efficient storage (DWH, Data Lake, NoSQL)
3. Data Transformation → Clean, enrich, aggregate data
4. Pipeline Building → Automate data flow
5. Quality Assurance → Ensure accuracy and consistency

Types of Data

Structured Data

  • Data saved in rows, columns and tables
  • Fixed schema (pre-defined structure)

Easily managed by RDBMS (SQL)

Examples:

  • Bank transaction records (ID, Date, Amount, Account)
  • Employee records (EmpID, Name, Salary, Department)
  • Student grades table
ID  | Name    | Age | Salary
----|---------|-----|-------
001 | Alice   | 28  | 55000
002 | Bob     | 35  | 72000

Semi-Structured Data

  • Flexible schema — structure varies per record
  • Has some organization (tags, keys) but not fixed table format
  • Serialization formats: JSON, XML, CSV, YAML

Examples:

{
  "post": "...",
  "location": "Mumbai",
  "likes": 240,
  "tags": ["travel", "tech"],
  "image": "photo.jpg",
  "feeling": "happy"
}
  • Facebook post, Twitter tweet, MongoDB document, API response

Unstructured Data

  • No fixed format — no predefined schema
  • Cannot be stored in traditional rows/columns
  • Examples: images (JPG/PNG), videos, audio, ML model files
Images  → JPG, PNG, BMP
Video   → MP4, AVI
Audio   → MP3, WAV
Text    → Free-form articles, emails
PDFs    → Documents
Emails  → Mixed content

Comparison

FeatureStructuredSemi-StructuredUnstructured
SchemaFixed tableFlexible (JSON/XML)None
StorageRDBMS (SQL)NoSQL, FilesFile systems, Object store
SearchabilityEasy (SQL)Moderate (JSONPath)Hard (AI/ML needed)
ExamplesBank recordsJSON API dataImages, videos
% of total data~20%~20%~80%

RDBMS — Relational Database Management System

What is RDBMS?

RDBMS (Relational Database Management System) is software that:

  • Stores data in tables (rows and columns)
  • Based on E.F. Codd's relational model (1970)
  • Every enterprise application needs to manage data through RDBMS

Key Concepts:

TermDefinition
Table / RelationOrganized collection of rows (tuples) and columns (attributes)
Row / Tuple / RecordSingle entry in a table
Column / Attribute / FieldA property/characteristic of the entity
Primary KeyUnique identifier for each row; NOT NULL
Foreign KeyReferences primary key of another table; enforces referential integrity
IndexData structure for fast data retrieval
ViewVirtual table based on a SELECT query
SchemaLogical structure/blueprint of the database

Popular RDBMS Products

ProductVendorType
Oracle DatabaseOracleEnterprise (paid)
SQL ServerMicrosoftEnterprise (paid)
MySQLOracle (open source)Very popular, web apps
PostgreSQLOpen sourceAdvanced, enterprise grade
SQLiteOpen sourceEmbedded, mobile apps
MariaDBCommunity fork of MySQLMySQL-compatible open source
Db2IBMEnterprise mainframe

E.F. Codd's 12 Rules for RDBMS

Key rules (important ones):

  1. Information Rule — All information in tables

Guaranteed Access Rule — Every value accessible by table+column+primary key 3. NULL Values — NULL supported for missing information 4. Active Online Catalog — Database schema stored in same database 5. Data Independence — Logical and physical independence

CAP Theorem

What is CAP Theorem?

CAP Theorem (Brewer's Theorem, 2000): A distributed system can guarantee at most 2 out of 3 properties simultaneously:

PropertyDescription
C — ConsistencyEvery read receives the most recent write or an error
A — AvailabilityEvery request receives a response (not necessarily latest)
P — Partition ToleranceSystem continues operating despite network partitions
         C (Consistency)
             /\
            /  \
           /    \
          /      \
CA ------+--------+ CP
        /          \
       /            \
      /      AP      \
     +________________+
A (Availability)   P (Partition Tolerance)

Network partitions ALWAYS happen in distributed systems → you must choose between C and A.

Database Classification by CAP

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.