Big Data Technologies

MapReduce Data Types, Formats, Partitioners, Combiners and Counters

PGCP-BDA

Writable type

Writable is Hadoop’s mutable binary serialization contract; WritableComparable adds ordering required for shuffle keys.

InputFormat and OutputFormat

InputFormat creates logical splits and record readers, while OutputFormat validates output and supplies record writers and commit behavior.

partitioner

A MapReduce partitioner maps each intermediate key to a reducer number and therefore controls grouping distribution and possible skew.

combiner

A combiner is an optional local aggregation that may run zero or more times, so it is valid only when partial aggregation preserves the final result.

counter

A named MapReduce metric incremented by tasks to record events such as processed records, invalid input or framework activity.

distributed cache

The distributed cache localizes read-only files, archives or jars beside tasks so every task need not fetch common resources independently.

secondary sort

Secondary sort designs composite keys, partitioning and grouping comparators so reducer values arrive ordered by a field other than the grouping key.

custom key

An application-defined key type that implements the serialization, equality and ordering behavior required by a distributed framework.

map-side join

A map-side join loads or co-partitions compatible inputs so records can be joined without sending both full inputs through reducers.

reduce-side join

A reduce-side join tags records from each input with a common key, shuffles them together and joins grouped values in the reducer.

Cloud Computing Fundamentals

What is Cloud Computing?

Cloud computing is the delivery of computing services (servers, storage, databases, networking, software) over the internet on a pay-as-you-go basis.

Key Concepts

ConceptDefinition
VirtualizationCreating virtual versions of hardware, OS, storage or network (multiple VMs on one physical server)
Scaling / ElasticityAutomatically increase/decrease resources based on demand
Horizontal ScalingAdd more machines (scale out)
Vertical ScalingAdd more power to existing machine (scale up)

Cloud Service Models

+-------------------------------------------------------------+
|   IaaS                                                       |
|   Infrastructure as a Service                               |
|   You manage: OS, runtime, apps, data                       |
|   Provider manages: servers, storage, networking           |
|   Example: AWS EC2, Azure VMs, Google Compute Engine       |
+-------------------------------------------------------------+
              +-------------------------------------------+
              |   PaaS                                     |
              |   Platform as a Service                   |
              |   You manage: app, data                   |
              |   Provider manages: OS + infrastructure   |
              |   Example: Heroku, Google App Engine      |
              +-------------------------------------------+
                          +------------------------+
                          |   SaaS                  |
                          |   Software as a Service |
                          |   You just USE it       |
                          |   Example: Gmail,       |
                          |   Salesforce, Office365 |
                          +------------------------+

Cloud Providers

ProviderKey Services
AWS (Amazon)EC2, S3, RDS, Redshift, EMR, Kinesis, Glue
Azure (Microsoft)Virtual Machines, Blob Storage, Azure SQL, Synapse
GCP (Google)GCE, GCS, BigQuery, Dataflow, Pub/Sub, Dataproc

Virtualization

Physical Server (64 cores, 512GB RAM)
    ↓ Hypervisor (VMware, KVM, Hyper-V)
    ├── VM 1: 8 cores, 64GB (Ubuntu, web server)
    ├── VM 2: 16 cores, 128GB (Windows, databases)
    ├── VM 3: 8 cores, 64GB (CentOS, analytics)
    └── VM 4: 4 cores, 32GB (Test environment)

MapReduce Types and Formats

Hadoop serialization uses Writable types such as IntWritable, LongWritable, Text and BytesWritable. A mapper’s declared output key and value types must match what it emits. InputFormat chooses splits and creates RecordReaders; it does not parse every record itself. Text input commonly reports byte offset as the key and line text as the value. Sequence files store binary key-value records and avoid repeated text parsing.

An output format controls record representation and output commit behavior. Compression reduces disk and network traffic. Splittable formats permit independent tasks to begin at defined boundaries. Many small files increase metadata and scheduling overhead even when their combined size is modest.

Partitioning, Combining and Counters

The default hash partitioner selects a reducer from the key hash. A custom partitioner can group keys by domain or range but must remain consistent and distribute load. A combiner performs local aggregation before shuffle. It cannot be required for correctness because the framework does not guarantee its invocation. Operations such as sum, minimum and maximum often combine safely while average requires carrying both sum and count.

Counters collect distributed integer measurements such as malformed rows or domain totals. Built-in counters describe input, output, bytes and task behavior. Custom counters are useful for bounded categories but excessive dynamic counter names create overhead. Counters diagnose a completed job; they should complement rather than replace validation of output data.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.