Big Data Technologies

HBase Architecture, Regions, Storage Model and Installation

PGCP-BDA

HBase

Apache HBase is a distributed wide-column database on HDFS that provides low-latency random reads and writes over sparse, versioned rows ordered by row key.

HBase data model

An HBase table contains rows identified by byte-array keys, column families defined at schema time, dynamic qualifiers.

row key

The HBase row key uniquely identifies and sorts a row; its design determines region distribution, scan locality and hotspot risk.

column family and qualifier

An HBase column is addressed by family and qualifier; family properties control physical storage while qualifiers can vary independently by row.

cell timestamp and version

An HBase cell is identified by row, family, qualifier and timestamp, permitting several versions of one logical field.

region

An HBase region is a contiguous row-key range of one table served by one RegionServer and split as it grows.

RegionServer

An HBase RegionServer serves reads and writes for assigned regions and manages write-ahead logs, MemStores, HFiles, flushes and compactions.

HMaster

The HBase HMaster assigns regions, coordinates schema and administrative operations and balances cluster load while clients normally access RegionServers.

ZooKeeper

ZooKeeper provides coordinated metadata, membership and leader or master discovery using a replicated hierarchical namespace with ordered updates.

MemStore and HFile

HBase writes are recorded in a write-ahead log and MemStore; a flush creates immutable HFiles that later compactions reorganize.

Relational and MongoDB Vocabulary

Approximate analogies are:

These are learning aids, not exact equivalences. Documents can nest arrays and objects, while relational rows are organized into declared columns and relationships.

MongoDB

MongoDB is a document database. It stores records as BSON documents grouped into collections. A database contains collections and a deployment can contain several databases.

The document model supports nested objects, arrays and varied fields. Design should still define required structure, types, identifiers, relationships and evolution rules.

MongoDB provides a query language, aggregation pipeline, indexes, replication, sharding, validation and transactions. Behavior depends on server version, topology, read concern, write concern and session options.

Indexes

Create an index:

db.products.createIndex({status: 1, price: -1})

This can support equality on status followed by a compatible price range or order. Compound index prefixes matter, much like ordered relational indexes.

MongoDB index types include single-field, compound, multikey for arrays, text, geospatial, hashed, wildcard, sparse, partial, TTL and unique variants.

Every index consumes storage and write work. Build indexes from measured access patterns.

HBase Architecture and Storage

HBase stores sparse rows in tables ordered lexicographically by row key. A cell is identified by row key, column family, qualifier and timestamp. Column families are declared in the schema while qualifiers can vary by row. Values are byte arrays, so applications define encoding. Row operations are atomic within one row.

A table is divided into regions containing contiguous row-key ranges. RegionServers serve regions and split them as they grow. The master coordinates assignments and administrative operations while ZooKeeper assists coordination. Writes enter a write-ahead log and an in-memory MemStore before later flush to immutable HFiles. Compaction merges files and removes obsolete versions or deleted cells according to policy. Reads consult memory and files with indexes and Bloom filters reducing unnecessary access.

Row-Key Design and Operations

Row-key design determines distribution and access. A monotonically increasing key can direct new writes to one region and create a hotspot. Salting, hashing or reversed components spread load but can make range scans harder. Related columns should share a family only when their access and retention patterns align because each family has separate files and flush behavior.

Installation requires compatible Hadoop, coordination and Java versions plus correct filesystem and network configuration. Tables define column families and properties such as versions and time-to-live. Operational checks cover region distribution, compaction backlog, write-ahead logs, disk use and unavailable servers. HBase suits keyed lookups and range scans at scale; it does not provide relational joins or arbitrary secondary queries by default.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.