Big Data Technologies

HBase Shell, Java APIs, CRUD, Scans, Administration and Security

PGCP-BDA

HBase shell

An interactive command-line client for creating, inspecting, reading, changing and administering HBase objects.

HBase put get scan delete

Put writes cell values, Get reads one row, Scan streams an ordered row-key range and Delete records tombstones for selected cells or versions.

HBase Java API

The client interfaces used by Java programs to configure connections and perform HBase table, row, scan and administrative operations.

scan filter

An HBase scan filter removes unwanted rows or cells after a scan range is chosen; a narrow row-key range usually avoids more I/O.

HBase table design

HBase table design begins with access patterns and chooses row keys and column families to keep common reads contiguous while distributing writes.

hotspot avoidance

Hotspot avoidance prevents sequential or low-cardinality row-key prefixes from directing most requests to one HBase region.

HBase compaction

HBase compaction merges immutable HFiles, removes eligible deleted or expired versions and reduces read amplification at an I/O cost.

HBase administration

Creation and management of namespaces, tables, column families, regions, permissions, backups and cluster health.

HBase security

Authentication, authorization, encryption and auditing controls that protect HBase clusters, namespaces, tables and cells.

HBase consistency

HBase provides strongly consistent reads and writes for a single row; operations spanning rows are not one atomic transaction.

NoSQL Databases

What is NoSQL?

NoSQL (Not Only SQL) databases are designed for:

  • Large-scale distributed data storage
  • Flexible schemas — no fixed table structure
  • High performance at scale (reads and writes)
  • Horizontal scaling across commodity servers

Why NoSQL?

Limitation of RDBMSNoSQL Solution
Fixed schema hard to changeDynamic/flexible schemas
Vertical scaling limitHorizontal scaling (add nodes)
Joins become slow at huge scaleDenormalized data; no joins needed
Complex transactionsSimpler CRUD operations at scale

Types of NoSQL Databases

8.1 Key-Value Stores

Key → Value
"user:1001" → {"name":"Alice","age":28,"city":"Mumbai"}

Operations: GET(key), PUT(key, value), DELETE(key) — O(1)

Examples: Redis, Amazon DynamoDB, Memcached Use cases: Session management, caching, shopping cart

8.2 Document Databases

{
  "_id": "abc123",
  "name": "Alice",
  "address": {
    "street": "123 Main St",
    "city": "Mumbai"
  },
  "orders": [
    {"product": "Laptop", "amount": 50000},
    {"product": "Mouse", "amount": 1000}
  ]
}

Examples: MongoDB, CouchDB, Amazon DocumentDB Use cases: Content management, user profiles, e-commerce catalogs

8.3 Column-Family (Wide-Column) Stores

Row Key → Column Family → Columns
user_001 → profile: {name, email, age}
           orders:  {order_1, order_2}
           stats:   {login_count, last_seen}

Examples: Apache Cassandra, Google Bigtable, HBase (Hadoop) Use cases: IoT time series, analytics, messaging systems

8.4 Graph Databases

(Alice) --[FRIENDS]-- (Bob)
   |                    |
[LIKES]              [BOUGHT]
   |                    |
(Product)           (Laptop)

Examples: Neo4j, Amazon Neptune, ArangoDB Use cases: Social networks, recommendation engines, fraud detection

NoSQL vs RDBMS Comparison

FeatureRDBMSNoSQL
SchemaFixedFlexible
ScalingVertical (more RAM/CPU)Horizontal (more nodes)
ConsistencyStrong (ACID)Eventual (BASE)
JoinsSupportedLimited/Not supported
Query LanguageSQLVarious (MongoQL, CQL, etc.)
TransactionsFull ACIDLimited
Best forStructured data, OLTPLarge scale, varied data

Reliable MongoDB Design

For every update, verify whether one or many documents may match. Use atomic operators rather than read-modify-write cycles. Define deterministic sorting and keyset pagination. Inspect explain output and account for index write cost.

MongoDB's flexible document structure is valuable when paired with explicit contracts. Flexibility without validation, type consistency, bounded growth or migration rules merely moves schema problems into every application.

Databases, Collections, Shell and Compass

In mongosh:

use training

The command selects a database context. A database or collection may be created lazily on the first write or a collection can be created explicitly with options:

db.createCollection("products")

MongoDB Compass is a graphical tool for connecting, browsing documents, constructing queries and aggregation pipelines, inspecting schema patterns and managing indexes.

Shell and Compass are clients. Authentication and server authorization govern what each connected account can do.

Atomicity and Transactions

An update to one MongoDB document is atomic, even when it changes several fields or embedded values. This is one reason to embed data that forms one consistency unit.

MongoDB also supports multi-document transactions in supported replica-set and sharded configurations. They add coordination cost and should not compensate for a poor aggregate model.

Use conditional filters and inspect matched counts to detect concurrent conflicts. Configure read and write concerns according to durability and consistency needs.

Equality and Comparison Filters

Direct field equality:

db.products.find({status: "ACTIVE"})

Comparison operators include:

  • $eq and $ne;
  • $gt and $gte;
  • $lt and $lte;
  • $in and $nin.

db.products.find({ price: {$gte: NumberDecimal("100.00"), $lte: NumberDecimal("500.00")} })

$in compares against a list:

{status: {$in: ["NEW", "PAID"]}}

Use consistent BSON types in both stored data and query values.

BASE Properties

What is BASE?

BASE is the consistency model used by NoSQL databases (opposite of ACID):

LetterPropertyDefinition
BABasically AvailableSystem guarantees availability (not every request will get the latest data, but a response)
SSoft StateThe state of the system may change over time (even without new inputs — due to eventual consistency)
EEventually ConsistentSystem will become consistent over time (not immediately; all replicas eventually agree)

Example of Eventual Consistency:

User posts a photo on Instagram (write to primary replica)
Immediately visible on their phone
After 1-2 seconds: visible globally (replicas catch up)
Eventually consistent — not immediately consistent

Explain and Query Plans

Use explain to inspect access:

db.products.find({ status: "ACTIVE", price: {$gte: NumberDecimal("100.00")} }).explain("executionStats")

Examine index scans, collection scans, keys examined, documents examined, returned rows and execution stages.

An index name alone does not prove efficiency. A broad scan of index keys followed by many document fetches can still be expensive.

Reading with find and findOne

find returns a cursor over matching documents:

db.products.find({stock: {$gt: 0}})

findOne returns one matching document or null:

db.products.findOne({sku: "BK-100"})

Without a sort, which matching document findOne returns is not a stable business rule. Use a unique filter or an explicit ordered query when identity matters.

An empty filter matches all documents:

db.products.find({})

Documents and the _id Field

Every standard collection document has a unique _id field. If an inserted document omits it, the driver or server normally generates an ObjectId.

{ _id: ObjectId("..."), sku: "BK-100", name: "Database Design", price: NumberDecimal("599.00"), tags: ["database", "design"], publisher: { name: "Example Press", city: "Pune" } }

MongoDB automatically creates a unique index on _id. Choose a custom _id only when its uniqueness, immutability, size and distribution suit the workload.

Deleting Documents

deleteOne removes at most one match:

db.products.deleteOne({sku: "BK-100"})

deleteMany removes all matches:

db.session.deleteMany({expiresAt: {$lt: new Date()}})

An empty deleteMany filter removes every document in the collection. Verify filters and deleted counts. Dropping a collection additionally removes its indexes and metadata and is not the same as deleting documents.

Positional Array Updates

MongoDB supplies positional operators for matched array elements. The exact operator depends on whether the update targets the first match, all elements or elements satisfying array filters.

db.orders.updateOne( {_id: orderId}, {$set: {"items.$[item].status": "BACKORDERED"}}, {arrayFilters: [{"item.productId": productId}]} )

Test filters carefully so the intended elements are changed. Complex arrays often signal that aggregate boundaries need review.

OLTP vs OLAP

OLTP — Online Transaction Processing

FeatureOLTP
PurposeDay-to-day operations; transactions
DataCurrent, operational data
QueriesSimple CRUD; INSERT, UPDATE, DELETE
Response timeMilliseconds
UsersThousands of concurrent users (clerks, customers)
Data sizeGB
NormalizedHighly normalized (3NF)

Examples: Banking transactions, e-commerce orders, inventory management

OLAP — Online Analytical Processing

Examples: Sales trend analysis, customer behavior analysis, financial reporting

OLTP vs OLAP Comparison

Logical Operators

Multiple fields in one filter are implicitly combined with AND:

{status: "ACTIVE", stock: {$gt: 0}}

Explicit operators include $and, $or, $nor and $not:

{ $or: [ {category: "BOOK"}, {price: {$lt: NumberDecimal("100.00")}} ] }

Use $and explicitly when the same field needs conditions that cannot be combined in one object or when generated query structure requires it.

Inserting Documents

Insert one:

db.products.insertOne({ sku: "BK-100", name: "Database Design", price: NumberDecimal("599.00"), stock: 20 })

Insert several:

db.products.insertMany([ {sku: "BK-101", name: "SQL", stock: 10}, {sku: "BK-102", name: "MongoDB", stock: 15} ])

The result reports acknowledged status and inserted identifiers according to write concern. Ordered and unordered bulk behavior determines whether later operations continue after an error.

Aggregation Pipeline

The aggregation pipeline passes documents through stages:

db.orders.aggregate([ {$match: {status: "PAID"}}, {$unwind: "$items"}, {$group: { _id: "$items.productId", quantity: {$sum: "$items.quantity"} }}, {$sort: {quantity: -1}} ])

$match filters, $project reshapes, $unwind emits one document per array element, $group aggregates, $sort orders and $lookup can join collection data.

Place selective $match stages early when semantics permit and inspect multiplication caused by $unwind or $lookup.

Multikey Indexes

Indexing an array field creates a multikey index with entries derived from array elements. This supports array membership queries.

Compound multikey indexes have restrictions when more than one indexed path is an array. Arrays with many elements can produce many index entries and high write cost.

Use $elemMatch where predicates must apply to the same array element. Index design and query semantics must agree.

Unique, Partial and TTL Indexes

A unique index enforces uniqueness:

db.products.createIndex({sku: 1}, {unique: true})

Missing and null behavior needs careful testing, particularly with sparse or partial options.

A partial index contains documents satisfying a filter and can reduce size when queries use the same condition.

A TTL index allows background expiration based on a date field. Deletion is asynchronous rather than exactly at the expiration instant, so it should not be used as a precise scheduler.

Sorting, Limiting and Skipping

db.products.find({status: "ACTIVE"}) .sort({price: -1, _id: 1}) .limit(20)

1 requests ascending order and -1 descending. _id provides a unique tie-breaker.

skip supports offset pagination but becomes expensive for large offsets and can shift under concurrent writes. Keyset pagination applies a filter after the last seen sort key and uses a compatible index.

Sort without a suitable index may require an in-memory or disk-assisted blocking sort and is subject to server limits and options.

Upsert

An upsert updates a matching document or inserts a document when no match exists:

db.products.updateOne( {sku: "BK-103"}, { $set: {name: "Algorithms", price: NumberDecimal("499.00")}, $setOnInsert: {createdAt: new Date()} }, {upsert: true} )

The filter contributes equality fields to the inserted document under defined rules. Use a unique index on the logical identity to prevent concurrent upserts from creating duplicates.

$setOnInsert applies only to the insert branch.

Replacement

replaceOne substitutes an entire document except for immutable _id:

db.products.replaceOne( {sku: "BK-100"}, { sku: "BK-100", name: "Database Design", price: NumberDecimal("649.00"), stock: 20 } )

Fields omitted from the replacement disappear. Use update operators for partial changes and replacement only when the caller intentionally supplies the complete new document.

Update Operators

Common operators include:

  • $set assigns or creates fields;
  • $unset removes fields;
  • $inc atomically increments numeric values;
  • $mul multiplies;
  • $min and $max change a value conditionally;
  • $rename renames a field;
  • $currentDate assigns a current date or timestamp.

db.products.updateOne( {sku: "BK-100", stock: {$gte: 2}}, {$inc: {stock: -2}} )

The combined filter and increment form an atomic conditional update on one document, preventing a separate read-then-write stock race.

Updating Documents

updateOne changes the first matching document, while updateMany changes all matches:

db.products.updateOne( {sku: "BK-100"}, {$set: {status: "ACTIVE"}} )

The result distinguishes matched and modified counts. A document can match even when the requested value already equals the stored value, producing a matched count without a modification.

Use a unique filter for updateOne when one known entity is intended. Otherwise “first” is not a stable identity rule.

Projection

Projection controls returned fields:

db.products.find( {status: "ACTIVE"}, {name: 1, price: 1} )

In inclusion projection, _id remains included unless explicitly excluded:

{name: 1, price: 1, _id: 0}

Except for _id, inclusion and exclusion styles generally should not be mixed. Projection reduces network transfer and can support covered queries when an index contains all required filter and result fields.

Arrays

A scalar equality condition can match an array containing that value:

db.products.find({tags: "database"})

$all requires the array to contain every specified value:

{tags: {$all: ["database", "design"]}}

$size tests exact array length. $elemMatch requires one array element to satisfy several conditions:

{ reviews: { $elemMatch: {rating: {$gte: 4}, verified: true} } }

Without $elemMatch, separate dotted predicates might be satisfied by different array elements.

JSON and BSON

MongoDB interfaces display documents in a JSON-like syntax, while the stored binary representation is BSON.

BSON supports types beyond plain JSON, including ObjectId, Date, binary data, Decimal128, regular expression and multiple numeric types.

Type choice matters:

Two visually similar values of different BSON types may not compare as expected.

Schema Validation

A collection validator can require fields and BSON types:

db.createCollection("products", { validator: { $jsonSchema: { bsonType: "object", required: ["sku", "name", "price"], properties: { sku: {bsonType: "string"}, name: {bsonType: "string"}, price: {bsonType: "decimal"} } } } })

Validation can reject or warn according to settings. It provides a shared boundary while allowing intentional document variation.

Array Updates

$push appends an array element:

{$push: {tags: "featured"}}

$addToSet adds only when an equal element is absent. $pull removes matching elements. $pop removes one end.

Modifiers with $push can add several values, limit length, sort and position insertion:

{ $push: { scores: { $each: [90, 95], $sort: -1, $slice: 10 } } }

Arrays that grow without a bound can make documents large and contentious. Move unbounded events to another collection.

Nested Fields

Dot notation addresses nested paths:

db.products.find({"publisher.city": "Pune"})

An exact equality comparison against an embedded document is sensitive to its complete value and field order. Dot-path predicates are usually better for individual nested properties.

Updates use the same path:

{$set: {"publisher.city": "Mumbai"}}

Without dot notation, setting publisher would replace the entire embedded object.

Missing Fields and NULL

MongoDB distinguishes a field that is absent from one whose value is null, but some query forms can match both:

{middleName: null}

For explicit presence tests:

{middleName: {$exists: false}} {middleName: {$exists: true}}

Combine $exists and type/value conditions when the distinction matters. Schema validation can require fields and prevent inconsistent absence conventions.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.