Big Data Technologies
HBase Shell, Java APIs, CRUD, Scans, Administration and Security
PGCP-BDA
HBase shell
An interactive command-line client for creating, inspecting, reading, changing and administering HBase objects.
HBase put get scan delete
Put writes cell values, Get reads one row, Scan streams an ordered row-key range and Delete records tombstones for selected cells or versions.
HBase Java API
The client interfaces used by Java programs to configure connections and perform HBase table, row, scan and administrative operations.
scan filter
An HBase scan filter removes unwanted rows or cells after a scan range is chosen; a narrow row-key range usually avoids more I/O.
HBase table design
HBase table design begins with access patterns and chooses row keys and column families to keep common reads contiguous while distributing writes.
hotspot avoidance
Hotspot avoidance prevents sequential or low-cardinality row-key prefixes from directing most requests to one HBase region.
HBase compaction
HBase compaction merges immutable HFiles, removes eligible deleted or expired versions and reduces read amplification at an I/O cost.
HBase administration
Creation and management of namespaces, tables, column families, regions, permissions, backups and cluster health.
HBase security
Authentication, authorization, encryption and auditing controls that protect HBase clusters, namespaces, tables and cells.
HBase consistency
HBase provides strongly consistent reads and writes for a single row; operations spanning rows are not one atomic transaction.
NoSQL Databases
What is NoSQL?
NoSQL (Not Only SQL) databases are designed for:
- Large-scale distributed data storage
- Flexible schemas — no fixed table structure
- High performance at scale (reads and writes)
- Horizontal scaling across commodity servers
Why NoSQL?
| Limitation of RDBMS | NoSQL Solution |
|---|---|
| Fixed schema hard to change | Dynamic/flexible schemas |
| Vertical scaling limit | Horizontal scaling (add nodes) |
| Joins become slow at huge scale | Denormalized data; no joins needed |
| Complex transactions | Simpler CRUD operations at scale |
Types of NoSQL Databases
8.1 Key-Value Stores
Key → Value
"user:1001" → {"name":"Alice","age":28,"city":"Mumbai"}
Operations: GET(key), PUT(key, value), DELETE(key) — O(1)
Examples: Redis, Amazon DynamoDB, Memcached Use cases: Session management, caching, shopping cart
8.2 Document Databases
{
"_id": "abc123",
"name": "Alice",
"address": {
"street": "123 Main St",
"city": "Mumbai"
},
"orders": [
{"product": "Laptop", "amount": 50000},
{"product": "Mouse", "amount": 1000}
]
}
Examples: MongoDB, CouchDB, Amazon DocumentDB Use cases: Content management, user profiles, e-commerce catalogs
8.3 Column-Family (Wide-Column) Stores
Row Key → Column Family → Columns
user_001 → profile: {name, email, age}
orders: {order_1, order_2}
stats: {login_count, last_seen}
Examples: Apache Cassandra, Google Bigtable, HBase (Hadoop) Use cases: IoT time series, analytics, messaging systems
8.4 Graph Databases
(Alice) --[FRIENDS]-- (Bob)
| |
[LIKES] [BOUGHT]
| |
(Product) (Laptop)
Examples: Neo4j, Amazon Neptune, ArangoDB Use cases: Social networks, recommendation engines, fraud detection
NoSQL vs RDBMS Comparison
| Feature | RDBMS | NoSQL |
|---|---|---|
| Schema | Fixed | Flexible |
| Scaling | Vertical (more RAM/CPU) | Horizontal (more nodes) |
| Consistency | Strong (ACID) | Eventual (BASE) |
| Joins | Supported | Limited/Not supported |
| Query Language | SQL | Various (MongoQL, CQL, etc.) |
| Transactions | Full ACID | Limited |
| Best for | Structured data, OLTP | Large scale, varied data |
Reliable MongoDB Design
For every update, verify whether one or many documents may match. Use atomic operators rather than read-modify-write cycles. Define deterministic sorting and keyset pagination. Inspect explain output and account for index write cost.
MongoDB's flexible document structure is valuable when paired with explicit contracts. Flexibility without validation, type consistency, bounded growth or migration rules merely moves schema problems into every application.
Databases, Collections, Shell and Compass
In mongosh:
use training
The command selects a database context. A database or collection may be created lazily on the first write or a collection can be created explicitly with options:
db.createCollection("products")
MongoDB Compass is a graphical tool for connecting, browsing documents, constructing queries and aggregation pipelines, inspecting schema patterns and managing indexes.
Shell and Compass are clients. Authentication and server authorization govern what each connected account can do.
Atomicity and Transactions
An update to one MongoDB document is atomic, even when it changes several fields or embedded values. This is one reason to embed data that forms one consistency unit.
MongoDB also supports multi-document transactions in supported replica-set and sharded configurations. They add coordination cost and should not compensate for a poor aggregate model.
Use conditional filters and inspect matched counts to detect concurrent conflicts. Configure read and write concerns according to durability and consistency needs.
Equality and Comparison Filters
Direct field equality:
db.products.find({status: "ACTIVE"})
Comparison operators include:
- $eq and $ne;
- $gt and $gte;
- $lt and $lte;
- $in and $nin.
db.products.find({ price: {$gte: NumberDecimal("100.00"), $lte: NumberDecimal("500.00")} })
$in compares against a list:
{status: {$in: ["NEW", "PAID"]}}
Use consistent BSON types in both stored data and query values.
BASE Properties
What is BASE?
BASE is the consistency model used by NoSQL databases (opposite of ACID):
| Letter | Property | Definition |
|---|---|---|
| BA | Basically Available | System guarantees availability (not every request will get the latest data, but a response) |
| S | Soft State | The state of the system may change over time (even without new inputs — due to eventual consistency) |
| E | Eventually Consistent | System will become consistent over time (not immediately; all replicas eventually agree) |
Example of Eventual Consistency:
User posts a photo on Instagram (write to primary replica)
Immediately visible on their phone
After 1-2 seconds: visible globally (replicas catch up)
Eventually consistent — not immediately consistent
Explain and Query Plans
Use explain to inspect access:
db.products.find({ status: "ACTIVE", price: {$gte: NumberDecimal("100.00")} }).explain("executionStats")
Examine index scans, collection scans, keys examined, documents examined, returned rows and execution stages.
An index name alone does not prove efficiency. A broad scan of index keys followed by many document fetches can still be expensive.
Reading with find and findOne
find returns a cursor over matching documents:
db.products.find({stock: {$gt: 0}})
findOne returns one matching document or null:
db.products.findOne({sku: "BK-100"})
Without a sort, which matching document findOne returns is not a stable business rule. Use a unique filter or an explicit ordered query when identity matters.
An empty filter matches all documents:
db.products.find({})
Documents and the _id Field
Every standard collection document has a unique _id field. If an inserted document omits it, the driver or server normally generates an ObjectId.
{ _id: ObjectId("..."), sku: "BK-100", name: "Database Design", price: NumberDecimal("599.00"), tags: ["database", "design"], publisher: { name: "Example Press", city: "Pune" } }
MongoDB automatically creates a unique index on _id. Choose a custom _id only when its uniqueness, immutability, size and distribution suit the workload.
Deleting Documents
deleteOne removes at most one match:
db.products.deleteOne({sku: "BK-100"})
deleteMany removes all matches:
db.session.deleteMany({expiresAt: {$lt: new Date()}})
An empty deleteMany filter removes every document in the collection. Verify filters and deleted counts. Dropping a collection additionally removes its indexes and metadata and is not the same as deleting documents.
Positional Array Updates
MongoDB supplies positional operators for matched array elements. The exact operator depends on whether the update targets the first match, all elements or elements satisfying array filters.
db.orders.updateOne( {_id: orderId}, {$set: {"items.$[item].status": "BACKORDERED"}}, {arrayFilters: [{"item.productId": productId}]} )
Test filters carefully so the intended elements are changed. Complex arrays often signal that aggregate boundaries need review.
OLTP vs OLAP
OLTP — Online Transaction Processing
| Feature | OLTP |
|---|---|
| Purpose | Day-to-day operations; transactions |
| Data | Current, operational data |
| Queries | Simple CRUD; INSERT, UPDATE, DELETE |
| Response time | Milliseconds |
| Users | Thousands of concurrent users (clerks, customers) |
| Data size | GB |
| Normalized | Highly normalized (3NF) |
Examples: Banking transactions, e-commerce orders, inventory management
OLAP — Online Analytical Processing
Examples: Sales trend analysis, customer behavior analysis, financial reporting
OLTP vs OLAP Comparison
Logical Operators
Multiple fields in one filter are implicitly combined with AND:
{status: "ACTIVE", stock: {$gt: 0}}
Explicit operators include $and, $or, $nor and $not:
{ $or: [ {category: "BOOK"}, {price: {$lt: NumberDecimal("100.00")}} ] }
Use $and explicitly when the same field needs conditions that cannot be combined in one object or when generated query structure requires it.
Inserting Documents
Insert one:
db.products.insertOne({ sku: "BK-100", name: "Database Design", price: NumberDecimal("599.00"), stock: 20 })
Insert several:
db.products.insertMany([ {sku: "BK-101", name: "SQL", stock: 10}, {sku: "BK-102", name: "MongoDB", stock: 15} ])
The result reports acknowledged status and inserted identifiers according to write concern. Ordered and unordered bulk behavior determines whether later operations continue after an error.
Aggregation Pipeline
The aggregation pipeline passes documents through stages:
db.orders.aggregate([ {$match: {status: "PAID"}}, {$unwind: "$items"}, {$group: { _id: "$items.productId", quantity: {$sum: "$items.quantity"} }}, {$sort: {quantity: -1}} ])
$match filters, $project reshapes, $unwind emits one document per array element, $group aggregates, $sort orders and $lookup can join collection data.
Place selective $match stages early when semantics permit and inspect multiplication caused by $unwind or $lookup.
Multikey Indexes
Indexing an array field creates a multikey index with entries derived from array elements. This supports array membership queries.
Compound multikey indexes have restrictions when more than one indexed path is an array. Arrays with many elements can produce many index entries and high write cost.
Use $elemMatch where predicates must apply to the same array element. Index design and query semantics must agree.
Unique, Partial and TTL Indexes
A unique index enforces uniqueness:
db.products.createIndex({sku: 1}, {unique: true})
Missing and null behavior needs careful testing, particularly with sparse or partial options.
A partial index contains documents satisfying a filter and can reduce size when queries use the same condition.
A TTL index allows background expiration based on a date field. Deletion is asynchronous rather than exactly at the expiration instant, so it should not be used as a precise scheduler.
Sorting, Limiting and Skipping
db.products.find({status: "ACTIVE"}) .sort({price: -1, _id: 1}) .limit(20)
1 requests ascending order and -1 descending. _id provides a unique tie-breaker.
skip supports offset pagination but becomes expensive for large offsets and can shift under concurrent writes. Keyset pagination applies a filter after the last seen sort key and uses a compatible index.
Sort without a suitable index may require an in-memory or disk-assisted blocking sort and is subject to server limits and options.
Upsert
An upsert updates a matching document or inserts a document when no match exists:
db.products.updateOne( {sku: "BK-103"}, { $set: {name: "Algorithms", price: NumberDecimal("499.00")}, $setOnInsert: {createdAt: new Date()} }, {upsert: true} )
The filter contributes equality fields to the inserted document under defined rules. Use a unique index on the logical identity to prevent concurrent upserts from creating duplicates.
$setOnInsert applies only to the insert branch.
Replacement
replaceOne substitutes an entire document except for immutable _id:
db.products.replaceOne( {sku: "BK-100"}, { sku: "BK-100", name: "Database Design", price: NumberDecimal("649.00"), stock: 20 } )
Fields omitted from the replacement disappear. Use update operators for partial changes and replacement only when the caller intentionally supplies the complete new document.
Update Operators
Common operators include:
- $set assigns or creates fields;
- $unset removes fields;
- $inc atomically increments numeric values;
- $mul multiplies;
- $min and $max change a value conditionally;
- $rename renames a field;
- $currentDate assigns a current date or timestamp.
db.products.updateOne( {sku: "BK-100", stock: {$gte: 2}}, {$inc: {stock: -2}} )
The combined filter and increment form an atomic conditional update on one document, preventing a separate read-then-write stock race.
Updating Documents
updateOne changes the first matching document, while updateMany changes all matches:
db.products.updateOne( {sku: "BK-100"}, {$set: {status: "ACTIVE"}} )
The result distinguishes matched and modified counts. A document can match even when the requested value already equals the stored value, producing a matched count without a modification.
Use a unique filter for updateOne when one known entity is intended. Otherwise “first” is not a stable identity rule.
Projection
Projection controls returned fields:
db.products.find( {status: "ACTIVE"}, {name: 1, price: 1} )
In inclusion projection, _id remains included unless explicitly excluded:
{name: 1, price: 1, _id: 0}
Except for _id, inclusion and exclusion styles generally should not be mixed. Projection reduces network transfer and can support covered queries when an index contains all required filter and result fields.
Arrays
A scalar equality condition can match an array containing that value:
db.products.find({tags: "database"})
$all requires the array to contain every specified value:
{tags: {$all: ["database", "design"]}}
$size tests exact array length. $elemMatch requires one array element to satisfy several conditions:
{ reviews: { $elemMatch: {rating: {$gte: 4}, verified: true} } }
Without $elemMatch, separate dotted predicates might be satisfied by different array elements.
JSON and BSON
MongoDB interfaces display documents in a JSON-like syntax, while the stored binary representation is BSON.
BSON supports types beyond plain JSON, including ObjectId, Date, binary data, Decimal128, regular expression and multiple numeric types.
Type choice matters:
Two visually similar values of different BSON types may not compare as expected.
Schema Validation
A collection validator can require fields and BSON types:
db.createCollection("products", { validator: { $jsonSchema: { bsonType: "object", required: ["sku", "name", "price"], properties: { sku: {bsonType: "string"}, name: {bsonType: "string"}, price: {bsonType: "decimal"} } } } })
Validation can reject or warn according to settings. It provides a shared boundary while allowing intentional document variation.
Array Updates
$push appends an array element:
{$push: {tags: "featured"}}
$addToSet adds only when an equal element is absent. $pull removes matching elements. $pop removes one end.
Modifiers with $push can add several values, limit length, sort and position insertion:
{ $push: { scores: { $each: [90, 95], $sort: -1, $slice: 10 } } }
Arrays that grow without a bound can make documents large and contentious. Move unbounded events to another collection.
Nested Fields
Dot notation addresses nested paths:
db.products.find({"publisher.city": "Pune"})
An exact equality comparison against an embedded document is sensitive to its complete value and field order. Dot-path predicates are usually better for individual nested properties.
Updates use the same path:
{$set: {"publisher.city": "Mumbai"}}
Without dot notation, setting publisher would replace the entire embedded object.
Missing Fields and NULL
MongoDB distinguishes a field that is absent from one whose value is null, but some query forms can match both:
{middleName: null}
For explicit presence tests:
{middleName: {$exists: false}} {middleName: {$exists: true}}
Combine $exists and type/value conditions when the distinction matters. Schema validation can require fields and prevent inconsistent absence conventions.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.