Data Collection and DBMS
MongoDB Aggregation Pipelines and Performance
PGCP-BDA
MongoDB aggregation pipeline
An aggregation pipeline passes documents through ordered stages that filter, reshape, join, unwind, group, sort and calculate results.
match project group
In a MongoDB pipeline, $match filters documents, $project reshapes fields and $group aggregates documents by a key.
unwind
MongoDB $unwind emits one output document for each element of an array field.
lookup
MongoDB $lookup performs a left outer join with documents from another collection and stores matches in an array field.
sort limit skip
MongoDB $sort orders documents, $skip discards an initial count and $limit restricts the number returned.
accumulator
An aggregation operator such as $sum or $avg that combines values while documents in a group are processed.
pipeline optimization
Reordering, combining or pushing down pipeline stages to reduce documents, fields, memory and work without changing results.
aggregation memory
Working memory used by blocking aggregation stages; large workloads may require indexes, smaller inputs or permitted disk spilling.
aggregation result validation
Checking pipeline output against expected counts, keys, totals, samples and edge cases to confirm semantic correctness.
Aggregation Pipeline
The aggregation pipeline passes documents through stages:
db.orders.aggregate([ {$match: {status: "PAID"}}, {$unwind: "$items"}, {$group: { _id: "$items.productId", quantity: {$sum: "$items.quantity"} }}, {$sort: {quantity: -1}} ])
$match filters, $project reshapes, $unwind emits one document per array element, $group aggregates, $sort orders and $lookup can join collection data.
Place selective $match stages early when semantics permit and inspect multiplication caused by $unwind or $lookup.
Sorting, Limiting and Skipping
db.products.find({status: "ACTIVE"}) .sort({price: -1, _id: 1}) .limit(20)
1 requests ascending order and -1 descending. _id provides a unique tie-breaker.
skip supports offset pagination but becomes expensive for large offsets and can shift under concurrent writes. Keyset pagination applies a filter after the last seen sort key and uses a compatible index.
Sort without a suitable index may require an in-memory or disk-assisted blocking sort and is subject to server limits and options.
Schema Validation
A collection validator can require fields and BSON types:
db.createCollection("products", { validator: { $jsonSchema: { bsonType: "object", required: ["sku", "name", "price"], properties: { sku: {bsonType: "string"}, name: {bsonType: "string"}, price: {bsonType: "decimal"} } } } })
Validation can reject or warn according to settings. It provides a shared boundary while allowing intentional document variation.
Missing Fields and NULL
MongoDB distinguishes a field that is absent from one whose value is null, but some query forms can match both:
{middleName: null}
For explicit presence tests:
{middleName: {$exists: false}} {middleName: {$exists: true}}
Combine $exists and type/value conditions when the distinction matters. Schema validation can require fields and prevent inconsistent absence conventions.
Aggregation Pipeline Execution
A MongoDB aggregation pipeline passes documents through ordered stages. $match filters, $project selects or computes fields, $unwind emits one document per array element and $group aggregates by a key. $sort, $limit and $skip control order and range. $lookup joins another collection. Stage order affects meaning and cost. An early selective match can reduce later work while grouping before filtering may process unnecessary documents.
Expressions refer to fields with a dollar prefix. Accumulators such as $sum, $avg, $min, $max, $push and $addToSet operate within groups. Missing fields, null values and arrays need deliberate handling. A pipeline may produce a schema entirely different from its input.
Indexes and Performance
An index stores ordered keys with document references. A compound index supports prefixes of its key order. Equality fields commonly precede sort or range fields for a query pattern. Multikey indexes support arrays. A covered query obtains required fields from the index without fetching full documents.
explain reports the selected plan, examined keys, examined documents and returned records. A collection scan may be reasonable for a small or unselective collection, but many examined documents for few returned results indicates waste. Place selective stages early and discard large unused fields before expensive stages. Sorting and grouping consume memory and may spill to disk. $lookup benefits from indexed join fields and bounded cardinality. Every index improves selected reads at the cost of storage and write maintenance.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.