Data Collection and DBMS

MongoDB Aggregation Pipelines and Performance

PGCP-BDA

MongoDB aggregation pipeline

An aggregation pipeline passes documents through ordered stages that filter, reshape, join, unwind, group, sort and calculate results.

match project group

In a MongoDB pipeline, $match filters documents, $project reshapes fields and $group aggregates documents by a key.

unwind

MongoDB $unwind emits one output document for each element of an array field.

lookup

MongoDB $lookup performs a left outer join with documents from another collection and stores matches in an array field.

sort limit skip

MongoDB $sort orders documents, $skip discards an initial count and $limit restricts the number returned.

accumulator

An aggregation operator such as $sum or $avg that combines values while documents in a group are processed.

pipeline optimization

Reordering, combining or pushing down pipeline stages to reduce documents, fields, memory and work without changing results.

aggregation memory

Working memory used by blocking aggregation stages; large workloads may require indexes, smaller inputs or permitted disk spilling.

aggregation result validation

Checking pipeline output against expected counts, keys, totals, samples and edge cases to confirm semantic correctness.

Aggregation Pipeline

The aggregation pipeline passes documents through stages:

db.orders.aggregate([ {$match: {status: "PAID"}}, {$unwind: "$items"}, {$group: { _id: "$items.productId", quantity: {$sum: "$items.quantity"} }}, {$sort: {quantity: -1}} ])

$match filters, $project reshapes, $unwind emits one document per array element, $group aggregates, $sort orders and $lookup can join collection data.

Place selective $match stages early when semantics permit and inspect multiplication caused by $unwind or $lookup.

Sorting, Limiting and Skipping

db.products.find({status: "ACTIVE"}) .sort({price: -1, _id: 1}) .limit(20)

1 requests ascending order and -1 descending. _id provides a unique tie-breaker.

skip supports offset pagination but becomes expensive for large offsets and can shift under concurrent writes. Keyset pagination applies a filter after the last seen sort key and uses a compatible index.

Sort without a suitable index may require an in-memory or disk-assisted blocking sort and is subject to server limits and options.

Schema Validation

A collection validator can require fields and BSON types:

db.createCollection("products", { validator: { $jsonSchema: { bsonType: "object", required: ["sku", "name", "price"], properties: { sku: {bsonType: "string"}, name: {bsonType: "string"}, price: {bsonType: "decimal"} } } } })

Validation can reject or warn according to settings. It provides a shared boundary while allowing intentional document variation.

Missing Fields and NULL

MongoDB distinguishes a field that is absent from one whose value is null, but some query forms can match both:

{middleName: null}

For explicit presence tests:

{middleName: {$exists: false}} {middleName: {$exists: true}}

Combine $exists and type/value conditions when the distinction matters. Schema validation can require fields and prevent inconsistent absence conventions.

Aggregation Pipeline Execution

A MongoDB aggregation pipeline passes documents through ordered stages. $match filters, $project selects or computes fields, $unwind emits one document per array element and $group aggregates by a key. $sort, $limit and $skip control order and range. $lookup joins another collection. Stage order affects meaning and cost. An early selective match can reduce later work while grouping before filtering may process unnecessary documents.

Expressions refer to fields with a dollar prefix. Accumulators such as $sum, $avg, $min, $max, $push and $addToSet operate within groups. Missing fields, null values and arrays need deliberate handling. A pipeline may produce a schema entirely different from its input.

Indexes and Performance

An index stores ordered keys with document references. A compound index supports prefixes of its key order. Equality fields commonly precede sort or range fields for a query pattern. Multikey indexes support arrays. A covered query obtains required fields from the index without fetching full documents.

explain reports the selected plan, examined keys, examined documents and returned records. A collection scan may be reasonable for a small or unselective collection, but many examined documents for few returned results indicates waste. Place selective stages early and discard large unused fields before expensive stages. Sorting and grouping consume memory and may spill to disk. $lookup benefits from indexed join fields and bounded cardinality. Every index improves selected reads at the cost of storage and write maintenance.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.