Database Technologies
NoSQL Models, CAP, BASE and Data Design
PGCP-AC
1. NoSQL Databases
NoSQL is a broad label for database systems whose primary data model is not the traditional relational table model. The term includes key-value, document, wide-column and graph systems with different query languages and guarantees.
NoSQL does not mean that SQL can never be supported, that schemas do not exist or that transactions are impossible. Modern products overlap. The exact database version, operation, topology and configuration determine actual behavior.
Select a model from data shape, access patterns, consistency, scale, latency and operational requirements rather than from category alone.
2. Structured, Semi-Structured and Unstructured Data
Structured data follows a defined organization of fields and types, such as relational rows.
Semi-structured data contains recognizable fields and nesting but allows variation among records. JSON documents, event messages and XML are common examples.
Unstructured data lacks a convenient fixed field model for its primary content, such as images, audio and free-form documents. Metadata around that content can still be structured.
These descriptions are a spectrum. A JSON collection can enforce strict validation, while a relational database can store flexible JSON fields.
3. Key-Value Stores
A key-value database maps a unique key to a value:
cart:customer:42 -> encoded cart data
Its central operation is direct retrieval by known key. Values may be opaque bytes or structured objects depending on the product.
Key-value systems fit caches, sessions, counters, feature state and objects naturally addressed by one identifier. They are less natural for ad hoc predicates across arbitrary value fields unless secondary indexing or additional structures are provided.
Key design controls distribution, locality and access. A poorly chosen hot key can concentrate load.
4. Document Databases
A document database stores self-contained structured documents, commonly in JSON-like or BSON form:
{
"_id": 42,
"customer": {
"name": "Anita Rao",
"email": "anita@example.com"
},
"items": [
{"productId": 8, "quantity": 2, "price": 120.00}
],
"status": "PAID"
}
Documents can contain nested objects and arrays. Fields can vary across documents, while validation rules can still enforce required structure and types.
Document stores fit aggregates commonly read and changed together.
5. Wide-Column Databases
Wide-column systems organize data around partition keys, clustering or sort keys and flexible column sets. Their design is strongly query-driven.
Rows sharing a partition key are stored together conceptually, while clustering keys order data inside the partition. Efficient queries usually specify the partition key and compatible ranges.
They suit high-volume distributed workloads with known access paths, such as time-series events by device and time. They are not merely relational tables with many nullable columns.
Partition size and hot-partition risk must be planned.
6. Graph Databases
A property graph contains vertices, edges, labels and properties. Relationships are stored as first-class structures rather than reconstructed mainly through foreign-key joins.
Graph databases fit variable-depth traversal and relationship-centered questions:
- paths between users;
- dependency chains;
- fraud rings;
- authorization relationships;
- network topology.
A graph is not automatically superior for every connected dataset. Simple one-hop relationships and aggregate transactions may remain clearer in a relational or document model.
7. Relational and NoSQL Differences
Relational design commonly normalizes facts into relations, enforces constraints centrally and reconstructs related results with joins.
NoSQL designs often organize data around aggregates and known access patterns, duplicate selected facts and distribute data through application-visible keys.
Relational databases can scale and store JSON. NoSQL systems can support transactions, validation, indexes and SQL-like querying. The useful comparison is between concrete guarantees and operational models, not stereotypes.
8. Schema Flexibility
A schemaless write interface does not mean data has no schema. Applications still expect field names, types, versions and relationships.
Without database validation, schema responsibility moves to writers, readers, migration jobs and monitoring. Old and new document shapes may coexist and every reader must handle that evolution.
Validation at the database boundary reduces malformed data. Store an explicit schema version when documents evolve significantly and provide tested migrations or compatible readers.
9. Aggregate-Oriented Design
An aggregate is a cluster of data treated as one consistency and access unit. An order document may contain its line items because they are created, read and owned with the order.
Aggregate boundaries affect:
- atomic update scope;
- document size;
- duplication;
- distribution;
- query patterns;
- contention.
An aggregate should not grow without bound. A customer document containing every lifetime event may exceed limits and create one hot record.
10. Embedding
Embedding places related data inside one parent document. It can provide one-read retrieval and atomic updates within the document.
Good candidates are bounded child data with the same lifecycle and dominant read pattern as the parent, such as an order's shipping address snapshot and line items.
Embedding is less suitable for large unbounded collections, independently shared entities or child data updated frequently across many parents.
11. Referencing
Referencing stores another entity's identifier:
{
"orderId": 9001,
"customerId": 42
}
It avoids duplicating current customer state and supports independently changing entities. Reading both may require another query, aggregation pipeline, application join or cache.
References do not automatically enforce referential integrity in every document product. Applications or transactions may need to prevent dangling identifiers.
12. Historical Snapshots and Current Facts
Duplication has different meanings. An order's customer name and address at purchase time may be an intentional historical snapshot. It should not change when the customer later edits a profile.
A copied current customer email used in several active views is a cache and may need propagation after every update.
Distinguish immutable event-time facts from duplicated current facts. The former preserves history; the latter creates synchronization work.
13. Denormalization
NoSQL designs commonly duplicate data to answer important queries without expensive distributed joins. Denormalization can reduce read latency and cross-partition work.
It increases:
- storage;
- number of write locations;
- partial-failure paths;
- stale-read possibilities;
- migration effort;
- repair requirements.
Document the authoritative copy, propagation mechanism, acceptable lag, conflict policy and reconciliation process.
14. Partitioning
Partitioning distributes records across nodes according to a partition or shard key. A good key spreads storage and requests while supporting common queries.
A monotonically increasing or low-cardinality key can concentrate writes. A highly distributed key may scatter a query that needs a related range.
Changing a partition key is often expensive because data must move. Evaluate cardinality, access locality, growth, hot tenants and rebalancing before committing to it.
15. Replication
Replication keeps copies of data on several nodes for availability, read scaling and fault tolerance.
Synchronous coordination can provide stronger immediate agreement but adds latency and can reject work when nodes cannot communicate. Asynchronous replication improves local responsiveness but permits lag and stale reads.
Replication is not backup. A mistaken deletion can propagate to every replica. Independent backups and recovery history remain necessary.
16. CAP Theorem
CAP concerns a distributed read/write data system during a network partition. It states that the system cannot simultaneously guarantee both:
- Consistency: operations behave like one up-to-date linearizable copy;
- Availability: every request to a nonfailed node receives a non-error response;
while also continuing across the partition.
Partition tolerance means the distributed system continues according to its chosen policy despite lost or delayed communication between node groups.
17. CAP Consistency
Consistency in CAP means linearizability, not merely valid schema or eventual replica convergence. A completed write should appear to take effect at one instant and later reads should observe it according to real-time ordering.
This differs from ACID consistency, which concerns preservation of application invariants. The same word names different concepts.
A system may offer several consistency levels per operation, such as local, quorum or linearizable access. Avoid labeling an entire product with one simplistic letter pair.
18. CAP Availability
CAP availability means every request received by a nonfailed node eventually returns a successful response, though it need not contain the newest value if consistency is sacrificed.
Returning a timeout or explicit failure to preserve consistency does not meet CAP availability for that operation during the partition.
Ordinary uptime percentage and low latency are important operational measures but are not the precise theorem definition.
19. Partition Tradeoffs
During a partition, a system prioritizing linearizable consistency may reject or delay operations in a side that cannot coordinate safely.
A system prioritizing availability may accept operations on separated sides, creating divergent versions that require later convergence or conflict resolution.
Outside partitions, systems still face latency, durability, cost and failure tradeoffs. CAP is not “choose any two forever”; partitioned systems must decide how particular operations behave when communication is disrupted.
20. Consistency Models
Consistency forms a spectrum:
- strong or linearizable reads observe real-time ordered writes;
- sequential consistency preserves one global operation order without full real-time constraints;
- causal consistency preserves cause-before-effect relationships;
- read-your-writes lets a client observe its own completed changes;
- monotonic reads prevent one client from moving backward to an older version;
- eventual consistency promises convergence under stated propagation assumptions.
Applications should choose the minimum sufficient guarantee per operation. Account balances and social-feed counters have different risks.
21. Eventual Consistency
Eventual consistency does not promise immediate replica agreement. Broadly, if updates cease and communication and repair assumptions hold, replicas converge.
It does not specify how long convergence takes, which value a read sees during convergence or how concurrent conflicts are resolved. Those are product and configuration contracts.
Permanent partition or failed propagation can prevent convergence. Monitor replication lag and repair failures rather than treating “eventual” as automatic.
22. Conflict Resolution
Available writes on separated replicas can conflict. Resolution strategies include:
- last-write-wins using timestamps or logical clocks;
- application merges;
- version vectors;
- conflict-free replicated data types for supported operations;
- manual review.
Last-write-wins is simple but can discard a legitimate concurrent update and depends on time semantics. Choose conflict handling from domain meaning and test concurrent scenarios.
23. BASE
BASE is commonly expanded as Basically Available, Soft state, Eventually consistent. It describes a family of distributed design choices emphasizing availability and asynchronous convergence.
Soft state means observed state can change as background propagation or reconciliation occurs, even without a new local user action.
BASE is not a formal opposite of all ACID behavior. A system can provide ACID transactions within an aggregate while replicating results asynchronously across regions.
24. Transactions in NoSQL Systems
NoSQL systems vary widely. Some provide atomicity only within one key or document. Others support multi-document or distributed transactions at additional cost and with restrictions.
Before relying on a transaction, verify:
- operation and object scope;
- isolation model;
- conflict behavior;
- timeout and retry semantics;
- durability and acknowledgment level;
- partition behavior.
Do not infer these guarantees from the NoSQL label.
25. Polyglot Persistence
One system can use different databases for different bounded responsibilities: a relational database for orders, a document store for content, a search engine for text discovery and a cache for short-lived responses.
This can align models with workloads but increases operational complexity, data movement, security surfaces and consistency coordination.
Use another database only when its benefit exceeds the cost of operating and reconciling another source of state.
26. Designing a NoSQL Model
List required reads and writes with expected volume, latency and consistency. Define aggregate boundaries and maximum sizes. Choose partition keys from distribution and locality. Decide what to embed, reference or duplicate.
For every duplicate, name the authority and repair path. For every cross-aggregate rule, state whether it uses a transaction, conditional write, asynchronous process or compensation.
Then test hot keys, network partitions, stale reads, concurrent writers, retries and schema evolution. A flexible data model works only when its consistency and lifecycle rules are precise.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.