Data Collection and DBMS

NoSQL Models, Storage Architectures and Schema Evolution

PGCP-BDA

NoSQL

NoSQL describes nonrelational data stores designed around models such as key-value, document, wide-column or graph and selected distribution tradeoffs.

key-value database

A database that retrieves an opaque value by a unique key and usually favors simple, horizontally scalable access.

document database

A database that stores self-describing hierarchical documents and queries fields within those documents.

wide-column database

A distributed store that organizes sparse rows by row key and groups related columns into column families.

graph database

A database that represents entities as vertices and relationships as edges and supports relationship traversal.

CAP theorem

During a network partition, a distributed data service cannot guarantee both linearizable consistency and availability for every request.

What is CAP Theorem?

CAP Theorem (Brewer's Theorem, 2000): A distributed system can guarantee at most 2 out of 3 properties simultaneously:

         C (Consistency)
             /\
            /  \
           /    \
          /      \
CA ------+--------+ CP
        /          \
       /            \
      /      AP      \
     +________________+
A (Availability)   P (Partition Tolerance)

Network partitions ALWAYS happen in distributed systems → you must choose between C and A.

Database Classification by CAP

TypePrioritizesExamples
CPConsistency + Partition Tolerance (sacrifice Availability)HBase, MongoDB, Redis
APAvailability + Partition Tolerance (sacrifice Consistency)Cassandra, DynamoDB, CouchDB
CAConsistency + Availability (only on single node — no partition)Traditional RDBMS

CAP concerns a distributed read/write data system during a network partition. It states that the system cannot simultaneously guarantee both:

  • Consistency: operations behave like one up-to-date linearizable copy;
  • Availability: every request to a nonfailed node receives a non-error response;

while also continuing across the partition.

Partition tolerance means the distributed system continues according to its chosen policy despite lost or delayed communication between node groups.

BASE

Basically Available, Soft state and Eventual consistency describes systems that may accept temporary replica divergence so availability can be favored.

BASE is commonly expanded as Basically Available, Soft state, Eventually consistent. It describes a family of distributed design choices emphasizing availability and asynchronous convergence.

Soft state means observed state can change as background propagation or reconciliation occurs, even without a new local user action.

BASE is not a formal opposite of all ACID behavior. A system can provide ACID transactions within an aggregate while replicating results asynchronously across regions.

schema evolution

Controlled change to a data structure while preserving compatibility, migration safety and the meaning of existing data.

polyglot persistence

Using different database models for different application workloads according to their access and consistency needs.

One system can use different databases for different bounded responsibilities: a relational database for orders, a document store for content, a search engine for text discovery and a cache for short-lived responses.

This can align models with workloads but increases operational complexity, data movement, security surfaces and consistency coordination.

Use another database only when its benefit exceeds the cost of operating and reconciling another source of state.

NoSQL data modeling

Designing records around access patterns, partition keys, denormalization and the consistency behavior of a NoSQL store.

NoSQL Databases

NoSQL is a broad label for database systems whose primary data model is not the traditional relational table model. The term includes key-value, document, wide-column and graph systems with different query languages and guarantees.

NoSQL does not mean that SQL can never be supported, that schemas do not exist or that transactions are impossible. Modern products overlap. The exact database version, operation, topology and configuration determine actual behavior.

Select a model from data shape, access patterns, consistency, scale, latency and operational requirements rather than from category alone.

Graph Databases

A property graph contains vertices, edges, labels and properties. Relationships are stored as first-class structures rather than reconstructed mainly through foreign-key joins.

Graph databases fit variable-depth traversal and relationship-centered questions:

  • paths between users;
  • dependency chains;
  • fraud rings;
  • authorization relationships;
  • network topology.

A graph is not automatically superior for every connected dataset. Simple one-hop relationships and aggregate transactions may remain clearer in a relational or document model.

Schema Flexibility

A schemaless write interface does not mean data has no schema. Applications still expect field names, types, versions and relationships.

Without database validation, schema responsibility moves to writers, readers, migration jobs and monitoring. Old and new document shapes may coexist and every reader must handle that evolution.

Validation at the database boundary reduces malformed data. Store an explicit schema version when documents evolve significantly and provide tested migrations or compatible readers.

Designing a NoSQL Model

List required reads and writes with expected volume, latency and consistency. Define aggregate boundaries and maximum sizes. Choose partition keys from distribution and locality. Decide what to embed, reference or duplicate.

For every duplicate, name the authority and repair path. For every cross-aggregate rule, state whether it uses a transaction, conditional write, asynchronous process or compensation.

Then test hot keys, network partitions, stale reads, concurrent writers, retries and schema evolution. A flexible data model works only when its consistency and lifecycle rules are precise.

BASE Properties

What is BASE?

BASE is the consistency model used by NoSQL databases (opposite of ACID):

Document Databases

A document database stores self-contained structured documents, commonly in JSON-like or BSON form:

{ "_id": 42, "customer": { "name": "Anita Rao", "email": "anita@example.com" }, "items": [ {"productId": 8, "quantity": 2, "price": 120.00} ], "status": "PAID" }

Documents can contain nested objects and arrays. Fields can vary across documents, while validation rules can still enforce required structure and types.

Document stores fit aggregates commonly read and changed together.

Key-Value Stores

A key-value database maps a unique key to a value:

cart:customer:42 -> encoded cart data

Its central operation is direct retrieval by known key. Values may be opaque bytes or structured objects depending on the product.

Key-value systems fit caches, sessions, counters, feature state and objects naturally addressed by one identifier. They are less natural for ad hoc predicates across arbitrary value fields unless secondary indexing or additional structures are provided.

Key design controls distribution, locality and access. A poorly chosen hot key can concentrate load.

CAP Availability

CAP availability means every request received by a nonfailed node eventually returns a successful response, though it need not contain the newest value if consistency is sacrificed.

Returning a timeout or explicit failure to preserve consistency does not meet CAP availability for that operation during the partition.

Ordinary uptime percentage and low latency are important operational measures but are not the precise theorem definition.

Transactions in NoSQL Systems

NoSQL systems vary widely. Some provide atomicity only within one key or document. Others support multi-document or distributed transactions at additional cost and with restrictions.

Before relying on a transaction, verify:

  • operation and object scope;
  • isolation model;
  • conflict behavior;
  • timeout and retry semantics;
  • durability and acknowledgment level;
  • partition behavior.

Do not infer these guarantees from the NoSQL label.

Relational and NoSQL Differences

Relational design commonly normalizes facts into relations, enforces constraints centrally and reconstructs related results with joins.

NoSQL designs often organize data around aggregates and known access patterns, duplicate selected facts and distribute data through application-visible keys.

Relational databases can scale and store JSON. NoSQL systems can support transactions, validation, indexes and SQL-like querying. The useful comparison is between concrete guarantees and operational models, not stereotypes.

Structured, Semi-Structured and Unstructured Data

Structured data follows a defined organization of fields and types, such as relational rows.

Semi-structured data contains recognizable fields and nesting but allows variation among records. JSON documents, event messages and XML are common examples.

Unstructured data lacks a convenient fixed field model for its primary content, such as images, audio and free-form documents. Metadata around that content can still be structured.

These descriptions are a spectrum. A JSON collection can enforce strict validation, while a relational database can store flexible JSON fields.

Denormalization

NoSQL designs commonly duplicate data to answer important queries without expensive distributed joins. Denormalization can reduce read latency and cross-partition work.

It increases:

  • storage;
  • number of write locations;
  • partial-failure paths;
  • stale-read possibilities;
  • migration effort;
  • repair requirements.

Document the authoritative copy, propagation mechanism, acceptable lag, conflict policy and reconciliation process.

Aggregate-Oriented Design

An aggregate is a cluster of data treated as one consistency and access unit. An order document may contain its line items because they are created, read and owned with the order.

Aggregate boundaries affect:

  • atomic update scope;
  • document size;
  • duplication;
  • distribution;
  • query patterns;
  • contention.

An aggregate should not grow without bound. A customer document containing every lifetime event may exceed limits and create one hot record.

Embedding

Embedding places related data inside one parent document. It can provide one-read retrieval and atomic updates within the document.

Good candidates are bounded child data with the same lifecycle and dominant read pattern as the parent, such as an order's shipping address snapshot and line items.

Embedding is less suitable for large unbounded collections, independently shared entities or child data updated frequently across many parents.

Partitioning

Partitioning distributes records across nodes according to a partition or shard key. A good key spreads storage and requests while supporting common queries.

A monotonically increasing or low-cardinality key can concentrate writes. A highly distributed key may scatter a query that needs a related range.

Changing a partition key is often expensive because data must move. Evaluate cardinality, access locality, growth, hot tenants and rebalancing before committing to it.

Referencing

Referencing stores another entity's identifier:

{ "orderId": 9001, "customerId": 42 }

It avoids duplicating current customer state and supports independently changing entities. Reading both may require another query, aggregation pipeline, application join or cache.

References do not automatically enforce referential integrity in every document product. Applications or transactions may need to prevent dangling identifiers.

Conflict Resolution

Available writes on separated replicas can conflict. Resolution strategies include:

  • last-write-wins using timestamps or logical clocks;
  • application merges;
  • version vectors;
  • conflict-free replicated data types for supported operations;
  • manual review.

Last-write-wins is simple but can discard a legitimate concurrent update and depends on time semantics. Choose conflict handling from domain meaning and test concurrent scenarios.

Historical Snapshots and Current Facts

Duplication has different meanings. An order's customer name and address at purchase time may be an intentional historical snapshot. It should not change when the customer later edits a profile.

A copied current customer email used in several active views is a cache and may need propagation after every update.

Distinguish immutable event-time facts from duplicated current facts. The former preserves history; the latter creates synchronization work.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.