Data Collection and DBMS

Cassandra Architecture, Data Modeling, CQL and Operations

PGCP-BDA

Cassandra cluster

A peer-to-peer set of nodes that partitions rows by token and replicates them across failure domains without a permanent master.

peer-to-peer architecture

A Cassandra design in which nodes have equal roles and coordinate requests without a permanent master.

partitioner and token ring

The Cassandra partitioner hashes a partition key to a token; token ranges assigned to nodes determine data placement around the logical ring.

replication factor

The replication factor is the desired number of HDFS block copies

consistency level

A Cassandra consistency level specifies how many replicas must acknowledge a read or write.

partition key and clustering columns

The partition key selects the Cassandra partition and nodes, while clustering columns order rows within that partition and support bounded range access.

CQL

Cassandra Query Language defines keyspaces and tables and reads or changes partitioned Cassandra data using SQL-like syntax.

tombstone

A Cassandra tombstone is a timestamped deletion marker retained until replicas have had time to learn the deletion and compaction can safely discard it.

compaction

The background merging of immutable Cassandra SSTables to discard obsolete versions, reclaim space and improve reads.

repair and operations

Cassandra maintenance that reconciles replicas, monitors health, manages capacity and safely replaces or removes nodes.

Eventual Consistency

Eventual consistency does not promise immediate replica agreement. Broadly, if updates cease and communication and repair assumptions hold, replicas converge.

It does not specify how long convergence takes, which value a read sees during convergence or how concurrent conflicts are resolved. Those are product and configuration contracts.

Permanent partition or failed propagation can prevent convergence. Monitor replication lag and repair failures rather than treating “eventual” as automatic.

CAP Consistency

Consistency in CAP means linearizability, not merely valid schema or eventual replica convergence. A completed write should appear to take effect at one instant and later reads should observe it according to real-time ordering.

This differs from ACID consistency, which concerns preservation of application invariants. The same word names different concepts.

A system may offer several consistency levels per operation, such as local, quorum or linearizable access. Avoid labeling an entire product with one simplistic letter pair.

Wide-Column Databases

Wide-column systems organize data around partition keys, clustering or sort keys and flexible column sets. Their design is strongly query-driven.

Rows sharing a partition key are stored together conceptually, while clustering keys order data inside the partition. Efficient queries usually specify the partition key and compatible ranges.

They suit high-volume distributed workloads with known access paths, such as time-series events by device and time. They are not merely relational tables with many nullable columns.

Partition size and hot-partition risk must be planned.

Partition Tradeoffs

During a partition, a system prioritizing linearizable consistency may reject or delay operations in a side that cannot coordinate safely.

A system prioritizing availability may accept operations on separated sides, creating divergent versions that require later convergence or conflict resolution.

Outside partitions, systems still face latency, durability, cost and failure tradeoffs. CAP is not “choose any two forever”; partitioned systems must decide how particular operations behave when communication is disrupted.

Consistency Models

Consistency forms a spectrum:

  • strong or linearizable reads observe real-time ordered writes;
  • sequential consistency preserves one global operation order without full real-time constraints;
  • causal consistency preserves cause-before-effect relationships;
  • read-your-writes lets a client observe its own completed changes;
  • monotonic reads prevent one client from moving backward to an older version;
  • eventual consistency promises convergence under stated propagation assumptions.

Applications should choose the minimum sufficient guarantee per operation. Account balances and social-feed counters have different risks.

Replication

Replication keeps copies of data on several nodes for availability, read scaling and fault tolerance.

Synchronous coordination can provide stronger immediate agreement but adds latency and can reject work when nodes cannot communicate. Asynchronous replication improves local responsiveness but permits lag and stale reads.

Replication is not backup. A mistaken deletion can propagate to every replica. Independent backups and recovery history remain necessary.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.