Big Data Technologies
HDFS Read and Write Paths, Replication, Rack Awareness and High Availability
PGCP-BDA
HDFS read path
An HDFS client obtains block locations from the NameNode, reads replicas directly from nearby DataNodes.
HDFS write pipeline
An HDFS client requests block targets, sends packets through an ordered DataNode pipeline and receives acknowledgements in reverse after replicas persist.
block placement
HDFS block placement chooses replica nodes using available space, load and rack topology so common failures do not remove every copy.
rack awareness
Rack awareness lets Hadoop place replicas and schedule work using network topology so a rack failure does not remove all copies and cross-rack traffic is.
checksum
HDFS stores checksums for chunks of block data and verifies them during transfer to detect corruption independently of a successful disk or network read.
replication recovery
When replicas are missing or corrupt, the NameNode schedules copying from a healthy replica
NameNode high availability
HDFS high availability uses active and standby NameNodes sharing edits through JournalNodes.
JournalNode
A JournalNode is a member of the quorum that durably stores shared namespace edits used to keep an HDFS standby NameNode synchronized.
Standby NameNode
A standby NameNode continuously applies shared edits and block reports so it can take over the namespace after controlled or automatic failover.
federation
HDFS federation uses several independent NameNodes and block pools to scale namespace throughput while DataNodes store blocks for all pools.
Index Write Costs
Each INSERT must add entries to every relevant index. UPDATE must change indexes containing modified keys. DELETE must remove entries. More and wider indexes increase:
- write I/O and CPU;
- transaction duration;
- buffer-pool use;
- lock and latch pressure;
- storage and backup size;
- schema-change time.
Duplicate or unused indexes should be identified through evidence and removed carefully after checking constraint roles and workload history.
Primary, Unique and Ordinary Indexes
A primary-key index enforces primary identity. A unique index enforces uniqueness of its key values under MySQL's NULL rules. An ordinary index supplies an access path without a uniqueness rule.
Constraint and performance roles overlap but are not identical. Do not remove a required unique constraint merely because no current query appears to use its index. It protects data integrity.
Foreign-key columns often benefit from indexes because parent checks, child lookups, joins and referential actions use them. MySQL also imposes index requirements for enforced foreign keys.
The Query Optimizer
The optimizer considers access methods, indexes, join orders, join algorithms, predicate placement, sorting, grouping and materialization. It estimates costs using metadata and statistics.
Cost estimates can be wrong when statistics are stale, data is skewed, predicates are correlated or parameter values vary widely.
The optimizer does not understand unstated business assumptions. Correct keys and constraints can improve both integrity and available optimization information.
An Index Design Method
For an important query:
- state result correctness and row grain;
- identify equality predicates and join keys;
- identify range predicates;
- identify ordering and grouping;
- decide whether coverage is worthwhile;
- propose the narrowest reusable composite index;
- examine EXPLAIN and actual execution;
- test write impact and other queries.
Index design is workload design. The best index is not the one with the most columns; it is the smallest maintainable structure that measurably supports important correct queries.
Indexes as Access Structures
An index is an auxiliary structure that helps the DBMS locate rows without examining every table record. It stores indexed key values with information that leads to corresponding rows.
CREATE INDEX idx_employee_department ON employee(department_id);
The index can support queries that search, join, group or order by department_id. It does not change relational meaning and does not guarantee output order without ORDER BY.
An index trades faster access for storage, maintenance during writes, memory use and administrative work.
Prefix Indexes
MySQL can index a prefix of a long string:
CREATE INDEX idx_customer_email_prefix ON customer(email(20));
This reduces index size but can decrease selectivity and cannot enforce full-value uniqueness unless the indexed prefix itself is guaranteed unique.
Choose prefix length from data-distribution measurements, not guesswork. Modern limits depend on engine, row format, character set and version.
Query Shape Before Indexes
Indexes cannot repair an incorrect query. First verify join conditions, result grain, NULL behavior and predicates.
Reduce unnecessary columns and rows, pre-aggregate many-side data when appropriate, avoid repeated correlated work and use EXISTS for existence. Then index the resulting stable access patterns.
Adding indexes to compensate for SELECT *, accidental cross joins or non-sargable conversions creates maintenance cost without addressing the root problem.
Join Indexes
For a join:
FROM order_header AS o JOIN order_line AS l ON l.order_id = o.order_id
an index on order_line(order_id) supports finding lines for each order. The order_header primary key already supports the opposite lookup.
Composite indexes can include additional join filters:
order_line(order_id, product_id)
Choose direction from likely driving tables and predicates. Indexing every join column separately may be inferior to a composite index matching the full access pattern.
Sargable Predicates
A sargable predicate can be translated into a direct index search argument. Equality, bounded ranges and prefix LIKE patterns often qualify.
Common obstacles include:
- wrapping the indexed column in a function;
- implicit conversion between incompatible types;
- arithmetic on the indexed column;
- a leading wildcard such as LIKE '%text';
- conditions not aligned with a composite prefix.
Rewrite without changing semantics. A computed indexed column may be appropriate when the transformed value is a frequent search key.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.