Big Data Technologies

Hadoop Cluster Setup, Configuration, Security, Administration and Monitoring

PGCP-BDA

Hadoop cluster planning

Sizing and topology decisions for masters, workers, storage, memory, network, replication, workload growth and failure tolerance.

core-site hdfs-site mapred-site yarn-site

Hadoop configuration files for common settings, HDFS, MapReduce and YARN resource-management behavior respectively.

worker node

A cluster machine that stores data blocks, runs assigned tasks or performs both roles under coordination by master services.

cluster startup

The ordered launch and verification of storage, resource-management, history and supporting services before jobs are accepted.

security and Kerberos

Kerberos provides strong service and user authentication through tickets; Hadoop authorization then controls permitted operations.

commission and decommission

Commissioning safely adds a node; decommissioning migrates or replicates its data and work before the node leaves service.

capacity and quota

Capacity defines available storage or compute resources, while quotas limit the amount a user, directory or queue may consume.

monitoring

Continuous collection and interpretation of health, capacity, latency, throughput, error and resource metrics.

log diagnosis

Correlating service and task logs with timestamps, identifiers and metrics to locate the component and condition that caused a failure.

cluster administration

Installation, configuration, security, capacity management, upgrades, monitoring, recovery and routine operation of cluster services.

Covering Indexes

An index covers a query when it contains every column needed for filtering and output:

CREATE INDEX idx_order_customer_date_total ON order_header(customer_id, ordered_at, total_amount);

For a compatible query, InnoDB may answer from secondary index entries without fetching full clustered rows. EXPLAIN may report an index-only access indicator such as “Using index.”

Coverage is query-specific. Adding every output column makes indexes wide and expensive to maintain, so covering design should target important measured workloads.

Selectivity

Selectivity describes how strongly a predicate narrows candidates. A unique identifier is highly selective; a Boolean status may match half the table.

The optimizer estimates selectivity from statistics and data distribution. “This column has an index” does not imply that using it is cheapest.

The Leftmost-Prefix Principle

For index (a, b, c), useful ordered prefixes are:

(a) (a, b) (a, b, c)

A query using a and c but not b may use a effectively but cannot usually use c as the next continuous ordered component.

Equality on earlier columns followed by a range on the next column is a strong pattern. After a range component, later columns may still help filtering or covering but often do not extend the same search range.

Column order should reflect important predicates, joins, ordering, selectivity and reuse across real queries.

Composite Indexes

A composite index contains several columns in a defined order:

CREATE INDEX idx_employee_dept_salary ON employee(department_id, salary);

It naturally orders first by department_id and then by salary within each department. It can support:

WHERE department_id = 10

and:

WHERE department_id = 10 AND salary >= 50000

It usually cannot provide the same direct ordered search for salary alone because salary is not the leading column.

Clustered InnoDB Storage

InnoDB normally organizes table records by primary key. The clustered index's leaf records contain the row data.

Secondary index entries contain their secondary key plus the primary-key value used to locate the clustered record. Therefore, a wide primary key enlarges every secondary index.

Choose a primary key that is stable, unique, non-NULL and reasonably compact. Randomly distributed keys can increase page splitting and reduce locality, while monotonically increasing identifiers can concentrate inserts. Workload and distributed-generation needs determine the tradeoff.

Statistics

Statistics estimate table sizes, value distributions and selectivity. MySQL can refresh statistics through analysis operations and newer versions support histograms for selected columns.

Maintenance should respond to evidence of stale or inaccurate estimates. Refreshing statistics can change plans, so important queries should be observed afterward.

Statistics describe individual distributions imperfectly and may not capture correlation between columns. Composite indexes and schema constraints can provide more useful structure.

B-Tree-Style Indexes

MySQL commonly uses B-tree-style indexes. Their balanced ordered structure supports:

  • equality lookup;
  • ordered range lookup;
  • prefix traversal of composite keys;
  • minimum and maximum access;
  • ordered scans that may satisfy ORDER BY;
  • grouping in compatible key order.

A range such as salary BETWEEN 50000 AND 70000 maps naturally to an interval of ordered keys. Hash-like structures suit equality but do not naturally support ordered ranges.

EXPLAIN ANALYZE

Where supported, EXPLAIN ANALYZE executes the query and reports actual timing and row counts alongside estimates.

Large differences between estimated and actual rows reveal poor statistics, skew or model limitations. Because the query executes, use care with cost and with statements that can modify data.

Run tests with representative parameter values and data volume. A plan fast for one selective customer may be poor for another customer owning millions of rows.

Functional and Generated-Column Indexing

A predicate that applies a function may not use an ordinary index on the original column:

WHERE LOWER(email) = 'a@example.com'

Supported MySQL versions can index expressions or generated columns. Another solution is a collation whose comparison rules already match the requirement.

For dates, a range is often clearer:

WHERE ordered_at >= '2026-01-01' AND ordered_at < '2027-01-01'

rather than YEAR(ordered_at) = 2026.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.