Big Data Technologies

Hive Architecture, Tables, Partitions, Buckets and Storage Formats

PGCP-BDA

Hive

Apache Hive provides SQL-like querying and metadata over data stored in distributed files, translating logical queries into distributed execution plans.

Hive architecture

Hive parses HiveQL, performs semantic analysis with metastore metadata, optimizes a logical plan and submits physical work to a distributed engine.

metastore

The Hive metastore stores table, column, partition, location and statistics metadata in a relational service used during query planning.

managed and external table

Hive normally controls metadata and data deletion for a managed table; an external table leaves the underlying data under external ownership.

schema on read

Schema on read interprets stored bytes using a schema when queried, allowing inexpensive ingestion but moving validation and conversion responsibility to.

partition

A data partition divides a dataset by selected values or rules so operations can process only relevant subsets and work can run in parallel.

bucket

A Hive bucket assigns rows to a fixed number of files using a hash of bucket columns, which can assist sampling and compatible joins.

SerDe

A Hive serializer/deserializer converts between stored record bytes and Hive's column representation according to the table format.

ORC and Parquet

ORC and Parquet are columnar formats that store values by column with metadata, compression and statistics for projection and predicate pruning.

Hive storage format

The physical encoding used for Hive table data, such as Text, ORC or Parquet, which affects compression, pruning and read performance.

Apache Hive

What is Hive?

Apache Hive is a data warehouse built on top of Hadoop that provides:

  • HiveQL — SQL-like query language for querying HDFS data
  • Converts HiveQL to MapReduce/Tez/Spark jobs
HiveQL Query → Hive → MapReduce job → HDFS result
-- Create external table pointing to HDFS data
CREATE EXTERNAL TABLE sales (
    date STRING,
    product STRING,
    amount DECIMAL(10,2),
    quantity INT
)
ROW FORMAT DELIMITED
FIELDS TERMINATED BY ','
LOCATION '/user/hive/warehouse/sales/';

-- Query the data
SELECT product, SUM(amount) as total_sales
FROM sales
WHERE date BETWEEN '2024-01-01' AND '2024-12-31'
GROUP BY product
ORDER BY total_sales DESC;

Hive vs RDBMS

FeatureHiveRDBMS
SchemaSchema on readSchema on write
Data stored inHDFSLocal storage
LatencyHigh (minutes)Low (ms)
ACIDLimitedFull
ScalabilityMassiveLimited
UpdatesLimitedFull
Best forAnalytics (OLAP)Transactions (OLTP)

Hive Architecture and Tables

Hive accepts SQL-like statements and compiles them into execution plans for engines such as Tez or Spark. The driver manages a query, the compiler performs semantic analysis and optimization and the metastore holds database, table, column, partition and storage metadata. Data normally remains in distributed or object storage.

A managed table gives Hive lifecycle control over its data location. An external table manages metadata while files have an independent lifecycle. Dropping behavior depends on table type and platform configuration. Schema-on-read means stored bytes are interpreted using table metadata when queried. A SerDe converts between stored records and Hive’s row representation.

Partitions, Buckets and Formats

Partition columns map values to directory divisions. A predicate on partition columns allows partition pruning so unrelated directories are not scanned. Too many tiny partitions burden the metastore and filesystem. Bucketing hashes a column into a fixed number of files. It can support sampling and selected joins but does not replace partition pruning.

Text formats are easy to inspect but weak in typing and scan efficiency. Avro stores row-oriented records with a schema. Parquet and ORC are columnar and support compression, predicate pushdown and reading only required columns. File size matters: many tiny files create task and metadata overhead while extremely large files may limit parallelism. Storage format, compression and layout should follow the main query patterns.

Continue learning

Related notes

Put this topic into timed practice

Open mock tests when you want full-exam pacing, or keep drilling in practice mode.