Big Data Technologies
Hive Architecture, Tables, Partitions, Buckets and Storage Formats
PGCP-BDA
Hive
Apache Hive provides SQL-like querying and metadata over data stored in distributed files, translating logical queries into distributed execution plans.
Hive architecture
Hive parses HiveQL, performs semantic analysis with metastore metadata, optimizes a logical plan and submits physical work to a distributed engine.
metastore
The Hive metastore stores table, column, partition, location and statistics metadata in a relational service used during query planning.
managed and external table
Hive normally controls metadata and data deletion for a managed table; an external table leaves the underlying data under external ownership.
schema on read
Schema on read interprets stored bytes using a schema when queried, allowing inexpensive ingestion but moving validation and conversion responsibility to.
partition
A data partition divides a dataset by selected values or rules so operations can process only relevant subsets and work can run in parallel.
bucket
A Hive bucket assigns rows to a fixed number of files using a hash of bucket columns, which can assist sampling and compatible joins.
SerDe
A Hive serializer/deserializer converts between stored record bytes and Hive's column representation according to the table format.
ORC and Parquet
ORC and Parquet are columnar formats that store values by column with metadata, compression and statistics for projection and predicate pruning.
Hive storage format
The physical encoding used for Hive table data, such as Text, ORC or Parquet, which affects compression, pruning and read performance.
Apache Hive
What is Hive?
Apache Hive is a data warehouse built on top of Hadoop that provides:
- HiveQL — SQL-like query language for querying HDFS data
- Converts HiveQL to MapReduce/Tez/Spark jobs
HiveQL Query → Hive → MapReduce job → HDFS result
-- Create external table pointing to HDFS data
CREATE EXTERNAL TABLE sales (
date STRING,
product STRING,
amount DECIMAL(10,2),
quantity INT
)
ROW FORMAT DELIMITED
FIELDS TERMINATED BY ','
LOCATION '/user/hive/warehouse/sales/';
-- Query the data
SELECT product, SUM(amount) as total_sales
FROM sales
WHERE date BETWEEN '2024-01-01' AND '2024-12-31'
GROUP BY product
ORDER BY total_sales DESC;
Hive vs RDBMS
| Feature | Hive | RDBMS |
|---|---|---|
| Schema | Schema on read | Schema on write |
| Data stored in | HDFS | Local storage |
| Latency | High (minutes) | Low (ms) |
| ACID | Limited | Full |
| Scalability | Massive | Limited |
| Updates | Limited | Full |
| Best for | Analytics (OLAP) | Transactions (OLTP) |
Hive Architecture and Tables
Hive accepts SQL-like statements and compiles them into execution plans for engines such as Tez or Spark. The driver manages a query, the compiler performs semantic analysis and optimization and the metastore holds database, table, column, partition and storage metadata. Data normally remains in distributed or object storage.
A managed table gives Hive lifecycle control over its data location. An external table manages metadata while files have an independent lifecycle. Dropping behavior depends on table type and platform configuration. Schema-on-read means stored bytes are interpreted using table metadata when queried. A SerDe converts between stored records and Hive’s row representation.
Partitions, Buckets and Formats
Partition columns map values to directory divisions. A predicate on partition columns allows partition pruning so unrelated directories are not scanned. Too many tiny partitions burden the metastore and filesystem. Bucketing hashes a column into a fixed number of files. It can support sampling and selected joins but does not replace partition pruning.
Text formats are easy to inspect but weak in typing and scan efficiency. Avro stores row-oriented records with a schema. Parquet and ORC are columnar and support compression, predicate pushdown and reading only required columns. File size matters: many tiny files create task and metadata overhead while extremely large files may limit parallelism. Storage format, compression and layout should follow the main query patterns.
Continue learning
Related notes
Put this topic into timed practice
Open mock tests when you want full-exam pacing, or keep drilling in practice mode.