SuDo

Parquet Unpacked

Understanding Apache Parquet: From Row Groups to Page Indexes

Parquet becomes much easier to understand once you stop thinking of it as simply “a columnar file format” and instead learn its physical structure one layer at a time.

The core hierarchy is:

PARQUET FILE
    │
    └── ROW GROUP
          │
          └── COLUMN CHUNK
                  │
                  └── PAGE

Everything else—dictionary encoding, statistics, Bloom filters, page indexes, schema annotations—fits around this structure.


1. Why Parquet Is Columnar

Consider this table:

┌────┬────────┬─────┬────────┐
│ id │ name   │ age │ salary │
├────┼────────┼─────┼────────┤
│ 1  │ Alice  │ 28  │ 100k   │
│ 2  │ Bob    │ 35  │ 120k   │
│ 3  │ Carol  │ 42  │ 150k   │
└────┴────────┴─────┴────────┘

A row-oriented format conceptually stores:

1, Alice, 28, 100k
2, Bob,   35, 120k
3, Carol, 42, 150k

Parquet organizes values by column:

id      → 1, 2, 3
name    → Alice, Bob, Carol
age     → 28, 35, 42
salary  → 100k, 120k, 150k

Why?

Consider:

SELECT AVG(salary)
FROM employees;

A columnar reader can conceptually do:

id       → SKIP
name     → SKIP
age      → SKIP

salary   → READ
             │
             ▼
           AVG()

This is column pruning.

Parquet is therefore especially useful when queries examine:

MANY ROWS
   +
FEW COLUMNS

2. Row Groups

Parquet does not necessarily store an entire file’s column as one giant contiguous unit.

First, the logical table is divided horizontally into row groups.

┌────┬────────┬─────┐
│ id │ name   │ age │
├────┼────────┼─────┤
│ 1  │ Alice  │ 28  │  ┐
│ 2  │ Bob    │ 35  │  ├── Row Group 1
│ 3  │ Carol  │ 42  │  ┘
├────┼────────┼─────┤  ← horizontal boundary
│ 4  │ Dan    │ 25  │  ┐
│ 5  │ Eve    │ 31  │  ├── Row Group 2
│ 6  │ Frank  │ 39  │  ┘
└────┴────────┴─────┘

“Horizontal” describes how we divide the logical table.

Inside each row group, the data is still organized by column:

                ROW GROUP 1
                     │
          ┌──────────┼──────────┐
          ▼          ▼          ▼
         id         name       age

      [1,2,3]    [A,B,C]   [28,35,42]

Row Group: A horizontal batch of rows whose contents are stored column-by-column.


3. Column Chunks

Inside a row group, each column becomes a column chunk.

                    ROW GROUP 1
                         │
          ┌──────────────┼──────────────┐
          ▼              ▼              ▼
      id chunk       name chunk      age chunk
       [1,2,3]        [A,B,C]       [28,35,42]

Column Chunk: One column’s data within one row group.

Therefore:

3 Row Groups
×
4 Columns
────────────
12 Column Chunks

The same logical column has a different chunk in every row group:

                  age
                   │
        ┌──────────┼──────────┐
        ▼          ▼          ▼
       RG1        RG2        RG3
        │          │          │
    age chunk  age chunk  age chunk

4. Pages

Column chunks can still be large, so Parquet divides them into pages.

age COLUMN CHUNK
        │
        ├── Page 1
        ├── Page 2
        ├── Page 3
        └── Page 4

Our hierarchy is now:

PARQUET FILE
    │
    └── ROW GROUP
          │
          └── COLUMN CHUNK
                  │
                  ├── PAGE
                  ├── PAGE
                  └── PAGE

Mental model:

Row Group
    ↓
batch of rows

Column Chunk
    ↓
one column from that batch

Page
    ↓
smaller unit inside the column chunk

5. Data Pages and Dictionary Pages

Two important page concepts are:

PAGE
 │
 ├── Data Page
 │
 └── Dictionary Page

A data page contains encoded column data.

Suppose:

country

US
US
UK
US
IN
UK
US

Conceptually:

country Column Chunk
        │
        ├── Data Page 1
        │     US
        │     US
        │     UK
        │     US
        │
        └── Data Page 2
              IN
              UK
              US

6. Dictionary Encoding

Repeated values can be expensive to store repeatedly.

Instead of:

US, US, UK, US, IN, UK, US

Parquet can construct a dictionary:

0 → US
1 → UK
2 → IN

The data becomes:

US  US  UK  US  IN  UK  US
             ↓
0   0   1   0   2   1   0

The column chunk now conceptually looks like:

country COLUMN CHUNK
        │
        ├── Dictionary Page
        │      0 → US
        │      1 → UK
        │      2 → IN
        │
        ├── Data Page
        │      0 0 1 0
        │
        └── Data Page
               2 1 0

Remember:

Dictionary Page
       ↓
dictionary values


Data Page
       ↓
encoded references to those values

Dictionary encoding is particularly effective for low-cardinality or highly repetitive data.


7. Plain-Encoding Fallback

Dictionary encoding is not always efficient.

Consider unique identifiers:

a81f...
b72c...
c93e...
d11a...
e57b...

A dictionary might become almost as large as the original data.

A writer can therefore stop dictionary encoding and switch later pages to another encoding, commonly PLAIN.

COLUMN CHUNK
│
├── Dictionary Page
│
├── Data Page
│      dictionary encoded
│
├── Data Page
│      dictionary encoded
│
│   dictionary becomes inefficient
│
├── Data Page
│      PLAIN
│
└── Data Page
       PLAIN

Do not assume one column chunk must use exactly one encoding for all its data pages.


At the end of a Parquet file is its file metadata/footer.

┌──────────────────────────────┐
│         PARQUET FILE         │
├──────────────────────────────┤
│ Row Group 1                  │
│   column chunks + pages      │
├──────────────────────────────┤
│ Row Group 2                  │
│   column chunks + pages      │
├──────────────────────────────┤
│ Row Group 3                  │
│   column chunks + pages      │
├──────────────────────────────┤
│                              │
│ FILE METADATA / FOOTER       │
│                              │
│ • schema                     │
│ • row-group metadata         │
│ • column metadata            │
│ • statistics                 │
│ • offsets                    │
│ • encodings                  │
│ • compression information   │
└──────────────────────────────┘

Think of the footer as the reader’s map of the file.


9. Column-Chunk Statistics

Parquet can store statistics describing a column chunk.

age Column Chunk

values:
[18, 22, 25, 31, 40]

statistics:
min         = 18
max         = 40
null_count  = 0
value_count = 5

Now imagine:

             AGE

RG1       min=18  max=40
RG2       min=41  max=65
RG3       min=66  max=90

Query:

WHERE age > 70

The reader can reason:

RG1   18 ─── 40      SKIP

RG2   41 ─── 65      SKIP

RG3   66 ───── 90    MAY MATCH
                        │
                        ▼
                       READ

This allows row-group pruning.

Important Statistics

StatisticMeaning
minLower bound / minimum
maxUpper bound / maximum
null_countNumber of nulls
value_countNumber of encoded values in the column chunk

For nested/repeated schemas, don’t automatically interpret value_count as simply the number of non-null SQL values.


10. Statistics Truncation

Numbers have small min/max representations:

min = 18
max = 90

Strings or binary values can be enormous.

Instead of putting giant values into metadata, writers may use bounded/truncated statistics.

Conceptually:

actual value:

"watermelon-some-extremely-long-value........"

                 ↓

metadata bound:

"water..."

The actual truncation rules must preserve safe comparison bounds. It is not simply “cut both strings after N characters.”

The goal is:

keep metadata small
       +
retain useful skipping information

11. Bloom Filters

Min/max statistics have a limitation.

Suppose:

user_id chunk

min = 100
max = 9,000,000

Query:

WHERE user_id = 472938

Because:

100 ───────── 472938 ───────── 9,000,000

min/max cannot eliminate the chunk.

A Bloom filter can help.

Column Chunk values
        │
        ▼
┌────────────────┐
│  Bloom Filter  │
└───────┬────────┘
        │
        ▼
     472938?

It can answer:

Definitely NOT present
        │
        ▼
       SKIP

or:

MAYBE present
        │
        ▼
continue checking

Bloom filters can produce false positives:

Bloom says MAYBE
      │
      ├── actually exists
      │
      └── actually absent

But a correct Bloom filter does not produce false negatives:

Bloom says NO
      │
      ▼
definitely absent

Memory trick:

min/max
   ↓
range filtering


Bloom filter
   ↓
equality / point lookups

12. Page Index

Column-chunk statistics operate at a coarse level.

Suppose:

age Column Chunk

overall:
min=10
max=50

Query:

WHERE age >= 45

The entire chunk cannot be skipped.

But its pages might look like:

Page 1    10–20
Page 2    21–30
Page 3    31–40
Page 4    41–50

Only Page 4 could contain values ≥45.

This is where the Page Index helps:

             PAGE INDEX
                 │
        ┌────────┴────────┐
        ▼                 ▼
  Column Index       Offset Index

Memory trick:

Column Index
     ↓
WHAT is in the pages?


Offset Index
     ↓
WHERE are the pages?

13. Column Index

The Column Index provides page-level statistics/information.

age COLUMN CHUNK

┌────────┬─────┬─────┐
│        │ min │ max │
├────────┼─────┼─────┤
│ Page 1 │ 10  │ 20  │
│ Page 2 │ 21  │ 30  │
│ Page 3 │ 31  │ 40  │
│ Page 4 │ 41  │ 50  │
└────────┴─────┴─────┘

For:

WHERE age >= 45

the reader gets:

Page 1   10–20   SKIP
Page 2   21–30   SKIP
Page 3   31–40   SKIP
Page 4   41–50   MAY MATCH

So the Column Index answers:

Which pages might contain useful data?


14. Offset Index

After deciding that Page 4 matters, the reader needs to locate it.

That’s the Offset Index’s job.

Conceptually:

Page 1 → offset 10,000
Page 2 → offset 14,000
Page 3 → offset 18,000
Page 4 → offset 22,000

Page locations contain information such as:

offset
compressed_page_size
first_row_index

The relationship:

Column Index
      │
      ▼
"Page 4 might match"
      │
      ▼
Offset Index
      │
      ▼
"Page 4 is here"
      │
      ▼
READ PAGE 4

Therefore:

Chunk statistics → skip large chunk/row group

Column Index      → choose pages

Offset Index      → locate pages

15. Physical Types vs Logical Types

So far we’ve discussed where data lives.

Now:

What does the stored data mean?

Parquet has primitive/physical storage types such as:

BOOLEAN
INT32
INT64
FLOAT
DOUBLE
BYTE_ARRAY
FIXED_LEN_BYTE_ARRAY
...

Applications use richer concepts:

DATE
TIMESTAMP
STRING
DECIMAL
UUID
JSON
...

Parquet combines physical representation with logical annotations.

For example:

birth_date
│
├── Physical type: INT32
│
└── Logical type: DATE

Conceptually:

2026-09-16
     │
     ▼
    DATE
     │
     ▼
   INT32
     │
     ▼
days since 1970-01-01

The physical type answers:

How is this stored?

The logical annotation answers:

How should I interpret it?


16. Logical-Type Examples

String

name
│
├── Physical: BYTE_ARRAY
└── Logical:  STRING
"Alice"
   ↓
STRING
   ↓
BYTE_ARRAY

Decimal

price
│
├── Physical: INT64
└── Logical:
      DECIMAL
      precision=10
      scale=2

A physical integer:

12345

with scale 2 represents:

123.45

Remember:

Physical type
      ↓
storage


Logical type
      ↓
meaning

17. OriginalType vs LogicalTypeAnnotation

This distinction mostly comes from Parquet’s evolution.

Older representations used converted types, commonly represented in APIs as OriginalType.

OLDER

INT32
  │
  └── OriginalType.DATE

Modern Parquet uses richer logical-type annotations:

NEWER

INT32
  │
  └── LogicalTypeAnnotation.DATE

The newer representation can express richer semantics.

For example:

TIMESTAMP
│
├── unit
│    ├── MILLIS
│    ├── MICROS
│    └── NANOS
│
└── isAdjustedToUTC

There aren’t separate physical and logical values.

There is one physical representation plus schema metadata describing its meaning:

stored INT32
     │
     ▼
    20712
     │
     │ schema says DATE
     ▼
 interpreted as
 2026-09-16

18. Field IDs

Schema fields normally have names:

id
name
age

But names can change.

Version 1:

name

Version 2:

full_name

How can a system know this was a rename rather than a completely different field?

A Parquet schema field can have a Field ID.

name
├── Field ID = 2
└── STRING

After renaming:

full_name
├── Field ID = 2
└── STRING

Therefore:

OLD                         NEW

name                        full_name
Field ID = 2  ───────────►  Field ID = 2

The name changed.

The identity remained stable.

This is particularly useful for schema evolution.

Nested Example

user
│
├── name          [ID 1]
│
├── address       [ID 2]
│    ├── city     [ID 3]
│    └── zip      [ID 4]
│
└── age           [ID 5]

A nested field can be renamed while preserving its identity.

Field IDs are schema metadata. They are not repeated alongside every value.


19. Putting Everything Together

Consider:

SELECT name
FROM users
WHERE age = 35;

The reader can progressively eliminate unnecessary work:

                    QUERY
                 age = 35
                     │
                     ▼
                FILE FOOTER
                     │
                     ▼
             inspect schema
                     │
                     ▼
             column pruning
              age + name only
                     │
                     ▼
          Column Chunk Statistics
                     │
              Which Row Groups?
                     │
           ┌─────────┼─────────┐
           ▼         ▼         ▼
          RG1       RG2       RG3
          SKIP      KEEP      SKIP
                     │
                     ▼
                Bloom Filter
                     │
                  MAYBE
                     │
                     ▼
                Column Index
                     │
              Which pages?
                     │
                     ▼
                  Page 3
                     │
                     ▼
                Offset Index
                     │
              Where is Page 3?
                     │
                     ▼
                 Read Page
                     │
                     ▼
          decompress + decode
                     │
                     ▼
                age = 35
                     │
                     ▼
            matching name values

20. The Optimization Ladder

1. COLUMN PRUNING
        ↓
Don't read unnecessary columns

2. COLUMN-CHUNK STATISTICS
        ↓
Don't read unnecessary row groups

3. BLOOM FILTER
        ↓
Prove searched values are absent

4. COLUMN INDEX
        ↓
Don't read unnecessary pages

5. OFFSET INDEX
        ↓
Locate useful pages

6. ENCODING + COMPRESSION
        ↓
Store/read fewer bytes

Not every Parquet file contains Bloom filters or page indexes, and not every reader uses every optional optimization.

The fundamental structure remains:

FILE
 │
 ├── ROW GROUP
 │      │
 │      └── COLUMN CHUNK
 │              │
 │              └── PAGE
 │
 └── FOOTER / METADATA

Final Mental Model

If you remember only one diagram, use this:

                       PARQUET FILE
                            │
              ┌─────────────┴─────────────┐
              │                           │
             DATA                       FOOTER
              │                           │
          ROW GROUP                       ├── Schema
              │                           │     │
         COLUMN CHUNK                     │     ├── name
              │                           │     ├── Field ID
       ┌──────┴──────┐                    │     ├── physical type
       │             │                    │     └── logical type
       ▼             ▼                    │
 Dictionary       Data Pages              └── Row Group /
    Page                                      Column metadata
                                                   │
                                                   ├── min
                                                   ├── max
                                                   ├── null_count
                                                   └── value_count

                    COLUMN CHUNK
                          │
               ┌──────────┼──────────┐
               ▼          ▼          ▼
            Statistics   Bloom    Page Index
                                    │
                             ┌──────┴──────┐
                             ▼             ▼
                       Column Index   Offset Index
                             │             │
                             ▼             ▼
                       WHAT pages?    WHERE pages?

Cheat Sheet

ConceptRemember it as
Row GroupHorizontal batch of rows
Column ChunkOne column within one row group
PageSmaller storage unit inside a column chunk
Data PageEncoded column data
Dictionary PageDictionary values
Dictionary EncodingReplace repeated values with compact dictionary IDs
PLAINDirect representation; useful when dictionary encoding isn’t appropriate
FooterMap/schema/metadata describing the file
min / maxRange-based skipping
null_countNumber of nulls represented by statistics
Bloom Filter“Definitely absent” or “maybe present”
Column IndexWhat is inside individual pages?
Offset IndexWhere are those pages?
Physical TypeHow bytes are represented
Logical TypeWhat those bytes mean
OriginalTypeOlder converted/logical type representation
LogicalTypeAnnotationNewer, richer logical-type representation
Field IDStable schema-field identity

The One Idea Behind Parquet

Most of Parquet’s design ultimately helps a query engine answer:

"Can I avoid reading this?"
        FILE
         │
         ▼
      COLUMN? ── no ──► SKIP
         │
        yes
         ▼
    ROW GROUP? ─ no ──► SKIP
         │
        yes
         ▼
       PAGE? ─── no ──► SKIP
         │
        yes
         ▼
       READ
         │
         ▼
      DECODE
         │
         ▼
      RESULT

That is the core mental model:

Parquet organizes data and metadata so analytical engines can read as little data as possible.

NOTE: Written by AI, Prompted by Human