Files
seaweedfs/weed/worker
Chris LuandDevin 0305e837fd iceberg maintenance: keep compacted files prunable (stats, bound order, row groups) (#11654)
* iceberg maintenance: record column statistics on compacted files

A compacted file's manifest entry was built with no column_sizes,
value_counts, null_value_counts, lower_bounds, upper_bounds or
split_offsets, so no reader could skip a compacted file on any
predicate. parquet-go already writes exact per-chunk min/max and null
counts into the footer; read that footer back after the merge and
record it on the data file, bounds as the spec's single-value
serialization with string and binary truncated as truncate(16). A
column whose bounds cannot be converted exactly gets none, and a
statistics failure is logged while the compaction commits anyway:
metrics are an optimization, not a correctness requirement.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: merge bins in bound order so compacted files stay prunable

A bin's files were concatenated in manifest order, or largest-first
when a partition was split under the target size, so inputs disjoint
on a column came out as outputs that overlapped on it. Order each
bin's files by their bounds on one column before merging and split an
oversized partition into runs of consecutive files, so every output
covers one contiguous range. The column is the first identity field
of the table's sort order when it declares one (a descending order
sorts by upper bound), otherwise the first schema column every
candidate file has bounds for; detection resolves the same order so
it plans the bins execution builds. Files without bounds keep the old
behavior.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: cap compacted files' row groups from table config

Neither merge writer set a row-group limit (parquet-go's default is
unlimited rows), so every compacted file was a single row group and
readers could not skip inside it either. Rows per row group now come
from the table's write.parquet.row-group-limit and
write.parquet.row-group-size-bytes, defaulting to PyIceberg's
1 048 576 rows and Iceberg's 128 MiB, with the byte size turned into
rows from the bin's inputs' compressed bytes per row and a floor of
1 024 rows. Both writers take the cap, and the statistics the entry
records list one split offset per row group.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg: keep compaction order eligibility per group, rescue stranded runs

The merge order resolved over all candidates, so one oversized or
non-Parquet file without bounds disabled ordering for files that could
participate. Resolve it per partition group over the eligible entries.

Ordered runs too short to merge were dropped entirely. Runs from an
inferred bounds order now fall back to size-based packing — ordering is
a preference there — while runs under a declared sort order are still
left for later passes so the sort contract holds.

An explicit write.parquet.row-group-limit is a cap, not a floor: values
below the estimate floor are now honored instead of being raised.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg: only use the declared sort order when its bounds are complete

An entry without bounds on the sort column sorted to the tail and merged
into an output claiming an order it cannot verify. Fall back to
bound inference instead of ordering by a later sort field alone.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* iceberg maintenance: keep ordered runs when the full repack yields nothing

The bestEffort leftover fallback removed the ordered runs before checking
whether repacking the whole bin produced any bins, discarding valid
compaction work.

Generated with [Devin](https://devin.ai)

Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-10-08 22:06:44 +08:00
..