mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-09 07:47:54 +02:00
* iceberg maintenance: record column statistics on compacted files A compacted file's manifest entry was built with no column_sizes, value_counts, null_value_counts, lower_bounds, upper_bounds or split_offsets, so no reader could skip a compacted file on any predicate. parquet-go already writes exact per-chunk min/max and null counts into the footer; read that footer back after the merge and record it on the data file, bounds as the spec's single-value serialization with string and binary truncated as truncate(16). A column whose bounds cannot be converted exactly gets none, and a statistics failure is logged while the compaction commits anyway: metrics are an optimization, not a correctness requirement. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: merge bins in bound order so compacted files stay prunable A bin's files were concatenated in manifest order, or largest-first when a partition was split under the target size, so inputs disjoint on a column came out as outputs that overlapped on it. Order each bin's files by their bounds on one column before merging and split an oversized partition into runs of consecutive files, so every output covers one contiguous range. The column is the first identity field of the table's sort order when it declares one (a descending order sorts by upper bound), otherwise the first schema column every candidate file has bounds for; detection resolves the same order so it plans the bins execution builds. Files without bounds keep the old behavior. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: cap compacted files' row groups from table config Neither merge writer set a row-group limit (parquet-go's default is unlimited rows), so every compacted file was a single row group and readers could not skip inside it either. Rows per row group now come from the table's write.parquet.row-group-limit and write.parquet.row-group-size-bytes, defaulting to PyIceberg's 1 048 576 rows and Iceberg's 128 MiB, with the byte size turned into rows from the bin's inputs' compressed bytes per row and a floor of 1 024 rows. Both writers take the cap, and the statistics the entry records list one split offset per row group. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg: keep compaction order eligibility per group, rescue stranded runs The merge order resolved over all candidates, so one oversized or non-Parquet file without bounds disabled ordering for files that could participate. Resolve it per partition group over the eligible entries. Ordered runs too short to merge were dropped entirely. Runs from an inferred bounds order now fall back to size-based packing — ordering is a preference there — while runs under a declared sort order are still left for later passes so the sort contract holds. An explicit write.parquet.row-group-limit is a cap, not a floor: values below the estimate floor are now honored instead of being raised. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg: only use the declared sort order when its bounds are complete An entry without bounds on the sort column sorted to the tail and merged into an output claiming an order it cannot verify. Fall back to bound inference instead of ordering by a later sort field alone. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> * iceberg maintenance: keep ordered runs when the full repack yields nothing The bestEffort leftover fallback removed the ordered runs before checking whether repacking the whole bin produced any bins, discarding valid compaction work. Generated with [Devin](https://devin.ai) Co-Authored-By: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Devin <158243242+devin-ai-integration[bot]@users.noreply.github.com>