From 98bb2c947b8faaba4c0bd38ee02d0f383bad26b4 Mon Sep 17 00:00:00 2001 From: Chris Lu Date: Sun, 21 Jun 2026 14:02:21 -0700 Subject: [PATCH] S3 lifecycle docs: recipes page, accurate metrics, drop stale knobs - add S3-Lifecycle-Recipes with copy-paste configs and verification - link the Architecture/Monitoring/Operator/Troubleshooting/Recipes sub-pages from the sidebar; they were orphaned from nav - fix Prometheus metric names to the real SeaweedFS_ namespace; the bare s3_lifecycle_ names matched nothing in a scrape - main page: raw XML wire format, 1 MiB / error-code validation table, Transition rejected (not ignored), GET legacy directory-TTL fallback - document events_total and schedule_depth; run-shard flag table - drop max_runtime_minutes (removed); per-pass cap is now the scheduler Execution Timeout, default effectively unbounded - Days < 1 is client-side only; server stores Days=0 as a no-op --- S3-Lifecycle-Monitoring.md | 41 +++--- S3-Lifecycle-Operator-Guide.md | 34 ++++- S3-Lifecycle-Recipes.md | 251 ++++++++++++++++++++++++++++++++ S3-Lifecycle-Troubleshooting.md | 18 +-- S3-Lifecycle.md | 68 ++++++++- _Sidebar.md | 5 + 6 files changed, 383 insertions(+), 34 deletions(-) create mode 100644 S3-Lifecycle-Recipes.md diff --git a/S3-Lifecycle-Monitoring.md b/S3-Lifecycle-Monitoring.md index d507617..f759f13 100644 --- a/S3-Lifecycle-Monitoring.md +++ b/S3-Lifecycle-Monitoring.md @@ -4,38 +4,43 @@ This page lists the Prometheus signals the worker exposes and how to read the he ## Prometheus metrics -All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem. +All metrics are registered in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem. The name a scrape sees is the namespace + subsystem + metric — e.g. the `cursor_min_ts_ns` gauge is exposed as `SeaweedFS_s3_lifecycle_cursor_min_ts_ns`. The fully-qualified names are used throughout this page; query them exactly as written. ### Per-shard gauges | Metric | Labels | What | |---|---|---| -| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event on this shard for which all matches dispatched successfully | -| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire | +| `SeaweedFS_s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event on this shard for which all matches dispatched successfully | +| `SeaweedFS_s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire | Derived queries: ```promql # Per-shard replay lag in seconds -(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 +(time() * 1e9 - SeaweedFS_s3_lifecycle_cursor_min_ts_ns) / 1e9 # Per-shard walker freshness in seconds -(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 +(time() * 1e9 - SeaweedFS_s3_lifecycle_daily_run_last_walked_ns) / 1e9 # Worst-shard lag across the cluster -max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 +max(time() * 1e9 - SeaweedFS_s3_lifecycle_cursor_min_ts_ns) / 1e9 ``` Zero values mean "not started yet" — distinct from "0s caught up". The heartbeat line uses `cold` as the marker for that state. +`SeaweedFS_s3_lifecycle_schedule_depth{shard}` is still registered but is a leftover from the prior streaming dispatcher, which kept an in-memory match heap. The daily-run worker has no such heap, so this gauge stays at zero — don't alert on it. + ### Counters | Metric | Labels | What | |---|---|---| -| `s3_lifecycle_dispatch_total` | `bucket`, `kind`, `outcome` | Per-bucket dispatch counter, partitioned by action kind and server outcome | -| `s3_lifecycle_daily_run_events_scanned_total` | `shard` | Meta-log events `drainShardEvents` processed | -| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter | -| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path | +| `SeaweedFS_s3_lifecycle_dispatch_total` | `bucket`, `kind`, `outcome` | Per-bucket dispatch counter, partitioned by action kind and server outcome | +| `SeaweedFS_s3_lifecycle_events_total` | `shard` | Meta-log events the reader emitted to the router | +| `SeaweedFS_s3_lifecycle_daily_run_events_scanned_total` | `shard` | Meta-log events `drainShardEvents` processed | +| `SeaweedFS_s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter | +| `SeaweedFS_s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path. `rule_hash` is the hex-encoded rule identity; on clusters with many rules, drop it via Prometheus relabeling to control cardinality | + +`events_total` counts what the reader fanned out; `daily_run_events_scanned_total` counts what each shard actually drained. In steady state they track together — a persistent gap means events are being filtered before drain (out-of-window TsNs) rather than processed. `outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED` is the proto zero-value — a healthy worker / server pair should never emit it; a non-zero count there indicates an internal error or a version mismatch between worker and server. @@ -43,8 +48,8 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe | Metric | What | |---|---| -| `s3_lifecycle_daily_run_shard_duration_seconds{shard}` | Wall-clock per shard pass. p95 climbing toward `max_runtime_minutes` means the shard is brushing its budget. | -| `s3_lifecycle_dispatch_limiter_wait_seconds` | Time spent waiting on the cluster rate limiter before issuing `LifecycleDelete`. Near-zero = cap not binding; long-tail at `1/rate` = cap is the active throttle. | +| `SeaweedFS_s3_lifecycle_daily_run_shard_duration_seconds{shard}` | Wall-clock per shard pass. Rising p95 means shards are taking longer; if you've set a finite Execution Timeout for the job type, p95 approaching it means passes risk being cut short (the default is effectively unbounded). | +| `SeaweedFS_s3_lifecycle_dispatch_limiter_wait_seconds` | Time spent waiting on the cluster rate limiter before issuing `LifecycleDelete`. Near-zero = cap not binding; long-tail at `1/rate` = cap is the active throttle. | ## Heartbeat log line @@ -78,10 +83,10 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag | Pattern | Meaning | What to do | |---|---|---| | `cursor_lag_max` grows unbounded | Stuck cursor; head-of-line blocking on some shard | See [Troubleshooting → Stuck cursor](S3-Lifecycle-Troubleshooting#stuck-cursor) | -| `walked_max_age` exceeds `walker_interval_minutes × 2` | Walker isn't firing as configured | Check `errors=N` in heartbeat and `s3_lifecycle_dispatch_total{outcome="RPC_ERROR"}` | +| `walked_max_age` exceeds `walker_interval_minutes × 2` | Walker isn't firing as configured | Check `errors=N` in heartbeat and `SeaweedFS_s3_lifecycle_dispatch_total{outcome="RPC_ERROR"}` | | `errors=16` (all shards) on every pass | Filer is unreachable or returning errors | Check filer health | -| `s3_lifecycle_dispatch_total{outcome="RETRY_LATER"}` rising fast | Server rate-limited or filer overloaded | Lower `cluster_deletes_per_second` or add capacity | -| `s3_lifecycle_dispatch_total{outcome="BLOCKED"}` non-zero | Programmatic event content error | Check worker logs for `FATAL_EVENT_ERROR` | +| `SeaweedFS_s3_lifecycle_dispatch_total{outcome="RETRY_LATER"}` rising fast | Server rate-limited or filer overloaded | Lower `cluster_deletes_per_second` or add capacity | +| `SeaweedFS_s3_lifecycle_dispatch_total{outcome="BLOCKED"}` non-zero | Programmatic event content error | Check worker logs for `FATAL_EVENT_ERROR` | | `duration=Ns` ramping up across passes | Walker is firing too often | Set `walker_interval_minutes` | ## Suggested alerts @@ -90,14 +95,14 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag # `> 0` filters out shards whose gauge is still the proto zero (never # started); without it, every fresh-install heartbeat triggers the alert. - alert: S3LifecycleCursorLagHigh - expr: max(time() * 1e9 - (s3_lifecycle_cursor_min_ts_ns > 0)) / 1e9 > 3600 + expr: max(time() * 1e9 - (SeaweedFS_s3_lifecycle_cursor_min_ts_ns > 0)) / 1e9 > 3600 for: 30m annotations: summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}" runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor - alert: S3LifecycleWalkerStuck - expr: max(time() * 1e9 - (s3_lifecycle_daily_run_last_walked_ns > 0)) / 1e9 > 86400 + expr: max(time() * 1e9 - (SeaweedFS_s3_lifecycle_daily_run_last_walked_ns > 0)) / 1e9 > 86400 for: 1h annotations: summary: "S3 lifecycle walker hasn't run in > 24h" @@ -105,7 +110,7 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag - alert: S3LifecycleDispatchFailures expr: | - rate(s3_lifecycle_dispatch_total{outcome=~"RETRY_LATER|BLOCKED|RPC_ERROR"}[5m]) > 0.1 + rate(SeaweedFS_s3_lifecycle_dispatch_total{outcome=~"RETRY_LATER|BLOCKED|RPC_ERROR"}[5m]) > 0.1 for: 15m annotations: summary: "S3 lifecycle delete failure rate > 0.1/s" diff --git a/S3-Lifecycle-Operator-Guide.md b/S3-Lifecycle-Operator-Guide.md index 2465047..1063f97 100644 --- a/S3-Lifecycle-Operator-Guide.md +++ b/S3-Lifecycle-Operator-Guide.md @@ -17,15 +17,24 @@ All keys are set through the admin UI's plugin config for `s3_lifecycle`. | `meta_log_retention_days` | int64 | `0` (treated as unbounded) | Stock SeaweedFS doesn't GC the meta-log, so the default is fine. Set positive if your deployment manually trims `/topics/.system/log` — then rules with TTL > retention will route through the walker. | | `walker_interval_minutes` | int64 | `0` (fire every pass) | **Important.** See "Walker interval" below — most production deployments should set this to a positive value. | -### Worker config +### Per-pass runtime cap -| Key | Default | What | -|---|---|---| -| `max_runtime_minutes` | 60 | Wall-clock cap per `dailyrun.Run` invocation. The pass returns early if it hits this. | +There is no separate worker knob for this. The earlier "Per-Run Time Limit (minutes)" form was removed because it duplicated the admin scheduler's **Execution Timeout** — both capped the same `Execute` call and had to be kept in agreement. The single source of truth is now the scheduler's runtime settings for the `s3_lifecycle` job type. + +For `s3_lifecycle` the shipped defaults set both **Execution Timeout** and **Job Type Max Runtime** to `math.MaxInt32` seconds — effectively no cap. Lifecycle is a batch job whose natural duration is "as long as today's events take"; the scheduler's generic 90s default would kill every real run, and a finite estimate (1h? 8h?) tends to either truncate a large-bucket pass or be meaningless. Set a finite Execution Timeout in the admin UI only if you specifically want passes bounded. ### Detection schedule -The admin scheduler's `DetectionIntervalMinutes` for `s3_lifecycle` is `1440` by default — once per day. Each detection produces one execution. Change in the admin UI's runtime defaults for the job type. +The admin scheduler's runtime defaults for `s3_lifecycle`: + +| Setting | Default | What | +|---|---|---| +| `DetectionIntervalMinutes` | `1440` (daily) | How often a detection runs; each detection produces one execution (`MaxJobsPerDetection` = 1). | +| `DetectionTimeoutSeconds` | `60` | Cap on the detection step itself. | +| Execution Timeout | effectively unbounded | Per-pass wall-clock cap; see above. | +| Enabled | `true` | Worker is on by default — a bucket with rules but no worker silently retains data past its declared expiration. | + +Change any of these in the admin UI's runtime defaults for the job type. ## Walker interval @@ -94,7 +103,7 @@ After applying a rule: 1. Wait at least one detection interval (default 24h). - **Replay-eligible rules** (`Expiration.Days`, `NoncurrentVersionExpiration.NoncurrentDays`, `AbortIncompleteMultipartUpload`) dispatch as soon as the next pass runs — the walker throttle doesn't apply. - **Walker-only rules** (`Expiration.Date`, `Expiration.ExpiredObjectDeleteMarker`, `NoncurrentVersionExpiration.NewerNoncurrentVersions`) additionally wait up to one `walker_interval_minutes` window before the next walker fire. -2. Read `s3_lifecycle_dispatch_total{bucket="my-bucket"}` — the counter should advance. +2. Read `SeaweedFS_s3_lifecycle_dispatch_total{bucket="my-bucket"}` — the counter should advance. 3. Verify a target object is gone: `aws s3 head-object --bucket my-bucket --key ` should return 404. For testing without waiting, the `weed shell` command supports manual invocation: @@ -104,6 +113,19 @@ weed shell -master > s3.lifecycle.run-shard -shards 0-15 -s3 -refresh 1s -runtime 30s ``` +This runs the same `dailyrun.Run` code path as the scheduled worker, driven from your shell. Flags: + +| Flag | Default | What | +|---|---|---| +| `-shard` | `-1` | Single shard id in `[0, 16)`. Mutually exclusive with `-shards`. | +| `-shards` | "" | Shard range `"lo-hi"` (inclusive) or comma list `"0,3,7"`. | +| `-s3` | "" | S3 server gRPC endpoint `host:port` (required). | +| `-events` | `1000` | Max in-shard events drained per pass. `0` = drain to now. | +| `-runtime` | `0` | Wall-clock cap on the whole run. `0` = no timeout. | +| `-refresh` | `0` | Inter-pass interval. `0` = single pass, then exit. | + +The legacy `-dispatch`, `-checkpoint`, and `-bootstrap-interval` flags are accepted but ignored (they belonged to the removed streaming worker). + This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for examples. ## Rate limit allocation diff --git a/S3-Lifecycle-Recipes.md b/S3-Lifecycle-Recipes.md new file mode 100644 index 0000000..94e15ca --- /dev/null +++ b/S3-Lifecycle-Recipes.md @@ -0,0 +1,251 @@ +# S3 Lifecycle — Recipes + +Worked, end-to-end configurations for the common lifecycle scenarios, each with how to apply it and how to confirm it ran. For the feature reference and validation rules see [S3 Lifecycle](S3-Lifecycle); for config knobs see the [Operator Guide](S3-Lifecycle-Operator-Guide). + +All examples use the AWS CLI against the S3 endpoint. Set it once: + +```bash +export S3_ENDPOINT=http://localhost:8333 +``` + +Each rule needs a `Status` and a `Filter` (use `{}` to match the whole bucket). The CLI accepts the JSON shown here; the bytes on the wire are the equivalent `LifecycleConfiguration` XML. + +## Verifying any rule + +Two independent signals tell you a rule worked: + +1. **The dispatch counter advances.** Each delete increments `SeaweedFS_s3_lifecycle_dispatch_total{bucket,kind,outcome}`. A successful expiry shows up as `outcome="DONE"`. The `kind` label is one of: + + | `kind` | Set by | + |---|---| + | `expiration_days` | `Expiration.Days` | + | `expiration_date` | `Expiration.Date` | + | `noncurrent_days` | `NoncurrentVersionExpiration.NoncurrentDays` | + | `newer_noncurrent` | `NoncurrentVersionExpiration.NewerNoncurrentVersions` (stand-alone) | + | `abort_mpu` | `AbortIncompleteMultipartUpload` | + | `expired_delete_marker` | `Expiration.ExpiredObjectDeleteMarker` | + + ```promql + sum by (kind, outcome) (SeaweedFS_s3_lifecycle_dispatch_total{bucket="my-bucket"}) + ``` + +2. **The object is gone.** `aws s3api head-object` returns `404`, or `list-object-versions` no longer shows the version. + +**Timing.** Replay-eligible rules (`Expiration.Days`, `NoncurrentDays`, `AbortIncompleteMultipartUpload`) fire on the next worker pass. Version-list-aware rules (`Expiration.Date`, `NewerNoncurrentVersions`, `ExpiredObjectDeleteMarker`) fire on the next walker pass — additionally up to one `walker_interval_minutes` window. With the default daily schedule, budget up to 24h. To check a rule immediately without waiting, drive the worker by hand with `s3.lifecycle.run-shard` (see the [Operator Guide](S3-Lifecycle-Operator-Guide#verifying-a-rule-is-working)). + +## Expire objects under a prefix + +Delete everything under `logs/` 30 days after it was written. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "expire-logs-30d", + "Status": "Enabled", + "Filter": { "Prefix": "logs/" }, + "Expiration": { "Days": 30 } + }] + }' +``` + +Replay path (`kind="expiration_days"`). The 30-day clock starts at the object's latest-version PUT. On a versioned bucket this writes a delete marker rather than removing data — pair it with the cleanup recipes below. + +Verify: + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api head-object --bucket my-bucket --key logs/old.log +# expected: An error occurred (404) +``` + +## Abort incomplete multipart uploads + +Reclaim parts from uploads that were started but never completed, 7 days after initiation. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "abort-stuck-mpu", + "Status": "Enabled", + "Filter": {}, + "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } + }] + }' +``` + +Replay path (`kind="abort_mpu"`). The clock starts at the `CreateMultipartUpload` time. Safe to run on every bucket — it only touches in-flight uploads, never completed objects. + +Verify: + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api list-multipart-uploads --bucket my-bucket +# expected: no Uploads older than 7 days +``` + +## Expire objects by tag + +Delete objects tagged `temp=true`, one day after write, regardless of prefix. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "expire-temp-tagged", + "Status": "Enabled", + "Filter": { "Tag": { "Key": "temp", "Value": "true" } }, + "Expiration": { "Days": 1 } + }] + }' +``` + +Replay path. The tag is read from the live object at evaluation time, so re-tagging an object changes whether it matches. (This mutability is also why the [TTL fast path](S3-Lifecycle-vs-Volume-TTL#the-ttl-fast-path-opt-in) refuses to stamp tag-filtered rules.) + +## Expire only large objects + +Combine predicates with `And`. Here: delete objects under `tmp/` that are larger than 5 MiB. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "expire-large-tmp", + "Status": "Enabled", + "Filter": { + "And": { + "Prefix": "tmp/", + "ObjectSizeGreaterThan": 5242880 + } + }, + "Expiration": { "Days": 7 } + }] + }' +``` + +`And` is required whenever a filter has more than one predicate. Size bounds are strict: `ObjectSizeGreaterThan` matches objects *strictly larger* than the value, `ObjectSizeLessThan` *strictly smaller*. Set both to target a size band. + +## Versioned bucket: keep the N newest versions + +On a versioned bucket, retain the 5 most recent noncurrent versions and expire older noncurrent versions 30 days after they were superseded. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "prune-noncurrent", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "NoncurrentVersionExpiration": { + "NoncurrentDays": 30, + "NewerNoncurrentVersions": 5 + } + }] + }' +``` + +A noncurrent version is removed only when **both** conditions hold: it is older than `NoncurrentDays` *and* there are at least `NewerNoncurrentVersions` newer noncurrent versions ahead of it. The five newest noncurrent versions are always kept, however old they get; everything behind them ages out at 30 days. The noncurrent clock starts at the PUT that demoted the version, not the version's own mtime. + +Because the retain-N cap needs the full version list, this combination is evaluated on the walker pass, so it also waits up to one `walker_interval_minutes` window. + +Use `NoncurrentDays` alone for a pure age policy, or `NewerNoncurrentVersions` alone for a pure count cap (keep exactly the N newest, no age requirement). + +Verify: + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api list-object-versions --bucket my-bucket --prefix my-key +# expected: at most 5 noncurrent versions remain +``` + +## Versioned bucket: clean up fully after deletes + +A `DELETE` on a versioned bucket leaves a delete marker, and previous versions remain as noncurrent. To reclaim everything for deleted objects, use two rules — one to expire the old versions, one to remove the leftover delete marker once it is the only thing left: + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [ + { + "ID": "expire-noncurrent", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "NoncurrentVersionExpiration": { "NoncurrentDays": 30 } + }, + { + "ID": "expire-orphan-delete-markers", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "Expiration": { "ExpiredObjectDeleteMarker": true } + } + ] + }' +``` + +`ExpiredObjectDeleteMarker` removes a delete marker only when it is the sole remaining version of the key (no noncurrent versions behind it). With the first rule clearing the noncurrent versions, the second eventually finds each delete marker orphaned and removes it. Both run on the walker pass. + +## One-time cleanup at a fixed date + +Delete everything under `archive/2024/` on a specific date, rather than relative to each object's age. + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [{ + "ID": "drop-2024-archive", + "Status": "Enabled", + "Filter": { "Prefix": "archive/2024/" }, + "Expiration": { "Date": "2026-07-01T00:00:00Z" } + }] + }' +``` + +Walker path (`kind="expiration_date"`). The rule fires on the first walker pass at or after the date; a date already in the past triggers on the next walk. Remove the rule afterward — it stays active and would expire anything later written under the prefix. + +## A combined production config + +Rules are independent and evaluated together, so one document can carry the whole policy for a bucket: + +```bash +aws --endpoint-url "$S3_ENDPOINT" s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [ + { + "ID": "expire-logs-30d", + "Status": "Enabled", + "Filter": { "Prefix": "logs/" }, + "Expiration": { "Days": 30 } + }, + { + "ID": "abort-stuck-mpu", + "Status": "Enabled", + "Filter": {}, + "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } + }, + { + "ID": "prune-noncurrent", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "NoncurrentVersionExpiration": { "NoncurrentDays": 30, "NewerNoncurrentVersions": 5 } + }, + { + "ID": "expire-orphan-delete-markers", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "Expiration": { "ExpiredObjectDeleteMarker": true } + } + ] + }' +``` + +`GET` returns the stored document verbatim. To temporarily disable a single rule without deleting it, set its `Status` to `Disabled` and re-`PUT`; the worker reads the change on its next pass. + +## See also + +[[S3 Lifecycle]] · [S3 Lifecycle Operator Guide](S3-Lifecycle-Operator-Guide) · [S3 Lifecycle Monitoring](S3-Lifecycle-Monitoring) · [S3 Lifecycle Troubleshooting](S3-Lifecycle-Troubleshooting) · [[S3 Lifecycle vs Volume TTL]] · [[S3 Object Versioning]] diff --git a/S3-Lifecycle-Troubleshooting.md b/S3-Lifecycle-Troubleshooting.md index bccd849..b6257f7 100644 --- a/S3-Lifecycle-Troubleshooting.md +++ b/S3-Lifecycle-Troubleshooting.md @@ -4,7 +4,7 @@ Incident-response playbook for the S3 lifecycle worker. For monitoring backgroun ## Stuck cursor -**Symptom:** `s3_lifecycle_cursor_min_ts_ns{shard=N}` is not advancing. Heartbeat shows `cursor_lag_max` growing unbounded. +**Symptom:** `SeaweedFS_s3_lifecycle_cursor_min_ts_ns{shard=N}` is not advancing. Heartbeat shows `cursor_lag_max` growing unbounded. **Cause:** The cursor advance is gated on every match from the event dispatching successfully (`DONE`, `NOOP_RESOLVED`, or `SKIPPED_OBJECT_LOCK`). Any unresolved outcome (`RETRY_LATER`, `BLOCKED`, transport error after in-run retries) halts the run for that shard and persists the cursor at the last fully-processed event. Head-of-line blocking is intentional — it surfaces a real problem rather than silently retrying forever. @@ -12,10 +12,10 @@ Incident-response playbook for the S3 lifecycle worker. For monitoring backgroun ```promql # Which shard is stuck? -time() * 1e9 - s3_lifecycle_cursor_min_ts_ns +time() * 1e9 - SeaweedFS_s3_lifecycle_cursor_min_ts_ns # What outcomes are being returned? -sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m])) +sum by (outcome) (rate(SeaweedFS_s3_lifecycle_dispatch_total[5m])) ``` Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`: @@ -39,7 +39,7 @@ The worker doesn't auto-skip past a stuck event. If you've verified the event is ## Walker stuck (no progress on walker-only rules) -**Symptom:** `s3_lifecycle_daily_run_last_walked_ns{shard=N}` is not advancing. Rules like `Expiration.Date`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` aren't firing on objects that should be due. +**Symptom:** `SeaweedFS_s3_lifecycle_daily_run_last_walked_ns{shard=N}` is not advancing. Rules like `Expiration.Date`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` aren't firing on objects that should be due. **Causes:** @@ -51,7 +51,7 @@ The worker doesn't auto-skip past a stuck event. If you've verified the event is ```promql # Walker age per shard -(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 +(time() * 1e9 - SeaweedFS_s3_lifecycle_daily_run_last_walked_ns) / 1e9 ``` Check worker config: `walker_interval_minutes` should be ≤ the daily worker schedule interval. @@ -60,7 +60,7 @@ Check worker config: `walker_interval_minutes` should be ≤ the daily worker sc ## Test PUT a file with a 1-day rule, didn't expire -The S3 API rejects `Expiration.Days < 1`, so the smallest "expire after N days" you can configure is 1 day. The worker runs once per day by default. Object PUT + 1-day rule + waiting one day is the minimum scenario. +The AWS CLI and SDKs reject `Expiration.Days < 1` client-side, so through them the smallest "expire after N days" you can configure is 1 day. (SeaweedFS itself doesn't validate the value — a raw `PUT` with `Days` 0 or omitted stores fine but compiles to a no-op, which is one way to end up with a rule that never fires.) The worker runs once per day by default. Object PUT + 1-day rule + waiting one day is the minimum real scenario. For testing, the in-repo integration suite uses a trick: backdate the entry's `Mtime` via `filer_pb.UpdateEntry` to 30+ days ago. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for the pattern. @@ -81,7 +81,7 @@ This runs the same code path as the scheduled worker, but driven from your shell 1. **Filer unreachable.** Subscription fails on every shard. Check filer health and gRPC connectivity. 2. **passCtx timeout from `-refresh` loop.** If `-refresh` is less than the pass cap, the timeout fires before the drain completes. This is now treated as "clean end-of-pass" — it should not show as errors=N. If you see this on a build before #9481, upgrade. -3. **Bucket walker is timing out.** Big bucket, walker hits ctx deadline. Increase `max_runtime_minutes`. +3. **Bucket walker is timing out.** Big bucket, walker hits a ctx deadline. This only happens if you've set a finite **Execution Timeout** for the `s3_lifecycle` job type — the shipped default is effectively unbounded, so a pass isn't cut short by a runtime cap out of the box. Raise (or clear) the Execution Timeout in the admin UI's runtime defaults. ## Some objects expired, others didn't (same rule) @@ -90,9 +90,9 @@ This runs the same code path as the scheduled worker, but driven from your shell **Common causes:** 1. **The non-deleted object's mtime is wrong.** Check the entry's mtime — it might be more recent than you expect (e.g., a recent metadata update bumped it). -2. **The objects are on different shards** and one shard has a stuck cursor while the other doesn't. Check `s3_lifecycle_cursor_min_ts_ns` per shard. +2. **The objects are on different shards** and one shard has a stuck cursor while the other doesn't. Check `SeaweedFS_s3_lifecycle_cursor_min_ts_ns` per shard. 3. **`Filter` doesn't match what you think.** A prefix-only filter requires the object key to start with that prefix; a tag filter requires the matching tag. Verify with `aws s3api head-object` (returns tags via `--query`). -4. **Object lock or retention.** The dispatcher returns `SKIPPED_OBJECT_LOCK` for protected objects. Check `s3_lifecycle_dispatch_total{outcome="SKIPPED_OBJECT_LOCK"}`. +4. **Object lock or retention.** The dispatcher returns `SKIPPED_OBJECT_LOCK` for protected objects. Check `SeaweedFS_s3_lifecycle_dispatch_total{outcome="SKIPPED_OBJECT_LOCK"}`. ## How to read the cursor files diff --git a/S3-Lifecycle.md b/S3-Lifecycle.md index 8bae445..62c998e 100644 --- a/S3-Lifecycle.md +++ b/S3-Lifecycle.md @@ -18,7 +18,7 @@ This page is the operator-facing entry point. Developers and architecture reader | `Filter.Tag` | Yes | | | `Filter.ObjectSizeGreaterThan` / `ObjectSizeLessThan` | Yes | | | `Filter.And` (composite) | Yes | | -| `Transition` / `NoncurrentVersionTransition` | No | SeaweedFS doesn't model storage class tiers | +| `Transition` / `NoncurrentVersionTransition` | Rejected | SeaweedFS doesn't model storage class tiers. A `PUT` whose enabled rules contain either is rejected with `NotImplemented` — the config is not stored. See [Limits and validation](#limits-and-validation). | ## API endpoints @@ -28,6 +28,52 @@ GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration DELETE /{bucket}?lifecycle # DeleteBucketLifecycle ``` +The CLI examples below are the convenient way to drive these. The PUT body on the wire is XML — the same `LifecycleConfiguration` document AWS uses: + +```xml + + + expire-logs + Enabled + + + logs/ + classtemp + 4096 + + + 30 + + 7 + 3 + + + 7 + + + +``` + +Unknown elements are skipped on decode. `Filter` with a single `Prefix` (or a single `Tag`) needs no `And` wrapper; `And` is only required when combining more than one predicate. The stored XML is returned verbatim by `GET` — SeaweedFS does not rewrite or canonicalize it. + +## Limits and validation + +The `PUT` handler validates the whole document before touching any state, so a rejected config never half-applies: + +| Condition | Result | +|---|---| +| Body larger than **1 MiB** | `EntityTooLarge` | +| Body not valid lifecycle XML | `MalformedXML` | +| An **enabled** rule contains `Transition` or `NoncurrentVersionTransition` | `NotImplemented` (whole config rejected) | +| Filer / backing-store error | `InternalError` | +| `GET` when no lifecycle is configured | `NoSuchLifecycleConfiguration` | + +Notes: + +- Only **enabled** rules are checked for transitions — a `Disabled` rule carrying a `Transition` is ignored rather than rejected, matching how a disabled rule is otherwise inert. +- There is no enforced cap on the number of rules, rule-ID length, or tag count beyond the 1 MiB document size. +- `GET` has a legacy fallback: if a bucket has no lifecycle XML but does have directory TTLs configured via `fs.configure -ttl` (see [S3 Lifecycle vs Volume TTL](S3-Lifecycle-vs-Volume-TTL)), `GET` synthesizes `Expiration.Days` rules from those TTLs instead of returning `NoSuchLifecycleConfiguration`. + ## Example: AWS CLI ```bash @@ -102,6 +148,12 @@ The lifecycle worker is a scheduled job (default daily). Each invocation: The worker exits when the pass is done. The admin scheduler invokes it on a daily cadence by default; operators can change that via the standard plugin scheduler config. +### Timing + +Expiration is not event-time accurate. The lag from the triggering PUT to the actual delete is bounded by the worker invocation cadence (default 24h) plus, for walker-only rules, up to one `walker_interval_minutes` window. Budget for "up to a day," not seconds. + +`Days` count in 24-hour units from the relevant clock (the latest-version PUT for `Expiration.Days`, the demoting PUT for `NoncurrentDays`, the MPU initiation for `AbortIncompleteMultipartUpload`). The AWS CLI and SDKs reject `Days < 1` client-side, so one day is the smallest window you'll set through them. SeaweedFS does **not** validate this server-side — a raw `PUT` with `Days` set to `0` (or omitted) is stored and returned as-is, but the rule compiles to a no-op and deletes nothing. The in-repo integration tests compile with a build tag that shortens one "day" to 10 seconds so a full lifecycle can be exercised in CI; released binaries always use 24 hours. + ## Versioning integration Lifecycle rules interact with [S3 Object Versioning](S3-Object-Versioning): @@ -110,8 +162,22 @@ Lifecycle rules interact with [S3 Object Versioning](S3-Object-Versioning): - **`ExpiredObjectDeleteMarker`** removes delete markers that are the sole remaining version of an object (no non-current versions behind them). - **`Expiration.Days`** on a versioned bucket creates a delete marker when the current version expires; it does not permanently delete the object. +## Disk reclamation and the TTL fast path + +By default the worker deletes objects individually and disk is freed later by [volume vacuum](Volume-Management). For high-churn, non-versioned buckets with a fixed `Expiration.Days` retention, an opt-in **TTL fast path** stamps the expiry as a volume TTL at write time, so disk is reclaimed by dropping the whole volume and the worker only removes the metadata entry. It is off by default and applies only to non-versioned, non-Object-Lock buckets with prefix/size `Expiration.Days` rules; everything else falls back to the worker automatically. + +```text +weed shell -master +> s3.bucket.lifecycle.fastpath -name my-bucket # show current state +> s3.bucket.lifecycle.fastpath -name my-bucket -enable +> s3.bucket.lifecycle.fastpath -name my-bucket -disable +``` + +Trade-offs and when to choose it are covered in [S3 Lifecycle vs Volume TTL](S3-Lifecycle-vs-Volume-TTL). + ## Quick references +- **[Recipes](S3-Lifecycle-Recipes)** — copy-paste configs for the common scenarios, with how to verify each - **[Operator Guide](S3-Lifecycle-Operator-Guide)** — config knobs, defaults, when to change each - **[Monitoring](S3-Lifecycle-Monitoring)** — Prometheus metrics, heartbeat log line, what a healthy run looks like - **[Troubleshooting](S3-Lifecycle-Troubleshooting)** — stuck cursor, missing deletes, head-of-line blocking diff --git a/_Sidebar.md b/_Sidebar.md index 7e75bc6..dbb908d 100644 --- a/_Sidebar.md +++ b/_Sidebar.md @@ -83,6 +83,11 @@ * [[Amazon S3 API]] * [[Supported APIs vs Minio]] * [[S3 Lifecycle]] + * [[S3 Lifecycle Recipes]] + * [[S3 Lifecycle Operator Guide]] + * [[S3 Lifecycle Monitoring]] + * [[S3 Lifecycle Troubleshooting]] + * [[S3 Lifecycle Architecture]] * [[S3 Lifecycle vs Volume TTL]] * [[S3 Conditional Operations]] * [[S3 CORS]]