From 29aa6c5c5f059453abc1a1a7ee5e3e77ce8f7f18 Mon Sep 17 00:00:00 2001 From: Chris Lu Date: Wed, 13 May 2026 15:12:22 -0700 Subject: [PATCH] S3 Lifecycle: refresh + add operator, monitoring, troubleshooting, architecture pages Synced from seaweedfs/seaweedfs:docs/wiki/s3-lifecycle/ (PR #9491). The existing S3-Lifecycle page described the pre-redesign two-tier TTL+scan architecture. Replace the architecture and config-knob sections with the current as-built daily-replay model. Keep the useful AWS CLI / Terraform examples and the rule-support table. Add four new pages: - S3-Lifecycle-Operator-Guide: config knobs, walker interval recommendations by cluster size - S3-Lifecycle-Monitoring: Prometheus metric reference, heartbeat tokens, suggested PromQL alerts - S3-Lifecycle-Troubleshooting: stuck cursor playbook, failure outcomes, cursor schema for manual inspection - S3-Lifecycle-Architecture: high-level overview between Home and the in-tree DESIGN.md --- S3-Lifecycle-Architecture.md | 121 ++++++++++++++++++++++++ S3-Lifecycle-Monitoring.md | 112 ++++++++++++++++++++++ S3-Lifecycle-Operator-Guide.md | 135 +++++++++++++++++++++++++++ S3-Lifecycle-Troubleshooting.md | 151 ++++++++++++++++++++++++++++++ S3-Lifecycle.md | 158 +++++++++++++------------------- 5 files changed, 584 insertions(+), 93 deletions(-) create mode 100644 S3-Lifecycle-Architecture.md create mode 100644 S3-Lifecycle-Monitoring.md create mode 100644 S3-Lifecycle-Operator-Guide.md create mode 100644 S3-Lifecycle-Troubleshooting.md diff --git a/S3-Lifecycle-Architecture.md b/S3-Lifecycle-Architecture.md new file mode 100644 index 0000000..0ae217d --- /dev/null +++ b/S3-Lifecycle-Architecture.md @@ -0,0 +1,121 @@ +# S3 Lifecycle — Architecture + +High-level overview of the lifecycle worker. For implementation detail, see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md). + +## At a glance + +The lifecycle worker runs as a scheduled job. Each invocation: + +``` + ┌──────────────────────────────────────────┐ + │ dailyrun.Run (one filer subscription) │ + │ │ + meta-log ──→ │ reader ──→ fan-out ──→ per-shard │ + │ channels │ + │ │ + │ ┌──────────────────────────────────┐ │ + │ │ 16 shard goroutines │ │ + │ │ ┌──────────────────────────┐ │ │ + │ │ │ walker(view, shardID)? │ │ │ + │ │ │ drainShardEvents │ │ │ + │ │ │ saveCursorAndPublish │ │ │ + │ │ └──────────────────────────┘ │ │ + │ └──────────────────────────────────┘ │ + │ │ + │ summary heartbeat + exit │ + └──────────────────────────────────────────┘ +``` + +One filer `SubscribeMetadata` stream covers every shard in this worker's set. A fan-out goroutine routes events to per-shard channels by `ev.ShardID = sha256(bucket || "/" || key) >> 252`. Each shard's goroutine independently runs the walker (when due), drains events, and persists its cursor. + +Once every shard's goroutine returns, the worker tears down the subscription, emits a summary heartbeat, and exits. + +## Per-shard state + +Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`: + +``` +TsNs — last meta-log event whose matches all dispatched successfully +RuleSetHash — ReplayContentHash of the rule set when this cursor was written +PromotedHash — PromotedHash(retentionWindow) at write time +LastWalkedNs — wall-clock of the last successful walker fire +``` + +The two hashes together detect every situation that invalidates the cursor: a replay-rule edit (`RuleSetHash` changes) or a partition flip (`PromotedHash` changes). On mismatch, the next pass triggers a recovery walk over `RecoveryView(snap)` to catch already-due objects across the full rule set, then rewinds the cursor. + +## Replay vs walker + +The lifecycle rule space splits two ways: + +| Path | Action kinds | Why this path | +|---|---|---| +| Replay (meta-log) | `ExpirationDays`, `NoncurrentDays`, `AbortMPU` | DueTime is monotonic in event TsNs. The `done` early-stop works. | +| Walker (bucket list) | `ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` | DueTime depends on current sibling/version state, not event age. | + +The engine's `RulesForShard(shardID, retentionWindow)` returns two snapshot views (`replay`, `walk`); each is a clone of the base snapshot with the action map masked to the right partition. `router.Route` consumes the `replay` view per event; the walker consumes the `walk` view per bucket. + +A rule promoted to scan-only because its TTL exceeds meta-log retention moves from `replay` to `walk` — visible via `PromotedHash`. + +## Cadence layers + +Three independent cadences shape worker behavior: + +| Cadence | Set by | Default | +|---|---|---| +| Worker invocation | Admin scheduler `DetectionIntervalMinutes` | 1440 (daily) | +| Walker fire | `walker_interval_minutes` admin config | 0 (every invocation) | +| Cursor save | After each `runShard` | n/a | + +The walker throttle decouples walker firing from invocation rate. CI invokes the worker every 2s; production invokes once per day. Both can use the same code with appropriate `walker_interval_minutes`. + +## Failure model + +- **Worker crash mid-run.** Cursor only advances past events whose matches all succeeded. On restart, the next pass resumes at the same cursor. Identity-CAS makes redundant deletes no-ops. +- **Transient delete failure.** Pass halts at the failing event, cursor persists. Next pass retries from there. Head-of-line blocking is intentional — surfaces real problems instead of silently retrying forever. +- **Rule edits.** Replay-rule edits trigger one-time recovery walk over `RecoveryView`. Walker-only rule edits don't change either hash; walker reads the new rules on its next steady-state fire. +- **Object overwritten between event and delete.** `LifecycleDelete` RPC's identity-CAS returns `NOOP_RESOLVED`; cursor advances normally. + +## Rate limiting + +Cluster-wide cap allocated per worker at job dispatch: + +``` +per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers) +``` + +Each worker shares one `rate.Limiter` across all shard goroutines. `dispatchWithRetry` calls `limiter.Wait(ctx)` before each `LifecycleDelete` RPC. + +## Components + +| Path | Role | +|---|---| +| `engine/` | Rule compilation, partition views (`RulesForShard`, `RecoveryView`) | +| `evaluate.go` | Per-event rule evaluation (`EvaluateAction`) | +| `due_at.go` | Per-(rule, kind, info) due-time computation | +| `router/router.go` | Per-event match emission (calls engine.Action and EvaluateAction) | +| `reader/reader.go` | Meta-log subscribe with `ShardPredicate` | +| `bootstrap/walker.go` | Bucket-walker with `RunForShard` filter | +| `dailyrun/run.go` | Main orchestrator: subscription, fan-out, per-shard runShard | +| `dailyrun/cursor.go` | Cursor type + filer JSON serializer | +| `dailyrun/walker_dispatcher.go` | Walker-to-`LifecycleDelete` adapter | + +## What it's not + +- **Not a streaming dispatcher.** The earlier model kept a long-running goroutine per shard with an in-memory match heap. That code is gone. Worker is now "start, do today's work, stop." +- **Not event-time accurate.** Latency from PUT to delete is bounded by the worker invocation cadence plus the walker interval — typically up to 24h, not seconds. +- **Not a general-purpose scheduler.** The two action paths (replay, walker) are specific to lifecycle semantics. Don't add new event sources or actions without thinking through which path they belong on. + +## Why this shape + +Each design choice points back to a specific failure mode of the prior streaming worker: + +| Choice | Replaces | +|---|---| +| Per-pass run + exit | Long-running goroutines with ticker drift, leak risk, restart pain | +| Cursor file per shard | Per-key freeze state, retry counters, in-memory heap on every restart | +| Identity-CAS at dispatch time | Pre-dispatch consistency checks at schedule time, racing object updates | +| Recovery branch over `RecoveryView` | Implicit "is this rule new" tracking with bookkeeping flags | +| Walker throttle independent of invocation | Walker hammering filer when test driver invokes every 2s | +| Single subscription per pass | 16x filer load with 16 per-shard subscriptions | + +The result is a worker the operator can reason about by reading 2 metrics and a heartbeat line, with a state machine small enough to fit in one design doc. diff --git a/S3-Lifecycle-Monitoring.md b/S3-Lifecycle-Monitoring.md new file mode 100644 index 0000000..fbb8fc7 --- /dev/null +++ b/S3-Lifecycle-Monitoring.md @@ -0,0 +1,112 @@ +# S3 Lifecycle — Monitoring + +This page lists the Prometheus signals the worker exposes and how to read the heartbeat log line. For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting). + +## Prometheus metrics + +All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem. + +### Per-shard gauges + +| Metric | Labels | What | +|---|---|---| +| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard | +| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire | + +Derived queries: + +```promql +# Per-shard replay lag in seconds +(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 + +# Per-shard walker freshness in seconds +(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 + +# Worst-shard lag across the cluster +max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 +``` + +Zero values mean "not started yet" — distinct from "0s caught up". The heartbeat line uses `cold` as the marker for that state. + +### Counters + +| Metric | Labels | What | +|---|---|---| +| `s3_lifecycle_dispatch_total` | `bucket`, `kind`, `outcome` | Per-bucket dispatch counter, partitioned by action kind and server outcome | +| `s3_lifecycle_daily_run_events_scanned_total` | `shard` | Meta-log events `drainShardEvents` processed | +| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter | +| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path | + +`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. + +### Histograms + +| Metric | What | +|---|---| +| `s3_lifecycle_daily_run_shard_duration_seconds{shard}` | Wall-clock per shard pass. p95 climbing toward `max_runtime_minutes` means the shard is brushing its budget. | +| `s3_lifecycle_dispatch_limiter_wait_seconds` | Time spent waiting on the cluster rate limiter before issuing `LifecycleDelete`. Near-zero = cap not binding; long-tail at `1/rate` = cap is the active throttle. | + +## Heartbeat log line + +Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity): + +``` +daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m +``` + +Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stable across versions: + +| Token | Meaning | +|---|---| +| `status=ok` or `status=error` | Whether any shard returned an error | +| `shards=N` | Number of shards processed this pass | +| `errors=N` | Per-shard error count | +| `duration=Ns` | Wall-clock for the whole pass | +| `cursor_lag_max=...` | Worst per-shard replay lag, or `cold` if no shard has a persisted cursor yet | +| `walked_max_age=...` | Worst per-shard walker age, or `cold` if no shard has walked yet | + +A healthy production heartbeat looks like: + +``` +daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m +``` + +Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag is 45 seconds behind real-time; the oldest walker fire on any shard is 58 minutes ago (so `walker_interval_minutes=60` is roughly honored). + +## Anti-patterns to alert on + +| Pattern | Meaning | What to do | +|---|---|---| +| `cursor_lag_max` grows unbounded | Stuck cursor; head-of-line blocking on some shard | See [Troubleshooting → Stuck cursor](S3-Lifecycle-Troubleshooting#stuck-cursor) | +| `walked_max_age` exceeds `walker_interval_minutes × 2` | Walker isn't firing as configured | Check `errors=N` in heartbeat and `s3_lifecycle_dispatch_total{outcome="RPC_ERROR"}` | +| `errors=16` (all shards) on every pass | Filer is unreachable or returning errors | Check filer health | +| `s3_lifecycle_dispatch_total{outcome="RETRY_LATER"}` rising fast | Server rate-limited or filer overloaded | Lower `cluster_deletes_per_second` or add capacity | +| `s3_lifecycle_dispatch_total{outcome="BLOCKED"}` non-zero | Programmatic event content error | Check worker logs for `FATAL_EVENT_ERROR` | +| `duration=Ns` ramping up across passes | Walker is firing too often | Set `walker_interval_minutes` | + +## Suggested alerts + +```yaml +- alert: S3LifecycleCursorLagHigh + expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600 + for: 30m + annotations: + summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}" + runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor + +- alert: S3LifecycleWalkerStuck + expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400 + for: 1h + annotations: + summary: "S3 lifecycle walker hasn't run in > 24h" + runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#walker-stuck + +- alert: S3LifecycleDispatchFailures + expr: | + rate(s3_lifecycle_dispatch_total{outcome=~"RETRY_LATER|BLOCKED|RPC_ERROR"}[5m]) > 0.1 + for: 15m + annotations: + summary: "S3 lifecycle delete failure rate > 0.1/s" +``` + +Adjust thresholds to your cluster's normal levels — these are starting points. diff --git a/S3-Lifecycle-Operator-Guide.md b/S3-Lifecycle-Operator-Guide.md new file mode 100644 index 0000000..35b12ce --- /dev/null +++ b/S3-Lifecycle-Operator-Guide.md @@ -0,0 +1,135 @@ +# S3 Lifecycle — Operator Guide + +This page covers the admin and worker config knobs for the S3 lifecycle worker, plus when to change each one. + +For monitoring guidance, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring). For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting). + +## Configuration + +All keys are set through the admin UI's plugin config for `s3_lifecycle`. + +### Admin config + +| Key | Type | Default | When to change | +|---|---|---|---| +| `cluster_deletes_per_second` | int64 | `0` (unlimited) | Set a positive value when lifecycle deletes are causing filer contention. Allocated evenly across active workers at job dispatch. | +| `cluster_deletes_burst` | int64 | `0` (= 2× rate) | Adjust if delete bursts overload the filer faster than the per-second rate allows. | +| `meta_log_retention_days` | int64 | `0` (treated as unbounded) | Stock SeaweedFS doesn't GC the meta-log, so the default is fine. Set positive if your deployment manually trims `/topics/.system/log` — then rules with TTL > retention will route through the walker. | +| `walker_interval_minutes` | int64 | `0` (fire every pass) | **Important.** See "Walker interval" below — most production deployments should set this to a positive value. | + +### Worker config + +| Key | Default | What | +|---|---|---| +| `max_runtime_minutes` | 60 | Wall-clock cap per `dailyrun.Run` invocation. The pass returns early if it hits this. | + +### Detection schedule + +The admin scheduler's `DetectionIntervalMinutes` for `s3_lifecycle` is `1440` by default — once per day. Each detection produces one execution. Change in the admin UI's runtime defaults for the job type. + +## Walker interval + +The walker is the part of the worker that lists bucket contents and evaluates them against rules. It fires: + +- **Always**, on cold start (no persisted cursor) and on rule changes — these are bounded events. +- **Periodically**, in steady state, gated by `walker_interval_minutes`. + +The throttle exists because walker cost is bucket size, not event rate. If the worker is scheduled at a tighter cadence than the desired walk frequency (CI, sub-hourly admin schedules, manual runs), the steady-state walker would crush the filer with a full subtree scan per invocation. + +Recommended values by cluster size: + +| Cluster | Recommended `walker_interval_minutes` | +|---|---| +| Small (≤1M objects, single bucket) | 60 | +| Medium (≤100M objects, mixed buckets) | 360 (6h) | +| Large (≥100M objects) | 1440 (24h) | +| Testing / CI | 0 (fire every pass) | + +Setting `0` keeps the prior behavior — fire on every invocation. That's appropriate when the worker is scheduled at exactly the desired walk frequency (e.g., once per day) and the in-repo integration tests rely on it. + +A negative value is rejected at worker start (loud error rather than silent fall-through to "walk every pass"). + +## Per-bucket lifecycle XML + +Set via the standard S3 API: + +```bash +aws --endpoint $S3_ENDPOINT s3api put-bucket-lifecycle-configuration \ + --bucket my-bucket \ + --lifecycle-configuration file://lifecycle.json +``` + +```jsonc +// lifecycle.json +{ + "Rules": [ + { + "ID": "expire-logs-after-30d", + "Status": "Enabled", + "Filter": { "Prefix": "logs/" }, + "Expiration": { "Days": 30 } + }, + { + "ID": "abort-stuck-mpu", + "Status": "Enabled", + "Filter": {}, + "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } + }, + { + "ID": "keep-3-versions", + "Status": "Enabled", + "Filter": { "Prefix": "versioned/" }, + "NoncurrentVersionExpiration": { "NewerNoncurrentVersions": 3 } + } + ] +} +``` + +The worker picks up rule changes on the next pass. Replay-eligible rule edits trigger a one-time recovery walk to catch already-due objects under the new rule. + +## Verifying a rule is working + +After applying a rule: + +1. Wait one detection interval (default 24h) plus one walker interval. +2. Read `s3_lifecycle_dispatch_total{bucket="my-bucket"}` — the counter should advance. +3. Verify a target object is gone: `aws s3 head-object --bucket my-bucket --key ` should return 404. + +For testing without waiting, the `weed shell` command supports manual invocation: + +``` +weed shell -master +> s3.lifecycle.run-shard -shards 0-15 -s3 -refresh 1s -runtime 30s +``` + +This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for examples. + +## Rate limit allocation + +The cluster delete cap is allocated per-worker at job dispatch: + +``` +per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers) +``` + +Brief over/undershoot during worker join/leave is acceptable for a bulk workload. The allocation is recomputed each detection cycle, so adding workers smooths out within one day. + +`cluster_deletes_burst` is divided the same way. `0` is treated as `2 × rate` (a token bucket reasonable default). + +## Common change patterns + +### Tightening a TTL (e.g., 60d → 30d) + +The new rule hash differs from the persisted hash — next pass triggers a recovery walk over `RecoveryView`. Already-due objects under the new shorter TTL get caught. The cursor rewinds to `runNow - new_maxTTL`. No operator action needed beyond updating the XML. + +### Adding a walker-only rule (`Expiration.Date`) + +`RuleSetHash` and `PromotedHash` are unchanged (walker-only rules aren't in the replay hash). The replay cursor stays put. The walker reads the updated rule set on its next steady-state fire (or immediately, if invoked manually). + +### Operator forgot to set `walker_interval_minutes` + +Symptom: heartbeat log line shows `duration` ramping up across passes (each pass does a full subtree walk). Mitigation: set the throttle to a value matching your scheduling cadence. + +### Filer is GC'ing meta-log (custom deployment) + +Set `meta_log_retention_days` to the GC retention. Rules with TTL > retention will route through the walker (`PromotedHash` will be non-empty). The first run after this change triggers a recovery walk. diff --git a/S3-Lifecycle-Troubleshooting.md b/S3-Lifecycle-Troubleshooting.md new file mode 100644 index 0000000..0209496 --- /dev/null +++ b/S3-Lifecycle-Troubleshooting.md @@ -0,0 +1,151 @@ +# S3 Lifecycle — Troubleshooting + +Incident-response playbook for the S3 lifecycle worker. For monitoring background, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring). + +## Stuck cursor + +**Symptom:** `s3_lifecycle_cursor_min_ts_ns{shard=N}` is not advancing. Heartbeat shows `cursor_lag_max` growing unbounded. + +**Cause:** The cursor advance is gated on every match from the event dispatching successfully (`DONE`, `NOOP_RESOLVED`, or `SKIPPED_OBJECT_LOCK`). Any unresolved outcome (`RETRY_LATER`, `BLOCKED`, transport error after in-run retries) halts the run for that shard and persists the cursor at the last fully-processed event. Head-of-line blocking is intentional — it surfaces a real problem rather than silently retrying forever. + +**Diagnostics:** + +```promql +# Which shard is stuck? +time() * 1e9 - s3_lifecycle_cursor_min_ts_ns + +# What outcomes are being returned? +sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m])) +``` + +Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`: + +``` +daily_run: RETRY_LATER on / EXPIRATION_DAYS +daily_run: BLOCKED on / NONCURRENT_DAYS +daily_run: transport error on / ABORT_MPU: +``` + +**Mitigations:** + +| Outcome | Root cause | Action | +|---|---|---| +| `RETRY_LATER` (high rate) | Filer or rate limiter is throttling | Lower `cluster_deletes_per_second` to give the filer headroom, or scale filer capacity | +| `BLOCKED FATAL_EVENT_ERROR` | A malformed event the server refuses to dispatch | Check log for the specific reason. May need a code fix; file an issue with the log line | +| `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. | +| `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability | + +The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker. + +## Walker stuck (no progress on walker-only rules) + +**Symptom:** `s3_lifecycle_daily_run_last_walked_ns{shard=N}` is not advancing. Rules like `Expiration.Date`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` aren't firing on objects that should be due. + +**Causes:** + +1. `walker_interval_minutes` is too long for your invocation cadence. Worker runs once per day but interval is set to 48h. +2. Walker is hitting an error mid-walk (filer listing failure). Look for `recovery walk:` or `steady walk:` errors in the heartbeat's `errors=N` count. +3. The bucket has only walker-bound rules and the empty-replay branch's throttle hasn't elapsed. + +**Diagnostics:** + +```promql +# Walker age per shard +(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 +``` + +Check worker config: `walker_interval_minutes` should be ≤ the daily worker schedule interval. + +**Mitigation:** lower `walker_interval_minutes`. Setting `0` temporarily forces every pass to walk. + +## Test PUT a file with a 1-day rule, didn't expire + +The S3 API rejects `Expiration.Days < 1`, so the smallest "expire after N days" you can configure is 1 day. The worker runs once per day by default. Object PUT + 1-day rule + waiting one day is the minimum scenario. + +For testing, the in-repo integration suite uses a trick: backdate the entry's `Mtime` via `filer_pb.UpdateEntry` to 30+ days ago. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for the pattern. + +For ad-hoc verification, invoke the worker manually: + +``` +weed shell -master +> s3.lifecycle.run-shard -shards 0-15 -s3 -refresh 1s -runtime 30s +``` + +This runs the same code path as the scheduled worker, but driven from your shell rather than the admin scheduler. + +## All shards report `errors=16` every pass + +**Symptom:** Heartbeat consistently shows `status=error shards=16 errors=16 duration=Ns`. + +**Common causes:** + +1. **Filer unreachable.** Subscription fails on every shard. Check filer health and gRPC connectivity. +2. **passCtx timeout from `-refresh` loop.** If `-refresh` is less than the pass cap, the timeout fires before the drain completes. This is now treated as "clean end-of-pass" — it should not show as errors=N. If you see this on a build before #9481, upgrade. +3. **Bucket walker is timing out.** Big bucket, walker hits ctx deadline. Increase `max_runtime_minutes`. + +## Some objects expired, others didn't (same rule) + +**Symptom:** Two objects matching the same rule with the same age — one is deleted, the other isn't. + +**Common causes:** + +1. **The non-deleted object's mtime is wrong.** Check the entry's mtime — it might be more recent than you expect (e.g., a recent metadata update bumped it). +2. **The objects are on different shards** and one shard has a stuck cursor while the other doesn't. Check `s3_lifecycle_cursor_min_ts_ns` per shard. +3. **`Filter` doesn't match what you think.** A prefix-only filter requires the object key to start with that prefix; a tag filter requires the matching tag. Verify with `aws s3api head-object` (returns tags via `--query`). +4. **Object lock or retention.** The dispatcher returns `SKIPPED_OBJECT_LOCK` for protected objects. Check `s3_lifecycle_dispatch_total{outcome="SKIPPED_OBJECT_LOCK"}`. + +## How to read the cursor files + +Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`: + +``` +weed shell -master +> fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json +``` + +Schema: + +```json +{ + "version": 1, + "shard_id": 0, + "ts_ns": 1715600000000000000, + "rule_set_hash": "", + "promoted_hash": "", + "last_walked_ns": 1715620000000000000 +} +``` + +- `ts_ns == 0` means "no replay progress" — either cold start or a bucket whose rules are all walker-only. +- `last_walked_ns == 0` (or absent) means "never walked steady-state". Next pass will walk. + +Manually editing the cursor is supported as an escape hatch but obviously breaks the invariant that "everything before persisted.TsNs has been processed under the same rules." Use sparingly. + +## Resetting a shard + +If a shard's cursor is corrupted or wedged in an unrecoverable state: + +``` +weed shell -master +> fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json +``` + +Next pass treats the shard as cold start: recovery walker fires over `RecoveryView(snap)`, then the cursor seeds at `runNow - maxTTL`. This re-replays a `maxTTL`-wide window of meta-log events. Identity-CAS on the server side makes redundant deletes no-ops, so re-replay is safe. + +## Suspending the worker + +The admin UI's plugin scheduler allows pausing the `s3_lifecycle` job type. The worker won't be invoked while paused; cursors are preserved, and on resume the next pass picks up at the persisted state. + +For a single-bucket "stop deleting from this bucket" without pausing the worker: + +```bash +aws --endpoint $S3_ENDPOINT s3api delete-bucket-lifecycle --bucket my-bucket +``` + +The worker reads the empty rule set on the next pass; `rsh == [32]byte{}` causes the empty-replay branch to run only the walker (which has nothing to walk for an empty rule set) and exit. No deletes. + +## Reverting + +If something goes very wrong, the streaming worker (the previous design) is no longer in the codebase. Revert is via downgrading the binary. The cursor format has a `version` field — versions `>1` would fail-loud on load by an older binary that only knows version `1`. Currently version `1` is the only version. + +For escape-hatch operations (delete all cursors, suspend worker globally), prefer pausing via the admin UI over force-killing the worker process; the worker exits cleanly between passes. diff --git a/S3-Lifecycle.md b/S3-Lifecycle.md index 251eac2..8972201 100644 --- a/S3-Lifecycle.md +++ b/S3-Lifecycle.md @@ -1,26 +1,26 @@ -# S3 Lifecycle Configuration +# S3 Lifecycle -SeaweedFS supports S3 bucket lifecycle configuration for automated object expiration, non-current version cleanup, delete marker removal, and incomplete multipart upload abortion. +SeaweedFS implements the S3 `PutBucketLifecycleConfiguration` API. Configured rules are evaluated and enforced by a worker that runs as a scheduled job and exits when each pass completes. -## Supported Features +This page is the operator-facing entry point. Developers and architecture readers should see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md). + +## Supported features | Feature | Status | Notes | |---|---|---| -| `Expiration.Days` | Supported | Fast path via TTL + RocksDB compaction filter | -| `Expiration.Date` | Supported | Evaluated by lifecycle worker at scan time | -| `ExpiredObjectDeleteMarker` | Supported | Removes delete markers that are the sole remaining version | -| `NoncurrentVersionExpiration.NoncurrentDays` | Supported | Requires bucket versioning enabled | -| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Supported | Keep N newest non-current versions | -| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Supported | Per-rule prefix scoping | -| `Filter.Prefix` | Supported | | -| `Filter.Tag` | Supported | Evaluated at scan time | -| `Filter.And` (Prefix + Tags + Size) | Supported | Evaluated at scan time | -| `Filter.ObjectSizeGreaterThan` | Supported | Evaluated at scan time | -| `Filter.ObjectSizeLessThan` | Supported | Evaluated at scan time | -| `Transition` | Not supported | Requires storage class tiers | -| `NoncurrentVersionTransition` | Not supported | Requires storage class tiers | +| `Expiration.Days` | Yes | Latest-version PUT clock | +| `Expiration.Date` | Yes | Walker path; fires once date is reached | +| `Expiration.ExpiredObjectDeleteMarker` | Yes | Walker path; sibling-aware | +| `NoncurrentVersionExpiration.NoncurrentDays` | Yes | Clock starts at the demoting PUT, not the entry's own mtime | +| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Yes | Walker path; version-list aware | +| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Yes | | +| `Filter.Prefix` | Yes | | +| `Filter.Tag` | Yes | | +| `Filter.ObjectSizeGreaterThan` / `ObjectSizeLessThan` | Yes | | +| `Filter.And` (composite) | Yes | | +| `Transition` / `NoncurrentVersionTransition` | No | SeaweedFS doesn't model storage class tiers | -## API Endpoints +## API endpoints ``` PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration @@ -28,6 +28,37 @@ GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration DELETE /{bucket}?lifecycle # DeleteBucketLifecycle ``` +## Example: AWS CLI + +```bash +# Set lifecycle configuration +aws s3api put-bucket-lifecycle-configuration \ + --endpoint-url http://localhost:8333 \ + --bucket my-bucket \ + --lifecycle-configuration '{ + "Rules": [ + { + "ID": "expire-old", + "Status": "Enabled", + "Filter": { "Prefix": "" }, + "Expiration": { "Days": 90 }, + "NoncurrentVersionExpiration": { "NoncurrentDays": 30 }, + "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 } + } + ] + }' + +# Get lifecycle configuration +aws s3api get-bucket-lifecycle-configuration \ + --endpoint-url http://localhost:8333 \ + --bucket my-bucket + +# Delete lifecycle configuration +aws s3api delete-bucket-lifecycle \ + --endpoint-url http://localhost:8333 \ + --bucket my-bucket +``` + ## Example: Terraform ```hcl @@ -47,7 +78,7 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" { } noncurrent_version_expiration { - noncurrent_days = 7 + noncurrent_days = 7 newer_noncurrent_versions = 2 } @@ -58,89 +89,30 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" { } ``` -## Example: AWS CLI +## How it works -```bash -# Set lifecycle configuration -aws s3api put-bucket-lifecycle-configuration \ - --endpoint-url http://localhost:8333 \ - --bucket my-bucket \ - --lifecycle-configuration '{ - "Rules": [ - { - "ID": "expire-old", - "Status": "Enabled", - "Filter": { "Prefix": "" }, - "Expiration": { "Days": 90 }, - "NoncurrentVersionExpiration": { - "NoncurrentDays": 30 - }, - "AbortIncompleteMultipartUpload": { - "DaysAfterInitiation": 7 - } - } - ] - }' +The lifecycle worker is a scheduled job (default daily). Each invocation: -# Get lifecycle configuration -aws s3api get-bucket-lifecycle-configuration \ - --endpoint-url http://localhost:8333 \ - --bucket my-bucket +1. Reads bucket lifecycle XML from each bucket's metadata. +2. Compiles rules into a per-shard partition (replay-eligible vs. walker-bound). +3. Subscribes to the filer meta-log — one stream covering all 16 shards in this worker process. +4. For replay-eligible actions (`ExpirationDays`, `NoncurrentDays`, `AbortMPU`), checks each event's DueTime and dispatches `LifecycleDelete` if elapsed. +5. For walker-bound rules (`ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent`, or anything promoted to scan-only), iterates the bucket and evaluates each entry against current state. +6. Persists per-shard cursors so the next pass resumes where this one left off. -# Delete lifecycle configuration -aws s3api delete-bucket-lifecycle \ - --endpoint-url http://localhost:8333 \ - --bucket my-bucket -``` +The worker exits when the pass is done. The admin scheduler invokes it on a daily cadence by default; operators can change that via the standard plugin scheduler config. -## How It Works - -### Two-Tier Expiration Architecture - -**Tier 1 — TTL fast path:** Simple `Expiration.Days` rules with prefix-only filters are translated into TTL entries in `filer.conf`. The RocksDB compaction filter automatically removes expired entries during normal compaction at zero additional cost. This is the most efficient path for the common case. - -**Tier 2 — Scan-time evaluation:** Rules with tag filters, size filters, date-based expiration, non-current version expiration, and delete marker cleanup are evaluated by the lifecycle plugin worker. The worker periodically scans buckets and evaluates each object against the stored lifecycle XML configuration. - -### Which rules use which tier? - -| Rule Type | Tier | Why | -|---|---|---| -| `Expiration.Days` (prefix only) | TTL fast path | Can be expressed as per-entry TTL | -| `Expiration.Days` (with tags/size) | Worker scan | TTL can't express tag/size constraints | -| `Expiration.Date` | Worker scan | Absolute date, not relative TTL | -| `NoncurrentVersionExpiration` | Worker scan | Requires version enumeration | -| `ExpiredObjectDeleteMarker` | Worker scan | Requires version counting | -| `AbortIncompleteMultipartUpload` | Worker scan | Scans `.uploads` directory | - -### Lifecycle Worker - -The lifecycle plugin worker runs as part of the SeaweedFS plugin system. It: - -1. **Detects** buckets with lifecycle rules (from stored lifecycle XML or filer.conf TTLs) -2. **Scans** bucket contents and evaluates lifecycle rules against each object -3. **Executes** actions: delete expired objects, remove old non-current versions, clean up delete markers, abort stale multipart uploads - -Worker configuration options: - -| Setting | Default | Description | -|---|---|---| -| `batch_size` | 1000 | Entries per filer listing page | -| `max_deletes_per_bucket` | 10000 | Max expired objects to delete per run | -| `dry_run` | false | Detect but don't delete | -| `delete_marker_cleanup` | true | Remove expired delete markers | -| `abort_mpu_days` | 7 | Fallback for buckets without lifecycle XML MPU rules | -| `bucket_filter` | (all) | Wildcard pattern to scope lifecycle to specific buckets | - -## Versioning Integration +## Versioning integration Lifecycle rules interact with [S3 Object Versioning](S3-Object-Versioning): -- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded. `NewerNoncurrentVersions` retains the N newest non-current versions. +- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded (the demoting PUT's TsNs, not the version's own mtime). `NewerNoncurrentVersions` retains the N newest non-current versions. - **`ExpiredObjectDeleteMarker`** removes delete markers that are the sole remaining version of an object (no non-current versions behind them). - **`Expiration.Days`** on a versioned bucket creates a delete marker when the current version expires; it does not permanently delete the object. -## Limitations +## Quick references -- **Transition rules** (`Transition`, `NoncurrentVersionTransition`) are not supported. SeaweedFS does not have S3-equivalent storage class tiers. -- **Detection interval**: The lifecycle worker scans on a configurable interval (default 5 minutes). Objects may persist slightly beyond their configured expiration until the next scan. -- **TTL fast path** applies only to `Expiration.Days` rules with prefix-only filters. Rules with tag or size constraints are always evaluated at scan time. +- **[Operator Guide](S3-Lifecycle-Operator-Guide)** — config knobs, defaults, when to change each +- **[Monitoring](S3-Lifecycle-Monitoring)** — Prometheus metrics, heartbeat log line, what a healthy run looks like +- **[Troubleshooting](S3-Lifecycle-Troubleshooting)** — stuck cursor, missing deletes, head-of-line blocking +- **[Architecture](S3-Lifecycle-Architecture)** — high-level overview of the worker, engine, and dispatch path