S3 Lifecycle: refresh + add operator, monitoring, troubleshooting, architecture pages

Synced from seaweedfs/seaweedfs:docs/wiki/s3-lifecycle/ (PR #9491).

The existing S3-Lifecycle page described the pre-redesign two-tier
TTL+scan architecture. Replace the architecture and config-knob
sections with the current as-built daily-replay model. Keep the
useful AWS CLI / Terraform examples and the rule-support table.

Add four new pages:
- S3-Lifecycle-Operator-Guide: config knobs, walker interval
  recommendations by cluster size
- S3-Lifecycle-Monitoring: Prometheus metric reference, heartbeat
  tokens, suggested PromQL alerts
- S3-Lifecycle-Troubleshooting: stuck cursor playbook, failure
  outcomes, cursor schema for manual inspection
- S3-Lifecycle-Architecture: high-level overview between Home and
  the in-tree DESIGN.md
Chris Lu committed 2026-05-13 15:12:22 -07:00
1 parent f8ec586b82
commit 29aa6c5c5f
5 files changed
+584 -93

No files matched your search

+121
@@ -0,0 +1,121 @@
# S3 Lifecycle — Architecture
High-level overview of the lifecycle worker. For implementation detail, see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md).
## At a glance
The lifecycle worker runs as a scheduled job. Each invocation:
```
┌──────────────────────────────────────────┐
│ dailyrun.Run (one filer subscription) │
│ │
meta-log ──→ │ reader ──→ fan-out ──→ per-shard │
│ channels │
│ │
│ ┌──────────────────────────────────┐ │
│ │ 16 shard goroutines │ │
│ │ ┌──────────────────────────┐ │ │
│ │ │ walker(view, shardID)? │ │ │
│ │ │ drainShardEvents │ │ │
│ │ │ saveCursorAndPublish │ │ │
│ │ └──────────────────────────┘ │ │
│ └──────────────────────────────────┘ │
│ │
│ summary heartbeat + exit │
└──────────────────────────────────────────┘
```
One filer `SubscribeMetadata` stream covers every shard in this worker's set. A fan-out goroutine routes events to per-shard channels by `ev.ShardID = sha256(bucket || "/" || key) >> 252`. Each shard's goroutine independently runs the walker (when due), drains events, and persists its cursor.
Once every shard's goroutine returns, the worker tears down the subscription, emits a summary heartbeat, and exits.
## Per-shard state
Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`:
```
TsNs — last meta-log event whose matches all dispatched successfully
RuleSetHash — ReplayContentHash of the rule set when this cursor was written
PromotedHash — PromotedHash(retentionWindow) at write time
LastWalkedNs — wall-clock of the last successful walker fire
```
The two hashes together detect every situation that invalidates the cursor: a replay-rule edit (`RuleSetHash` changes) or a partition flip (`PromotedHash` changes). On mismatch, the next pass triggers a recovery walk over `RecoveryView(snap)` to catch already-due objects across the full rule set, then rewinds the cursor.
## Replay vs walker
The lifecycle rule space splits two ways:
| Path | Action kinds | Why this path |
|---|---|---|
| Replay (meta-log) | `ExpirationDays`, `NoncurrentDays`, `AbortMPU` | DueTime is monotonic in event TsNs. The `done` early-stop works. |
| Walker (bucket list) | `ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` | DueTime depends on current sibling/version state, not event age. |
The engine's `RulesForShard(shardID, retentionWindow)` returns two snapshot views (`replay`, `walk`); each is a clone of the base snapshot with the action map masked to the right partition. `router.Route` consumes the `replay` view per event; the walker consumes the `walk` view per bucket.
A rule promoted to scan-only because its TTL exceeds meta-log retention moves from `replay` to `walk` — visible via `PromotedHash`.
## Cadence layers
Three independent cadences shape worker behavior:
| Cadence | Set by | Default |
|---|---|---|
| Worker invocation | Admin scheduler `DetectionIntervalMinutes` | 1440 (daily) |
| Walker fire | `walker_interval_minutes` admin config | 0 (every invocation) |
| Cursor save | After each `runShard` | n/a |
The walker throttle decouples walker firing from invocation rate. CI invokes the worker every 2s; production invokes once per day. Both can use the same code with appropriate `walker_interval_minutes`.
## Failure model
- **Worker crash mid-run.** Cursor only advances past events whose matches all succeeded. On restart, the next pass resumes at the same cursor. Identity-CAS makes redundant deletes no-ops.
- **Transient delete failure.** Pass halts at the failing event, cursor persists. Next pass retries from there. Head-of-line blocking is intentional — surfaces real problems instead of silently retrying forever.
- **Rule edits.** Replay-rule edits trigger one-time recovery walk over `RecoveryView`. Walker-only rule edits don't change either hash; walker reads the new rules on its next steady-state fire.
- **Object overwritten between event and delete.** `LifecycleDelete` RPC's identity-CAS returns `NOOP_RESOLVED`; cursor advances normally.
## Rate limiting
Cluster-wide cap allocated per worker at job dispatch:
```
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
```
Each worker shares one `rate.Limiter` across all shard goroutines. `dispatchWithRetry` calls `limiter.Wait(ctx)` before each `LifecycleDelete` RPC.
## Components
| Path | Role |
|---|---|
| `engine/` | Rule compilation, partition views (`RulesForShard`, `RecoveryView`) |
| `evaluate.go` | Per-event rule evaluation (`EvaluateAction`) |
| `due_at.go` | Per-(rule, kind, info) due-time computation |
| `router/router.go` | Per-event match emission (calls engine.Action and EvaluateAction) |
| `reader/reader.go` | Meta-log subscribe with `ShardPredicate` |
| `bootstrap/walker.go` | Bucket-walker with `RunForShard` filter |
| `dailyrun/run.go` | Main orchestrator: subscription, fan-out, per-shard runShard |
| `dailyrun/cursor.go` | Cursor type + filer JSON serializer |
| `dailyrun/walker_dispatcher.go` | Walker-to-`LifecycleDelete` adapter |
## What it's not
- **Not a streaming dispatcher.** The earlier model kept a long-running goroutine per shard with an in-memory match heap. That code is gone. Worker is now "start, do today's work, stop."
- **Not event-time accurate.** Latency from PUT to delete is bounded by the worker invocation cadence plus the walker interval — typically up to 24h, not seconds.
- **Not a general-purpose scheduler.** The two action paths (replay, walker) are specific to lifecycle semantics. Don't add new event sources or actions without thinking through which path they belong on.
## Why this shape
Each design choice points back to a specific failure mode of the prior streaming worker:
| Choice | Replaces |
|---|---|
| Per-pass run + exit | Long-running goroutines with ticker drift, leak risk, restart pain |
| Cursor file per shard | Per-key freeze state, retry counters, in-memory heap on every restart |
| Identity-CAS at dispatch time | Pre-dispatch consistency checks at schedule time, racing object updates |
| Recovery branch over `RecoveryView` | Implicit "is this rule new" tracking with bookkeeping flags |
| Walker throttle independent of invocation | Walker hammering filer when test driver invokes every 2s |
| Single subscription per pass | 16x filer load with 16 per-shard subscriptions |
The result is a worker the operator can reason about by reading 2 metrics and a heartbeat line, with a state machine small enough to fit in one design doc.
+112
@@ -0,0 +1,112 @@
# S3 Lifecycle — Monitoring
This page lists the Prometheus signals the worker exposes and how to read the heartbeat log line. For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting).
## Prometheus metrics
All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem.
### Per-shard gauges
| Metric | Labels | What |
|---|---|---|
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard |
| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire |
Derived queries:
```promql
# Per-shard replay lag in seconds
(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9
# Per-shard walker freshness in seconds
(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9
# Worst-shard lag across the cluster
max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9
```
Zero values mean "not started yet" — distinct from "0s caught up". The heartbeat line uses `cold` as the marker for that state.
### Counters
| Metric | Labels | What |
|---|---|---|
| `s3_lifecycle_dispatch_total` | `bucket`, `kind`, `outcome` | Per-bucket dispatch counter, partitioned by action kind and server outcome |
| `s3_lifecycle_daily_run_events_scanned_total` | `shard` | Meta-log events `drainShardEvents` processed |
| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter |
| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path |
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run.
### Histograms
| Metric | What |
|---|---|
| `s3_lifecycle_daily_run_shard_duration_seconds{shard}` | Wall-clock per shard pass. p95 climbing toward `max_runtime_minutes` means the shard is brushing its budget. |
| `s3_lifecycle_dispatch_limiter_wait_seconds` | Time spent waiting on the cluster rate limiter before issuing `LifecycleDelete`. Near-zero = cap not binding; long-tail at `1/rate` = cap is the active throttle. |
## Heartbeat log line
Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity):
```
daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m
```
Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stable across versions:
| Token | Meaning |
|---|---|
| `status=ok` or `status=error` | Whether any shard returned an error |
| `shards=N` | Number of shards processed this pass |
| `errors=N` | Per-shard error count |
| `duration=Ns` | Wall-clock for the whole pass |
| `cursor_lag_max=...` | Worst per-shard replay lag, or `cold` if no shard has a persisted cursor yet |
| `walked_max_age=...` | Worst per-shard walker age, or `cold` if no shard has walked yet |
A healthy production heartbeat looks like:
```
daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m
```
Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag is 45 seconds behind real-time; the oldest walker fire on any shard is 58 minutes ago (so `walker_interval_minutes=60` is roughly honored).
## Anti-patterns to alert on
| Pattern | Meaning | What to do |
|---|---|---|
| `cursor_lag_max` grows unbounded | Stuck cursor; head-of-line blocking on some shard | See [Troubleshooting → Stuck cursor](S3-Lifecycle-Troubleshooting#stuck-cursor) |
| `walked_max_age` exceeds `walker_interval_minutes × 2` | Walker isn't firing as configured | Check `errors=N` in heartbeat and `s3_lifecycle_dispatch_total{outcome="RPC_ERROR"}` |
| `errors=16` (all shards) on every pass | Filer is unreachable or returning errors | Check filer health |
| `s3_lifecycle_dispatch_total{outcome="RETRY_LATER"}` rising fast | Server rate-limited or filer overloaded | Lower `cluster_deletes_per_second` or add capacity |
| `s3_lifecycle_dispatch_total{outcome="BLOCKED"}` non-zero | Programmatic event content error | Check worker logs for `FATAL_EVENT_ERROR` |
| `duration=Ns` ramping up across passes | Walker is firing too often | Set `walker_interval_minutes` |
## Suggested alerts
```yaml
- alert: S3LifecycleCursorLagHigh
expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600
for: 30m
annotations:
summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}"
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor
- alert: S3LifecycleWalkerStuck
expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400
for: 1h
annotations:
summary: "S3 lifecycle walker hasn't run in > 24h"
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#walker-stuck
- alert: S3LifecycleDispatchFailures
expr: |
rate(s3_lifecycle_dispatch_total{outcome=~"RETRY_LATER|BLOCKED|RPC_ERROR"}[5m]) > 0.1
for: 15m
annotations:
summary: "S3 lifecycle delete failure rate > 0.1/s"
```
Adjust thresholds to your cluster's normal levels — these are starting points.
+135
@@ -0,0 +1,135 @@
# S3 Lifecycle — Operator Guide
This page covers the admin and worker config knobs for the S3 lifecycle worker, plus when to change each one.
For monitoring guidance, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring). For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting).
## Configuration
All keys are set through the admin UI's plugin config for `s3_lifecycle`.
### Admin config
| Key | Type | Default | When to change |
|---|---|---|---|
| `cluster_deletes_per_second` | int64 | `0` (unlimited) | Set a positive value when lifecycle deletes are causing filer contention. Allocated evenly across active workers at job dispatch. |
| `cluster_deletes_burst` | int64 | `0` (= 2× rate) | Adjust if delete bursts overload the filer faster than the per-second rate allows. |
| `meta_log_retention_days` | int64 | `0` (treated as unbounded) | Stock SeaweedFS doesn't GC the meta-log, so the default is fine. Set positive if your deployment manually trims `/topics/.system/log` — then rules with TTL > retention will route through the walker. |
| `walker_interval_minutes` | int64 | `0` (fire every pass) | **Important.** See "Walker interval" below — most production deployments should set this to a positive value. |
### Worker config
| Key | Default | What |
|---|---|---|
| `max_runtime_minutes` | 60 | Wall-clock cap per `dailyrun.Run` invocation. The pass returns early if it hits this. |
### Detection schedule
The admin scheduler's `DetectionIntervalMinutes` for `s3_lifecycle` is `1440` by default — once per day. Each detection produces one execution. Change in the admin UI's runtime defaults for the job type.
## Walker interval
The walker is the part of the worker that lists bucket contents and evaluates them against rules. It fires:
- **Always**, on cold start (no persisted cursor) and on rule changes — these are bounded events.
- **Periodically**, in steady state, gated by `walker_interval_minutes`.
The throttle exists because walker cost is bucket size, not event rate. If the worker is scheduled at a tighter cadence than the desired walk frequency (CI, sub-hourly admin schedules, manual runs), the steady-state walker would crush the filer with a full subtree scan per invocation.
Recommended values by cluster size:
| Cluster | Recommended `walker_interval_minutes` |
|---|---|
| Small (≤1M objects, single bucket) | 60 |
| Medium (≤100M objects, mixed buckets) | 360 (6h) |
| Large (≥100M objects) | 1440 (24h) |
| Testing / CI | 0 (fire every pass) |
Setting `0` keeps the prior behavior — fire on every invocation. That's appropriate when the worker is scheduled at exactly the desired walk frequency (e.g., once per day) and the in-repo integration tests rely on it.
A negative value is rejected at worker start (loud error rather than silent fall-through to "walk every pass").
## Per-bucket lifecycle XML
Set via the standard S3 API:
```bash
aws --endpoint $S3_ENDPOINT s3api put-bucket-lifecycle-configuration \
--bucket my-bucket \
--lifecycle-configuration file://lifecycle.json
```
```jsonc
// lifecycle.json
{
"Rules": [
{
"ID": "expire-logs-after-30d",
"Status": "Enabled",
"Filter": { "Prefix": "logs/" },
"Expiration": { "Days": 30 }
},
{
"ID": "abort-stuck-mpu",
"Status": "Enabled",
"Filter": {},
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
},
{
"ID": "keep-3-versions",
"Status": "Enabled",
"Filter": { "Prefix": "versioned/" },
"NoncurrentVersionExpiration": { "NewerNoncurrentVersions": 3 }
}
]
}
```
The worker picks up rule changes on the next pass. Replay-eligible rule edits trigger a one-time recovery walk to catch already-due objects under the new rule.
## Verifying a rule is working
After applying a rule:
1. Wait one detection interval (default 24h) plus one walker interval.
2. Read `s3_lifecycle_dispatch_total{bucket="my-bucket"}` — the counter should advance.
3. Verify a target object is gone: `aws s3 head-object --bucket my-bucket --key <expected-expired>` should return 404.
For testing without waiting, the `weed shell` command supports manual invocation:
```
weed shell -master <addr>
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
```
This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for examples.
## Rate limit allocation
The cluster delete cap is allocated per-worker at job dispatch:
```
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
```
Brief over/undershoot during worker join/leave is acceptable for a bulk workload. The allocation is recomputed each detection cycle, so adding workers smooths out within one day.
`cluster_deletes_burst` is divided the same way. `0` is treated as `2 × rate` (a token bucket reasonable default).
## Common change patterns
### Tightening a TTL (e.g., 60d → 30d)
The new rule hash differs from the persisted hash — next pass triggers a recovery walk over `RecoveryView`. Already-due objects under the new shorter TTL get caught. The cursor rewinds to `runNow - new_maxTTL`. No operator action needed beyond updating the XML.
### Adding a walker-only rule (`Expiration.Date`)
`RuleSetHash` and `PromotedHash` are unchanged (walker-only rules aren't in the replay hash). The replay cursor stays put. The walker reads the updated rule set on its next steady-state fire (or immediately, if invoked manually).
### Operator forgot to set `walker_interval_minutes`
Symptom: heartbeat log line shows `duration` ramping up across passes (each pass does a full subtree walk). Mitigation: set the throttle to a value matching your scheduling cadence.
### Filer is GC'ing meta-log (custom deployment)
Set `meta_log_retention_days` to the GC retention. Rules with TTL > retention will route through the walker (`PromotedHash` will be non-empty). The first run after this change triggers a recovery walk.
+151
@@ -0,0 +1,151 @@
# S3 Lifecycle — Troubleshooting
Incident-response playbook for the S3 lifecycle worker. For monitoring background, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring).
## Stuck cursor
**Symptom:** `s3_lifecycle_cursor_min_ts_ns{shard=N}` is not advancing. Heartbeat shows `cursor_lag_max` growing unbounded.
**Cause:** The cursor advance is gated on every match from the event dispatching successfully (`DONE`, `NOOP_RESOLVED`, or `SKIPPED_OBJECT_LOCK`). Any unresolved outcome (`RETRY_LATER`, `BLOCKED`, transport error after in-run retries) halts the run for that shard and persists the cursor at the last fully-processed event. Head-of-line blocking is intentional — it surfaces a real problem rather than silently retrying forever.
**Diagnostics:**
```promql
# Which shard is stuck?
time() * 1e9 - s3_lifecycle_cursor_min_ts_ns
# What outcomes are being returned?
sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m]))
```
Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`:
```
daily_run: RETRY_LATER on <bucket>/<key> EXPIRATION_DAYS
daily_run: BLOCKED on <bucket>/<key> NONCURRENT_DAYS
daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
```
**Mitigations:**
| Outcome | Root cause | Action |
|---|---|---|
| `RETRY_LATER` (high rate) | Filer or rate limiter is throttling | Lower `cluster_deletes_per_second` to give the filer headroom, or scale filer capacity |
| `BLOCKED FATAL_EVENT_ERROR` | A malformed event the server refuses to dispatch | Check log for the specific reason. May need a code fix; file an issue with the log line |
| `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. |
| `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability |
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker.
## Walker stuck (no progress on walker-only rules)
**Symptom:** `s3_lifecycle_daily_run_last_walked_ns{shard=N}` is not advancing. Rules like `Expiration.Date`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` aren't firing on objects that should be due.
**Causes:**
1. `walker_interval_minutes` is too long for your invocation cadence. Worker runs once per day but interval is set to 48h.
2. Walker is hitting an error mid-walk (filer listing failure). Look for `recovery walk:` or `steady walk:` errors in the heartbeat's `errors=N` count.
3. The bucket has only walker-bound rules and the empty-replay branch's throttle hasn't elapsed.
**Diagnostics:**
```promql
# Walker age per shard
(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9
```
Check worker config: `walker_interval_minutes` should be ≤ the daily worker schedule interval.
**Mitigation:** lower `walker_interval_minutes`. Setting `0` temporarily forces every pass to walk.
## Test PUT a file with a 1-day rule, didn't expire
The S3 API rejects `Expiration.Days < 1`, so the smallest "expire after N days" you can configure is 1 day. The worker runs once per day by default. Object PUT + 1-day rule + waiting one day is the minimum scenario.
For testing, the in-repo integration suite uses a trick: backdate the entry's `Mtime` via `filer_pb.UpdateEntry` to 30+ days ago. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for the pattern.
For ad-hoc verification, invoke the worker manually:
```
weed shell -master <addr>
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
```
This runs the same code path as the scheduled worker, but driven from your shell rather than the admin scheduler.
## All shards report `errors=16` every pass
**Symptom:** Heartbeat consistently shows `status=error shards=16 errors=16 duration=Ns`.
**Common causes:**
1. **Filer unreachable.** Subscription fails on every shard. Check filer health and gRPC connectivity.
2. **passCtx timeout from `-refresh` loop.** If `-refresh` is less than the pass cap, the timeout fires before the drain completes. This is now treated as "clean end-of-pass" — it should not show as errors=N. If you see this on a build before #9481, upgrade.
3. **Bucket walker is timing out.** Big bucket, walker hits ctx deadline. Increase `max_runtime_minutes`.
## Some objects expired, others didn't (same rule)
**Symptom:** Two objects matching the same rule with the same age — one is deleted, the other isn't.
**Common causes:**
1. **The non-deleted object's mtime is wrong.** Check the entry's mtime — it might be more recent than you expect (e.g., a recent metadata update bumped it).
2. **The objects are on different shards** and one shard has a stuck cursor while the other doesn't. Check `s3_lifecycle_cursor_min_ts_ns` per shard.
3. **`Filter` doesn't match what you think.** A prefix-only filter requires the object key to start with that prefix; a tag filter requires the matching tag. Verify with `aws s3api head-object` (returns tags via `--query`).
4. **Object lock or retention.** The dispatcher returns `SKIPPED_OBJECT_LOCK` for protected objects. Check `s3_lifecycle_dispatch_total{outcome="SKIPPED_OBJECT_LOCK"}`.
## How to read the cursor files
Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`:
```
weed shell -master <addr>
> fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json
```
Schema:
```json
{
"version": 1,
"shard_id": 0,
"ts_ns": 1715600000000000000,
"rule_set_hash": "<base64 32 bytes>",
"promoted_hash": "<base64 32 bytes>",
"last_walked_ns": 1715620000000000000
}
```
- `ts_ns == 0` means "no replay progress" — either cold start or a bucket whose rules are all walker-only.
- `last_walked_ns == 0` (or absent) means "never walked steady-state". Next pass will walk.
Manually editing the cursor is supported as an escape hatch but obviously breaks the invariant that "everything before persisted.TsNs has been processed under the same rules." Use sparingly.
## Resetting a shard
If a shard's cursor is corrupted or wedged in an unrecoverable state:
```
weed shell -master <addr>
> fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json
```
Next pass treats the shard as cold start: recovery walker fires over `RecoveryView(snap)`, then the cursor seeds at `runNow - maxTTL`. This re-replays a `maxTTL`-wide window of meta-log events. Identity-CAS on the server side makes redundant deletes no-ops, so re-replay is safe.
## Suspending the worker
The admin UI's plugin scheduler allows pausing the `s3_lifecycle` job type. The worker won't be invoked while paused; cursors are preserved, and on resume the next pass picks up at the persisted state.
For a single-bucket "stop deleting from this bucket" without pausing the worker:
```bash
aws --endpoint $S3_ENDPOINT s3api delete-bucket-lifecycle --bucket my-bucket
```
The worker reads the empty rule set on the next pass; `rsh == [32]byte{}` causes the empty-replay branch to run only the walker (which has nothing to walk for an empty rule set) and exit. No deletes.
## Reverting
If something goes very wrong, the streaming worker (the previous design) is no longer in the codebase. Revert is via downgrading the binary. The cursor format has a `version` field — versions `>1` would fail-loud on load by an older binary that only knows version `1`. Currently version `1` is the only version.
For escape-hatch operations (delete all cursors, suspend worker globally), prefer pausing via the admin UI over force-killing the worker process; the worker exits cleanly between passes.
+65 -93
@@ -1,26 +1,26 @@
# S3 Lifecycle Configuration
# S3 Lifecycle
SeaweedFS supports S3 bucket lifecycle configuration for automated object expiration, non-current version cleanup, delete marker removal, and incomplete multipart upload abortion.
SeaweedFS implements the S3 `PutBucketLifecycleConfiguration` API. Configured rules are evaluated and enforced by a worker that runs as a scheduled job and exits when each pass completes.
## Supported Features
This page is the operator-facing entry point. Developers and architecture readers should see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md).
## Supported features
| Feature | Status | Notes |
|---|---|---|
| `Expiration.Days` | Supported | Fast path via TTL + RocksDB compaction filter |
| `Expiration.Date` | Supported | Evaluated by lifecycle worker at scan time |
| `ExpiredObjectDeleteMarker` | Supported | Removes delete markers that are the sole remaining version |
| `NoncurrentVersionExpiration.NoncurrentDays` | Supported | Requires bucket versioning enabled |
| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Supported | Keep N newest non-current versions |
| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Supported | Per-rule prefix scoping |
| `Filter.Prefix` | Supported | |
| `Filter.Tag` | Supported | Evaluated at scan time |
| `Filter.And` (Prefix + Tags + Size) | Supported | Evaluated at scan time |
| `Filter.ObjectSizeGreaterThan` | Supported | Evaluated at scan time |
| `Filter.ObjectSizeLessThan` | Supported | Evaluated at scan time |
| `Transition` | Not supported | Requires storage class tiers |
| `NoncurrentVersionTransition` | Not supported | Requires storage class tiers |
| `Expiration.Days` | Yes | Latest-version PUT clock |
| `Expiration.Date` | Yes | Walker path; fires once date is reached |
| `Expiration.ExpiredObjectDeleteMarker` | Yes | Walker path; sibling-aware |
| `NoncurrentVersionExpiration.NoncurrentDays` | Yes | Clock starts at the demoting PUT, not the entry's own mtime |
| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Yes | Walker path; version-list aware |
| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Yes | |
| `Filter.Prefix` | Yes | |
| `Filter.Tag` | Yes | |
| `Filter.ObjectSizeGreaterThan` / `ObjectSizeLessThan` | Yes | |
| `Filter.And` (composite) | Yes | |
| `Transition` / `NoncurrentVersionTransition` | No | SeaweedFS doesn't model storage class tiers |
## API Endpoints
## API endpoints
```
PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration
@@ -28,6 +28,37 @@ GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration
DELETE /{bucket}?lifecycle # DeleteBucketLifecycle
```
## Example: AWS CLI
```bash
# Set lifecycle configuration
aws s3api put-bucket-lifecycle-configuration \
--endpoint-url http://localhost:8333 \
--bucket my-bucket \
--lifecycle-configuration '{
"Rules": [
{
"ID": "expire-old",
"Status": "Enabled",
"Filter": { "Prefix": "" },
"Expiration": { "Days": 90 },
"NoncurrentVersionExpiration": { "NoncurrentDays": 30 },
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
}
]
}'
# Get lifecycle configuration
aws s3api get-bucket-lifecycle-configuration \
--endpoint-url http://localhost:8333 \
--bucket my-bucket
# Delete lifecycle configuration
aws s3api delete-bucket-lifecycle \
--endpoint-url http://localhost:8333 \
--bucket my-bucket
```
## Example: Terraform
```hcl
@@ -47,7 +78,7 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" {
}
noncurrent_version_expiration {
noncurrent_days = 7
noncurrent_days = 7
newer_noncurrent_versions = 2
}
@@ -58,89 +89,30 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" {
}
```
## Example: AWS CLI
## How it works
```bash
# Set lifecycle configuration
aws s3api put-bucket-lifecycle-configuration \
--endpoint-url http://localhost:8333 \
--bucket my-bucket \
--lifecycle-configuration '{
"Rules": [
{
"ID": "expire-old",
"Status": "Enabled",
"Filter": { "Prefix": "" },
"Expiration": { "Days": 90 },
"NoncurrentVersionExpiration": {
"NoncurrentDays": 30
},
"AbortIncompleteMultipartUpload": {
"DaysAfterInitiation": 7
}
}
]
}'
The lifecycle worker is a scheduled job (default daily). Each invocation:
# Get lifecycle configuration
aws s3api get-bucket-lifecycle-configuration \
--endpoint-url http://localhost:8333 \
--bucket my-bucket
1. Reads bucket lifecycle XML from each bucket's metadata.
2. Compiles rules into a per-shard partition (replay-eligible vs. walker-bound).
3. Subscribes to the filer meta-log — one stream covering all 16 shards in this worker process.
4. For replay-eligible actions (`ExpirationDays`, `NoncurrentDays`, `AbortMPU`), checks each event's DueTime and dispatches `LifecycleDelete` if elapsed.
5. For walker-bound rules (`ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent`, or anything promoted to scan-only), iterates the bucket and evaluates each entry against current state.
6. Persists per-shard cursors so the next pass resumes where this one left off.
# Delete lifecycle configuration
aws s3api delete-bucket-lifecycle \
--endpoint-url http://localhost:8333 \
--bucket my-bucket
```
The worker exits when the pass is done. The admin scheduler invokes it on a daily cadence by default; operators can change that via the standard plugin scheduler config.
## How It Works
### Two-Tier Expiration Architecture
**Tier 1 — TTL fast path:** Simple `Expiration.Days` rules with prefix-only filters are translated into TTL entries in `filer.conf`. The RocksDB compaction filter automatically removes expired entries during normal compaction at zero additional cost. This is the most efficient path for the common case.
**Tier 2 — Scan-time evaluation:** Rules with tag filters, size filters, date-based expiration, non-current version expiration, and delete marker cleanup are evaluated by the lifecycle plugin worker. The worker periodically scans buckets and evaluates each object against the stored lifecycle XML configuration.
### Which rules use which tier?
| Rule Type | Tier | Why |
|---|---|---|
| `Expiration.Days` (prefix only) | TTL fast path | Can be expressed as per-entry TTL |
| `Expiration.Days` (with tags/size) | Worker scan | TTL can't express tag/size constraints |
| `Expiration.Date` | Worker scan | Absolute date, not relative TTL |
| `NoncurrentVersionExpiration` | Worker scan | Requires version enumeration |
| `ExpiredObjectDeleteMarker` | Worker scan | Requires version counting |
| `AbortIncompleteMultipartUpload` | Worker scan | Scans `.uploads` directory |
### Lifecycle Worker
The lifecycle plugin worker runs as part of the SeaweedFS plugin system. It:
1. **Detects** buckets with lifecycle rules (from stored lifecycle XML or filer.conf TTLs)
2. **Scans** bucket contents and evaluates lifecycle rules against each object
3. **Executes** actions: delete expired objects, remove old non-current versions, clean up delete markers, abort stale multipart uploads
Worker configuration options:
| Setting | Default | Description |
|---|---|---|
| `batch_size` | 1000 | Entries per filer listing page |
| `max_deletes_per_bucket` | 10000 | Max expired objects to delete per run |
| `dry_run` | false | Detect but don't delete |
| `delete_marker_cleanup` | true | Remove expired delete markers |
| `abort_mpu_days` | 7 | Fallback for buckets without lifecycle XML MPU rules |
| `bucket_filter` | (all) | Wildcard pattern to scope lifecycle to specific buckets |
## Versioning Integration
## Versioning integration
Lifecycle rules interact with [S3 Object Versioning](S3-Object-Versioning):
- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded. `NewerNoncurrentVersions` retains the N newest non-current versions.
- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded (the demoting PUT's TsNs, not the version's own mtime). `NewerNoncurrentVersions` retains the N newest non-current versions.
- **`ExpiredObjectDeleteMarker`** removes delete markers that are the sole remaining version of an object (no non-current versions behind them).
- **`Expiration.Days`** on a versioned bucket creates a delete marker when the current version expires; it does not permanently delete the object.
## Limitations
## Quick references
- **Transition rules** (`Transition`, `NoncurrentVersionTransition`) are not supported. SeaweedFS does not have S3-equivalent storage class tiers.
- **Detection interval**: The lifecycle worker scans on a configurable interval (default 5 minutes). Objects may persist slightly beyond their configured expiration until the next scan.
- **TTL fast path** applies only to `Expiration.Days` rules with prefix-only filters. Rules with tag or size constraints are always evaluated at scan time.
- **[Operator Guide](S3-Lifecycle-Operator-Guide)** — config knobs, defaults, when to change each
- **[Monitoring](S3-Lifecycle-Monitoring)** — Prometheus metrics, heartbeat log line, what a healthy run looks like
- **[Troubleshooting](S3-Lifecycle-Troubleshooting)** — stuck cursor, missing deletes, head-of-line blocking
- **[Architecture](S3-Lifecycle-Architecture)** — high-level overview of the worker, engine, and dispatch path