mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-11 16:57:45 +02:00
S3 Lifecycle: refresh + add operator, monitoring, troubleshooting, architecture pages
Synced from seaweedfs/seaweedfs:docs/wiki/s3-lifecycle/ (PR #9491). The existing S3-Lifecycle page described the pre-redesign two-tier TTL+scan architecture. Replace the architecture and config-knob sections with the current as-built daily-replay model. Keep the useful AWS CLI / Terraform examples and the rule-support table. Add four new pages: - S3-Lifecycle-Operator-Guide: config knobs, walker interval recommendations by cluster size - S3-Lifecycle-Monitoring: Prometheus metric reference, heartbeat tokens, suggested PromQL alerts - S3-Lifecycle-Troubleshooting: stuck cursor playbook, failure outcomes, cursor schema for manual inspection - S3-Lifecycle-Architecture: high-level overview between Home and the in-tree DESIGN.md
1 parent
f8ec586b82
commit
29aa6c5c5f
5 files changed
+584
-93
No files matched your search
@@ -0,0 +1,121 @@
|
||||
# S3 Lifecycle — Architecture
|
||||
|
||||
High-level overview of the lifecycle worker. For implementation detail, see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md).
|
||||
|
||||
## At a glance
|
||||
|
||||
The lifecycle worker runs as a scheduled job. Each invocation:
|
||||
|
||||
```
|
||||
┌──────────────────────────────────────────┐
|
||||
│ dailyrun.Run (one filer subscription) │
|
||||
│ │
|
||||
meta-log ──→ │ reader ──→ fan-out ──→ per-shard │
|
||||
│ channels │
|
||||
│ │
|
||||
│ ┌──────────────────────────────────┐ │
|
||||
│ │ 16 shard goroutines │ │
|
||||
│ │ ┌──────────────────────────┐ │ │
|
||||
│ │ │ walker(view, shardID)? │ │ │
|
||||
│ │ │ drainShardEvents │ │ │
|
||||
│ │ │ saveCursorAndPublish │ │ │
|
||||
│ │ └──────────────────────────┘ │ │
|
||||
│ └──────────────────────────────────┘ │
|
||||
│ │
|
||||
│ summary heartbeat + exit │
|
||||
└──────────────────────────────────────────┘
|
||||
```
|
||||
|
||||
One filer `SubscribeMetadata` stream covers every shard in this worker's set. A fan-out goroutine routes events to per-shard channels by `ev.ShardID = sha256(bucket || "/" || key) >> 252`. Each shard's goroutine independently runs the walker (when due), drains events, and persists its cursor.
|
||||
|
||||
Once every shard's goroutine returns, the worker tears down the subscription, emits a summary heartbeat, and exits.
|
||||
|
||||
## Per-shard state
|
||||
|
||||
Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`:
|
||||
|
||||
```
|
||||
TsNs — last meta-log event whose matches all dispatched successfully
|
||||
RuleSetHash — ReplayContentHash of the rule set when this cursor was written
|
||||
PromotedHash — PromotedHash(retentionWindow) at write time
|
||||
LastWalkedNs — wall-clock of the last successful walker fire
|
||||
```
|
||||
|
||||
The two hashes together detect every situation that invalidates the cursor: a replay-rule edit (`RuleSetHash` changes) or a partition flip (`PromotedHash` changes). On mismatch, the next pass triggers a recovery walk over `RecoveryView(snap)` to catch already-due objects across the full rule set, then rewinds the cursor.
|
||||
|
||||
## Replay vs walker
|
||||
|
||||
The lifecycle rule space splits two ways:
|
||||
|
||||
| Path | Action kinds | Why this path |
|
||||
|---|---|---|
|
||||
| Replay (meta-log) | `ExpirationDays`, `NoncurrentDays`, `AbortMPU` | DueTime is monotonic in event TsNs. The `done` early-stop works. |
|
||||
| Walker (bucket list) | `ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` | DueTime depends on current sibling/version state, not event age. |
|
||||
|
||||
The engine's `RulesForShard(shardID, retentionWindow)` returns two snapshot views (`replay`, `walk`); each is a clone of the base snapshot with the action map masked to the right partition. `router.Route` consumes the `replay` view per event; the walker consumes the `walk` view per bucket.
|
||||
|
||||
A rule promoted to scan-only because its TTL exceeds meta-log retention moves from `replay` to `walk` — visible via `PromotedHash`.
|
||||
|
||||
## Cadence layers
|
||||
|
||||
Three independent cadences shape worker behavior:
|
||||
|
||||
| Cadence | Set by | Default |
|
||||
|---|---|---|
|
||||
| Worker invocation | Admin scheduler `DetectionIntervalMinutes` | 1440 (daily) |
|
||||
| Walker fire | `walker_interval_minutes` admin config | 0 (every invocation) |
|
||||
| Cursor save | After each `runShard` | n/a |
|
||||
|
||||
The walker throttle decouples walker firing from invocation rate. CI invokes the worker every 2s; production invokes once per day. Both can use the same code with appropriate `walker_interval_minutes`.
|
||||
|
||||
## Failure model
|
||||
|
||||
- **Worker crash mid-run.** Cursor only advances past events whose matches all succeeded. On restart, the next pass resumes at the same cursor. Identity-CAS makes redundant deletes no-ops.
|
||||
- **Transient delete failure.** Pass halts at the failing event, cursor persists. Next pass retries from there. Head-of-line blocking is intentional — surfaces real problems instead of silently retrying forever.
|
||||
- **Rule edits.** Replay-rule edits trigger one-time recovery walk over `RecoveryView`. Walker-only rule edits don't change either hash; walker reads the new rules on its next steady-state fire.
|
||||
- **Object overwritten between event and delete.** `LifecycleDelete` RPC's identity-CAS returns `NOOP_RESOLVED`; cursor advances normally.
|
||||
|
||||
## Rate limiting
|
||||
|
||||
Cluster-wide cap allocated per worker at job dispatch:
|
||||
|
||||
```
|
||||
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
|
||||
```
|
||||
|
||||
Each worker shares one `rate.Limiter` across all shard goroutines. `dispatchWithRetry` calls `limiter.Wait(ctx)` before each `LifecycleDelete` RPC.
|
||||
|
||||
## Components
|
||||
|
||||
| Path | Role |
|
||||
|---|---|
|
||||
| `engine/` | Rule compilation, partition views (`RulesForShard`, `RecoveryView`) |
|
||||
| `evaluate.go` | Per-event rule evaluation (`EvaluateAction`) |
|
||||
| `due_at.go` | Per-(rule, kind, info) due-time computation |
|
||||
| `router/router.go` | Per-event match emission (calls engine.Action and EvaluateAction) |
|
||||
| `reader/reader.go` | Meta-log subscribe with `ShardPredicate` |
|
||||
| `bootstrap/walker.go` | Bucket-walker with `RunForShard` filter |
|
||||
| `dailyrun/run.go` | Main orchestrator: subscription, fan-out, per-shard runShard |
|
||||
| `dailyrun/cursor.go` | Cursor type + filer JSON serializer |
|
||||
| `dailyrun/walker_dispatcher.go` | Walker-to-`LifecycleDelete` adapter |
|
||||
|
||||
## What it's not
|
||||
|
||||
- **Not a streaming dispatcher.** The earlier model kept a long-running goroutine per shard with an in-memory match heap. That code is gone. Worker is now "start, do today's work, stop."
|
||||
- **Not event-time accurate.** Latency from PUT to delete is bounded by the worker invocation cadence plus the walker interval — typically up to 24h, not seconds.
|
||||
- **Not a general-purpose scheduler.** The two action paths (replay, walker) are specific to lifecycle semantics. Don't add new event sources or actions without thinking through which path they belong on.
|
||||
|
||||
## Why this shape
|
||||
|
||||
Each design choice points back to a specific failure mode of the prior streaming worker:
|
||||
|
||||
| Choice | Replaces |
|
||||
|---|---|
|
||||
| Per-pass run + exit | Long-running goroutines with ticker drift, leak risk, restart pain |
|
||||
| Cursor file per shard | Per-key freeze state, retry counters, in-memory heap on every restart |
|
||||
| Identity-CAS at dispatch time | Pre-dispatch consistency checks at schedule time, racing object updates |
|
||||
| Recovery branch over `RecoveryView` | Implicit "is this rule new" tracking with bookkeeping flags |
|
||||
| Walker throttle independent of invocation | Walker hammering filer when test driver invokes every 2s |
|
||||
| Single subscription per pass | 16x filer load with 16 per-shard subscriptions |
|
||||
|
||||
The result is a worker the operator can reason about by reading 2 metrics and a heartbeat line, with a state machine small enough to fit in one design doc.
|
||||
@@ -0,0 +1,112 @@
|
||||
# S3 Lifecycle — Monitoring
|
||||
|
||||
This page lists the Prometheus signals the worker exposes and how to read the heartbeat log line. For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting).
|
||||
|
||||
## Prometheus metrics
|
||||
|
||||
All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem.
|
||||
|
||||
### Per-shard gauges
|
||||
|
||||
| Metric | Labels | What |
|
||||
|---|---|---|
|
||||
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard |
|
||||
| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire |
|
||||
|
||||
Derived queries:
|
||||
|
||||
```promql
|
||||
# Per-shard replay lag in seconds
|
||||
(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9
|
||||
|
||||
# Per-shard walker freshness in seconds
|
||||
(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9
|
||||
|
||||
# Worst-shard lag across the cluster
|
||||
max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9
|
||||
```
|
||||
|
||||
Zero values mean "not started yet" — distinct from "0s caught up". The heartbeat line uses `cold` as the marker for that state.
|
||||
|
||||
### Counters
|
||||
|
||||
| Metric | Labels | What |
|
||||
|---|---|---|
|
||||
| `s3_lifecycle_dispatch_total` | `bucket`, `kind`, `outcome` | Per-bucket dispatch counter, partitioned by action kind and server outcome |
|
||||
| `s3_lifecycle_daily_run_events_scanned_total` | `shard` | Meta-log events `drainShardEvents` processed |
|
||||
| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter |
|
||||
| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path |
|
||||
|
||||
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run.
|
||||
|
||||
### Histograms
|
||||
|
||||
| Metric | What |
|
||||
|---|---|
|
||||
| `s3_lifecycle_daily_run_shard_duration_seconds{shard}` | Wall-clock per shard pass. p95 climbing toward `max_runtime_minutes` means the shard is brushing its budget. |
|
||||
| `s3_lifecycle_dispatch_limiter_wait_seconds` | Time spent waiting on the cluster rate limiter before issuing `LifecycleDelete`. Near-zero = cap not binding; long-tail at `1/rate` = cap is the active throttle. |
|
||||
|
||||
## Heartbeat log line
|
||||
|
||||
Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity):
|
||||
|
||||
```
|
||||
daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m
|
||||
```
|
||||
|
||||
Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stable across versions:
|
||||
|
||||
| Token | Meaning |
|
||||
|---|---|
|
||||
| `status=ok` or `status=error` | Whether any shard returned an error |
|
||||
| `shards=N` | Number of shards processed this pass |
|
||||
| `errors=N` | Per-shard error count |
|
||||
| `duration=Ns` | Wall-clock for the whole pass |
|
||||
| `cursor_lag_max=...` | Worst per-shard replay lag, or `cold` if no shard has a persisted cursor yet |
|
||||
| `walked_max_age=...` | Worst per-shard walker age, or `cold` if no shard has walked yet |
|
||||
|
||||
A healthy production heartbeat looks like:
|
||||
|
||||
```
|
||||
daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m
|
||||
```
|
||||
|
||||
Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag is 45 seconds behind real-time; the oldest walker fire on any shard is 58 minutes ago (so `walker_interval_minutes=60` is roughly honored).
|
||||
|
||||
## Anti-patterns to alert on
|
||||
|
||||
| Pattern | Meaning | What to do |
|
||||
|---|---|---|
|
||||
| `cursor_lag_max` grows unbounded | Stuck cursor; head-of-line blocking on some shard | See [Troubleshooting → Stuck cursor](S3-Lifecycle-Troubleshooting#stuck-cursor) |
|
||||
| `walked_max_age` exceeds `walker_interval_minutes × 2` | Walker isn't firing as configured | Check `errors=N` in heartbeat and `s3_lifecycle_dispatch_total{outcome="RPC_ERROR"}` |
|
||||
| `errors=16` (all shards) on every pass | Filer is unreachable or returning errors | Check filer health |
|
||||
| `s3_lifecycle_dispatch_total{outcome="RETRY_LATER"}` rising fast | Server rate-limited or filer overloaded | Lower `cluster_deletes_per_second` or add capacity |
|
||||
| `s3_lifecycle_dispatch_total{outcome="BLOCKED"}` non-zero | Programmatic event content error | Check worker logs for `FATAL_EVENT_ERROR` |
|
||||
| `duration=Ns` ramping up across passes | Walker is firing too often | Set `walker_interval_minutes` |
|
||||
|
||||
## Suggested alerts
|
||||
|
||||
```yaml
|
||||
- alert: S3LifecycleCursorLagHigh
|
||||
expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600
|
||||
for: 30m
|
||||
annotations:
|
||||
summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}"
|
||||
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor
|
||||
|
||||
- alert: S3LifecycleWalkerStuck
|
||||
expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400
|
||||
for: 1h
|
||||
annotations:
|
||||
summary: "S3 lifecycle walker hasn't run in > 24h"
|
||||
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#walker-stuck
|
||||
|
||||
- alert: S3LifecycleDispatchFailures
|
||||
expr: |
|
||||
rate(s3_lifecycle_dispatch_total{outcome=~"RETRY_LATER|BLOCKED|RPC_ERROR"}[5m]) > 0.1
|
||||
for: 15m
|
||||
annotations:
|
||||
summary: "S3 lifecycle delete failure rate > 0.1/s"
|
||||
```
|
||||
|
||||
Adjust thresholds to your cluster's normal levels — these are starting points.
|
||||
@@ -0,0 +1,135 @@
|
||||
# S3 Lifecycle — Operator Guide
|
||||
|
||||
This page covers the admin and worker config knobs for the S3 lifecycle worker, plus when to change each one.
|
||||
|
||||
For monitoring guidance, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring). For incident response, see [S3-Lifecycle-Troubleshooting](S3-Lifecycle-Troubleshooting).
|
||||
|
||||
## Configuration
|
||||
|
||||
All keys are set through the admin UI's plugin config for `s3_lifecycle`.
|
||||
|
||||
### Admin config
|
||||
|
||||
| Key | Type | Default | When to change |
|
||||
|---|---|---|---|
|
||||
| `cluster_deletes_per_second` | int64 | `0` (unlimited) | Set a positive value when lifecycle deletes are causing filer contention. Allocated evenly across active workers at job dispatch. |
|
||||
| `cluster_deletes_burst` | int64 | `0` (= 2× rate) | Adjust if delete bursts overload the filer faster than the per-second rate allows. |
|
||||
| `meta_log_retention_days` | int64 | `0` (treated as unbounded) | Stock SeaweedFS doesn't GC the meta-log, so the default is fine. Set positive if your deployment manually trims `/topics/.system/log` — then rules with TTL > retention will route through the walker. |
|
||||
| `walker_interval_minutes` | int64 | `0` (fire every pass) | **Important.** See "Walker interval" below — most production deployments should set this to a positive value. |
|
||||
|
||||
### Worker config
|
||||
|
||||
| Key | Default | What |
|
||||
|---|---|---|
|
||||
| `max_runtime_minutes` | 60 | Wall-clock cap per `dailyrun.Run` invocation. The pass returns early if it hits this. |
|
||||
|
||||
### Detection schedule
|
||||
|
||||
The admin scheduler's `DetectionIntervalMinutes` for `s3_lifecycle` is `1440` by default — once per day. Each detection produces one execution. Change in the admin UI's runtime defaults for the job type.
|
||||
|
||||
## Walker interval
|
||||
|
||||
The walker is the part of the worker that lists bucket contents and evaluates them against rules. It fires:
|
||||
|
||||
- **Always**, on cold start (no persisted cursor) and on rule changes — these are bounded events.
|
||||
- **Periodically**, in steady state, gated by `walker_interval_minutes`.
|
||||
|
||||
The throttle exists because walker cost is bucket size, not event rate. If the worker is scheduled at a tighter cadence than the desired walk frequency (CI, sub-hourly admin schedules, manual runs), the steady-state walker would crush the filer with a full subtree scan per invocation.
|
||||
|
||||
Recommended values by cluster size:
|
||||
|
||||
| Cluster | Recommended `walker_interval_minutes` |
|
||||
|---|---|
|
||||
| Small (≤1M objects, single bucket) | 60 |
|
||||
| Medium (≤100M objects, mixed buckets) | 360 (6h) |
|
||||
| Large (≥100M objects) | 1440 (24h) |
|
||||
| Testing / CI | 0 (fire every pass) |
|
||||
|
||||
Setting `0` keeps the prior behavior — fire on every invocation. That's appropriate when the worker is scheduled at exactly the desired walk frequency (e.g., once per day) and the in-repo integration tests rely on it.
|
||||
|
||||
A negative value is rejected at worker start (loud error rather than silent fall-through to "walk every pass").
|
||||
|
||||
## Per-bucket lifecycle XML
|
||||
|
||||
Set via the standard S3 API:
|
||||
|
||||
```bash
|
||||
aws --endpoint $S3_ENDPOINT s3api put-bucket-lifecycle-configuration \
|
||||
--bucket my-bucket \
|
||||
--lifecycle-configuration file://lifecycle.json
|
||||
```
|
||||
|
||||
```jsonc
|
||||
// lifecycle.json
|
||||
{
|
||||
"Rules": [
|
||||
{
|
||||
"ID": "expire-logs-after-30d",
|
||||
"Status": "Enabled",
|
||||
"Filter": { "Prefix": "logs/" },
|
||||
"Expiration": { "Days": 30 }
|
||||
},
|
||||
{
|
||||
"ID": "abort-stuck-mpu",
|
||||
"Status": "Enabled",
|
||||
"Filter": {},
|
||||
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
|
||||
},
|
||||
{
|
||||
"ID": "keep-3-versions",
|
||||
"Status": "Enabled",
|
||||
"Filter": { "Prefix": "versioned/" },
|
||||
"NoncurrentVersionExpiration": { "NewerNoncurrentVersions": 3 }
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The worker picks up rule changes on the next pass. Replay-eligible rule edits trigger a one-time recovery walk to catch already-due objects under the new rule.
|
||||
|
||||
## Verifying a rule is working
|
||||
|
||||
After applying a rule:
|
||||
|
||||
1. Wait one detection interval (default 24h) plus one walker interval.
|
||||
2. Read `s3_lifecycle_dispatch_total{bucket="my-bucket"}` — the counter should advance.
|
||||
3. Verify a target object is gone: `aws s3 head-object --bucket my-bucket --key <expected-expired>` should return 404.
|
||||
|
||||
For testing without waiting, the `weed shell` command supports manual invocation:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
|
||||
```
|
||||
|
||||
This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for examples.
|
||||
|
||||
## Rate limit allocation
|
||||
|
||||
The cluster delete cap is allocated per-worker at job dispatch:
|
||||
|
||||
```
|
||||
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
|
||||
```
|
||||
|
||||
Brief over/undershoot during worker join/leave is acceptable for a bulk workload. The allocation is recomputed each detection cycle, so adding workers smooths out within one day.
|
||||
|
||||
`cluster_deletes_burst` is divided the same way. `0` is treated as `2 × rate` (a token bucket reasonable default).
|
||||
|
||||
## Common change patterns
|
||||
|
||||
### Tightening a TTL (e.g., 60d → 30d)
|
||||
|
||||
The new rule hash differs from the persisted hash — next pass triggers a recovery walk over `RecoveryView`. Already-due objects under the new shorter TTL get caught. The cursor rewinds to `runNow - new_maxTTL`. No operator action needed beyond updating the XML.
|
||||
|
||||
### Adding a walker-only rule (`Expiration.Date`)
|
||||
|
||||
`RuleSetHash` and `PromotedHash` are unchanged (walker-only rules aren't in the replay hash). The replay cursor stays put. The walker reads the updated rule set on its next steady-state fire (or immediately, if invoked manually).
|
||||
|
||||
### Operator forgot to set `walker_interval_minutes`
|
||||
|
||||
Symptom: heartbeat log line shows `duration` ramping up across passes (each pass does a full subtree walk). Mitigation: set the throttle to a value matching your scheduling cadence.
|
||||
|
||||
### Filer is GC'ing meta-log (custom deployment)
|
||||
|
||||
Set `meta_log_retention_days` to the GC retention. Rules with TTL > retention will route through the walker (`PromotedHash` will be non-empty). The first run after this change triggers a recovery walk.
|
||||
@@ -0,0 +1,151 @@
|
||||
# S3 Lifecycle — Troubleshooting
|
||||
|
||||
Incident-response playbook for the S3 lifecycle worker. For monitoring background, see [S3-Lifecycle-Monitoring](S3-Lifecycle-Monitoring).
|
||||
|
||||
## Stuck cursor
|
||||
|
||||
**Symptom:** `s3_lifecycle_cursor_min_ts_ns{shard=N}` is not advancing. Heartbeat shows `cursor_lag_max` growing unbounded.
|
||||
|
||||
**Cause:** The cursor advance is gated on every match from the event dispatching successfully (`DONE`, `NOOP_RESOLVED`, or `SKIPPED_OBJECT_LOCK`). Any unresolved outcome (`RETRY_LATER`, `BLOCKED`, transport error after in-run retries) halts the run for that shard and persists the cursor at the last fully-processed event. Head-of-line blocking is intentional — it surfaces a real problem rather than silently retrying forever.
|
||||
|
||||
**Diagnostics:**
|
||||
|
||||
```promql
|
||||
# Which shard is stuck?
|
||||
time() * 1e9 - s3_lifecycle_cursor_min_ts_ns
|
||||
|
||||
# What outcomes are being returned?
|
||||
sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m]))
|
||||
```
|
||||
|
||||
Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`:
|
||||
|
||||
```
|
||||
daily_run: RETRY_LATER on <bucket>/<key> EXPIRATION_DAYS
|
||||
daily_run: BLOCKED on <bucket>/<key> NONCURRENT_DAYS
|
||||
daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
|
||||
```
|
||||
|
||||
**Mitigations:**
|
||||
|
||||
| Outcome | Root cause | Action |
|
||||
|---|---|---|
|
||||
| `RETRY_LATER` (high rate) | Filer or rate limiter is throttling | Lower `cluster_deletes_per_second` to give the filer headroom, or scale filer capacity |
|
||||
| `BLOCKED FATAL_EVENT_ERROR` | A malformed event the server refuses to dispatch | Check log for the specific reason. May need a code fix; file an issue with the log line |
|
||||
| `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. |
|
||||
| `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability |
|
||||
|
||||
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker.
|
||||
|
||||
## Walker stuck (no progress on walker-only rules)
|
||||
|
||||
**Symptom:** `s3_lifecycle_daily_run_last_walked_ns{shard=N}` is not advancing. Rules like `Expiration.Date`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent` aren't firing on objects that should be due.
|
||||
|
||||
**Causes:**
|
||||
|
||||
1. `walker_interval_minutes` is too long for your invocation cadence. Worker runs once per day but interval is set to 48h.
|
||||
2. Walker is hitting an error mid-walk (filer listing failure). Look for `recovery walk:` or `steady walk:` errors in the heartbeat's `errors=N` count.
|
||||
3. The bucket has only walker-bound rules and the empty-replay branch's throttle hasn't elapsed.
|
||||
|
||||
**Diagnostics:**
|
||||
|
||||
```promql
|
||||
# Walker age per shard
|
||||
(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9
|
||||
```
|
||||
|
||||
Check worker config: `walker_interval_minutes` should be ≤ the daily worker schedule interval.
|
||||
|
||||
**Mitigation:** lower `walker_interval_minutes`. Setting `0` temporarily forces every pass to walk.
|
||||
|
||||
## Test PUT a file with a 1-day rule, didn't expire
|
||||
|
||||
The S3 API rejects `Expiration.Days < 1`, so the smallest "expire after N days" you can configure is 1 day. The worker runs once per day by default. Object PUT + 1-day rule + waiting one day is the minimum scenario.
|
||||
|
||||
For testing, the in-repo integration suite uses a trick: backdate the entry's `Mtime` via `filer_pb.UpdateEntry` to 30+ days ago. See [test/s3/lifecycle/](https://github.com/seaweedfs/seaweedfs/tree/master/test/s3/lifecycle) for the pattern.
|
||||
|
||||
For ad-hoc verification, invoke the worker manually:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
|
||||
```
|
||||
|
||||
This runs the same code path as the scheduled worker, but driven from your shell rather than the admin scheduler.
|
||||
|
||||
## All shards report `errors=16` every pass
|
||||
|
||||
**Symptom:** Heartbeat consistently shows `status=error shards=16 errors=16 duration=Ns`.
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **Filer unreachable.** Subscription fails on every shard. Check filer health and gRPC connectivity.
|
||||
2. **passCtx timeout from `-refresh` loop.** If `-refresh` is less than the pass cap, the timeout fires before the drain completes. This is now treated as "clean end-of-pass" — it should not show as errors=N. If you see this on a build before #9481, upgrade.
|
||||
3. **Bucket walker is timing out.** Big bucket, walker hits ctx deadline. Increase `max_runtime_minutes`.
|
||||
|
||||
## Some objects expired, others didn't (same rule)
|
||||
|
||||
**Symptom:** Two objects matching the same rule with the same age — one is deleted, the other isn't.
|
||||
|
||||
**Common causes:**
|
||||
|
||||
1. **The non-deleted object's mtime is wrong.** Check the entry's mtime — it might be more recent than you expect (e.g., a recent metadata update bumped it).
|
||||
2. **The objects are on different shards** and one shard has a stuck cursor while the other doesn't. Check `s3_lifecycle_cursor_min_ts_ns` per shard.
|
||||
3. **`Filter` doesn't match what you think.** A prefix-only filter requires the object key to start with that prefix; a tag filter requires the matching tag. Verify with `aws s3api head-object` (returns tags via `--query`).
|
||||
4. **Object lock or retention.** The dispatcher returns `SKIPPED_OBJECT_LOCK` for protected objects. Check `s3_lifecycle_dispatch_total{outcome="SKIPPED_OBJECT_LOCK"}`.
|
||||
|
||||
## How to read the cursor files
|
||||
|
||||
Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
> fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json
|
||||
```
|
||||
|
||||
Schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"version": 1,
|
||||
"shard_id": 0,
|
||||
"ts_ns": 1715600000000000000,
|
||||
"rule_set_hash": "<base64 32 bytes>",
|
||||
"promoted_hash": "<base64 32 bytes>",
|
||||
"last_walked_ns": 1715620000000000000
|
||||
}
|
||||
```
|
||||
|
||||
- `ts_ns == 0` means "no replay progress" — either cold start or a bucket whose rules are all walker-only.
|
||||
- `last_walked_ns == 0` (or absent) means "never walked steady-state". Next pass will walk.
|
||||
|
||||
Manually editing the cursor is supported as an escape hatch but obviously breaks the invariant that "everything before persisted.TsNs has been processed under the same rules." Use sparingly.
|
||||
|
||||
## Resetting a shard
|
||||
|
||||
If a shard's cursor is corrupted or wedged in an unrecoverable state:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
> fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json
|
||||
```
|
||||
|
||||
Next pass treats the shard as cold start: recovery walker fires over `RecoveryView(snap)`, then the cursor seeds at `runNow - maxTTL`. This re-replays a `maxTTL`-wide window of meta-log events. Identity-CAS on the server side makes redundant deletes no-ops, so re-replay is safe.
|
||||
|
||||
## Suspending the worker
|
||||
|
||||
The admin UI's plugin scheduler allows pausing the `s3_lifecycle` job type. The worker won't be invoked while paused; cursors are preserved, and on resume the next pass picks up at the persisted state.
|
||||
|
||||
For a single-bucket "stop deleting from this bucket" without pausing the worker:
|
||||
|
||||
```bash
|
||||
aws --endpoint $S3_ENDPOINT s3api delete-bucket-lifecycle --bucket my-bucket
|
||||
```
|
||||
|
||||
The worker reads the empty rule set on the next pass; `rsh == [32]byte{}` causes the empty-replay branch to run only the walker (which has nothing to walk for an empty rule set) and exit. No deletes.
|
||||
|
||||
## Reverting
|
||||
|
||||
If something goes very wrong, the streaming worker (the previous design) is no longer in the codebase. Revert is via downgrading the binary. The cursor format has a `version` field — versions `>1` would fail-loud on load by an older binary that only knows version `1`. Currently version `1` is the only version.
|
||||
|
||||
For escape-hatch operations (delete all cursors, suspend worker globally), prefer pausing via the admin UI over force-killing the worker process; the worker exits cleanly between passes.
|
||||
+65
-93
@@ -1,26 +1,26 @@
|
||||
# S3 Lifecycle Configuration
|
||||
# S3 Lifecycle
|
||||
|
||||
SeaweedFS supports S3 bucket lifecycle configuration for automated object expiration, non-current version cleanup, delete marker removal, and incomplete multipart upload abortion.
|
||||
SeaweedFS implements the S3 `PutBucketLifecycleConfiguration` API. Configured rules are evaluated and enforced by a worker that runs as a scheduled job and exits when each pass completes.
|
||||
|
||||
## Supported Features
|
||||
This page is the operator-facing entry point. Developers and architecture readers should see [`weed/s3api/s3lifecycle/DESIGN.md`](https://github.com/seaweedfs/seaweedfs/blob/master/weed/s3api/s3lifecycle/DESIGN.md).
|
||||
|
||||
## Supported features
|
||||
|
||||
| Feature | Status | Notes |
|
||||
|---|---|---|
|
||||
| `Expiration.Days` | Supported | Fast path via TTL + RocksDB compaction filter |
|
||||
| `Expiration.Date` | Supported | Evaluated by lifecycle worker at scan time |
|
||||
| `ExpiredObjectDeleteMarker` | Supported | Removes delete markers that are the sole remaining version |
|
||||
| `NoncurrentVersionExpiration.NoncurrentDays` | Supported | Requires bucket versioning enabled |
|
||||
| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Supported | Keep N newest non-current versions |
|
||||
| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Supported | Per-rule prefix scoping |
|
||||
| `Filter.Prefix` | Supported | |
|
||||
| `Filter.Tag` | Supported | Evaluated at scan time |
|
||||
| `Filter.And` (Prefix + Tags + Size) | Supported | Evaluated at scan time |
|
||||
| `Filter.ObjectSizeGreaterThan` | Supported | Evaluated at scan time |
|
||||
| `Filter.ObjectSizeLessThan` | Supported | Evaluated at scan time |
|
||||
| `Transition` | Not supported | Requires storage class tiers |
|
||||
| `NoncurrentVersionTransition` | Not supported | Requires storage class tiers |
|
||||
| `Expiration.Days` | Yes | Latest-version PUT clock |
|
||||
| `Expiration.Date` | Yes | Walker path; fires once date is reached |
|
||||
| `Expiration.ExpiredObjectDeleteMarker` | Yes | Walker path; sibling-aware |
|
||||
| `NoncurrentVersionExpiration.NoncurrentDays` | Yes | Clock starts at the demoting PUT, not the entry's own mtime |
|
||||
| `NoncurrentVersionExpiration.NewerNoncurrentVersions` | Yes | Walker path; version-list aware |
|
||||
| `AbortIncompleteMultipartUpload.DaysAfterInitiation` | Yes | |
|
||||
| `Filter.Prefix` | Yes | |
|
||||
| `Filter.Tag` | Yes | |
|
||||
| `Filter.ObjectSizeGreaterThan` / `ObjectSizeLessThan` | Yes | |
|
||||
| `Filter.And` (composite) | Yes | |
|
||||
| `Transition` / `NoncurrentVersionTransition` | No | SeaweedFS doesn't model storage class tiers |
|
||||
|
||||
## API Endpoints
|
||||
## API endpoints
|
||||
|
||||
```
|
||||
PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration
|
||||
@@ -28,6 +28,37 @@ GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration
|
||||
DELETE /{bucket}?lifecycle # DeleteBucketLifecycle
|
||||
```
|
||||
|
||||
## Example: AWS CLI
|
||||
|
||||
```bash
|
||||
# Set lifecycle configuration
|
||||
aws s3api put-bucket-lifecycle-configuration \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket \
|
||||
--lifecycle-configuration '{
|
||||
"Rules": [
|
||||
{
|
||||
"ID": "expire-old",
|
||||
"Status": "Enabled",
|
||||
"Filter": { "Prefix": "" },
|
||||
"Expiration": { "Days": 90 },
|
||||
"NoncurrentVersionExpiration": { "NoncurrentDays": 30 },
|
||||
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 7 }
|
||||
}
|
||||
]
|
||||
}'
|
||||
|
||||
# Get lifecycle configuration
|
||||
aws s3api get-bucket-lifecycle-configuration \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket
|
||||
|
||||
# Delete lifecycle configuration
|
||||
aws s3api delete-bucket-lifecycle \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket
|
||||
```
|
||||
|
||||
## Example: Terraform
|
||||
|
||||
```hcl
|
||||
@@ -47,7 +78,7 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" {
|
||||
}
|
||||
|
||||
noncurrent_version_expiration {
|
||||
noncurrent_days = 7
|
||||
noncurrent_days = 7
|
||||
newer_noncurrent_versions = 2
|
||||
}
|
||||
|
||||
@@ -58,89 +89,30 @@ resource "aws_s3_bucket_lifecycle_configuration" "example" {
|
||||
}
|
||||
```
|
||||
|
||||
## Example: AWS CLI
|
||||
## How it works
|
||||
|
||||
```bash
|
||||
# Set lifecycle configuration
|
||||
aws s3api put-bucket-lifecycle-configuration \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket \
|
||||
--lifecycle-configuration '{
|
||||
"Rules": [
|
||||
{
|
||||
"ID": "expire-old",
|
||||
"Status": "Enabled",
|
||||
"Filter": { "Prefix": "" },
|
||||
"Expiration": { "Days": 90 },
|
||||
"NoncurrentVersionExpiration": {
|
||||
"NoncurrentDays": 30
|
||||
},
|
||||
"AbortIncompleteMultipartUpload": {
|
||||
"DaysAfterInitiation": 7
|
||||
}
|
||||
}
|
||||
]
|
||||
}'
|
||||
The lifecycle worker is a scheduled job (default daily). Each invocation:
|
||||
|
||||
# Get lifecycle configuration
|
||||
aws s3api get-bucket-lifecycle-configuration \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket
|
||||
1. Reads bucket lifecycle XML from each bucket's metadata.
|
||||
2. Compiles rules into a per-shard partition (replay-eligible vs. walker-bound).
|
||||
3. Subscribes to the filer meta-log — one stream covering all 16 shards in this worker process.
|
||||
4. For replay-eligible actions (`ExpirationDays`, `NoncurrentDays`, `AbortMPU`), checks each event's DueTime and dispatches `LifecycleDelete` if elapsed.
|
||||
5. For walker-bound rules (`ExpirationDate`, `ExpiredObjectDeleteMarker`, `NewerNoncurrent`, or anything promoted to scan-only), iterates the bucket and evaluates each entry against current state.
|
||||
6. Persists per-shard cursors so the next pass resumes where this one left off.
|
||||
|
||||
# Delete lifecycle configuration
|
||||
aws s3api delete-bucket-lifecycle \
|
||||
--endpoint-url http://localhost:8333 \
|
||||
--bucket my-bucket
|
||||
```
|
||||
The worker exits when the pass is done. The admin scheduler invokes it on a daily cadence by default; operators can change that via the standard plugin scheduler config.
|
||||
|
||||
## How It Works
|
||||
|
||||
### Two-Tier Expiration Architecture
|
||||
|
||||
**Tier 1 — TTL fast path:** Simple `Expiration.Days` rules with prefix-only filters are translated into TTL entries in `filer.conf`. The RocksDB compaction filter automatically removes expired entries during normal compaction at zero additional cost. This is the most efficient path for the common case.
|
||||
|
||||
**Tier 2 — Scan-time evaluation:** Rules with tag filters, size filters, date-based expiration, non-current version expiration, and delete marker cleanup are evaluated by the lifecycle plugin worker. The worker periodically scans buckets and evaluates each object against the stored lifecycle XML configuration.
|
||||
|
||||
### Which rules use which tier?
|
||||
|
||||
| Rule Type | Tier | Why |
|
||||
|---|---|---|
|
||||
| `Expiration.Days` (prefix only) | TTL fast path | Can be expressed as per-entry TTL |
|
||||
| `Expiration.Days` (with tags/size) | Worker scan | TTL can't express tag/size constraints |
|
||||
| `Expiration.Date` | Worker scan | Absolute date, not relative TTL |
|
||||
| `NoncurrentVersionExpiration` | Worker scan | Requires version enumeration |
|
||||
| `ExpiredObjectDeleteMarker` | Worker scan | Requires version counting |
|
||||
| `AbortIncompleteMultipartUpload` | Worker scan | Scans `.uploads` directory |
|
||||
|
||||
### Lifecycle Worker
|
||||
|
||||
The lifecycle plugin worker runs as part of the SeaweedFS plugin system. It:
|
||||
|
||||
1. **Detects** buckets with lifecycle rules (from stored lifecycle XML or filer.conf TTLs)
|
||||
2. **Scans** bucket contents and evaluates lifecycle rules against each object
|
||||
3. **Executes** actions: delete expired objects, remove old non-current versions, clean up delete markers, abort stale multipart uploads
|
||||
|
||||
Worker configuration options:
|
||||
|
||||
| Setting | Default | Description |
|
||||
|---|---|---|
|
||||
| `batch_size` | 1000 | Entries per filer listing page |
|
||||
| `max_deletes_per_bucket` | 10000 | Max expired objects to delete per run |
|
||||
| `dry_run` | false | Detect but don't delete |
|
||||
| `delete_marker_cleanup` | true | Remove expired delete markers |
|
||||
| `abort_mpu_days` | 7 | Fallback for buckets without lifecycle XML MPU rules |
|
||||
| `bucket_filter` | (all) | Wildcard pattern to scope lifecycle to specific buckets |
|
||||
|
||||
## Versioning Integration
|
||||
## Versioning integration
|
||||
|
||||
Lifecycle rules interact with [S3 Object Versioning](S3-Object-Versioning):
|
||||
|
||||
- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded. `NewerNoncurrentVersions` retains the N newest non-current versions.
|
||||
- **`NoncurrentVersionExpiration`** only applies to versioned buckets. Non-current versions are deleted after `NoncurrentDays` days since they were superseded (the demoting PUT's TsNs, not the version's own mtime). `NewerNoncurrentVersions` retains the N newest non-current versions.
|
||||
- **`ExpiredObjectDeleteMarker`** removes delete markers that are the sole remaining version of an object (no non-current versions behind them).
|
||||
- **`Expiration.Days`** on a versioned bucket creates a delete marker when the current version expires; it does not permanently delete the object.
|
||||
|
||||
## Limitations
|
||||
## Quick references
|
||||
|
||||
- **Transition rules** (`Transition`, `NoncurrentVersionTransition`) are not supported. SeaweedFS does not have S3-equivalent storage class tiers.
|
||||
- **Detection interval**: The lifecycle worker scans on a configurable interval (default 5 minutes). Objects may persist slightly beyond their configured expiration until the next scan.
|
||||
- **TTL fast path** applies only to `Expiration.Days` rules with prefix-only filters. Rules with tag or size constraints are always evaluated at scan time.
|
||||
- **[Operator Guide](S3-Lifecycle-Operator-Guide)** — config knobs, defaults, when to change each
|
||||
- **[Monitoring](S3-Lifecycle-Monitoring)** — Prometheus metrics, heartbeat log line, what a healthy run looks like
|
||||
- **[Troubleshooting](S3-Lifecycle-Troubleshooting)** — stuck cursor, missing deletes, head-of-line blocking
|
||||
- **[Architecture](S3-Lifecycle-Architecture)** — high-level overview of the worker, engine, and dispatch path
|
||||
Reference in new issue
Block a user