diff --git a/S3-Lifecycle-Architecture.md b/S3-Lifecycle-Architecture.md index 0ae217d..6690748 100644 --- a/S3-Lifecycle-Architecture.md +++ b/S3-Lifecycle-Architecture.md @@ -6,7 +6,7 @@ High-level overview of the lifecycle worker. For implementation detail, see [`we The lifecycle worker runs as a scheduled job. Each invocation: -``` +```text ┌──────────────────────────────────────────┐ │ dailyrun.Run (one filer subscription) │ │ │ @@ -34,8 +34,8 @@ Once every shard's goroutine returns, the worker tears down the subscription, em Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`: -``` -TsNs — last meta-log event whose matches all dispatched successfully +```text +TsNs — last meta-log event for which all matches dispatched successfully RuleSetHash — ReplayContentHash of the rule set when this cursor was written PromotedHash — PromotedHash(retentionWindow) at write time LastWalkedNs — wall-clock of the last successful walker fire @@ -79,7 +79,7 @@ The walker throttle decouples walker firing from invocation rate. CI invokes the Cluster-wide cap allocated per worker at job dispatch: -``` +```text per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers) ``` diff --git a/S3-Lifecycle-Monitoring.md b/S3-Lifecycle-Monitoring.md index fbb8fc7..d507617 100644 --- a/S3-Lifecycle-Monitoring.md +++ b/S3-Lifecycle-Monitoring.md @@ -10,7 +10,7 @@ All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem. | Metric | Labels | What | |---|---|---| -| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard | +| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event on this shard for which all matches dispatched successfully | | `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire | Derived queries: @@ -37,7 +37,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe | `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter | | `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path | -`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. +`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED` is the proto zero-value — a healthy worker / server pair should never emit it; a non-zero count there indicates an internal error or a version mismatch between worker and server. ### Histograms @@ -50,7 +50,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity): -``` +```text daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m ``` @@ -67,7 +67,7 @@ Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stab A healthy production heartbeat looks like: -``` +```text daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m ``` @@ -87,15 +87,17 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag ## Suggested alerts ```yaml +# `> 0` filters out shards whose gauge is still the proto zero (never +# started); without it, every fresh-install heartbeat triggers the alert. - alert: S3LifecycleCursorLagHigh - expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600 + expr: max(time() * 1e9 - (s3_lifecycle_cursor_min_ts_ns > 0)) / 1e9 > 3600 for: 30m annotations: summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}" runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor - alert: S3LifecycleWalkerStuck - expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400 + expr: max(time() * 1e9 - (s3_lifecycle_daily_run_last_walked_ns > 0)) / 1e9 > 86400 for: 1h annotations: summary: "S3 lifecycle walker hasn't run in > 24h" diff --git a/S3-Lifecycle-Operator-Guide.md b/S3-Lifecycle-Operator-Guide.md index 35b12ce..a51df2d 100644 --- a/S3-Lifecycle-Operator-Guide.md +++ b/S3-Lifecycle-Operator-Guide.md @@ -97,8 +97,8 @@ After applying a rule: For testing without waiting, the `weed shell` command supports manual invocation: -``` -weed shell -master +```text +weed shell -master > s3.lifecycle.run-shard -shards 0-15 -s3 -refresh 1s -runtime 30s ``` @@ -108,7 +108,7 @@ This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](htt The cluster delete cap is allocated per-worker at job dispatch: -``` +```text per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers) ``` diff --git a/S3-Lifecycle-Troubleshooting.md b/S3-Lifecycle-Troubleshooting.md index 0209496..bccd849 100644 --- a/S3-Lifecycle-Troubleshooting.md +++ b/S3-Lifecycle-Troubleshooting.md @@ -20,7 +20,7 @@ sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m])) Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`: -``` +```text daily_run: RETRY_LATER on / EXPIRATION_DAYS daily_run: BLOCKED on / NONCURRENT_DAYS daily_run: transport error on / ABORT_MPU: @@ -35,7 +35,7 @@ daily_run: transport error on / ABORT_MPU: | `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. | | `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability | -The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker. +The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it: first **pause the `s3_lifecycle` job in the admin UI** so the worker isn't running mid-edit, then edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Resume the job. Editing while the worker is active would race with the worker's own save and either lose your change or overwrite the persisted progress. ## Walker stuck (no progress on walker-only rules) @@ -66,8 +66,8 @@ For testing, the in-repo integration suite uses a trick: backdate the entry's `M For ad-hoc verification, invoke the worker manually: -``` -weed shell -master +```text +weed shell -master > s3.lifecycle.run-shard -shards 0-15 -s3 -refresh 1s -runtime 30s ``` @@ -98,8 +98,8 @@ This runs the same code path as the scheduled worker, but driven from your shell Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`: -``` -weed shell -master +```text +weed shell -master > fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json ``` @@ -125,8 +125,8 @@ Manually editing the cursor is supported as an escape hatch but obviously breaks If a shard's cursor is corrupted or wedged in an unrecoverable state: -``` -weed shell -master +```text +weed shell -master > fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json ``` diff --git a/S3-Lifecycle.md b/S3-Lifecycle.md index 8972201..8bae445 100644 --- a/S3-Lifecycle.md +++ b/S3-Lifecycle.md @@ -22,7 +22,7 @@ This page is the operator-facing entry point. Developers and architecture reader ## API endpoints -``` +```text PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration DELETE /{bucket}?lifecycle # DeleteBucketLifecycle