mirror of
https://github.com/seaweedfs/seaweedfs.git
synced 2026-10-08 07:17:48 +02:00
S3 Lifecycle: address review feedback (alert filters, fence tags, weed shell flag, naming)
Pulls the post-review version from seaweedfs/seaweedfs#9491: - Alerts filter zero-valued gauges (avoids fresh-install false positives) - LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED documented as proto zero-value - weed shell -master flag clarified as host:http_port.grpc_port - Manual cursor edits gated on admin-UI pause of the s3_lifecycle job - Fenced code blocks tagged with 'text' for markdownlint compliance - S3 spec rule names used consistently (ExpiredObjectDeleteMarker, NewerNoncurrentVersions, AbortIncompleteMultipartUpload)
1 parent
29aa6c5c5f
commit
9dc49c9596
5 files changed
+24
-22
No files matched your search
@@ -6,7 +6,7 @@ High-level overview of the lifecycle worker. For implementation detail, see [`we
|
||||
|
||||
The lifecycle worker runs as a scheduled job. Each invocation:
|
||||
|
||||
```
|
||||
```text
|
||||
┌──────────────────────────────────────────┐
|
||||
│ dailyrun.Run (one filer subscription) │
|
||||
│ │
|
||||
@@ -34,8 +34,8 @@ Once every shard's goroutine returns, the worker tears down the subscription, em
|
||||
|
||||
Each shard owns a cursor file on the filer at `/etc/s3/lifecycle/daily-cursors/shard-NN.json`:
|
||||
|
||||
```
|
||||
TsNs — last meta-log event whose matches all dispatched successfully
|
||||
```text
|
||||
TsNs — last meta-log event for which all matches dispatched successfully
|
||||
RuleSetHash — ReplayContentHash of the rule set when this cursor was written
|
||||
PromotedHash — PromotedHash(retentionWindow) at write time
|
||||
LastWalkedNs — wall-clock of the last successful walker fire
|
||||
@@ -79,7 +79,7 @@ The walker throttle decouples walker firing from invocation rate. CI invokes the
|
||||
|
||||
Cluster-wide cap allocated per worker at job dispatch:
|
||||
|
||||
```
|
||||
```text
|
||||
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
|
||||
```
|
||||
|
||||
|
||||
@@ -10,7 +10,7 @@ All labels are in `weed/stats/metrics.go` under the `s3_lifecycle` subsystem.
|
||||
|
||||
| Metric | Labels | What |
|
||||
|---|---|---|
|
||||
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event whose matches all dispatched successfully on this shard |
|
||||
| `s3_lifecycle_cursor_min_ts_ns` | `shard` | UnixNano of the last meta-log event on this shard for which all matches dispatched successfully |
|
||||
| `s3_lifecycle_daily_run_last_walked_ns` | `shard` | UnixNano of the most recent successful walker fire |
|
||||
|
||||
Derived queries:
|
||||
@@ -37,7 +37,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe
|
||||
| `s3_lifecycle_bootstrap_dispatch_total` | `bucket`, `kind` | Walker dispatch counter |
|
||||
| `s3_lifecycle_metadata_only_total` | `bucket`, `rule_hash` | Successful deletes that took the metadata-only path |
|
||||
|
||||
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run.
|
||||
`outcome` values: `DONE`, `NOOP_RESOLVED`, `SKIPPED_OBJECT_LOCK`, `RETRY_LATER`, `BLOCKED`, `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED`, `RPC_ERROR`. The first three are success outcomes that advance the cursor; the others halt the run. `LIFECYCLE_DELETE_OUTCOME_UNSPECIFIED` is the proto zero-value — a healthy worker / server pair should never emit it; a non-zero count there indicates an internal error or a version mismatch between worker and server.
|
||||
|
||||
### Histograms
|
||||
|
||||
@@ -50,7 +50,7 @@ Zero values mean "not started yet" — distinct from "0s caught up". The heartbe
|
||||
|
||||
Emitted at the end of every `dailyrun.Run` invocation, at `glog.V(0)` (default verbosity):
|
||||
|
||||
```
|
||||
```text
|
||||
daily_run: status=ok shards=16 errors=0 duration=7s cursor_lag_max=2m walked_max_age=3m
|
||||
```
|
||||
|
||||
@@ -67,7 +67,7 @@ Tokens are space-separated `key=value` for grep / log-aggregator filtering. Stab
|
||||
|
||||
A healthy production heartbeat looks like:
|
||||
|
||||
```
|
||||
```text
|
||||
daily_run: status=ok shards=16 errors=0 duration=12.3s cursor_lag_max=45s walked_max_age=58m
|
||||
```
|
||||
|
||||
@@ -87,15 +87,17 @@ Read it as: 16 shards finished cleanly in 12 seconds; the worst-case replay lag
|
||||
## Suggested alerts
|
||||
|
||||
```yaml
|
||||
# `> 0` filters out shards whose gauge is still the proto zero (never
|
||||
# started); without it, every fresh-install heartbeat triggers the alert.
|
||||
- alert: S3LifecycleCursorLagHigh
|
||||
expr: max(time() * 1e9 - s3_lifecycle_cursor_min_ts_ns) / 1e9 > 3600
|
||||
expr: max(time() * 1e9 - (s3_lifecycle_cursor_min_ts_ns > 0)) / 1e9 > 3600
|
||||
for: 30m
|
||||
annotations:
|
||||
summary: "S3 lifecycle replay lag > 1h on shard {{ $labels.shard }}"
|
||||
runbook: https://github.com/seaweedfs/seaweedfs/wiki/S3-Lifecycle-Troubleshooting#stuck-cursor
|
||||
|
||||
- alert: S3LifecycleWalkerStuck
|
||||
expr: max(time() * 1e9 - s3_lifecycle_daily_run_last_walked_ns) / 1e9 > 86400
|
||||
expr: max(time() * 1e9 - (s3_lifecycle_daily_run_last_walked_ns > 0)) / 1e9 > 86400
|
||||
for: 1h
|
||||
annotations:
|
||||
summary: "S3 lifecycle walker hasn't run in > 24h"
|
||||
|
||||
@@ -97,8 +97,8 @@ After applying a rule:
|
||||
|
||||
For testing without waiting, the `weed shell` command supports manual invocation:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
```text
|
||||
weed shell -master <host:http_port.grpc_port>
|
||||
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
|
||||
```
|
||||
|
||||
@@ -108,7 +108,7 @@ This is exactly what the CI integration suite uses. See [test/s3/lifecycle/](htt
|
||||
|
||||
The cluster delete cap is allocated per-worker at job dispatch:
|
||||
|
||||
```
|
||||
```text
|
||||
per_worker_rate = cluster_deletes_per_second / count(active_s3_lifecycle_workers)
|
||||
```
|
||||
|
||||
|
||||
@@ -20,7 +20,7 @@ sum by (outcome) (rate(s3_lifecycle_dispatch_total[5m]))
|
||||
|
||||
Look at worker log for the offending event. The dispatcher logs at `glog.V(1)`:
|
||||
|
||||
```
|
||||
```text
|
||||
daily_run: RETRY_LATER on <bucket>/<key> EXPIRATION_DAYS
|
||||
daily_run: BLOCKED on <bucket>/<key> NONCURRENT_DAYS
|
||||
daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
|
||||
@@ -35,7 +35,7 @@ daily_run: transport error on <bucket>/<key> ABORT_MPU: <err>
|
||||
| `BLOCKED SKIPPED_OBJECT_LOCK` | Object is locked (legal hold, retention) | Wait for lock to expire, or remove lock manually. Cursor advances normally — this isn't stuck. |
|
||||
| `RPC_ERROR` (sustained) | Transport / network issue | Check S3 server health and filer reachability |
|
||||
|
||||
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it, edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Restart the worker.
|
||||
The worker doesn't auto-skip past a stuck event. If you've verified the event is malformed and want to skip it: first **pause the `s3_lifecycle` job in the admin UI** so the worker isn't running mid-edit, then edit the cursor file directly (`/etc/s3/lifecycle/daily-cursors/shard-NN.json`), advancing `ts_ns` past the bad event's TsNs. Resume the job. Editing while the worker is active would race with the worker's own save and either lose your change or overwrite the persisted progress.
|
||||
|
||||
## Walker stuck (no progress on walker-only rules)
|
||||
|
||||
@@ -66,8 +66,8 @@ For testing, the in-repo integration suite uses a trick: backdate the entry's `M
|
||||
|
||||
For ad-hoc verification, invoke the worker manually:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
```text
|
||||
weed shell -master <host:http_port.grpc_port>
|
||||
> s3.lifecycle.run-shard -shards 0-15 -s3 <s3-host:port> -refresh 1s -runtime 30s
|
||||
```
|
||||
|
||||
@@ -98,8 +98,8 @@ This runs the same code path as the scheduled worker, but driven from your shell
|
||||
|
||||
Cursors live at `/etc/s3/lifecycle/daily-cursors/shard-NN.json` on the filer. Read with the filer's `read` API or `weed shell`:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
```text
|
||||
weed shell -master <host:http_port.grpc_port>
|
||||
> fs.cat /etc/s3/lifecycle/daily-cursors/shard-00.json
|
||||
```
|
||||
|
||||
@@ -125,8 +125,8 @@ Manually editing the cursor is supported as an escape hatch but obviously breaks
|
||||
|
||||
If a shard's cursor is corrupted or wedged in an unrecoverable state:
|
||||
|
||||
```
|
||||
weed shell -master <addr>
|
||||
```text
|
||||
weed shell -master <host:http_port.grpc_port>
|
||||
> fs.rm /etc/s3/lifecycle/daily-cursors/shard-07.json
|
||||
```
|
||||
|
||||
|
||||
+1
-1
@@ -22,7 +22,7 @@ This page is the operator-facing entry point. Developers and architecture reader
|
||||
|
||||
## API endpoints
|
||||
|
||||
```
|
||||
```text
|
||||
PUT /{bucket}?lifecycle # PutBucketLifecycleConfiguration
|
||||
GET /{bucket}?lifecycle # GetBucketLifecycleConfiguration
|
||||
DELETE /{bucket}?lifecycle # DeleteBucketLifecycle
|
||||
|
||||
Reference in new issue
Block a user