Deployment to Kubernetes: document volume scale-down evacuation and AdminScript

Chris Lu committed 2026-06-16 18:38:26 -07:00
1 parent 716489eafd
commit 02f3b09691
1 file changed
+65
+65
@@ -100,6 +100,30 @@ replicas never share a host directory. For rack/datacenter-aware placement
across multiple volume groups, see
[Topology Support](https://github.com/seaweedfs/seaweedfs-operator/blob/master/TOPOLOGY_SUPPORT.md).
## Scaling volume servers (graceful evacuation)
When you lower `volume.replicas` (or a `volumeTopology` group's `replicas`), the
Operator scales the volume servers down **gracefully** rather than deleting
their pods outright. Kubernetes removes the highest-ordinal pods first, so for
each volume server about to go away the Operator:
1. holds the `StatefulSet` at its current size so the doomed pod keeps serving,
2. runs `volumeServer.evacuate` to move that server's volumes and EC shards onto
the remaining servers, and
3. deletes the pod only once the master confirms the server holds no data.
Servers drain and pods are removed one at a time, top ordinal first. If a volume
cannot be moved safely — for example a replicated volume with no
replication-compliant destination left — the evacuation fails and the
scale-down is **held** (the pod is kept) rather than risk losing data. The
Operator records `VolumeServerEvacuating` and `VolumeServerEvacuationFailed`
events on the `Seaweed` resource, so you can follow progress with
`kubectl describe seaweed <name>`.
This is automatic and needs no configuration, and it applies to both the flat
`volume` group and per-`volumeTopology` groups. Scaling **up** is unchanged —
new volume servers simply join the cluster.
## Declarative Buckets and IAM
Besides provisioning the cluster itself, the Operator can manage S3 buckets,
@@ -192,6 +216,47 @@ For the full guide and field reference, see the
[CSI Support](https://github.com/seaweedfs/seaweedfs-operator/blob/master/CSI_SUPPORT.md)
documentation in the operator repository.
## Scheduled `weed shell` scripts (AdminScript)
The Operator can run `weed shell` administrative scripts on a cron schedule —
for recurring maintenance like `volume.balance`, `volume.fix.replication`, or
`volume.vacuum`. An `AdminScript` resource reconciles into a native Kubernetes
`CronJob` (same namespace, owned by the resource) whose pod runs
`weed shell -master=<cluster masters>` with the script piped to stdin. When the
referenced cluster has mTLS enabled, the job mounts the same
`security.toml`/TLS material so the shell authenticates to the masters over gRPC.
```yaml
apiVersion: seaweed.seaweedfs.com/v1
kind: AdminScript
metadata:
name: nightly-balance
namespace: default
spec:
clusterRef: { name: seaweed1 } # a Seaweed CR in the same namespace
schedule: "0 2 * * *" # daily at 02:00
script: |
lock
volume.balance -force
volume.fix.replication
unlock
```
`concurrencyPolicy` defaults to `Forbid` so runs never overlap; `suspend: true`
pauses scheduling. The usual CronJob/Job knobs (`timeZone`, history limits,
`backoffLimit`, `activeDeadlineSeconds`, `resources`, scheduling) pass through.
For the full field reference, see the
[AdminScript](https://github.com/seaweedfs/seaweedfs-operator#scheduled-admin-scripts-adminscript)
section of the operator README.
> **`AdminScript` vs the [Admin Script Plugin](Migrate-Maintenance-Scripts-to-Admin-Script-Plugin):**
> the `AdminScript` CRD schedules a standalone `weed shell` run as a Kubernetes
> `CronJob` — handy when you want maintenance expressed as plain, GitOps-managed
> manifests. The Admin Script Plugin instead runs maintenance on the admin
> server's `weed worker` processes with a `/plugin` UI, run history, and
> dedicated workers for erasure coding and volume balancing. If you run an admin
> server, the plugin is the richer option.
# Recommended: Helm Deployment
The SeaweedFS Helm chart is the most flexible and maintained way to deploy SeaweedFS on Kubernetes.