Under-replicated, writable and crowded volume counts are cluster-wide
values that only the leader maintains. Summing them across masters
double-counted them, and a demoted master keeps serving its last values,
so a leadership change could inflate the health alerts.
The request counters label "type" with the HTTP method, so filtering on
type=read/write matched nothing and left the throughput chart empty. A
live cluster reports type=GET/POST.
Aggregates the scraped series into per-page structs: overview health,
throughput, latency, errors and maintenance, plus per-server detail for
volume servers, filers, S3, masters and workers. Rates and counts are
summed across servers; latency quantiles take the worst server, since
quantiles cannot be summed. Disk usage is derived from the volume server
resource gauge's used/all mounts.