Skip to content

IAM Data Access (Runway) Service

The runbook annotation on this service’s alerts links here. All alerts are currently severity: s4 — Slack-only to #g_sscs_authentication, no paging.

AlertMeansFirst checks
IamDataGkeGrpcServiceIamLookupErrorSLOViolationLookupService is returning server-fault gRPC codes (Internal, Unavailable, DeadlineExceeded, ResourceExhausted, Unknown, DataLoss) above 1% on two burn-rate windows. Client-input codes are excluded, so this is our fault by construction.Errors-by-category panel on the overview dashboard to see which code dominates; then the DB Pool Contention row — Unavailable/DeadlineExceeded bursts usually trace back to Yugabyte or pool exhaustion, not to the RPC layer.
IamDataGkeGrpcServiceIamLookupApdexSLOViolationFewer than 99% of LookupService RPCs completed within 250 ms. The histogram has no grpc_code label, so failed RPCs count here too.Lookup RPC latency p50/p99 panels, then repository operation duration — if the DB layer is slow, the RPC apdex follows.
IamDataGkeGrpcServiceIamUpdateErrorSLOViolation / ...IamUpdateApdexSLOViolationSame, for UpdateService (control plane, 500 ms apdex threshold).As above. Update traffic is bursty and low-volume, so check the request rate panel before treating a percentile as a trend.
IamDataGkeGrpcServiceIamLookupTrafficCessationLookupService has served no requests for a sustained period.Whether the deployment is up, and whether Runway is still scraping (/-/metrics via .runway/fairway.yaml). Update has traffic cessation alerting disabled — it is legitimately idle.
component_saturation_slo_out_of_bounds:iam_data_db_poolChecked-out connections have exceeded 95% of the pool’s configured maximum for 15 minutes, on whichever of the lookup/update pools is worst. Measured fleet-wide, so the ceiling already accounts for HPA scaling.The DB Pool Contention row: a rising empty-acquire ratio alongside this confirms callers are waiting on the pool. Then decide between raising MaxConns and fixing whatever holds connections open — the mean acquire wait and repository operation duration panels distinguish those.

Not yet covered: repository operation duration and error rate. It is a deeper diagnostic signal largely covered by the gRPC SLIs above — read the Repository operation duration panel by hand. Tracked in gitlab-org/gitlab#607976.

Staging alerts go to the same channel as production. Runway labels its environments staging and production rather than gstg and gprd, so none of the usual non-prod handling applies to them — a dedicated route in alertmanager/alertmanager.jsonnet sends env=staging alerts for this service to the team channel and stops there, rather than letting them fall through to the generic #alerts channel. Staging alerts never page and never reach incident.io, since those routes match production environments only.

The two ops-rate anomaly alerts (service_ops_out_of_bounds_upper_5m and service_ops_out_of_bounds_lower_5m) are the exception: in staging a route ahead of the team-channel route sends them to the blackhole receiver. Staging traffic is a few short bursts per day, so the 3-sigma bounds flap on every burst and on every quiet hour that follows a burst from a previous week. The production copies of these alerts are not affected.