IAM Data Access (Runway) Service
- Service Overview
- Alerts: https://alerts.gitlab.net/#/alerts?filter=%7Btype%3D%22iam-data-gke-grpc%22%2C%20tier%3D%22sv%22%7D
- Label: gitlab-com/gl-infra/production~“Service::IAMDataAccess”
Logging
Section titled “Logging”Alerts
Section titled “Alerts”The runbook annotation on this service’s alerts links here. All alerts are
currently severity: s4 — Slack-only to #g_sscs_authentication, no paging.
| Alert | Means | First checks |
|---|---|---|
IamDataGkeGrpcServiceIamLookupErrorSLOViolation | LookupService is returning server-fault gRPC codes (Internal, Unavailable, DeadlineExceeded, ResourceExhausted, Unknown, DataLoss) above 1% on two burn-rate windows. Client-input codes are excluded, so this is our fault by construction. | Errors-by-category panel on the overview dashboard to see which code dominates; then the DB Pool Contention row — Unavailable/DeadlineExceeded bursts usually trace back to Yugabyte or pool exhaustion, not to the RPC layer. |
IamDataGkeGrpcServiceIamLookupApdexSLOViolation | Fewer than 99% of LookupService RPCs completed within 250 ms. The histogram has no grpc_code label, so failed RPCs count here too. | Lookup RPC latency p50/p99 panels, then repository operation duration — if the DB layer is slow, the RPC apdex follows. |
IamDataGkeGrpcServiceIamUpdateErrorSLOViolation / ...IamUpdateApdexSLOViolation | Same, for UpdateService (control plane, 500 ms apdex threshold). | As above. Update traffic is bursty and low-volume, so check the request rate panel before treating a percentile as a trend. |
IamDataGkeGrpcServiceIamLookupTrafficCessation | LookupService has served no requests for a sustained period. | Whether the deployment is up, and whether Runway is still scraping (/-/metrics via .runway/fairway.yaml). Update has traffic cessation alerting disabled — it is legitimately idle. |
component_saturation_slo_out_of_bounds:iam_data_db_pool | Checked-out connections have exceeded 95% of the pool’s configured maximum for 15 minutes, on whichever of the lookup/update pools is worst. Measured fleet-wide, so the ceiling already accounts for HPA scaling. | The DB Pool Contention row: a rising empty-acquire ratio alongside this confirms callers are waiting on the pool. Then decide between raising MaxConns and fixing whatever holds connections open — the mean acquire wait and repository operation duration panels distinguish those. |
Not yet covered: repository operation duration and error rate. It is a deeper diagnostic signal largely covered by the gRPC SLIs above — read the Repository operation duration panel by hand. Tracked in gitlab-org/gitlab#607976.
Staging alerts go to the same channel as production. Runway labels its
environments staging and production rather than gstg and gprd, so none
of the usual non-prod handling applies to them — a dedicated route in
alertmanager/alertmanager.jsonnet sends env=staging alerts for this service
to the team channel and stops there, rather than letting them fall through to
the generic #alerts channel. Staging alerts never page and never reach
incident.io, since those routes match production environments only.
The two ops-rate anomaly alerts (service_ops_out_of_bounds_upper_5m and
service_ops_out_of_bounds_lower_5m) are the exception: in staging a route ahead
of the team-channel route sends them to the blackhole receiver. Staging traffic
is a few short bursts per day, so the 3-sigma bounds flap on every burst and on
every quiet hour that follows a burst from a previous week. The production
copies of these alerts are not affected.