Markdown
302 words
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
# Q1 Platform Reliability Report *Prepared by the Infrastructure team — internal distribution only.* --- ## Executive summary Service availability across the core platform held at **99.94%** for the quarter, exceeding the 99.9% target. Two incidents accounted for the majority of downtime; both have remediation tracked in the next section. Mean time to recovery dropped from 38 minutes to 22 minutes quarter-over-quarter. ## Incident overview | Date | Service | Severity | Duration | Root cause | |------------|-------------|---------:|---------:|-------------------------| | 2026-01-14 | Auth API | S2 | 41 min | Expired upstream cert | | 2026-02-08 | Search | S1 | 78 min | Index shard saturation | | 2026-03-22 | Payments | S3 | 12 min | Deploy rollback | ### Auth API outage — 2026-01-14 A renewal job failed silently three days before the certificate expired. The on-call engineer was paged at 03:12 UTC. Recovery required manual issuance and a rolling restart of seven pods. > **Action item.** Migrate cert renewal to the centralized PKI service. Owner: Platform Security. Target: 2026-05-01. ### Search saturation — 2026-02-08 Query volume exceeded the per-shard ceiling during a regional traffic spike. The autoscaler did not react because shard count is fixed at deploy time. ```yaml # proposed change to search-cluster.yaml shards: min: 12 max: 32 # was: 12 (fixed) scale_metric: query_p99_latency_ms scale_threshold: 250 ``` ## Capacity outlook Forecast traffic for Q2 grows ~18% based on current cohort retention. Compute headroom is sufficient; storage on the analytics cluster will reach 80% by mid-May without intervention. 1. Provision an additional 40 TiB on the analytics tier. 2. Enable tiered storage for cold partitions older than 90 days. 3. Re-baseline alert thresholds after the storage migration. ## Recommendations - Consolidate certificate management under a single owner. - Move shard sizing from static config to a metric-driven autoscaler. - Schedule a game-day exercise for the auth path before end of Q2. --- *Questions? Reach out in #infra-reliability.*