Resilience Coffee/2026/08/21
Appearance
- Availability and Reliability:
- Orthogonal:
- Systems can be unreliable but high availability
- e.g. self-hosting may be more reliable because the change rate can be smaller or off-peak
- Not that the failure modes are reduce, but rather that change rate is reduced
- "How can you have a critical system and not control it?" e.g. GitHub
- Tax Season deployments for a tax preparer software
- https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed
- Scale of impact and duration of impact
- Object to self-hosted as more available: maybe!
- Steve McGee's Inverted pyramid model https://wiki.resilience-coffee.org/wiki/Inverted_reliability_pyramid
- Reliability beyond availability: e.g. incorrect search results
- Or LLMs hallucinating
- Should SRE be able to control e.g. feature flags?
- Help teams help themselves vs. separating responsibilities
- Connection to leadership maturity and "rowing in the same direction"
- SRE as limiting options:
- Paved path, guardrails, desire lines
- Use RFC process
- Can this happen post AI?
- Maybe in an ADR style approach where the LLMs pull the RFCs into context
- AI can make the "quality" of a RFC better, but social benefits lost: alignment of decision making process across a group of humans
- "Let people touch the hot stove" as a precursor to group learning
- Related to educational philosophies
- "Looking for trouble" meetings
- Categorization of incidents: is it valuable?
- re: https://direct.mit.edu/books/monograph/4738/Sorting-Things-OutClassification-and-Its