Jump to content

Resilience Coffee/2026/08/21

From Resilience Coffee
  • Availability and Reliability:
  • Orthogonal:
  • Systems can be unreliable but high availability
  • e.g. self-hosting may be more reliable because the change rate can be smaller or off-peak
  • Not that the failure modes are reduce, but rather that change rate is reduced
  • "How can you have a critical system and not control it?" e.g. GitHub
  • Tax Season deployments for a tax preparer software
  • https://incident.io/blog/we-turned-off-pub-sub-and-nobody-noticed
  • Scale of impact and duration of impact
  • Object to self-hosted as more available: maybe!
  • Steve McGee's Inverted pyramid model https://wiki.resilience-coffee.org/wiki/Inverted_reliability_pyramid
  • Reliability beyond availability: e.g. incorrect search results
  • Or LLMs hallucinating
  • Should SRE be able to control e.g. feature flags?
  • Help teams help themselves vs. separating responsibilities
  • Connection to leadership maturity and "rowing in the same direction"
  • SRE as limiting options:
  • Paved path, guardrails, desire lines
  • Use RFC process
  • Can this happen post AI?
  • Maybe in an ADR style approach where the LLMs pull the RFCs into context
  • AI can make the "quality" of a RFC better, but social benefits lost: alignment of decision making process across a group of humans
  • "Let people touch the hot stove" as a precursor to group learning
  • Related to educational philosophies
  • "Looking for trouble" meetings
  • Categorization of incidents: is it valuable?
  • re: https://direct.mit.edu/books/monograph/4738/Sorting-Things-OutClassification-and-Its​