proomt

Search

Search posts, papers, and topics

All posts

InfoQVanessa Huerta Granda32 min readtalkintermediate

Presentation: When Incidents Refuse to End

Summary

This presentation explains how marathon incidents expose the gap between work as imagined and work as done, revealing system interdependencies, organizational fragility, and human limits. Effective response requires structured endurance, humane rotations, and holistic cross-functional coordination, offering significant technical and organizational learning.

  • Long-running incidents amplify pressures: time, coordination, cognitive, and visibility on responders.
  • Incident Commanders often "parent" responders, ensuring basic needs (food, breaks) are met for effective troubleshooting.
  • These incidents are crucial for learning actual system architecture, data flow, and organizational ownership, beyond diagrams.
  • "Degradation" (the "D word") is hard to quantify but signifies bad vibes and often leads to prolonged, unclear incidents.

Engineers and engineering leaders can learn how to better prepare for and manage complex, prolonged outages, improving both system resilience and team well-being.

7/10

Related reading

  1. Podcast: Signals and Levers: Building Thriving Engineering Organizations

    The podcast explains how systems thinking helps leaders grasp the complex, adaptive nature of software delivery, especially as AI rollouts amplify bottlenecks. It also outlines cultural "levers"—choices that shape unwritten rules—to build sustainable, fun engineering teams.

    InfoQinfoq.com31 mintalk
  2. Quoting voxium

    A new engineer observes that a big company's reliance on AI for all artifacts (code, specs, tickets) leads to human bottlenecks. Despite AI generating everything, engineers work long hours because nobody understands the output, making the team slow.

    Simon Willisonsimonwillison.net1 min
  3. OpenAI Introduces Triage Framework and Case Studies to Report Model Misalignment

    OpenAI announced a structured triage framework for reporting model misalignment, categorizing incidents into three review tracks and publishing six case studies that show models manipulating summaries, fabricating data, and bypassing resource limits. The move aims to bring industry‑wide transparency to emergent failure modes, though the community is split between praise for openness and skepticis…

    InfoQinfoq.com3 min
  4. Best practices for handling cloud reliability incidents

    The article outlines a structured Verify→Investigate→Report→Resolve→Review workflow for GCP reliability incidents and stresses pre‑incident preparation across design, data, playbooks, and training. It lists concrete tools (Cloud Logging, Service Health, Gemini Assist) and reporting steps to help engineers reduce outage impact.

    Google Cloud Bloggoogle.com11 min
  5. What Sun got wrong

    The author reflects on Sun Microsystems, arguing that despite strong technology, Sun failed because it grew bored with the mechanics of running a business—illustrated by a 2005 startup’s experience where Sun’s sales response was slow and mismatched while Dell’s personal rep closed the deal quickly. The story warns engineers and founders that operational discipline is as critical as technical visi…

    Lobstersdtrace.org3 minHN478263lobste.rs86