Resources
Browse through videos, guides, and other educational resources that cover incident management, reliability, team culture, and more.
Videos
Ebook
11.9.2021
Walkthrough Configuring SLOs in Blameless
Service level objectives (SLOs) define what “success” means for engineering orgs, based on the user experience. It’s a whole new way of understanding your production services that helps keep customers satisfied and boosts development velocity. Another benefit to using SLOs is being able to align multiple teams on shared goals — from product to dev, ops, sales, and so on.
Videos
Ebook
11.8.2021
Retrospectives Do We Need Them & Why
A retrospective (or post-mortem) is a critical part of Incident Response. If you don’t document, share, and learn, everyone misses out on a critical learning opportunity. Retrospectives are designed to help teams collectively learn, iterate, and improve. It involves taking the time to story-tell exactly how an incident took place and discussing how to do better next time. The key is to shine new light on processes, tools, and systems so that you know how to fine-tune.
Blog
Ebook
11.2.2021
DevOps Culture: How to Build a Stronger Team
Trying to improve your DevOps team? We’ll explain what DevOps culture is, how it benefits your team, and how you can build it within your organization.
Podcasts
Ebook
10.27.2021
Resilience in Action E11: Applying SRE Principles to Other Aspects of Work and Life with Jennifer Petoff
Blog
Ebook
10.27.2021
What is DevOps CI/CD & Why Is It Important?
Curious about DevOps CI/CD? We explain the continuous integration and continuous deployment pipeline and why it’s important to the DevOps process.
Videos
Ebook
10.14.2021
Unsolved Problems in SRE
Every field of endeavor has its leading edge where the answers are unclear and active exploration is warranted. Although the phrase "here be dragons" might be an appropriate warning, this panel of intrepid adventurers will venture into that unknown territory.
Blog
Ebook
10.6.2021
Incident Response: A Step-by-Step Guide to Managing Incidents
Looking into Incident Response? We explain incident response, the end-to-end process, the teams involved, and steps to take to avoid friction and slow-down. The goal is to manage the incident as efficiently as possible in order to restore or resume the service to its expected operational state.
Videos
Ebook
10.5.2021
"She's not dead yet, Jim": Vulnerability and Retrospectives in Emergency Medicine
What can we learn about working through and analyzing IT incidents from the high tempo, very high consequence world of medical emergency rooms? What does vulnerability have to do with psychological safety, creativity, and team collaboration? It turns out that we can learn quite a lot. We can all benefit from drawing connections across different fields of study and this panel will explore some of these connections. DevOps Enterprise Summit Virtual - US 2021
Videos
Ebook
9.16.2021
How Everyone Learns From Incidents (If you’re on-call or not).
On call doesn’t have to suck. Right? Some vendors even have this as a tagline. Truth is some on-calls do suck (a little) especially if it’s a Sev0. Ever been in a chaotic on-call environment? Unclear on who’s doing what and how to run the playbook. Is there a playbook? This session shares approaches that simply work and more importantly how teams learn and improve over time. From living through inevitable incidents and reviewing retrospectives, everyone up-levels and learns. Happier DevOps teams, happier customers. Presented by Paul Chu, Head of Customer Success at Blameless.
Incident Impact Calculator
Find out how much you could save
Incidents can do real damage to companies that aren't sufficiently prepared them. Use our calculator to estimate the full cost of incidents for your team.
use the calculator