Dashboard for Mean Time to Recovery
Mean time to recovery, or MTTR, is the average time it takes to recover from a partial or total failure. This metric is used specifically for DevOps, giving insight into team stability and flow. MTTR covers the entire process of recovering from a failure from start to finish, with the “finish” meaning the service is fully operational again. Using MTTR is also an excellent way to compare how well your recovery times are against your competitors. While figuring out your MTTR doesn’t fully address everything happening when a failure occurs, it’s excellent for having a record of the speed your team addresses a failure and how much time the overall recovery process takes.
How to Measure Mean Time to Recovery
Learn about Mean Time to Recovery, including how to measure it, and leverage it in dashboards and visualizations with Metabase.
What is Mean Time to Recovery?
Mean time to recovery, or MTTR, is the average time it takes to recover from a partial or total failure. This metric is used specifically for DevOps, giving insight into team stability and flow. MTTR covers the entire process of recovering from a failure from start to finish, with the “finish” meaning the service is fully operational again. Using MTTR is also an excellent way to compare how well your recovery times are against your competitors. While figuring out your MTTR doesn’t fully address everything happening when a failure occurs, it’s excellent for having a record of the speed your team addresses a failure and how much time the overall recovery process takes.
Get Started
How to calculate Mean Time to Recovery
You’ll need to know the total downtime for every incident within a set period of time, like the average for a day, week, month, and so on. Then, you’ll take the total number of incidents that occurred in that timeframe. You’re going to divide the total number of minutes down by the number of incidents that occurred during the specified period of time. For example, if your service was down for a total of 2 hours (120 minutes) in a week and there were 3 separate incidents total, you would divide 120 by 3. Your mean time to recovery would then be 40 minutes.
Other KPIs to measure related to Mean Time to Recovery
- Deployment Frequency
- Change Failure Rate
- Downtime
- Uptime
- Online Application Performance
- Mean Time to Detect
- Lead Time For Changes
- Error Rate
- Automated Test Pass Percentage
Why build a dashboard for Mean Time to Recovery?

Everything in one place
Get everyone on the same page by collecting your most important metrics into a single view.


Share your perspective
Take your data wherever it needs to go by embedding it in your internal wikis, websites, and content.

Unlock exploration
Empower your team to measure their own progress and explore new paths to achieve their goals.
How to use Metabase to measure Mean Time to Recovery

Step 1.
Skip the custom quoteThat's right, no sales calls necessary—just sign up, and get running in under 5 minutes.


Step 2.
Plugin your databaseWe connect to the most popular production databases and data warehouses.

Step 3.
Build your KPI dashboardsInvite your team and start building dashboards—no SQL required.
Get started with Metabase
- Free, no-commitment trial
- Easy for everyone—no SQL required
- Up and running in 5 minutes