Internal IT SLA: how to define a 99.9% your operations team can actually meet (not just sell)
The internal SLA is the metric your operations team either can or cannot meet when it is asked to. The distance between what you sell to the customer and what your own team delivers on average is where technical trust breaks down. This article starts from one principle: if you cannot measure your SLA inside your own operation, you should not be able to promise it.
The 99.9% figure shows up in commercial documents almost as a badge of being “professional”. The problem is that it gets promised without having measured the starting point. The difference between 99.9% and 99.99% is not one more decimal: it is a decade of additional operational complexity.
What an SLA measures and what an SLO measures
An SLA (Service Level Agreement) is the contractual commitment with a third party, normally a customer or an internal business unit. An SLO (Service Level Objective) is the internal target the team sets for itself in order to deliver value. When SLAs are signed without SLOs backing them, the commercial team promises availability that operations cannot verify, and the first serious incident opens the question of who pays the service credit.
Before setting the SLA with the customer, define the internal SLO first. It is the only way to know whether you can promise anything at all. Google SRE and the ITIL 4 texts treat this in different and complementary ways: the SLO belongs to operations; the SLA belongs to legal and sales.
How 99.9% translates into minutes per period
99.9% monthly is 43 minutes 12 seconds of total tolerated unavailability. Over a quarter it is 2 hours 9 minutes 36 seconds. Over a year it is 8 hours 45 minutes 36 seconds. To reach 99.99% those numbers are divided by ten: 4.3 minutes per month, 26 minutes per year. To reach 99.999%: 26 seconds per month, 4 minutes 22 seconds per year.
Any SLA that adds one more decimal without operations being able to document the previous level is selling something it does not know how to deliver. And even when a provider offers magical figures, it pays to read the fine print: how unavailability is measured (what counts as downtime), what is excluded (scheduled maintenance, announced windows), and what the credits for non-compliance are.
The starting point: your real unavailability today
Before setting any SLA, measure the last year of operation with your team. You need three data points:
1. Total real unavailability time (the sum of incidents that affected the customer, not just scheduled maintenance windows).
2. Frequency (how many incidents per month, their typical duration, their most common root cause).
3. Recurring causes (the root cause that repeats the most — is it operations, equipment, or upstream such as the utility or the ISP?).
Three decisions before signing the SLA with the customer
First: define the internal SLO. It is the target your team commits to meeting before facing customers. Without an internal SLO, any customer SLA is an empty word.
Second: define what is excluded from the calculation. Scheduled maintenance with prior notice, technology change windows, upstream provider failures (utility, ISP, carrier). What you exclude reduces your measured unavailability and improves your SLA; what you do not exclude requires additional redundancy.
Third: define the credits for non-compliance. They are not a punishment for the team: they are the way operations commits to reducing incidents. An SLA without a service credit is an SLA the provider can ignore.
Why most enterprise SLAs end up at 99.9%
Three reasons. First, it is the figure the commercial team already uses because it sounds “professional” and they have no way to defend 99.99% without the infrastructure. Second, because 99.9% is achievable with basic redundancy and 24×7 monitoring. Third, because operationally 99.9% leaves 43 minutes per month to absorb real incidents without triggering credits: it is the number that fits into an SLA without needing to invest much.
For most mid-sized companies in Mexico with a physical data center, 99.9% is a realistic and defensible objective. For 99.99% promises, the customer must understand that they are paying for real redundancy (N+1 across every critical layer) and for formalized processes, not just for a bigger UPS.
How to make the SLA actually hold
Three practices make an SLA hold: continuous measurement of unavailability (24×7 monitoring with an auditable log), documented and up-to-date runbooks (what gets done, in what order, who calls whom), and a quarterly review where operations reports to the customer how many incidents it had, what caused them and how they were reduced. The last one is the one most companies skip: turning the SLA into an improvement tool, not a forgotten contract.
Sources
[1] Google SRE — Service Level Objectives (book) — https://sre.google/sre-book/service-level-objectives/
[2] ITIL 4 — Practice: Service Level Management — https://www.axelos.com/certifications/itil-service-management/itil-4-foundation
[3] ISO/IEC 20000-1 — Service management system requirements — https://www.iso.org/standard/70636.html
Want to master this?
Noxtel Academy →