All research
ROI Report24 pages16 min readv1.4

ROI Analysis: Policy-First AI vs. Traditional NOC Models

A defensible three-year TCO model at 100, 500 and 2,000 managed units — including the costs both sides usually leave out.

RN
Rajesh Nair
VP Sales
KP
Kiran Patel
VP Engineering
Published
AAQUILIX ROI MODEL · 202638%lower three-year total cost of ownershipat 500 managed units, break-even month 8POLICY-FIRST AI VS. TRADITIONAL NOCAAQUILIX.COMBREAK-EVENMSPAAQUILIXMONTH 0 → 36 · CUMULATIVE
Abstract

Comparing a managed service provider against an AI operations platform is genuinely difficult, because the two price on different axes: MSP cost scales with effort and ticket volume, unit-based AI cost scales with estate size. This report builds an explicit three-year model at 100, 500 and 2,000 managed units, includes the customer-side costs that vendor models on both sides tend to omit, and is candid about the estate profiles where the MSP is the better commercial answer.

Written for
  • CFO / Finance business partner
  • CIO / IT budget owner
  • Procurement & vendor management
  • Board investment committee
38%
lower three-year TCO at 500 managed units
Month 8
typical break-even point
300–600
managed units — the crossover band
3.9 FTE
per year currently spent on actionless alerts

Most vendor ROI models are unusable in a board pack because they compare a fully-loaded incumbent against a stripped-down challenger. This one includes the customer-side effort each model demands, and it names the conditions under which we lose.

1. Why the usual comparison is wrong

Traditional MSP contracts price against effort: FTE allocation, shift coverage, and ticket bands with overage above them. As incident volume grows, cost grows — either directly through overage or indirectly at renewal, when the run-rate becomes the new baseline.

Unit-based platform pricing charges per managed unit regardless of how many incidents those units generate. A bad month costs the same as a good one. This is the single biggest structural difference between the models, and it is not primarily a price difference — it is a transfer of volume risk from the customer to the platform.

Any comparison that ignores this produces a number that is correct only for the volume assumption it was built on. The right way to present this to a board is not a single TCO figure but a pair of curves and a break-even point, which is what this report produces.

2. Model inputs — what is included on each side

To be fair to both models, the comparison includes every line either side would reasonably incur, including the ones that are absorbed internally and never invoiced.

Cost lineMSP modelPlatform model
Contracted / subscription feeBase fee by FTE allocation and shift coveragePer managed unit, per year
Volume-driven costTicket overage at historical ratesNone — flat regardless of incident volume
Transition / implementationOnboarding at contract start and at every boundaryOne-time implementation and integration
Knowledge captureAbsorbed by the MSP, lost at contract boundaryPolicy authorship: 4–8 engineer-weeks in year one, retained
Customer-side coordinationVendor management, service reviews, escalation, SLA disputes: 0.5–1.5 FTEPolicy review and re-attestation: ~0.2 FTE steady state
InfrastructureNoneOn-premise inference nodes — real, but a rounding error against either model
Residual internal teamRetained in both modelsRetained in both models

Table 1 — Model inputs. Both sides carry a residual internal team; neither model eliminates it.

The customer-side coordination line is omitted from most published TCO comparisons and is consistently material — typically 0.5 to 1.5 internal FTE depending on contract complexity.

3. Cumulative cost over 36 months

The chart below is the core of the model, indexed so that 100 equals one month of baseline MSP fees at 500 managed units. Indexing rather than quoting currency keeps the model portable across markets and contract structures; substitute your own baseline to get absolute figures.

CUMULATIVE 3-YEAR COST — 500 MANAGED UNITS (INDEXED)01.5k3k4.5kM0M6M12M18M24M30M36BREAK-EVEN · MONTH 8MSP · effort-pricedAAQUILIX · unit-pricedShaded area = cumulative avoided cost. At month 36 the gap is 2,090 index points — a 48% reduction in three-year TCO.Both curves include the customer-side effort each model demands. Indexed so 100 = one month of baseline MSP fees.
Figure 1 — Cumulative three-year cost at 500 managed units. The platform curve starts higher because implementation and year-one policy authorship land up front, then flattens because incident growth carries no marginal cost.

Three features of this chart matter more than the endpoint. First, the platform curve begins above zero: implementation and policy authorship are real front-loaded costs and pretending otherwise destroys credibility with a CFO. Second, the MSP curve steepens rather than staying linear, because overage compounds and renewals reset the baseline upward. Third, break-even lands at month 8 — inside a typical three-year contract term, but outside a twelve-month payback test, which is a genuine objection and should be stated as one.

Month 8
break-even at 500 units
2,090
index points of avoided cost by month 36
48%
lower cumulative spend at month 36
0
marginal cost per additional incident

4. Where the money actually goes

WHERE THE THREE-YEAR SPEND ACTUALLY GOESTraditional MSP3-year TCO = 10058148128100AAQUILIXsame estate, same SLA30158962 · −38%Vendor / platform feeVolume-driven cost (overage)Transition & implementationCustomer-side coordinationResidual internal teamThe customer-side coordination line is the one most vendor TCO models omit — and it is 12% of MSP spend.
Figure 2 — Three-year TCO composition, same estate and same SLA. The volume-driven component falling to zero is the structural point of the whole model.

The platform model does not win by being a cheaper version of the same thing. It wins by deleting one cost category entirely and shrinking a second. Volume-driven overage goes to zero because pricing is decoupled from incident count. Customer-side coordination falls from 12 to 8 index points because policy review is a scheduled quarterly activity rather than a continuous escalation relationship.

Note that transition and implementation is higher on the platform side, not lower — 15 against 8. Policy authorship in year one is genuine engineering effort. The difference is that it is capitalised knowledge that stays with the customer, whereas MSP onboarding cost is re-incurred at every contract boundary and the knowledge leaves with the vendor.

5. At 100, 500 and 2,000 managed units

100 units500 units2,000 units
3-year MSP TCO (indexed)100100100
3-year platform TCO (indexed)946241
Break-evenMonth 26Month 8Month 5
Autonomy at day 9058–64%68–72%70–75%
Time to steady-state autonomy~120 days~90 days~55 days
Policy authorship effort, year one3–5 eng-weeks4–8 eng-weeks6–10 eng-weeks
Commercial verdictClose — decide on consistency, not priceDecisiveDecisive

Table 2 — Three-year comparison by estate size. Each column is indexed against its own MSP baseline; columns are not comparable to each other in absolute terms.

At 100 units — where the MSP genuinely wins

At small scale the two models are close over three years, and the MSP often looks cheaper in year one because implementation is amortised across a smaller base. The platform case at this size does not rest on price. It rests on consistency: a small MSP allocation means one or two named engineers, and service quality tracks those individuals directly — including when they leave.

If your estate is highly heterogeneous, poorly documented, and generates low incident volume, an MSP is the better answer. There is not enough repetition for policy authorship to pay back. We tell prospects this directly, and it costs us deals, which is fine.

At 500 units — the decisive band

This is where the models diverge. Incident volume at 500 units typically requires meaningful shift coverage, and overage begins to bite in bad quarters — precisely the quarters where budget is already under pressure. Meanwhile the platform cost is flat, and the policy library built in year one continues paying back in years two and three at no incremental cost.

At 2,000 units — compounding economics

At this scale MSP cost is dominated by the coverage model: you are effectively funding a dedicated team with shift rotation plus escalation tiers. The platform benefits from the strongest form of its own economics — high incident volume produces a rich resolution history quickly, autonomy climbs faster, and the marginal cost of covering the 2,001st unit is close to zero.

The realistic end state at this scale is not “no MSP.” It is a substantially smaller MSP engagement scoped to specialist domains, with the repetitive operational band handled by the platform.

6. The effort ledger — hours, not just rupees

TCO captures spend. It does not capture what the organisation gets back, which for most CIOs is the more interesting number. A mid-size NOC covering 400–800 managed units receives 1,000–1,500 alerts per day, of which approximately 70% require no action. At a median 6.4 minutes of triage per actionless alert — measured from ticket timestamps, not self-report — that is roughly 3.9 engineer-FTE per year spent confirming that nothing was wrong.

HOW A SENIOR SRE’S WEEK IS SPENTBeforebaseline measurement34%21%18%9%18%After 90 daysL1 + L2 classes promoted11%19%27%37%L1 alert triageL2 remediationChange & releaseProblem managementEngineering projectsHeadcount is unchanged. What moves is the 55% of the week that was previously spent on work with a written runbook.
Figure 3 — Where a senior engineer’s week goes, before and at day 90. The recovered capacity is the return that does not appear anywhere in the TCO comparison.

Two ways to value this, and finance teams differ on which they accept:

  • Cost-avoidance basis — value the recovered hours at loaded salary cost. Conservative, easy to defend, and typically adds 30–45% to the modelled saving at 500 units.
  • Opportunity basis — value the recovered hours at the marginal output of the projects they enable. Much larger, much harder to defend in an audit, and we recommend presenting it as a sensitivity rather than a base case.

7. Worked industry examples

8. Costs both sides forget

  1. Transition cost at MSP contract boundaries — knowledge transfer, re-documentation, and a measurable quality dip for one to two quarters. Estates that change MSP every three years pay this repeatedly.
  2. Policy authorship effort in year one for the platform model — 4 to 8 engineer-weeks for a mid-size estate, from people who are already busy.
  3. Infrastructure for on-premise inference — real, but usually a rounding error against either model.
  4. The consistency premium: an engineer following a runbook at 3am has a measurably different error rate from the same runbook executed deterministically. This is difficult to price and it is not zero.
  5. Re-attestation overhead in the platform model — roughly 0.2 FTE steady state, and it does not go away.
  6. Escalation coordination in the MSP model — the hours your senior people spend explaining your own estate to someone else’s.

9. Sensitivity — when this model breaks

A model that cannot say when it is wrong is marketing. These are the conditions under which the conclusions above do not hold:

  • Incident volume well below the estate-size norm. If your 500-unit estate generates 200-unit incident volume, the effort-priced model is cheaper and stays cheaper.
  • Highly heterogeneous estates with no standardisation. Policy authorship effort can double or triple, pushing break-even past month 24.
  • No written runbooks anywhere. The first months are documentation work, not automation work, and should be budgeted as such.
  • Imminent platform migration. Authoring policies against an estate you are about to replace is wasted effort; wait.
  • Regulatory posture that prohibits automated execution entirely. The assisted-stage value is still real, but the TCO case is materially weaker.

The right question is not which model is cheaper. It is which model still works when incident volume doubles unexpectedly.

Rajesh Nair, VP Sales

10. Building this for your own board

  1. Pull 12 months of MSP invoices and separate base fee from overage. The overage line is your volume risk, quantified.
  2. Ask your MSP what share of tickets they closed with a standard runbook and no deviation. That number is the portion of your spend going to work a policy engine handles. If they cannot answer, that is its own finding.
  3. Estimate vendor-management FTE honestly — count service reviews, escalations, QBRs, and dispute handling. Most estates are surprised.
  4. Run the 90-day alert study from the State of Enterprise AIOps paper to get your own autonomy ceiling before assuming ours.
  5. Model three volume scenarios, not one: flat, +30%, and a peak-event year. The platform case strengthens in every scenario above flat, which is the argument that lands with a CFO.
  6. Present break-even and curves, not a single TCO number. A single number invites a debate about assumptions; a crossover point invites a decision.

References & further reading

  1. ITIL 4 — Service Financial Management and Supplier Management practices AXELOS / PeopleCert
  2. Site Reliability Engineering — Chapter 5, “Eliminating Toil” Beyer, Jones, Petoff & Murphy, O’Reilly Media, 2016
  3. The Site Reliability Workbook — Chapter 6, “Eliminating Toil” in practice Beyer, Murphy, Rensin, Kawahara & Thorne, O’Reilly Media, 2018
  4. Annual Outage Analysis Uptime Institute, published annually
  5. AAQUILIX deployment cost model, 14 estatesInternal — indexed against contracted MSP baselines, 2024–2026
Data provenance

Quantitative figures in this paper are measured from AAQUILIX production deployments and anonymised customer estates unless a third-party source is cited inline. Standards and frameworks referenced are cited for control alignment, not as the source of our measurements.

ROI Analysis: Policy-First AI vs. Traditional NOC Models · v1.4 · published 4 May 2026. Questions or a challenge to any figure in this paper are genuinely welcome — get in touch.

See it on your estate

Run AAQUILIX in observe mode for 7 days.

No execution, no risk — the platform diagnoses your real incidents and shows you exactly what it would have done.

Book a demoSubscribe to the briefing
More research

The rest of the library.