Most vendor ROI models are unusable in a board pack because they compare a fully-loaded incumbent against a stripped-down challenger. This one includes the customer-side effort each model demands, and it names the conditions under which we lose.
1. Why the usual comparison is wrong
Traditional MSP contracts price against effort: FTE allocation, shift coverage, and ticket bands with overage above them. As incident volume grows, cost grows — either directly through overage or indirectly at renewal, when the run-rate becomes the new baseline.
Unit-based platform pricing charges per managed unit regardless of how many incidents those units generate. A bad month costs the same as a good one. This is the single biggest structural difference between the models, and it is not primarily a price difference — it is a transfer of volume risk from the customer to the platform.
Any comparison that ignores this produces a number that is correct only for the volume assumption it was built on. The right way to present this to a board is not a single TCO figure but a pair of curves and a break-even point, which is what this report produces.
2. Model inputs — what is included on each side
To be fair to both models, the comparison includes every line either side would reasonably incur, including the ones that are absorbed internally and never invoiced.
| Cost line | MSP model | Platform model |
|---|---|---|
| Contracted / subscription fee | Base fee by FTE allocation and shift coverage | Per managed unit, per year |
| Volume-driven cost | Ticket overage at historical rates | None — flat regardless of incident volume |
| Transition / implementation | Onboarding at contract start and at every boundary | One-time implementation and integration |
| Knowledge capture | Absorbed by the MSP, lost at contract boundary | Policy authorship: 4–8 engineer-weeks in year one, retained |
| Customer-side coordination | Vendor management, service reviews, escalation, SLA disputes: 0.5–1.5 FTE | Policy review and re-attestation: ~0.2 FTE steady state |
| Infrastructure | None | On-premise inference nodes — real, but a rounding error against either model |
| Residual internal team | Retained in both models | Retained in both models |
Table 1 — Model inputs. Both sides carry a residual internal team; neither model eliminates it.
The customer-side coordination line is omitted from most published TCO comparisons and is consistently material — typically 0.5 to 1.5 internal FTE depending on contract complexity.
3. Cumulative cost over 36 months
The chart below is the core of the model, indexed so that 100 equals one month of baseline MSP fees at 500 managed units. Indexing rather than quoting currency keeps the model portable across markets and contract structures; substitute your own baseline to get absolute figures.
Three features of this chart matter more than the endpoint. First, the platform curve begins above zero: implementation and policy authorship are real front-loaded costs and pretending otherwise destroys credibility with a CFO. Second, the MSP curve steepens rather than staying linear, because overage compounds and renewals reset the baseline upward. Third, break-even lands at month 8 — inside a typical three-year contract term, but outside a twelve-month payback test, which is a genuine objection and should be stated as one.
4. Where the money actually goes
The platform model does not win by being a cheaper version of the same thing. It wins by deleting one cost category entirely and shrinking a second. Volume-driven overage goes to zero because pricing is decoupled from incident count. Customer-side coordination falls from 12 to 8 index points because policy review is a scheduled quarterly activity rather than a continuous escalation relationship.
Note that transition and implementation is higher on the platform side, not lower — 15 against 8. Policy authorship in year one is genuine engineering effort. The difference is that it is capitalised knowledge that stays with the customer, whereas MSP onboarding cost is re-incurred at every contract boundary and the knowledge leaves with the vendor.
5. At 100, 500 and 2,000 managed units
| 100 units | 500 units | 2,000 units | |
|---|---|---|---|
| 3-year MSP TCO (indexed) | 100 | 100 | 100 |
| 3-year platform TCO (indexed) | 94 | 62 | 41 |
| Break-even | Month 26 | Month 8 | Month 5 |
| Autonomy at day 90 | 58–64% | 68–72% | 70–75% |
| Time to steady-state autonomy | ~120 days | ~90 days | ~55 days |
| Policy authorship effort, year one | 3–5 eng-weeks | 4–8 eng-weeks | 6–10 eng-weeks |
| Commercial verdict | Close — decide on consistency, not price | Decisive | Decisive |
Table 2 — Three-year comparison by estate size. Each column is indexed against its own MSP baseline; columns are not comparable to each other in absolute terms.
At 100 units — where the MSP genuinely wins
At small scale the two models are close over three years, and the MSP often looks cheaper in year one because implementation is amortised across a smaller base. The platform case at this size does not rest on price. It rests on consistency: a small MSP allocation means one or two named engineers, and service quality tracks those individuals directly — including when they leave.
If your estate is highly heterogeneous, poorly documented, and generates low incident volume, an MSP is the better answer. There is not enough repetition for policy authorship to pay back. We tell prospects this directly, and it costs us deals, which is fine.
At 500 units — the decisive band
This is where the models diverge. Incident volume at 500 units typically requires meaningful shift coverage, and overage begins to bite in bad quarters — precisely the quarters where budget is already under pressure. Meanwhile the platform cost is flat, and the policy library built in year one continues paying back in years two and three at no incremental cost.
At 2,000 units — compounding economics
At this scale MSP cost is dominated by the coverage model: you are effectively funding a dedicated team with shift rotation plus escalation tiers. The platform benefits from the strongest form of its own economics — high incident volume produces a rich resolution history quickly, autonomy climbs faster, and the marginal cost of covering the 2,001st unit is close to zero.
The realistic end state at this scale is not “no MSP.” It is a substantially smaller MSP engagement scoped to specialist domains, with the repetitive operational band handled by the platform.
6. The effort ledger — hours, not just rupees
TCO captures spend. It does not capture what the organisation gets back, which for most CIOs is the more interesting number. A mid-size NOC covering 400–800 managed units receives 1,000–1,500 alerts per day, of which approximately 70% require no action. At a median 6.4 minutes of triage per actionless alert — measured from ticket timestamps, not self-report — that is roughly 3.9 engineer-FTE per year spent confirming that nothing was wrong.
Two ways to value this, and finance teams differ on which they accept:
- Cost-avoidance basis — value the recovered hours at loaded salary cost. Conservative, easy to defend, and typically adds 30–45% to the modelled saving at 500 units.
- Opportunity basis — value the recovered hours at the marginal output of the projects they enable. Much larger, much harder to defend in an audit, and we recommend presenting it as a sensitivity rather than a base case.
7. Worked industry examples
8. Costs both sides forget
- Transition cost at MSP contract boundaries — knowledge transfer, re-documentation, and a measurable quality dip for one to two quarters. Estates that change MSP every three years pay this repeatedly.
- Policy authorship effort in year one for the platform model — 4 to 8 engineer-weeks for a mid-size estate, from people who are already busy.
- Infrastructure for on-premise inference — real, but usually a rounding error against either model.
- The consistency premium: an engineer following a runbook at 3am has a measurably different error rate from the same runbook executed deterministically. This is difficult to price and it is not zero.
- Re-attestation overhead in the platform model — roughly 0.2 FTE steady state, and it does not go away.
- Escalation coordination in the MSP model — the hours your senior people spend explaining your own estate to someone else’s.
9. Sensitivity — when this model breaks
A model that cannot say when it is wrong is marketing. These are the conditions under which the conclusions above do not hold:
- Incident volume well below the estate-size norm. If your 500-unit estate generates 200-unit incident volume, the effort-priced model is cheaper and stays cheaper.
- Highly heterogeneous estates with no standardisation. Policy authorship effort can double or triple, pushing break-even past month 24.
- No written runbooks anywhere. The first months are documentation work, not automation work, and should be budgeted as such.
- Imminent platform migration. Authoring policies against an estate you are about to replace is wasted effort; wait.
- Regulatory posture that prohibits automated execution entirely. The assisted-stage value is still real, but the TCO case is materially weaker.
“The right question is not which model is cheaper. It is which model still works when incident volume doubles unexpectedly.”
Rajesh Nair, VP Sales
10. Building this for your own board
- Pull 12 months of MSP invoices and separate base fee from overage. The overage line is your volume risk, quantified.
- Ask your MSP what share of tickets they closed with a standard runbook and no deviation. That number is the portion of your spend going to work a policy engine handles. If they cannot answer, that is its own finding.
- Estimate vendor-management FTE honestly — count service reviews, escalations, QBRs, and dispute handling. Most estates are surprised.
- Run the 90-day alert study from the State of Enterprise AIOps paper to get your own autonomy ceiling before assuming ours.
- Model three volume scenarios, not one: flat, +30%, and a peak-event year. The platform case strengthens in every scenario above flat, which is the argument that lands with a CFO.
- Present break-even and curves, not a single TCO number. A single number invites a debate about assumptions; a crossover point invites a decision.
References & further reading
- ITIL 4 — Service Financial Management and Supplier Management practices AXELOS / PeopleCert
- Site Reliability Engineering — Chapter 5, “Eliminating Toil” Beyer, Jones, Petoff & Murphy, O’Reilly Media, 2016
- The Site Reliability Workbook — Chapter 6, “Eliminating Toil” in practice Beyer, Murphy, Rensin, Kawahara & Thorne, O’Reilly Media, 2018
- Annual Outage Analysis Uptime Institute, published annually
- AAQUILIX deployment cost model, 14 estatesInternal — indexed against contracted MSP baselines, 2024–2026
Quantitative figures in this paper are measured from AAQUILIX production deployments and anonymised customer estates unless a third-party source is cited inline. Standards and frameworks referenced are cited for control alignment, not as the source of our measurements.
ROI Analysis: Policy-First AI vs. Traditional NOC Models · v1.4 · published 4 May 2026. Questions or a challenge to any figure in this paper are genuinely welcome — get in touch.