Hold The Line
Vendors underdelivering. Failed changes causing outages. Financial penalties stacking up. A managed services contract winding down with no governance to hold anyone accountable. Here is how I rebuilt vendor relationships from the ground up and kept the line clean.
The Suits problem
The situation
Vendor relationships managed by goodwill rather than governance. Performance issues raised in emails and forgotten. SLAs referenced in contracts but never enforced. Renewals happening automatically because nobody had built the case to challenge them.
One engagement had an unusual dimension: I was employed by the outgoing managed services vendor while simultaneously responsible for managing the transition to an in-house model. In effect, I was managing my own employer out of the contract.
In another, a multi-supplier environment running critical infrastructure needed real governance: multiple vendors to keep coordinated, none allowed to become a dependency, and a client that did not want to be the one doing the coordinating. That seat became mine.
Starting state
01Informal SLAs
Goodwill, not governance. Issues raised in emails and forgotten.
02Structured governance
SLAs enforced. Vendors coordinated. No single point of dependency.
03Clean transition
Managed my own employer out of the contract. No disruption. No dependency left behind.
The approach
Managing your own employer out
The dual-role transition
Being employed by the vendor you are managing out is professionally unusual. The only way it stays clean is complete transparency on both sides. I was accountable to my employer while also building the case for the client to operate without them. In practice that meant separating the waters in every meeting and managing egos on both sides: colleagues whose roles were winding down, and a client deciding how much to trust someone paid by the company on its way out. I could not afford to be anyone's advocate. The moment either side doubted that, the transition would have stalled.
- ServiceNow
- Workflows
- Ownership
- Governance
The ServiceNow rebuild became the way to scope everything before any decision was made. I dug through the CMDB, cross-checked documentation against monitoring data, and when the data was ambiguous I walked to the department and asked, mapping what was actually running, who really owned it, and what was politically pending. Only then was the operating model rebuilt: change templates, automated workflows, and real-time visibility teams had not had before. Process ownership was redefined on evidence, not assumption.
The transition completed without service disruption. My employer knew what I was doing and why throughout.
The team that made vendors optional
Governance kept partners honest; an internal team that could absorb the work is what made any single vendor expendable. Here is how that team was built from nothing.
Read Build What Lasts →Agile-ITIL hybrid and formal process ownership
This grew out of the same dual-role transition: ServiceNow gave us the platform, but the operating rhythm around it had to be rebuilt too. Governance without measurability is just procedure. Agile gave the rhythm, the culture of continuous improvement, and the discipline to break down the departmental silos that had kept teams working in parallel without talking. ITIL gave the structure and accountability baselines. The two were not in tension: Agile defines how teams improve, ITIL defines what they are accountable for improving. In practice, the hybrid cut the bureaucratic approval stages that added delay without adding control. Pre-templated change requests alone removed two to three minutes of planning per ticket, across roughly 125 tickets a week.
Agile practices
Rhythm, culture, improvement
Kanban
Visualised flow, limited WIP, blockers visible before incidents.
Backlog management
Prioritized and reviewed so work that mattered got done.
Standups
Daily alignment. Replaced email chains that let issues drift.
Retrospectives
Operational and executive level. Structured reflection at every layer.
Frequent delivery
Incremental work, no over-engineering, adjusted from data.
ITIL process ownership
Structure, accountability, baselines
Incident management
Response structure, escalation paths, post-incident review cadence.
Change management
Process Owner and CAB Manager.
Problem management
Root cause tracking to prevent recurring incidents.
Availability management
SLA thresholds defined, monitored, enforced against vendors.
Capacity management
Forward planning tied to real usage data, not vendor estimates.
Agile defines how teams improve. ITIL defines what they are accountable for improving.
The result was a process teams actually used instead of working around. Lighter, clearer gates let changes move without ceremony, and groups that had run in parallel started solving problems together instead of escalating past each other.
Rebuilding change governance to stop outages and eliminate penalties
Failed changes were causing outages and triggering financial penalties on both sides. When the client took a hit, the provider took one too. The exposure ran in both directions and compounded fast. For two clients, the downstream impact was not abstract:
Electricity distribution: a failed change could mean a household losing power.
Payment processing: a failed change could mean someone unable to pay at a grocery store.
Logistics: a failed change could mean a truck sitting at a warehouse because the system could not legally authorize the dispatch.
These were not IT incidents. They were real consequences for real people. The root causes were consistent:
No pre-change testing standard.
Unclear ownership at the point of failure.
No post-implementation review feeding back into the next cycle.
As CAB Manager and Process Owner, I rebuilt the governance model with mandatory pre-checks, rollback procedures, defined approval gates, and post-change reviews wired directly into Problem Management. The change advisory board got actual authority, not just a calendar invite.
testing
gates
authority
procedures
clarity
review
−75%
Failed changes
~200/week
−10%
Critical outages
4 quarters
$0
Penalties paid
one renegotiated into work
Failed changes dropped 75% across roughly 200 changes per week. Critical outages dropped 10% quarter over quarter for four consecutive quarters. Financial penalties, tied to strict deployment and P1/P2 incident SLAs, went to zero. Where one loomed, I renegotiated it into bundled project hours against the vendor's pending work, turning a punitive charge into delivered value.
The cost angle of the same governance work
The $150k infrastructure decommission and $27k/month contractor removal did not happen in isolation. Here is how the evidence was built and the cases were approved.
Read Stop The Bleed →Multi-supplier governance: no single vendor indispensable
Managing multiple vendors across critical infrastructure meant keeping everyone accountable without letting any single supplier become indispensable. Process ownership, escalation paths, and performance review cadences were defined for each relationship under ITIL and COBIT frameworks. The client did not want to sit in the middle of vendor coordination, so I took that seat: chasing responses, raising flags when commitments slipped, and treating each vendor as a strategic partner rather than a body to manage.
SLAs were renegotiated where they were not fit for purpose. Performance was tracked against agreed metrics, not just reported. Vendors that underperformed were challenged formally, with documentation and a defined consequence.
- Vendor A
- Vendor B
- Vendor C
- Vendor D
same metrics · same review cycle · no single point of dependency
Defined per relationship
Process ownership, escalation paths, review cadence.
Tracked, not just reported
Performance measured against agreed metrics; SLAs renegotiated where unfit.
Held under pressure
Some pushed back. Applied consistently to every vendor, every breach.
Some pushed back when I flagged underperformance. A few thought they could wait me out. They couldn't. The governance held because it was applied consistently, to every vendor, every review cycle, every SLA breach. The enforcement structure was designed so that pressure in the room could not override the process on paper.
Want to see how this works with development partners?
Same governance logic applied to a fragmented multi-product environment with four ticket systems and no shared standards.
Read Clear The Fog →A framework that makes accountability structural
Four service appendices, one per partner type. Each defines what is expected, how performance is measured, and what happens when it is not. The mechanics are concrete: monthly written status reports due 48 hours before each review, KPI thresholds for resolution time, reopen rate, and throughput, and a staged path from warning to corrective action plan to replacement when targets stay missed. Escalation clocks define who responds and how fast. Renewals are tied to performance, and the model was later proposed for expansion across all technical roles. The system carries the accountability so the relationship does not have to.
The partner relationship
no longer carrying the accountability itself
Development
- expectations
- metrics
- consequences
QA
- expectations
- metrics
- consequences
DevOps
- expectations
- metrics
- consequences
Design
- expectations
- metrics
- consequences
four service appendices: the structure everything else rests on
Once the standard lived in the appendix instead of the relationship, holding a vendor to it stopped being a confrontation and became routine. The same document onboarded a replacement just as cleanly, so no single partner was ever load-bearing.
The same standard, enforced by software
Governance here meant a published standard no vendor could slip past. The AI triage build applies the identical idea to intake: nothing reaches an engineer until it meets the bar.
Read The Pivot Point →The results
Failed Changes
-75%
Across 200+ changes per week.
Critical Outages
-10% QoQ
4 consecutive quarters.
Financial Penalties
None paid
One charge renegotiated into delivered work.
Vendors Under Governance
5
Held to SLAs and review cadence.
Dual-Role Transition
22 months
Own employer out, no disruption.
CAB Meetings Chaired
190
Local and global, up to 40 seats.
None of it depended on me staying in the room. The standard lived in the contract rather than any one relationship, so it held as partners changed and even as my own employer was transitioned out without disruption. Failed changes, outages, and penalties all moved the right way and stayed there.
Severance edition
What made it hard
The dual-role transition required a level of professional separation that most people avoid. You are building the case for the client to operate without your employer. That only works if you are honest about it with everyone and committed to the outcome regardless of where your paycheck comes from.
Getting engineers to treat governance gates as protection rather than bureaucracy took time and evidence. The first two change cycles were a negotiation. After that, the data from the post-implementation reviews did the argument for me.
The infrastructure decommission cases required political courage as much as technical evidence. Legacy systems survive because nobody wants ownership of the risk if something breaks during removal. The CMDB-evidenced approach removed that ambiguity. If the system had no active consumers and no justified cost, the case was watertight.
Vendor accountability is only durable if it is applied consistently. The moment you make an exception, you have reset the standard. That means having the same conversation more than once, with the same outcome, until the vendor accepts that the governance applies every time.
Vendors not delivering what they promised?
Let's talk about what good governance actually looks like.