Build What Lasts
I built a middleware team from nothing, watched the early model fail, and fixed it. 30 engineers across APAC, led from Portugal, running 24/7/365 with no inherited playbook. Then a ransomware attack hit and recovery fell to us.
Ted Lasso, but for middleware
The situation
A global energy provider running critical middleware infrastructure. I was brought in to build and lead a new team, a distributed APAC team run from Portugal, as the client moved off a vendor they had been with for 30 years, which made the transition painful. There was no documentation or runbooks in English, so I had to review everything that existed and make sure it was actually executable. The early model I ran failed fast: the engineers I assembled were specialists in silos.
Incidents happened daily. Overtime was constant. When any middleware engineer was unavailable, their incidents sat open until they came back online. The team I had built was the worst performing in the engagement, and I had built it.
Then a ransomware attack hit. 40 products affected. Recovery fell to us.
Starting state
No foundation
Siloed engineers. No docs, no runbooks, no shared visibility.
Worst in the engagement
Daily incidents. Overtime constant. And I had built it.
Took ownership
Restructured the model. Cross-trained. Documented everything.
Built to last
Ransomware hit. 40 products affected. Team held the line and recovered.
The approach
Building domain expertise from zero
My background was purely in Windows systems administration, not middleware, so at first I did not even know where to look. We could not do it alone, so we brought in specialist contractors to get the team started, learned alongside them, built the team around their feedback and structure, and turned those contractors into a vehicle for documentation and structured knowledge transfer to the offshore team.
Once the revamped middleware engineers could operate independently, the three specialist contractors were removed, taking $27k/month of external spend off the books. Over time I also replaced one underperforming FTE with the leveraged shared team. The financial case for these decisions, how it was built and approved, is covered separately.
How the cost case was built
The $27k saving did not come from cutting. It came from building internal capability first and removing the dependency once the evidence was there.
Read Stop The Bleed →T-shaped cross-training across the APAC team
The early model failed because each middleware engineer owned one area completely and nothing else. Coverage depended on individuals being available. When they weren't, incidents waited.
I redesigned the team around a T-shaped model. Every engineer kept their primary specialty and built working knowledge of adjacent areas. Runbooks were written. Knowledge transfer happened in structured sessions, not improvised during outages.
Siloed specialists
Single ownership. No coverage.
Coverage depends on individuals. When they aren't available, incidents wait.
T-shaped team
Shared coverage. Built to scale.
Cross-trained engineers. Depth in one area, working knowledge across others.
24/7 coverage
APAC team, led from Portugal
shift-based, no gaps
Incidents handled
regardless of availability
Overtime −80%
sustained reduction
Stronger team
broader skills, fewer bottlenecks
Incidents got picked up regardless of who was on call. On top of the cross-training, an adaptive ROTA model, built on historical incident and change probability, matched coverage to when issues actually happened, backed by an off-hours incentive scheme. Overtime dropped 80%.
More coverage, same team
Cross-training scaled what one team could cover. The same leverage instinct, structured workstreams and an internal triage tool, later expanded scope across fragmented development partners without adding a single hire.
Read Clear The Fog →Ransomware recovery: 3 months, up to 16 hours a day
A ransomware attack took down 40 products. Damage was contained to the Windows OS layer. I coordinated credential resets (middleware agents and service accounts) across affected systems, ran server-by-server damage assessments with my middleware engineer, and restored applications through either clean deployments from validated backups or full rebuilds from scratch.
Normal operations continued in parallel. At peak, 16-hour days sustained for weeks. The team held because ownership was clear, the documentation existed, and the cross-training had already happened.
- Detect Identify threats
- Analyze Understand impact
- Contain Isolate damage
- Recover Restore systems
- Lessons learned Review & improve
Full recovery in 3 months. No permanent outages. Automated recovery tasks that previously took 4 to 5 hours were rebuilt and reduced to under 1 hour, which let us repair disaster recovery and prove business continuity for real.
Structure that holds under pressure
The recovery held because the structure was already there, not improvised mid-crisis. The same principle ran the vendor side: governance that stayed enforceable because it was built before it was tested, not during the breach.
Read Hold The Line →On-premises to Azure migration
The ransomware attack accelerated a migration already in motion. With roughly half the infrastructure affected, spinning up replacement VMs in Azure on Windows Server 2019 became the faster path to recovery than restoring everything in place. Availability and disaster recovery were the drivers, not cost.
On-premises
~50% affected by ransomware
Microsoft Azure
Windows Server 2019 VMs
Availability + DR posture on-prem couldn't match
Full stack: installed, configured, validated
The cloud environment gave us a recovery posture that on-premises could not match, and running the assessment, recovery, rollback, and DR tests through it doubled as a full stress test.
What the team built on the cloud foundation
A cloud-ready, recoverable platform was the groundwork. The same team later built on it directly: a containerized, cloud-ready AI triage layer that turns a prepared ticket into a triaged fix for about four cents.
Read The Pivot Point →The results
On-call Incidents
-95%
Daily to roughly once a month.
Overtime
-80%
After cross-training rollout.
SLA Compliance
99.5%
Under 4h downtime per month.
Documentation
450
Runbooks and SOPs, from scratch.
Recovery Automation
4h → <1h
Disaster recovery, automated.
Azure Migration
200+ VMs
20 products, dev to prod.
What started as one middleware team became the template for five. After taking the middleware service line from worst to best performing, I was promoted to Operations Service Manager and rolled the same operating model across five teams (Middleware, Oracle & SQL Databases, Batch, SAP Basis, and Application Owners), this time spanning EMEA, LATAM, and APAC.
Band of Brothers edition
What made it hard
The team started poorly because I built it that way. There was no predecessor to blame. The silos, the individual dependencies, and the lack of documentation were consequences of the initial model I put in place and of my own inexperience. Fixing it meant owning that clearly and rebuilding without losing the team's confidence in the direction.
Leading distributed teams across EMEA, LATAM, and APAC meant constant cultural translation. How engineers raise blockers, take feedback, and handle escalation varies a lot from region to region. So to make one operating model land everywhere, I adapted the approach for each region, without ever diluting the standard.
The ransomware recovery compressed all of it. Long days, sustained for months, on systems that had just been damaged, with a distributed team and no precedent to follow. What made it survivable was that the structural work had been done before the attack arrived.
Cutting the contractors was the call with the least margin for error. The business case was straightforward. The timing was not.
Dealing with a team that isn't performing?
Let's talk about what it takes to fix it properly.