← Back to What I Do
Leadership Team Building

Build What Lasts

I built a middleware team from nothing, watched the early model fail, and fixed it. 30 engineers across APAC, led from Portugal, running 24/7/365 with no inherited playbook. Then a ransomware attack hit and recovery fell to us.

Work with me Request my CV

Ted Lasso, but for middleware

The situation

A global energy provider running critical middleware infrastructure. I was brought in to build and lead a new team, a distributed APAC team run from Portugal, as the client moved off a vendor they had been with for 30 years, which made the transition painful. There was no documentation or runbooks in English, so I had to review everything that existed and make sure it was actually executable. The early model I ran failed fast: the engineers I assembled were specialists in silos.

Incidents happened daily. Overtime was constant. When any middleware engineer was unavailable, their incidents sat open until they came back online. The team I had built was the worst performing in the engagement, and I had built it.

Then a ransomware attack hit. 40 products affected. Recovery fell to us.

Starting state

Team performance Worst in the engagement
Incident frequency Daily
Cross-team coverage None
Documentation None

No foundation

Siloed engineers. No docs, no runbooks, no shared visibility.

Worst in the engagement

Daily incidents. Overtime constant. And I had built it.

Took ownership

Restructured the model. Cross-trained. Documented everything.

Built to last

Ransomware hit. 40 products affected. Team held the line and recovered.


The approach

Building domain expertise from zero

My background was purely in Windows systems administration, not middleware, so at first I did not even know where to look. We could not do it alone, so we brought in specialist contractors to get the team started, learned alongside them, built the team around their feedback and structure, and turned those contractors into a vehicle for documentation and structured knowledge transfer to the offshore team.

Once the revamped middleware engineers could operate independently, the three specialist contractors were removed, taking $27k/month of external spend off the books. Over time I also replaced one underperforming FTE with the leveraged shared team. The financial case for these decisions, how it was built and approved, is covered separately.

The $27k saving did not come from cutting. It came from building internal capability first and removing the dependency once the evidence was there.

Read Stop The Bleed →

T-shaped cross-training across the APAC team

The early model failed because each middleware engineer owned one area completely and nothing else. Coverage depended on individuals being available. When they weren't, incidents waited.

I redesigned the team around a T-shaped model. Every engineer kept their primary specialty and built working knowledge of adjacent areas. Runbooks were written. Knowledge transfer happened in structured sessions, not improvised during outages.

Before

Siloed specialists

Single ownership. No coverage.

DB Apps Infra Integ

Coverage depends on individuals. When they aren't available, incidents wait.

After

T-shaped team

Shared coverage. Built to scale.

DB Apps Infra Integ

Cross-trained engineers. Depth in one area, working knowledge across others.

24/7 coverage

APAC team, led from Portugal
shift-based, no gaps

Incidents handled

regardless of availability

Overtime −80%

sustained reduction

Stronger team

broader skills, fewer bottlenecks

Incidents got picked up regardless of who was on call. On top of the cross-training, an adaptive ROTA model, built on historical incident and change probability, matched coverage to when issues actually happened, backed by an off-hours incentive scheme. Overtime dropped 80%.

Cross-training scaled what one team could cover. The same leverage instinct, structured workstreams and an internal triage tool, later expanded scope across fragmented development partners without adding a single hire.

Read Clear The Fog →

Ransomware recovery: 3 months, up to 16 hours a day

A ransomware attack took down 40 products. Damage was contained to the Windows OS layer. I coordinated credential resets (middleware agents and service accounts) across affected systems, ran server-by-server damage assessments with my middleware engineer, and restored applications through either clean deployments from validated backups or full rebuilds from scratch.

Normal operations continued in parallel. At peak, 16-hour days sustained for weeks. The team held because ownership was clear, the documentation existed, and the cross-training had already happened.

Continuous monitoring
  1. Detect Identify threats
  2. Analyze Understand impact
  3. Contain Isolate damage
  4. Recover Restore systems
  5. Lessons learned Review & improve

Full recovery

40 products · 3 months

Recovery tasks

4–5 hrs → under 1 hr

No permanent outages

Normal ops continued

Full recovery in 3 months. No permanent outages. Automated recovery tasks that previously took 4 to 5 hours were rebuilt and reduced to under 1 hour, which let us repair disaster recovery and prove business continuity for real.

The recovery held because the structure was already there, not improvised mid-crisis. The same principle ran the vendor side: governance that stayed enforceable because it was built before it was tested, not during the breach.

Read Hold The Line →

On-premises to Azure migration

The ransomware attack accelerated a migration already in motion. With roughly half the infrastructure affected, spinning up replacement VMs in Azure on Windows Server 2019 became the faster path to recovery than restoring everything in place. Availability and disaster recovery were the drivers, not cost.

On-premises

~50% affected by ransomware

LIFT · REBUILD · VALIDATE

Microsoft Azure

Windows Server 2019 VMs

Availability + DR posture on-prem couldn't match

Full stack: installed, configured, validated

Oracle Forms & Reports
IBM WebSphere
Oracle WebLogic
Apache Tomcat & HTTP
IIS & SQL Server Reporting
Informatica PowerCenter
SharePoint
Oracle Billing · FlexNet · ARIS

The cloud environment gave us a recovery posture that on-premises could not match, and running the assessment, recovery, rollback, and DR tests through it doubled as a full stress test.

A cloud-ready, recoverable platform was the groundwork. The same team later built on it directly: a containerized, cloud-ready AI triage layer that turns a prepared ticket into a triaged fix for about four cents.

Read The Pivot Point →

The results

On-call Incidents

-95%

Daily to roughly once a month.

Overtime

-80%

After cross-training rollout.

SLA Compliance

99.5%

Under 4h downtime per month.

Documentation

450

Runbooks and SOPs, from scratch.

Recovery Automation

4h → <1h

Disaster recovery, automated.

Azure Migration

200+ VMs

20 products, dev to prod.

What started as one middleware team became the template for five. After taking the middleware service line from worst to best performing, I was promoted to Operations Service Manager and rolled the same operating model across five teams (Middleware, Oracle & SQL Databases, Batch, SAP Basis, and Application Owners), this time spanning EMEA, LATAM, and APAC.


Band of Brothers edition

What made it hard

The team started poorly because I built it that way. There was no predecessor to blame. The silos, the individual dependencies, and the lack of documentation were consequences of the initial model I put in place and of my own inexperience. Fixing it meant owning that clearly and rebuilding without losing the team's confidence in the direction.

Leading distributed teams across EMEA, LATAM, and APAC meant constant cultural translation. How engineers raise blockers, take feedback, and handle escalation varies a lot from region to region. So to make one operating model land everywhere, I adapted the approach for each region, without ever diluting the standard.

The ransomware recovery compressed all of it. Long days, sustained for months, on systems that had just been damaged, with a distributed team and no precedent to follow. What made it survivable was that the structural work had been done before the attack arrived.

Cutting the contractors was the call with the least margin for error. The business case was straightforward. The timing was not.


Dealing with a team that isn't performing?

Let's talk about what it takes to fix it properly.