Set The Standard
Every department wanted AI. Nobody owned it. Here is how one team's build became a governed platform the whole company runs on: one gateway, a written charter, engineering standards, and a certified enablement lead in every department.
The gold rush problem
The situation
The pressure to adopt AI landed on every department at once, and several had already started. What they were building on was a LiteLLM deployment nobody relied on, with no engineering standards, no approval path, and no way to say what any of it cost.
That is how shadow tooling starts. Teams buy their own keys, wire them into whatever they already use, and the company ends up running a dozen small AI programs, none of them measurable and none of them safe to put in front of an auditor.
I had just finished proving the other half of this on my own team, where AI absorbed the repetitive load and a person still owned every customer outcome. This job was different. It meant turning one team's build into something every team could use without each of them writing its own rules.
Starting state
The approach
A charter the company runs on, not a strategy deck
Strategy that lives in a deck does not survive contact with a roadmap. So I wrote the program down as three pillars, each with its own mechanics, and everything since has been delivered against them. Govern sets the rules and the limits. Build turns ideas into tools somebody owns. Train makes sure the capability outlives any single team, including mine.
The company's AI program
no longer a collection of separate experiments
Govern
- one gateway
- budgets per team
- sso and logging
- approval policy
Build
- prioritized roadmap
- time-boxed pocs
- owner and budget
- measure or retire
Train
- curriculum
- certified leads
- quarterly rebuild
- open floor
three pillars: the structure every AI decision rests on
Writing it down changed the conversations. A team asking for a tool now knows what it has to bring, and a team asking to skip the approval path is arguing with a document rather than with me.
Where this came from
The charter grew out of one team's build, where I argued AI could not do the work, then proved where it could: heavy triage absorbed, a person still owning every customer outcome.
Read The Pivot Point →One gateway, replaced rather than inherited, then taken to production grade
The gateway I was handed was a poor LiteLLM deployment. Taking it over would have meant owning the state it was in, so I went looking for a replacement instead: a scan of the market, outreach to vendors and sellers, demos I requested and ran, and the procurement process that got the winner bought. That was Bifrost.
I deployed it on AWS, provisioned with Terraform, across two environments, staging and production, then took it from first rollout to production grade. That meant redundancy, so a node can drop without taking the gateway down; SSO in front, so access is tied to an identity rather than a shared key; secrets brokered rather than handed out, so almost no human ever touches a key; documentation the team can actually run from; and per-team usage tracking that turns spend into a report finance can read. Models from every major provider are available through it: Anthropic, OpenAI, Google, Mistral, and Meta. Consolidating onto it brought total spend down while putting roughly $40,000 a month of AI usage behind one set of budgets, with access, tracking, and cost visibility in one place.
Because every internal call routes through one gateway, it is where the standard gets set rather than just where the traffic goes: teams build on shared SDKs instead of raw keys, every call is observable and attributed to a cost center, and the choices that used to be ad hoc, which model to use, how prompts are managed, how a change is evaluated before it ships, get settled once and reused. It is also the safe place to try something, which is why nobody needs to stand up a shadow stack to experiment.
The same discipline, on the vendor side
Running a real search, holding vendors to a demo, and closing procurement properly is the same instinct that rebuilt vendor governance and took the penalties off the table.
Read Hold The Line →The rules: engineering standards, legal review, and a catalog per department
Standards cover the decisions every team would otherwise make separately and inconsistently: which model to select for a job, how prompts are managed and versioned, the security guardrails an integration has to carry, how a change is evaluated before it ships, and the operational practices that keep it running once it does.
The catalog side is less interesting and matters just as much. I took the legal and privacy notices through review, expanded the approved model list, and adapted what each department gets to what that department actually handles, including how personally identifiable information is treated. The test I held it to was whether someone who will never write code could use it without needing me in the room.
How an idea becomes a tool, or gets retired
Ideas arrive from everywhere, which is a good problem and an expensive one. The pipeline exists so enthusiasm turns into something owned rather than something abandoned, and so the answer to a weak idea is a process rather than my opinion. Six of the first ten initiatives cleared it into production, which is about the hit rate I want: high enough to be worth running, low enough that the bar is real.
1
Prioritize
Department enablement leads bring the demand into one roadmap.
2
Prove
Time-boxed proof of concept, success criteria written before work starts.
3
Deliver
What clears the bar gets an owner, a budget, and training.
4
Measure
Every delivered tool reports usage and cost. What nobody uses is retired.
The same four steps are how new models, providers, and emerging tools get assessed, which is the part that keeps the platform current without letting it churn. Anything new is evaluated against a real use case before it reaches production, so the decision rests on a measurement rather than on whoever read the announcement first.
Training the company, then certifying the people who carry it
A platform nobody knows how to use is shelfware with a budget line. I own what the company teaches about internal AI and the policies around it, which means the curriculum, how it is delivered, and the decision about what changes in it each quarter.
The program covers roughly 500 people, which it would not if I delivered all of it myself, so every department has an enablement lead who is trained, certified against provider courses, and then carries the work locally, along with the practices for using AI responsibly. The curriculum gets rebuilt each quarter from what usage actually shows rather than from what I assumed three months earlier, and a recurring open floor gives teams somewhere to show what they built and what did not work. Those sessions started quarterly and now run monthly, which is the clearest adoption signal I have.
Deploying agents, not one-off tools
What the pipeline produces is agents that run in the systems people already use, not demos that need a champion to stay alive. Each one came through the same four steps, has a named owner, and reports its own usage and cost. None of them is mine to maintain.
That inheritance is what makes deploying the next one cheap. A new agent needs no key of its own, no separate budget conversation, and no bespoke security review, because the rails it lands on already carry all three. The work left over is the part that is actually specific to the job.
Employee experience agents
Agents inside the tools people already work in, so the platform turns up in the working day rather than in a training deck.
Anti-piracy agent
Flags license abuse from the IP patterns of cloud instances running the product, so revenue leakage surfaces before a renewal conversation rather than after it.
Data warehouse agents
Adoption per team and per use case, demand signals, usage tied to business results, and projected spend forecast as adoption grows.
Across the program, delivery has spanned employee experience, customer retention, quality assurance, security, and engineering productivity. The list grows, which is the point of having a pipeline rather than a wishlist. The tools that proved the model in the first place, the triage engine, the intake gate, and the backlog agents, belong to the team where this started.
Adoption, cost, and return, reported rather than asserted
A governance model is only worth as much as its numbers. Every initiative clears a lightweight approval before it can spend, budgets are set per person and per team in the gateway with low-cost models as the default, and cost is attributed to the team that drives it rather than pooled where nobody feels it.
On top of that sits the reporting: platform usage, business impact, operational efficiency, and return on investment, with projected spend forecast as adoption grows. It means a new use case argues for itself with evidence, and the finance conversation stops being a surprise that arrives a quarter late. The same portfolio view, roadmap, spend, return, and risk, is what I take to executives and investors, so what the program is delivering and what it needs next are visible rather than asserted.
The same discipline, on the cost side
Making spend legible before anyone has to defend it is the same method that turned overspending into evidence-backed cost cases and saved hundreds of thousands without cutting headcount.
Read Stop The Bleed →The results
Employees enabled
~500
Every department onboarded onto the same governed platform, technical and non-technical alike.
Engineering workdays avoided
320
Work that came off the board rather than being added to somebody's plate.
Targeted workflows
2-4x
Faster on the specific engineering workflows the program set out to accelerate.
AI usage governed
$40k/mo
Centralized access, budgets, and usage tracking, with total spend down after the migration.
Initiatives to production
6 of 10
A framework that promotes what clears the bar and retires what does not.
Direct reports
0
Every program delivered through matrixed teams across five departments.
The program holds because the governed path is also the convenient one. Teams route through the gateway because it is faster than getting their own key, and usage gets reported because the tooling does it for them rather than because someone chases it. Nobody has needed to stand up a shadow stack, and the tools nobody opened have come back off the books.
None of it depends on me being in the room. The standards are written, the curriculum is owned, and every department has somebody certified to carry it, which is the only version of enablement that survives a reorganization.
Silicon Valley edition
What made it hard
Almost all of the delivery ran through people who did not report to me. Matrixed resources have their own priorities and their own managers, so the work moved at the speed of whatever I could make genuinely worth their time. That is a different skill from running a team, and I was better at the second one when I started.
Governing enthusiasm is harder than governing resistance. Telling a team its idea needs a success criterion before it gets a budget is not a popular message, and the approval path only survived because I kept it fast enough that going around it was not worth the effort.
The last risk was becoming the bottleneck I had just removed. Owning the curriculum, the standards, and the roadmap concentrates a lot in one person, which is exactly why certifying a lead in every department mattered more than any tool on this page.
Under pressure to adopt AI without losing control of it?
Let's talk about what governed AI enablement actually takes.