# Managing the LiteLLM Control Plane at Scale with Git-Driven Automation

# Managing the LiteLLM Control Plane at Scale with Git-Driven Automation

If your engineering teams are adopting generative AI, you are likely relying on proxies like LiteLLM to sit between your internal systems and the dozens of large language models (LLMs) your teams want to use. LiteLLM handles routing, API key management, and billing limits, among other features.

Using UI-driven configurations for this works initially, but at an enterprise scale, it quickly becomes unmanageable.

In this post, I will share the Git-driven architecture I designed for USC's **Trojan AI Gateway**. While we manage our *entire* AI Gateway infrastructure (AWS accounts, AWS resources, and the LiteLLM configuration) with infrastructure-as-code (IaC) to ensure it is auditable, secure, and developer-friendly, this blog post focuses exclusively on the **LiteLLM control plane** using OpenTofu and YAML.

To help others adopt this, I've open-sourced a generalized starter kit for the community based on this exact architecture!

## The Problem: UI ClickOps Does Not Scale

Tools like [LiteLLM](https://github.com/BerriAI/litellm) are excellent for proxying AI requests. However, managing hundreds of API keys, model routing rules, and strict budget limits via a web console inevitably leads to:

- **Lack of auditability**: "Who increased the budget for marketing-analysts by \$5,000?"
- **Configuration drift**: "Why does staging behave differently from production?"
- **Security risks**: Manual management of API keys makes offboarding complicated.

I needed to architect a system where every change was reviewed as code, checked against organizational policies automatically, and deployed reliably.

## The Git-Driven Solution

I mapped the AI proxy infrastructure configuration strictly to a Git repository. Using OpenTofu (the open-source Terraform fork) and the official LiteLLM Terraform provider, I abstracted the configuration behind simple, readable YAML files.

*Note: While purists argue true "GitOps" requires an in-cluster agent (like ArgoCD) for constant drift reconciliation, invoking OpenTofu via CI/CD acts as a highly effective "GitOps-style" or push-based automation that solves these auditability issues.*

### 1. Environment-Aware Configuration

I split the configurations cleanly by environment. *(Note: While the Trojan AI Gateway officially runs `test` and `prod` environments, the open-source starter repo generalizes these to `staging` and `prod`.)* For instance, `staging` is loosely governed and allows developers to freely test models. Meanwhile, `prod` has strict governance enforcing hard team budgets.

Inside a `prod/teams.yaml`, declaring team budgets looks exactly like this:

```yaml
teams:
  engineering-platform:
    team_alias: "Platform Engineering"
    organization: "engineering"
    models:
      - "openai/gpt-4o"
      - "openai/gpt-4o-mini"
    max_budget: 200
    budget_duration: "1mo"
```

Every PR that modifies this file is automatically planned and reviewed. Destructive changes (like entirely removing a team) are guarded by automated CI/CD checks before a human approval gate.

### 2. Zero Credentials in Code

With this setup, the master LLM API keys and proxy credentials never touch the repository. OpenTofu retrieves the necessary secrets directly from AWS Secrets Manager at runtime and maps them to the respective LiteLLM configurations securely.

## Identity vs. Governance

When building the Trojan AI Gateway, I had to solve a major problem: *How do we handle human accounts?*

Initially, I added about 20 early-adopter users directly into our IaC configuration. I quickly became uncomfortable with this approach, and a team member rightly pointed out that user provisioning should remain outside IaC. So, I came to a firm conclusion: **human user provisioning does not belong in IaC.**

Maintaining several individual YAML blocks for human users creates significant operational overhead and duplicates identity information already managed by the Identity Provider (IdP).

I considered three complementary mechanisms for user provisioning and access management:

1. **SCIM: IdP-driven lifecycle management**  
   With the appropriate LiteLLM Enterprise license and connectivity, SCIM allows an IdP such as Entra ID or Okta to automate user and team provisioning, updates, and deprovisioning. This is the preferred approach when the environment supports it.

2. **SSO with Just-in-Time Provisioning: Authentication and initial access**  
   Users can be created in LiteLLM when they first sign in through SSO, receiving configured defaults for their role and model access. This simplifies onboarding, but those defaults do not necessarily satisfy each approved access request. JIT provisioning alone also does not provide ongoing lifecycle synchronization or automated offboarding.

3. **Automated Workflows: Applying approved access**  
   A GitHub Actions workflow, backed by Python scripts calling LiteLLM’s administration APIs, creates or updates users and applies the model access specified in their approved onboarding requests. This complements SSO and supports requirements beyond the configured defaults.

Our implementation combines **SSO with workflow-based provisioning**. After an onboarding request is approved, we add the user to the appropriate Entra ID group and run the GitHub Actions workflow with the required inputs. The workflow provisions or updates the LiteLLM user, applies the approved model access, and sends a welcome email with SSO login instructions.

SSO remains the authentication mechanism. JIT creation remains available for users who do not yet have a LiteLLM record, while the onboarding workflow establishes requested access before their first login.

The separation of concerns is clear: **the IdP owns identity and authentication; IaC defines platform policy; workflows apply approved user access.**

## Get the Open-Source Starter Kit

The lessons from building USC’s Trojan AI Gateway can help other organizations bring their LiteLLM control plane under version control. To share that experience, I’ve published an open-source starter repository:

**[cybergavin/litellm-control-plane](https://github.com/cybergavin/litellm-control-plane)**

It includes OpenTofu code, JSON Schema validation, separate staging and production configurations, and GitHub Actions workflows for planning, applying, and validating changes.

The repository provides a foundation for managing models, organizations, teams, and service-account keys. You can adapt it to your environment and extend it to additional resources supported by the LiteLLM Terraform provider, including access groups and MCP servers.

Fork it, adapt it, and give Git-driven automation a try. Fewer manual changes. Clearer governance. A LiteLLM control plane you can review, version, and maintain.
