Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you avoid vendor lock-in when relying on one model provider?
What you need to know
Where lock-in actually comes from
The API call itself is easy to swap. The hard lock-in is elsewhere:
- Prompts tuned for one model family. They often behave differently on another model.
- Evals and traces stored only in a vendor's product.
- Proprietary features in the critical path: hosted vector stores, hosted conversation state, provider-specific tool or agent frameworks, fine-tunes that exist only on that platform.
- Contract and quota — committed spend, reserved throughput.
How to stay portable
Keep on your side
- Prompts and templates, in Git or your registry
- Eval datasets and results
- Retrieval corpus, embeddings pipeline and index
- Traces and feedback data
Keep behind an interface
- Model calls, via the gateway
- Vector search, via your retrieval service
- Tool definitions in a neutral schema
- Fine-tunes on open models you can host
- Standard formats help. OpenAI-compatible chat APIs and JSON-schema tool definitions are widely supported; the Model Context Protocol (MCP) makes tools portable across agent clients.
- Per-provider prompt variants with their own eval scores. Expect to maintain two versions of important prompts.
- A second provider in production, even at 1–5% of traffic, so the path is proven.
Be honest about the cost
Full portability means using fewer provider-specific features and doing eval work twice. For many products, the right balance is: portable core, with a few provider features wrapped behind interfaces where the benefit is large.
A real-life example
A state government's chatbot runs on one provider's model through an Indian cloud region. A new data-protection directive requires the service to be able to run fully within government-controlled infrastructure within six months, and the provider's regional price rises 30%.
Because the team built for portability, the move is a project, not a crisis:
- Prompts and the 600-case eval set (in 12 languages) are in their own repo. They run the eval set against two open-weight models self-hosted on a government cloud. The better one scores 4 points lower overall and 9 points lower in Odia.
- Retrieval uses their own pgvector database, so nothing moves there.
- They fine-tune a LoRA adapter on 5,000 reviewed Odia conversations, closing most of the gap, and route 5% of traffic to the new stack for a month before switching.
The migration takes seven weeks. A team that had used the provider's hosted assistant state and file search would have had to rebuild retrieval and conversation storage first.
Follow-up questions to expect
- "Isn't a gateway enough to avoid lock-in?" — It removes code lock-in, but not prompt behaviour or proprietary features. You also need your own evals and per-provider prompts.
- "Should we avoid provider features entirely?" — No; use them when the gain is large, but wrap them and know your exit cost.
- "How often do you test the second provider?" — Continuously with a small share of traffic, plus a planned failover drill every quarter.