LLMOps & Deployment

Course Content

LLMOps & Deployment

6 sections · 40 lessons

How do you avoid vendor lock-in when relying on one model provider?


What you need to know

Where lock-in actually comes from

The API call itself is easy to swap. The hard lock-in is elsewhere:

  • Prompts tuned for one model family. They often behave differently on another model.
  • Evals and traces stored only in a vendor's product.
  • Proprietary features in the critical path: hosted vector stores, hosted conversation state, provider-specific tool or agent frameworks, fine-tunes that exist only on that platform.
  • Contract and quota — committed spend, reserved throughput.

How to stay portable

Keep on your side

  • Prompts and templates, in Git or your registry
  • Eval datasets and results
  • Retrieval corpus, embeddings pipeline and index
  • Traces and feedback data

Keep behind an interface

  • Model calls, via the gateway
  • Vector search, via your retrieval service
  • Tool definitions in a neutral schema
  • Fine-tunes on open models you can host
  • Standard formats help. OpenAI-compatible chat APIs and JSON-schema tool definitions are widely supported; the Model Context Protocol (MCP) makes tools portable across agent clients.
  • Per-provider prompt variants with their own eval scores. Expect to maintain two versions of important prompts.
  • A second provider in production, even at 1–5% of traffic, so the path is proven.

Be honest about the cost

Full portability means using fewer provider-specific features and doing eval work twice. For many products, the right balance is: portable core, with a few provider features wrapped behind interfaces where the benefit is large.

A real-life example

A state government's chatbot runs on one provider's model through an Indian cloud region. A new data-protection directive requires the service to be able to run fully within government-controlled infrastructure within six months, and the provider's regional price rises 30%.

Because the team built for portability, the move is a project, not a crisis:

  • Prompts and the 600-case eval set (in 12 languages) are in their own repo. They run the eval set against two open-weight models self-hosted on a government cloud. The better one scores 4 points lower overall and 9 points lower in Odia.
  • Retrieval uses their own pgvector database, so nothing moves there.
  • They fine-tune a LoRA adapter on 5,000 reviewed Odia conversations, closing most of the gap, and route 5% of traffic to the new stack for a month before switching.

The migration takes seven weeks. A team that had used the provider's hosted assistant state and file search would have had to rebuild retrieval and conversation storage first.

Follow-up questions to expect

  • "Isn't a gateway enough to avoid lock-in?" — It removes code lock-in, but not prompt behaviour or proprietary features. You also need your own evals and per-provider prompts.
  • "Should we avoid provider features entirely?" — No; use them when the gain is large, but wrap them and know your exit cost.
  • "How often do you test the second provider?" — Continuously with a small share of traffic, plus a planned failover drill every quarter.