Course Content
Building with LLMs
4 sections · 10 lessons
Managing API Keys and Environment
A developer is debugging a failing call at half past eleven at night. To rule out an environment problem, they paste the key directly into the file:
client = OpenAI(api_key="sk-proj-8fT2...") # TODO: move back to env before commitIt works. The bug was elsewhere. They fix the real bug, commit everything, push, and go to bed.
Here is the timeline that follows, and every step of it is ordinary.
| Elapsed | What happens |
|---|---|
| 0 s | Push lands on a public repository |
| ~30 s | An automated scanner clones the repo — these run continuously against the public push firehose |
| ~2 min | The key is tested against the provider's API and confirmed live |
| ~6 min | The key is in use, typically resold or driving a proxy service |
| 9 hours | The developer wakes up. Usage dashboard shows 41,200 requests overnight. |
At a modest average of 3,000 input and 600 output tokens per request on a mid-tier model, that is 123.6 million input and 24.7 million output tokens. At 3 and 15 dollars per million: 370.80 plus 370.50, about 741 dollars in one night. On a large-tier model it would be several times that. And the money is the smaller problem — the key also had access to whatever else was on that account.
Then comes the part people get wrong. The developer deletes the line and pushes a fix. The key is still exposed. Git keeps every commit; the secret sits in the history, and it is already in the scanner's database regardless. The only fix that works is revoking the key at the provider.
A secret that has been committed is compromised permanently. Removing it from the current file changes nothing. Revoke it — that is the entire remediation, and everything else is theatre.
The baseline: environment variables
The principle is separation of configuration from code. Code says which setting it needs; the environment supplies the value. The same image runs in development, staging and production with different credentials and no code change.
import osapi_key = os.environ["OPENAI_API_KEY"] # raises KeyError if missingUse bracket access, not os.getenv(), for anything required. getenv returns None silently, so the failure surfaces later as an opaque 401 from the provider instead of a clear crash at startup naming the missing variable. Fail early and name the problem.
The .env file
Typing exports before every run is tedious, so keep local values in a file the loader reads into the environment:
1# .env — never committed2OPENAI_API_KEY=sk-proj-...3ANTHROPIC_API_KEY=sk-ant-...4DATABASE_URL=postgresql://localhost:5432/dev5ENVIRONMENT=development6LOG_LEVEL=DEBUGfrom dotenv import load_dotenvload_dotenv() # call once, at application entry, before anything reads configAlongside it, commit a .env.example with the same keys and no values. It is documentation that cannot drift, and it is how a new colleague knows what to fill in.
Keeping it out of git, properly
1# .gitignore2.env3.env.*4!.env.example5*.pem6*.key7secrets/Then check what you have already done, because .gitignore only affects untracked files — a file already committed keeps being tracked no matter what you add to the ignore list:
git ls-files | grep -E '\.env|\.pem|secret' # should print nothinggit log --all -p -S 'sk-' -- . | head -50 # search history for key-shaped stringsBetter still, make it impossible to commit one by accident. A pre-commit hook that scans staged changes takes ten minutes to set up and has caught more leaks than any policy document:
1# .git/hooks/pre-commit2#!/bin/sh3if git diff --cached | grep -Eq '(sk-[A-Za-z0-9_-]{20,}|AKIA[0-9A-Z]{16}|-----BEGIN [A-Z ]*PRIVATE KEY)'; then4 echo "Refusing to commit: a credential-shaped string is staged."5 exit 16fiFrom loose variables to typed settings
os.environ everywhere has four failure modes that show up as production bugs rather than errors: no validation, no types (everything is a string), no defaults, and no single place to see what the application needs. All four are fixed by one settings object validated at startup.
1from pydantic_settings import BaseSettings, SettingsConfigDict2from pydantic import Field, field_validator, SecretStr3from typing import Literal45class Settings(BaseSettings):6 model_config = SettingsConfigDict(7 env_file=".env",8 env_file_encoding="utf-8",9 case_sensitive=False,10 extra="ignore",11 )1213 openai_api_key: SecretStr14 anthropic_api_key: SecretStr | None = None15 database_url: str16 environment: Literal["development", "staging", "production"] = "development"17 request_timeout: float = Field(default=30.0, gt=0, le=300)18 max_retries: int = Field(default=3, ge=0, le=10)19 monthly_budget_usd: float = Field(default=500.0, gt=0)2021 @field_validator("openai_api_key")22 @classmethod23 def check_openai_prefix(cls, v: SecretStr) -> SecretStr:24 if not v.get_secret_value().startswith("sk-"):25 raise ValueError("OPENAI_API_KEY does not look like an OpenAI key")26 return v2728settings = Settings() # raises immediately if anything is missing or malformedWhat this buys, concretely:
Problem with raw os.environ | What the settings object does |
|---|---|
| Missing variable discovered at 3 a.m. on first use | Startup fails immediately, naming the field |
REQUEST_TIMEOUT=30 is the string "30" | Coerced to 30.0 and range-checked |
| A typo'd key produces a confusing 401 | Format validated before the first call |
| Nobody knows what config exists | One class is the complete inventory |
| Keys appear in tracebacks and logs | SecretStr renders as ********** |
SecretStr is the quietly valuable part. Printing or logging the settings object, or letting an exception include it in a traceback sent to an error tracker, shows asterisks. You must call .get_secret_value() deliberately to see the real value — which means the only places the secret is readable are places you wrote on purpose.
The version trap in copied code
Enormous amounts of tutorial code use the older Pydantic v1 style, and pasting it into a v2 project produces errors that do not obviously point at the cause. The differences that bite:
| Pydantic v1 | Pydantic v2 |
|---|---|
from pydantic import BaseSettings | from pydantic_settings import BaseSettings (separate package) |
class Config: inner class | model_config = SettingsConfigDict(...) |
@validator("field") | @field_validator("field") plus @classmethod |
.dict() / .json() | .model_dump() / .model_dump_json() |
parse_obj() | model_validate() |
The first row causes the most confusion: in v2, BaseSettings was moved out of pydantic entirely, so the import fails and the error message says nothing about versions. If you see that, install pydantic-settings rather than pinning yourself back to v1.
Production: managed secret stores
Environment variables on a server are a real improvement on files in a repository, and still weak in three ways: rotation requires a redeploy, there is no record of who read a secret, and the value sits in plaintext in whatever configured the process.
1import json, boto32from botocore.exceptions import ClientError3from functools import lru_cache45@lru_cache(maxsize=16)6def aws_secret(secret_id: str) -> dict:7 client = boto3.client("secretsmanager", region_name="eu-west-2")8 try:9 return json.loads(client.get_secret_value(SecretId=secret_id)["SecretString"])10 except ClientError as exc:11 code = exc.response["Error"]["Code"]12 if code == "ResourceNotFoundException":13 raise RuntimeError(f"Secret {secret_id} does not exist") from exc14 if code == "AccessDeniedException":15 raise RuntimeError(f"No permission to read {secret_id}") from exc16 raise1from google.cloud import secretmanager23def gcp_secret(project_id: str, name: str, version: str = "latest") -> str:4 client = secretmanager.SecretManagerServiceClient()5 path = f"projects/{project_id}/secrets/{name}/versions/{version}"6 return client.access_secret_version(name=path).payload.data.decode("utf-8")The cache is important for cost and latency both. Each fetch is a network call of 50–200 ms and is billed per API call; doing it inside a request handler multiplies both by your traffic. Load at startup, cache, refresh on a schedule.
Wire it into the settings object so the rest of the application never knows the difference:
1def load_settings() -> Settings:2 env = os.getenv("ENVIRONMENT", "development")3 if env == "production":4 secrets = aws_secret("prod/llm-service")5 return Settings(**secrets, environment=env)6 load_dotenv()7 return Settings()| Method | Rotation | Audit | Encrypted at rest | Use for |
|---|---|---|---|---|
| Hard-coded | Never | Public | No | Nothing |
.env file | Manual | None | No | Local development |
| Platform env vars | Redeploy | Deploy log | Usually | Small services, staging |
| Managed secret store | Automatic | Full access log | Yes | Production with real users |
Rotation, done without an outage
Rotation limits the damage window. A key rotated every 90 days is exposed for at most 90 days rather than forever. The naive approach — delete the old key, create a new one, redeploy — causes a brief outage and, worse, fails ugly if the redeploy is slow.
The pattern that avoids this is overlap: two keys valid at once, switch traffic, then retire the old one.
- Create key B at the provider. Both A and B are live.
- Write B into the secret store as a new version. Consumers refresh and start using B.
- Wait past your longest cache TTL and process lifetime — an hour is usually generous.
- Check the provider's usage dashboard: A should show zero requests.
- Revoke A.
Step 4 is the one people skip, and it is what turns a routine rotation into an incident. A batch job that runs nightly, or a worker that loaded its config a week ago, is still holding A. Confirm zero usage before revoking.
1import threading, time23class RotatingSecret:4 """Refreshes a secret in the background so long-lived processes pick up new versions."""5 def __init__(self, loader, interval=3600):6 self._loader, self._interval = loader, interval7 self._value, self._lock = loader(), threading.Lock()8 threading.Thread(target=self._refresh, daemon=True).start()910 def _refresh(self):11 while True:12 time.sleep(self._interval)13 try:14 new = self._loader()15 with self._lock:16 self._value = new17 except Exception:18 logging.exception("secret refresh failed; keeping previous value")1920 def get(self):21 with self._lock:22 return self._valueThe except that keeps the previous value on failure is deliberate: a transient failure to reach the secret store should not take down a running service that already holds a working credential.
Defence in depth
Assume every single control will fail once. Layer them so no single failure is fatal.
Validate formats at startup
1import re23PATTERNS = {4 "openai": r"^sk-[A-Za-z0-9_-]{20,}$",5 "anthropic": r"^sk-ant-[A-Za-z0-9_-]{20,}$",6 "aws_access": r"^AKIA[0-9A-Z]{16}$",7}89def validate(provider: str, key: str) -> bool:10 return bool(re.match(PATTERNS[provider], key))This catches the mundane and common failures: a key pasted with a trailing newline, two keys swapped between variables, a placeholder that was never replaced. Cheap, and it turns a confusing runtime 401 into a clear startup error.
Never log a secret in full
1def mask(secret: str, show: int = 4) -> str:2 if not secret or len(secret) <= show * 2:3 return "*" * 84 return f"{secret[:show]}{'*' * 12}{secret[-show:]}"56logging.info("using key %s", mask(api_key)) # sk-p************nQ4aShowing the first and last few characters keeps the log useful — you can tell which key is in use and confirm rotation happened — without printing the secret. Also add a filter so a secret cannot reach the logs by an indirect route, which is how it usually happens:
1class RedactFilter(logging.Filter):2 def __init__(self, secrets): 3 super().__init__(); self.secrets = [s for s in secrets if s]4 def filter(self, record):5 msg = record.getMessage()6 for s in self.secrets:7 msg = msg.replace(s, "[REDACTED]")8 record.msg, record.args = msg, ()9 return True1011logging.getLogger().addFilter(RedactFilter([settings.openai_api_key.get_secret_value()]))The indirect routes are the dangerous ones: an exception whose message includes the request headers, a debug dump of a config dictionary, an HTTP client logging the full request. A redaction filter catches all of them at the last moment.
Encrypt anything you must store yourself
If your application holds credentials on behalf of users — a customer's own API key, say — plaintext storage is not acceptable:
1from cryptography.fernet import Fernet23fernet = Fernet(os.environ["ENCRYPTION_KEY"].encode()) # 32 url-safe base64 bytes45def store_user_key(user_id: str, key: str):6 db.execute("UPDATE users SET provider_key = %s WHERE id = %s",7 (fernet.encrypt(key.encode()), user_id))89def load_user_key(user_id: str) -> str:10 row = db.execute("SELECT provider_key FROM users WHERE id = %s", (user_id,)).fetchone()11 return fernet.decrypt(row[0]).decode()Note that this moves the problem rather than eliminating it — now ENCRYPTION_KEY is the secret that matters, and it belongs in a managed store, never in the same database as the ciphertext. The gain is real nonetheless: a leaked database dump is useless without a key that lives somewhere else entirely.
Constrain what the credential can do
The layer that survives every other layer failing. If your provider supports scoped or project-limited keys, use a separate key per service with only the permissions that service needs, and set a spending cap on it. Then a leak costs you one capped project rather than the whole account. The overnight bill in the opening story would have stopped at the cap.
Multiple environments
Development, staging and production should differ in configuration and be identical in code. Make the differences explicit and validated:
1class Settings(BaseSettings):2 environment: Literal["development", "staging", "production"] = "development"3 model_name: str = "gpt-6-luna"4 debug: bool = False56 @property7 def is_production(self) -> bool:8 return self.environment == "production"910 def model_post_init(self, __ctx) -> None:11 if self.is_production:12 if self.debug:13 raise ValueError("debug must be off in production")14 if "localhost" in self.database_url:15 raise ValueError("production is pointed at a local database")Those two assertions look paranoid until the morning someone deploys with a staging .env still in the image and debug logging writes full request bodies — including whatever users typed — into a log aggregator retained for two years. A startup check that refuses to boot is a far better outcome than a service that runs happily in the wrong configuration.
Configuration errors are silent by default: the application starts, runs, and does the wrong thing. Validate at startup so the wrong configuration cannot boot at all.
Where people get this wrong
Thinking a deletion commit fixes a leak. It does not. Revoke.
One key for everything. One key across dev, CI, staging and production means one leak revokes all four, and you cannot tell from usage data which environment is responsible for a spike.
Reading secrets inside request handlers. Adds latency and secret-store bill to every request. Load once at startup.
Using os.getenv for required values. Converts a clear startup failure into a confusing runtime one.
Committing .env "just this once, it's only staging". Staging keys usually reach the same provider account and often the same data.
Putting secrets in Docker build arguments or image layers. Anyone who can pull the image can read them with docker history. Pass secrets at run time, never at build time.
No spending cap. The difference between a bad night and a bad quarter is whether a limit existed before the leak, not how fast you noticed.
What to do before your next deploy
This is a checklist because it is genuinely one, and because everything on it takes minutes while the failure it prevents takes days.
- Run
git ls-files | grep -E '\.env|\.pem|secret'andgit log --all -S 'sk-'. If either prints something, you have a key to revoke today. Do that before reading further. - Install a pre-commit secret scan. The single highest-value ten minutes in this lesson.
- Move to one validated settings object. One class, one import,
SecretStron every credential, validation that runs at startup. - Give every environment its own key, and every service its own key where the provider allows it. Scoped, capped, individually revocable.
- Set a hard spending limit at the provider. Not a billing alert — a cap. Alerts arrive after the money is gone.
- Write down the revocation procedure while nothing is wrong. Which dashboard, which account, who has access, how the new key reaches production. Discovering that only one person can rotate the key, and they are on a flight, is a bad way to spend an incident.
The engineer in the opening story was not careless in any unusual way. They did what almost everyone has done at least once, on a tired evening, with a comment promising to undo it. The systems that survive that evening are the ones where a hook refused the commit, or the key was scoped to one project with a cap, or both. Build those systems now, because the tired evening is coming.