Building with LLMs

Managing API Keys and Environment


A developer is debugging a failing call at half past eleven at night. To rule out an environment problem, they paste the key directly into the file:

Python
client = OpenAI(api_key="sk-proj-8fT2...")   # TODO: move back to env before commit

It works. The bug was elsewhere. They fix the real bug, commit everything, push, and go to bed.

Here is the timeline that follows, and every step of it is ordinary.

ElapsedWhat happens
0 sPush lands on a public repository
~30 sAn automated scanner clones the repo — these run continuously against the public push firehose
~2 minThe key is tested against the provider's API and confirmed live
~6 minThe key is in use, typically resold or driving a proxy service
9 hoursThe developer wakes up. Usage dashboard shows 41,200 requests overnight.

At a modest average of 3,000 input and 600 output tokens per request on a mid-tier model, that is 123.6 million input and 24.7 million output tokens. At 3 and 15 dollars per million: 370.80 plus 370.50, about 741 dollars in one night. On a large-tier model it would be several times that. And the money is the smaller problem — the key also had access to whatever else was on that account.

Then comes the part people get wrong. The developer deletes the line and pushes a fix. The key is still exposed. Git keeps every commit; the secret sits in the history, and it is already in the scanner's database regardless. The only fix that works is revoking the key at the provider.

A secret that has been committed is compromised permanently. Removing it from the current file changes nothing. Revoke it — that is the entire remediation, and everything else is theatre.

How a key gets out of the source filePasted in the file — one commit from a leakRead from an environment variableA .env file, git-ignored and never committedTyped settings that fail loudly at startupA managed secret store with rotation
Each rung removes one way the key can escape; the last one also lets you rotate it without an outage.

The baseline: environment variables

The principle is separation of configuration from code. Code says which setting it needs; the environment supplies the value. The same image runs in development, staging and production with different credentials and no code change.

Python
import osapi_key = os.environ["OPENAI_API_KEY"]      # raises KeyError if missing

Use bracket access, not os.getenv(), for anything required. getenv returns None silently, so the failure surfaces later as an opaque 401 from the provider instead of a clear crash at startup naming the missing variable. Fail early and name the problem.

The .env file

Typing exports before every run is tedious, so keep local values in a file the loader reads into the environment:

Bash
# .env  — never committedOPENAI_API_KEY=sk-proj-...ANTHROPIC_API_KEY=sk-ant-...DATABASE_URL=postgresql://localhost:5432/devENVIRONMENT=developmentLOG_LEVEL=DEBUG
Python
from dotenv import load_dotenvload_dotenv()          # call once, at application entry, before anything reads config

Alongside it, commit a .env.example with the same keys and no values. It is documentation that cannot drift, and it is how a new colleague knows what to fill in.

Keeping it out of git, properly

Bash
# .gitignore.env.env.*!.env.example*.pem*.keysecrets/

Then check what you have already done, because .gitignore only affects untracked files — a file already committed keeps being tracked no matter what you add to the ignore list:

Bash
git ls-files | grep -E '\.env|\.pem|secret'      # should print nothinggit log --all -p -S 'sk-' -- . | head -50        # search history for key-shaped strings

Better still, make it impossible to commit one by accident. A pre-commit hook that scans staged changes takes ten minutes to set up and has caught more leaks than any policy document:

Bash
# .git/hooks/pre-commit#!/bin/shif git diff --cached | grep -Eq '(sk-[A-Za-z0-9_-]{20,}|AKIA[0-9A-Z]{16}|-----BEGIN [A-Z ]*PRIVATE KEY)'; then  echo "Refusing to commit: a credential-shaped string is staged."  exit 1fi

From loose variables to typed settings

os.environ everywhere has four failure modes that show up as production bugs rather than errors: no validation, no types (everything is a string), no defaults, and no single place to see what the application needs. All four are fixed by one settings object validated at startup.

Python
from pydantic_settings import BaseSettings, SettingsConfigDictfrom pydantic import Field, field_validator, SecretStrfrom typing import Literalclass Settings(BaseSettings):    model_config = SettingsConfigDict(        env_file=".env",        env_file_encoding="utf-8",        case_sensitive=False,        extra="ignore",    )    openai_api_key: SecretStr    anthropic_api_key: SecretStr | None = None    database_url: str    environment: Literal["development", "staging", "production"] = "development"    request_timeout: float = Field(default=30.0, gt=0, le=300)    max_retries: int = Field(default=3, ge=0, le=10)    monthly_budget_usd: float = Field(default=500.0, gt=0)    @field_validator("openai_api_key")    @classmethod    def check_openai_prefix(cls, v: SecretStr) -> SecretStr:        if not v.get_secret_value().startswith("sk-"):            raise ValueError("OPENAI_API_KEY does not look like an OpenAI key")        return vsettings = Settings()      # raises immediately if anything is missing or malformed

What this buys, concretely:

Problem with raw os.environWhat the settings object does
Missing variable discovered at 3 a.m. on first useStartup fails immediately, naming the field
REQUEST_TIMEOUT=30 is the string "30"Coerced to 30.0 and range-checked
A typo'd key produces a confusing 401Format validated before the first call
Nobody knows what config existsOne class is the complete inventory
Keys appear in tracebacks and logsSecretStr renders as **********

SecretStr is the quietly valuable part. Printing or logging the settings object, or letting an exception include it in a traceback sent to an error tracker, shows asterisks. You must call .get_secret_value() deliberately to see the real value — which means the only places the secret is readable are places you wrote on purpose.

The version trap in copied code

Enormous amounts of tutorial code use the older Pydantic v1 style, and pasting it into a v2 project produces errors that do not obviously point at the cause. The differences that bite:

Pydantic v1Pydantic v2
from pydantic import BaseSettingsfrom pydantic_settings import BaseSettings (separate package)
class Config: inner classmodel_config = SettingsConfigDict(...)
@validator("field")@field_validator("field") plus @classmethod
.dict() / .json().model_dump() / .model_dump_json()
parse_obj()model_validate()

The first row causes the most confusion: in v2, BaseSettings was moved out of pydantic entirely, so the import fails and the error message says nothing about versions. If you see that, install pydantic-settings rather than pinning yourself back to v1.

Production: managed secret stores

Environment variables on a server are a real improvement on files in a repository, and still weak in three ways: rotation requires a redeploy, there is no record of who read a secret, and the value sits in plaintext in whatever configured the process.

Python
import json, boto3from botocore.exceptions import ClientErrorfrom functools import lru_cache@lru_cache(maxsize=16)def aws_secret(secret_id: str) -> dict:    client = boto3.client("secretsmanager", region_name="eu-west-2")    try:        return json.loads(client.get_secret_value(SecretId=secret_id)["SecretString"])    except ClientError as exc:        code = exc.response["Error"]["Code"]        if code == "ResourceNotFoundException":            raise RuntimeError(f"Secret {secret_id} does not exist") from exc        if code == "AccessDeniedException":            raise RuntimeError(f"No permission to read {secret_id}") from exc        raise
Python
from google.cloud import secretmanagerdef gcp_secret(project_id: str, name: str, version: str = "latest") -> str:    client = secretmanager.SecretManagerServiceClient()    path = f"projects/{project_id}/secrets/{name}/versions/{version}"    return client.access_secret_version(name=path).payload.data.decode("utf-8")

The cache is important for cost and latency both. Each fetch is a network call of 50–200 ms and is billed per API call; doing it inside a request handler multiplies both by your traffic. Load at startup, cache, refresh on a schedule.

Wire it into the settings object so the rest of the application never knows the difference:

Python
def load_settings() -> Settings:    env = os.getenv("ENVIRONMENT", "development")    if env == "production":        secrets = aws_secret("prod/llm-service")        return Settings(**secrets, environment=env)    load_dotenv()    return Settings()
MethodRotationAuditEncrypted at restUse for
Hard-codedNeverPublicNoNothing
.env fileManualNoneNoLocal development
Platform env varsRedeployDeploy logUsuallySmall services, staging
Managed secret storeAutomaticFull access logYesProduction with real users

Rotation, done without an outage

Rotation limits the damage window. A key rotated every 90 days is exposed for at most 90 days rather than forever. The naive approach — delete the old key, create a new one, redeploy — causes a brief outage and, worse, fails ugly if the redeploy is slow.

The pattern that avoids this is overlap: two keys valid at once, switch traffic, then retire the old one.

  1. Create key B at the provider. Both A and B are live.
  2. Write B into the secret store as a new version. Consumers refresh and start using B.
  3. Wait past your longest cache TTL and process lifetime — an hour is usually generous.
  4. Check the provider's usage dashboard: A should show zero requests.
  5. Revoke A.

Step 4 is the one people skip, and it is what turns a routine rotation into an incident. A batch job that runs nightly, or a worker that loaded its config a week ago, is still holding A. Confirm zero usage before revoking.

Python
import threading, timeclass RotatingSecret:    """Refreshes a secret in the background so long-lived processes pick up new versions."""    def __init__(self, loader, interval=3600):        self._loader, self._interval = loader, interval        self._value, self._lock = loader(), threading.Lock()        threading.Thread(target=self._refresh, daemon=True).start()    def _refresh(self):        while True:            time.sleep(self._interval)            try:                new = self._loader()                with self._lock:                    self._value = new            except Exception:                logging.exception("secret refresh failed; keeping previous value")    def get(self):        with self._lock:            return self._value

The except that keeps the previous value on failure is deliberate: a transient failure to reach the secret store should not take down a running service that already holds a working credential.

Defence in depth

Assume every single control will fail once. Layer them so no single failure is fatal.

Validate formats at startup

Python
import rePATTERNS = {    "openai":    r"^sk-[A-Za-z0-9_-]{20,}$",    "anthropic": r"^sk-ant-[A-Za-z0-9_-]{20,}$",    "aws_access": r"^AKIA[0-9A-Z]{16}$",}def validate(provider: str, key: str) -> bool:    return bool(re.match(PATTERNS[provider], key))

This catches the mundane and common failures: a key pasted with a trailing newline, two keys swapped between variables, a placeholder that was never replaced. Cheap, and it turns a confusing runtime 401 into a clear startup error.

Never log a secret in full

Python
def mask(secret: str, show: int = 4) -> str:    if not secret or len(secret) <= show * 2:        return "*" * 8    return f"{secret[:show]}{'*' * 12}{secret[-show:]}"logging.info("using key %s", mask(api_key))     # sk-p************nQ4a

Showing the first and last few characters keeps the log useful — you can tell which key is in use and confirm rotation happened — without printing the secret. Also add a filter so a secret cannot reach the logs by an indirect route, which is how it usually happens:

Python
class RedactFilter(logging.Filter):    def __init__(self, secrets):         super().__init__(); self.secrets = [s for s in secrets if s]    def filter(self, record):        msg = record.getMessage()        for s in self.secrets:            msg = msg.replace(s, "[REDACTED]")        record.msg, record.args = msg, ()        return Truelogging.getLogger().addFilter(RedactFilter([settings.openai_api_key.get_secret_value()]))

The indirect routes are the dangerous ones: an exception whose message includes the request headers, a debug dump of a config dictionary, an HTTP client logging the full request. A redaction filter catches all of them at the last moment.

Encrypt anything you must store yourself

If your application holds credentials on behalf of users — a customer's own API key, say — plaintext storage is not acceptable:

Python
from cryptography.fernet import Fernetfernet = Fernet(os.environ["ENCRYPTION_KEY"].encode())   # 32 url-safe base64 bytesdef store_user_key(user_id: str, key: str):    db.execute("UPDATE users SET provider_key = %s WHERE id = %s",               (fernet.encrypt(key.encode()), user_id))def load_user_key(user_id: str) -> str:    row = db.execute("SELECT provider_key FROM users WHERE id = %s", (user_id,)).fetchone()    return fernet.decrypt(row[0]).decode()

Note that this moves the problem rather than eliminating it — now ENCRYPTION_KEY is the secret that matters, and it belongs in a managed store, never in the same database as the ciphertext. The gain is real nonetheless: a leaked database dump is useless without a key that lives somewhere else entirely.

Constrain what the credential can do

The layer that survives every other layer failing. If your provider supports scoped or project-limited keys, use a separate key per service with only the permissions that service needs, and set a spending cap on it. Then a leak costs you one capped project rather than the whole account. The overnight bill in the opening story would have stopped at the cap.

Multiple environments

Development, staging and production should differ in configuration and be identical in code. Make the differences explicit and validated:

Python
class Settings(BaseSettings):    environment: Literal["development", "staging", "production"] = "development"    model_name: str = "gpt-6-luna"    debug: bool = False    @property    def is_production(self) -> bool:        return self.environment == "production"    def model_post_init(self, __ctx) -> None:        if self.is_production:            if self.debug:                raise ValueError("debug must be off in production")            if "localhost" in self.database_url:                raise ValueError("production is pointed at a local database")

Those two assertions look paranoid until the morning someone deploys with a staging .env still in the image and debug logging writes full request bodies — including whatever users typed — into a log aggregator retained for two years. A startup check that refuses to boot is a far better outcome than a service that runs happily in the wrong configuration.

Configuration errors are silent by default: the application starts, runs, and does the wrong thing. Validate at startup so the wrong configuration cannot boot at all.

Where people get this wrong

Thinking a deletion commit fixes a leak. It does not. Revoke.

One key for everything. One key across dev, CI, staging and production means one leak revokes all four, and you cannot tell from usage data which environment is responsible for a spike.

Reading secrets inside request handlers. Adds latency and secret-store bill to every request. Load once at startup.

Using os.getenv for required values. Converts a clear startup failure into a confusing runtime one.

Committing .env "just this once, it's only staging". Staging keys usually reach the same provider account and often the same data.

Putting secrets in Docker build arguments or image layers. Anyone who can pull the image can read them with docker history. Pass secrets at run time, never at build time.

No spending cap. The difference between a bad night and a bad quarter is whether a limit existed before the leak, not how fast you noticed.

What to do before your next deploy

This is a checklist because it is genuinely one, and because everything on it takes minutes while the failure it prevents takes days.

  1. Run git ls-files | grep -E '\.env|\.pem|secret' and git log --all -S 'sk-'. If either prints something, you have a key to revoke today. Do that before reading further.
  2. Install a pre-commit secret scan. The single highest-value ten minutes in this lesson.
  3. Move to one validated settings object. One class, one import, SecretStr on every credential, validation that runs at startup.
  4. Give every environment its own key, and every service its own key where the provider allows it. Scoped, capped, individually revocable.
  5. Set a hard spending limit at the provider. Not a billing alert — a cap. Alerts arrive after the money is gone.
  6. Write down the revocation procedure while nothing is wrong. Which dashboard, which account, who has access, how the new key reaches production. Discovering that only one person can rotate the key, and they are on a flight, is a bad way to spend an incident.

The engineer in the opening story was not careless in any unusual way. They did what almost everyone has done at least once, on a tired evening, with a comment promising to undo it. The systems that survive that evening are the ones where a hook refused the commit, or the key was scoped to one project with a cap, or both. Build those systems now, because the tired evening is coming.