Building AI Features in Python Backends

Logging without leaking personal data


ShipFast's customers send their phone numbers, home addresses, the names of family members who will receive the parcel and, sometimes, a photo of an ID "to prove it is me". Every one of those messages passes through your service and a model provider. The question for this lesson is where copies of that text end up, and who can read them.

In ShipFast's prototype, the answer was "everywhere". The first LLMClient logged the full prompt at INFO level for debugging. The error tracker captured exception messages, which included Pydantic validation errors, which included the invalid input values. The tracing tool recorded request bodies. Within a week, 60 engineers, a logging vendor and an error-tracking vendor had copies of about 250,000 customer messages, kept for 90 days by default. No one had decided that; it was the sum of defaults.

Every place a customer's message can end upCustomer messageProvider — checkretention termsDatabase — theone intended copyLogs — metadata onlyError tracker — nolocals or bodiesCache — hashed keysEval sets — redacted first
Most leaks come from defaults nobody chose, like validation errors and error-tracker capture, not from log lines someone wrote.

Where customer text goes

  1. The provider — the text is sent to the model. Check the provider's retention and training terms for API data, and your contract, before launch.
  2. Your database — the message and result are stored to run the service. This is the one copy you intend to keep.
  3. Application logs — anything you log, and anything a library logs for you.
  4. Error trackers and traces — exception messages, stack-frame variables and request bodies, often captured automatically.
  5. Caches — keys and values in Redis.
  6. Evaluation sets — real messages copied into files for testing prompts.

The goal is to make copy 2 the only full copy you control, and to keep every other copy free of personal data by design, not by hoping people are careful.

Log metadata, not messages

Look at the line LLMClient writes for every call:

Text
llm_call feature=classify model=claude-opus-5 in=413 out=38 ms=912 cost_usd=0.00302 stop=end_turn

It has no customer text, and it is enough to answer almost every operational question: what is slow, what is expensive, what is failing, which prompt changed. The triage service adds a few more safe fields when it logs a result: the message ID, the intent, the queue, the prompt versions and whether it degraded.

When you do need to connect several log lines about one customer, use a pseudonym: a salted hash of the customer ID. It is stable, so you can group by it, but it cannot be turned back into the ID without the salt, which lives in your secrets store, not in the logs.

Python
# shipfast/redact.pyimport hashlibimport loggingimport rePATTERNS = [    (re.compile(r"\b\d{4}\s?\d{4}\s?\d{4}\b"), "<aadhaar>"),          # 12-digit national ID    (re.compile(r"(?:\+91[\s-]?)?\b[6-9]\d{4}[\s-]?\d{5}\b"), "<phone>"),   # Indian mobiles    (re.compile(r"[\w.+-]+@[\w-]+(\.[\w-]+)+"), "<email>"),    (re.compile(r"\b[1-9]\d{5}\b"), "<pincode>"),]def redact(text: str) -> str:    for pattern, label in PATTERNS:        text = pattern.sub(label, text)    return textdef pseudonym(value: str, salt: str) -> str:    """Stable, non-reversible ID for joining log lines about the same customer."""    return hashlib.sha256((salt + value).encode()).hexdigest()[:16]class RedactingFilter(logging.Filter):    def filter(self, record: logging.LogRecord) -> bool:        record.msg = redact(record.getMessage())        record.args = None        return True

Redaction is a safety net, not the plan

RedactingFilter is attached to every log handler at startup, so any line that does contain a phone number or email is cleaned before it is written.

Python
import loggingfrom shipfast.redact import RedactingFilterhandler = logging.StreamHandler()handler.addFilter(RedactingFilter())logging.basicConfig(level=logging.INFO, handlers=[handler])logging.getLogger("shipfast").warning("customer says: call 98765 43210, pin 411014")# customer says: call <phone>, pin <pincode>

It catches the obvious patterns: phone numbers, emails, PIN codes and 12-digit ID numbers. It over-redacts sometimes (a six-digit amount looks like a PIN code), which is the right direction to be wrong in. But it cannot find names or street addresses, because those have no pattern: "Flat 12, Lake View, near the temple" passes straight through. That is why the plan is do not log message text at all, and the filter is only there for the line someone adds by mistake.

The leaks nobody writes on purpose

Most leaks come from code that never mentions logging.

Validation errors. Pydantic's ValidationError, turned into a string, includes the invalid input values. If call_structured raised it directly and an error tracker captured it, every failed extraction would send a customer's address to the tracker. ShipFast's short_errors builds messages from the field location and error text only, and StructuredOutputError carries those, never the raw reply.

Exception messages from your own code. LLMError messages contain the feature and the status code, never the prompt. Keep it that way in code review: an exception message is a log line that travels further.

Error-tracker defaults. Many error trackers capture local variables and request bodies by default. Turn that off for this service, or configure the tracker's own scrubbing, and test it by raising an exception with a fake phone number and checking what arrives.

Cache keys. ShipFast's cache keys are SHA-256 hashes, so Redis never stores message text as a key. The cached value is a label and a short reason, which the prompt limits to 20 words; check a sample of reasons to make sure they do not quote the customer.

Evaluation sets. Real messages are the best test data and the most dangerous files in the repository. ShipFast's evaluation sets are built with redact() plus a manual pass that replaces names and street names with invented ones, and they live in a repository with restricted access.

Check your understanding

0 of 3 answered

1.Why is a redaction filter not enough on its own?

2.A Pydantic ValidationError from an extraction is captured by the error tracker. What might it contain?

3.What is a pseudonym used for in ShipFast's logs?