We train on a mix of licensed clinical text, synthetic data generated under clinical supervision, and customer data — only when the customer has explicitly opted in and only after deidentification.
We do not train on data from one customer to improve a model serving another, unless both parties have opted in. This is contractual, not aspirational.
Every training data source is logged, with provenance, in our model registry.
