Anthropic
May 2026 — PresentMember of Technical Staff · San Francisco, CACurrent
- Work on model provenance and model governance — the systems that make it possible to say, precisely and after the fact, which datasets went into a given model and under what terms they were used.
- Build dataset lineage tracking across the training pipeline, so the composition of a training corpus is a recorded fact rather than institutional memory.
- Support training on data under strict privacy protections, including workloads that run inside secure enclaves where the data itself is never directly accessible.
- Build the data infrastructure that prepares training data: storage layout, bucket organisation, and reliable transfer across account boundaries at the scale frontier training demands.
- Own data deletion and retention as an enforced property of the platform — built on Apache Iceberg and Spark, so a deletion commitment resolves to jobs that actually run.