sasank.pagolu

Member of Technical Staff · Anthropic

The provenance layer beneath frontier models.

I build the systems that track what data a model was trained on — dataset lineage, model governance, and the infrastructure that prepares and moves training data across buckets, storage, and account boundaries before it ever reaches a training run.

Including the parts that are hard on purpose: training on sensitive data under strict privacy controls, and making deletion and retention real at petabyte scale.

San Francisco Bay Area Model governance · Dataset lineage · Data infrastructure

900+Citations across peer-reviewed publications
PB-scaleRetention and deletion enforced across the Lakehouse
22BMetrics per minute stored across Twitter's fleet
97% 99.99%Indexing service success rate after reliability work
01

Career highlights

Anthropic

May 2026 — Present

Member of Technical Staff · San Francisco, CACurrent

  • Work on model provenance and model governance — the systems that make it possible to say, precisely and after the fact, which datasets went into a given model and under what terms they were used.
  • Build dataset lineage tracking across the training pipeline, so the composition of a training corpus is a recorded fact rather than institutional memory.
  • Support training on data under strict privacy protections, including workloads that run inside secure enclaves where the data itself is never directly accessible.
  • Build the data infrastructure that prepares training data: storage layout, bucket organisation, and reliable transfer across account boundaries at the scale frontier training demands.
  • Own data deletion and retention as an enforced property of the platform — built on Apache Iceberg and Spark, so a deletion commitment resolves to jobs that actually run.

Apple

Sep 2022 — May 2026

Senior Software Engineer, Data & ML Platform · Cupertino, CATech Lead

  • Led design and implementation of the Data Retention Service across Apache Hive, Iceberg, and S3 datasets, using Spark-based enforcement jobs to hold data compliance at petabyte scale, with first-class integration into the Lakehouse platform.
  • Built AI-powered agents that give teams a conversational path through policy creation, onboarding, and troubleshooting — RAG pipelines over internal knowledge bases, Splunk logs, and MCP servers for automated diagnosis.
  • Developed a partition pattern detection system using LLMs to auto-suggest retention policies, accelerating governance onboarding and cutting manual configuration errors.
  • Provided technical leadership for Apache Iceberg adoption, shaping Lakehouse table format standards and cross-functional practice for large-scale data management.
  • Built a semantic search system on Ray, LangGraph, and a vector database for embedding-based dataset discovery, turning metadata retrieval into governance insight.

Twitter

Feb 2019 — Sep 2022

Senior Software Engineer, Observability · San Francisco, CA

  • Worked on the company-wide monitoring and alerting system, backed by an in-house timeseries database with aggressive compression that stores every metric across the fleet — 22 billion per minute — along with the read and write infrastructure around it.
  • Drove a key part of the company-wide hybrid cloud investment by extending the Observability stack to support Prometheus-style dimensional metrics with labels, including structured querying of dimensional and histogram metrics in the query service.
  • Shipped reliability improvements that took a stateful, distributed temporal key-value indexing service from 97% to 99.99% success, making container and metric indexing dependable across the data center.
  • Developed Kafka-based real-time indexing, cutting p99 delay for newly seen metrics from 30 minutes to 10 seconds.
  • Led the indexing service's lossless, zero-downtime migration from bare metal to shared cloud — a service holding 1 billion metrics in memory and serving 14 million Thrift RPCs per minute.
  • Extended the timeseries query language and scaled the query engine to 100M+ unique timeseries reads per minute, while reducing false alerts from 50/day to under 5/week.

Progression: joined as Software Engineer I and grew to Senior Software Engineer over three and a half years on the same team.

Qualcomm

Earlier

Safety application stack for connected autonomous vehicles (CV2X)

  • Built an intelligent transport systems stack for connected autonomous vehicles transmitting Basic Safety Messages over the IEEE 1609 protocol.
  • Wrote a C library for parsing and interpreting received safety messages, with detection algorithms driving real-time safety warnings.

Cisco

Earlier

Resource-based REST APIs for Unified Computing System Director

  • Built a workflow exposing virtual machine and cluster operations as REST APIs, using Java reflection and annotations to reshape the existing architecture into a single-pane-of-glass model.
  • Designed role-based access control for the generated APIs to guarantee secured access.
02

Publications

Google Scholar ↗

Semantic Loss Application to Entity Relation Recognition

Venkata Sasank Pagolu

arXiv:2006.04031 · cs.CL · 2020

Work from my master's at UCLA, applying semantic loss — a differentiable penalty for violating symbolic constraints — to entity and relation extraction, so a neural model is trained against the structural rules its outputs are supposed to obey.

Sentiment Analysis of Twitter Data for Predicting Stock Market Movements

743 citations

Venkata Sasank Pagolu, Kamal Nayan Reddy Challa, Ganapati Panda, Babita Majhi

IEEE SCOPES 2016 · Intl. Conf. on Signal Processing, Communication, Power and Embedded System

★  Best Paper Award — selected from ~500 submissions

First-author work testing whether public sentiment on Twitter carries signal about a company's stock movement — pairing sentiment features from tweets about a company with its daily price direction, and measuring how much of the movement is actually predictable from them. Now widely used as a baseline in financial-NLP research.

An Improved Approach for Prediction of Parkinson's Disease using Machine Learning Techniques

154 citations

Kamal Nayan Reddy Challa, Venkata Sasank Pagolu, Ganapati Panda, Babita Majhi

IEEE SCOPES 2016 · Intl. Conf. on Signal Processing, Communication, Power and Embedded System

Applying gradient-boosted and other supervised classifiers to clinical feature data for earlier, more reliable detection of Parkinson's disease than the baselines available at the time.

Citation counts via Google Scholar, June 2026 · live counts

03

Education

M.S. Computer Science

University of California, Los Angeles

2017 – 2018GPA 4.00 / 4.00

B.Tech. Computer Science

Indian Institute of Technology, Bhubaneswar

2013 – 2017GPA 9.38 / 10Institute Rank 5
04

Toolkit

Languages

JavaScalaPythonCSQL

Data & Storage

Apache IcebergApache HiveSparkHadoopKafkaS3MySQLRedisMemcached

Infrastructure

KubernetesZookeeperApache ThriftJVMMapReduceRESTCI/CD

ML & AI Systems

RayLangGraphRAG pipelinesVector databasesMCP
05

Contact

Let's talk

Happy to hear about data platform and governance work, reviewing and program committee invitations, or anything at the intersection of distributed systems and ML infrastructure.