Aniruddh Chandratre

Senior Engineer, ML Infrastructure · Tesla AI · Palo Alto

ML infrastructure for large-scale AI.

I build data platforms, schedulers, and cluster-efficiency systems for Tesla AI. My work supports data operations for FSD, Optimus, and Digital Optimus across large GPU fleets.

Internal data processing platform

Tesla AI · internal platform

Own a large-scale MapReduce system used for data acquisition, curation, mining, auto-labeling, training, and evaluation. It runs across 10 GPU clusters, processing 10B+ jobs per month and parallelizing an average of 3,000 years of compute.

  • 10B jobs / month
  • 3,000 yr compute / month
  • 10 GPU clusters

Scheduling & live defragmentation

Tesla AI · GPU infrastructure

Built placement and live-defragmentation systems that compact allocations into contiguous topology while clusters remain online. The scheduler operates across a heterogeneous GPU fleet.

  • 250K+ H100-equivalent GPUs
  • 98% effective utilization, fleet average
  • 24 → 88% usable capacity

Internet-scale data curation

Tesla AI · FSD · Optimus · Digital Optimus

Build pipelines that turn large-scale data acquisition into training-ready datasets, with lineage across artifacts, datasets, training jobs, and models.

  • end-to-end acquisition → training set

Core data infrastructure — Tesla Manufacturing

2023 – 2024

Built an edge-to-cloud data broker processing 10B+ manufacturing data points per day across Austin, Berlin, and Fremont. Developed a 5M writes/second durable queue and a distributed cache that cut storage by 80% across 50 deployments.

Graduate research — ASU, DARPA ARCOS

2021 – 2022

Researched formal verification for autonomous systems under DARPA ARCOS with Lockheed Martin. Co-authored PSY-TaLiRo and published eight peer-reviewed papers.

Founding engineer — Aftershoot

2020 – 2021

Built the image-ranking model at the core of AfterShoot's photo-culling product, now used by approximately 250,000 photographers.