Aniruddh Chandratre

Senior Engineer, ML Infrastructure · Tesla AI · Palo Alto

I keep large fleets of compute busy.

Standing up a fleet this size is hard, and it is only half the problem. The other half is making sure jobs can actually land on it — at scale, accelerators go idle behind fragmented placement, stalled pipelines, and data that arrives too late to use. I build the software in between.

Right now that means the data engine behind Tesla AI: the platform every training and evaluation pipeline for FSD, Optimus, and Digital Optimus runs through, across hundreds of thousands of GPUs. What that has been worth is below, in numbers someone else had to sign for.

Every figure here is attested in signed statements from Tesla's VP of AI Software and my direct manager, filed with USCIS in support of an EB-1A petition — so they were documented under oath and examined by an independent adjudicator, not just asserted here.
MeasureWasIsWhat that bought
Usable compute capacity 24%88% Hundreds of millions of dollars of AI compute, with no new hardware
Data items processed / month 940M15B+ 17× throughput through the same fleet
Dataset size per run 1M clips1B+ clips Three orders of magnitude — pipelines that simply could not run, now routine
Compute commanded $5M/mo~$200M/mo At public cloud list price
Engineering productivity $15M/yr Engineering time returned that used to go to waiting and reruns

Compute is not the bottleneck in AI research. Getting to the compute is. Every number above is the same trade: better software, same hardware, faster iteration for the people training models.

The data engine

Internal platform · fleet-wide · owner

Every data operation inside Tesla AI lands here — acquisition, curation, mining, auto-labeling, training, evaluation. It runs over 10 billion jobs and 3,000 years of compute a month across ten large GPU clusters, commanding roughly $200M/month measured at public cloud list prices. I rewrote it from the ground up: a lower-level core, distributed workers coordinating without a central service, and a memory model that survives datasets three orders of magnitude past what the original could hold.

  • 10B jobs / month
  • 3,000 yr compute / month
  • 10 GPU clusters
  • 3 AI programs, all data ops

Scheduling & live defragmentation

cortex · cortex2 · GB300 · fleet-wide

Big training jobs need contiguous topology, and a fleet left alone fragments until the scheduler can't place a job the hardware could obviously run. I build the placement and live-compaction layer that moves allocations back into clean blocks while the cluster stays hot — across heterogeneous racks, normalized to H100 equivalents. Effective utilization now averages 98% across the fleet, up from 24% usable capacity before.

  • 250K+ H100-equivalent GPUs
  • 98% effective utilization, fleet average
  • 24 → 88% usable capacity

Internet-scale data acquisition

FSD · Optimus · Digital Optimus

The pipelines that turn raw collection into training-ready corpora, for every program in the org. A model's quality ceiling is set long before a GPU is ever allocated.

  • 1B+ clips per run
  • 15B+ data items / month
  • end-to-end capture → training set

Core data infrastructure — Tesla Manufacturing

2023 – 2024

Built the data backbone for the gigafactories: an any-to-any broker sinking anything into anything — MQTT into InfluxDB, Kafka into ClickHouse. Underneath it, an mmap-backed durable queue sustaining 5M writes/second on NFS-backed Kubernetes volumes 100× slower than an SSD, and a distributed cache at 10M entries/second that deduplicated across 50 deployments and cut storage by 80%. Over 10 billion manufacturing data points a day, in production in Austin, Berlin, and Fremont.

Graduate research — ASU, DARPA ARCOS

2021 – 2022

Formal verification and falsification for autonomous systems, with Lockheed Martin under DARPA's Automated Rapid Certification of Software program. Co-authored PSY-TaLiRo, still in use in the CertGATE certification pipeline. Eight publications, 174 citations; first-author work at HSCC 2023 formalizing stealthy attacks on cyber-physical systems as temporal logic.

Founding engineer — AfterShoot

2020 – 2021

Built the image ranking model that is still the product's core six years later. Now used by ~250,000 professional photographers, nine billion images processed, $55M valuation.

Every system on this page is the same bet: the thing that decides how fast a team can move is the distance between having an idea and having a result. Factory telemetry, a scheduler, a data engine — different layers, one problem. Make the loop tighter and everything downstream of it gets better for free, without anyone having to buy more hardware or work more hours.

  • O-1A · current
  • EB-1A · approved

Both are extraordinary-ability classifications, granted on the record of the work above — the petitions were built on these systems and reviewed against them.

achandratre22@gmail.com