Engineered distributed training infrastructure for multi-billion parameter early-fusion models on NVIDIA H200 clusters using PyTorch FSDP and optimized multiprocessing.
Architected multi-million example datasets from raw video data to train and fine-tune world models.
Implemented efficient pipelines to train ensembled models across multiple cross-validation splits in PyTorch, classifying objects across 100GB+ of stellar light curve data.
Developed robust training infrastructure implementing FGSM and PGD attacks; improved model reliability from 50% to 90% while maintaining zero accuracy regression.
Engineered an adversarial benchmarking suite and integrated robustness evaluation into the ML production pipeline to automate large-scale experiment tracking.