Priya Goyal

AI Research & Engineering Leader | Agentic Systems, Evaluation & Synthetic Data | Building Learning Loop for AI Agents
LinkedIn, ex-Google Deepmind, Founding MTS DatologyAI, FAIR (Meta MSL)
Google Scholar (>70K citations)
LinkedIn
Twitter
Email
Github

My career has been focused on building increasingly capable learning systems. I started with large-scale representation learning at FAIR, moved into multimodal foundation models at Google DeepMind, then worked on data quality and evaluation at DatologyAI. Today I'm building agentic systems and, more importantly, the evaluation and data foundations that allows those agents to improve systematically. My current focus is essentially the learning loop for agents — how you generate tasks and data, measure behavior, identify failures, produce rewards, and use that feedback to make the system better.

My current research interests include agentic evaluations, synthetic data generation, data quality / curation, large scale ML.

Talks / Media coverage

TechCrunch article on ImageNet in 1-Hour.
CNBC article on SEER (training A.I. to "see").
NVIDIA Developer on ImageNet on 1-Hour.
Geekwire on ImageNet in 1-Hour.
NVIDIA Developer on Self-supervised learning beating SOTA Computer vision models.
WIRED article on AI Teaching Itself to See With Less Human Help.
CNET on training computers to learn like humans do.
ImageNet in 1-Hour at NeurIPS 2017 Supercomputing workshop.

Publications

Clip the bias: How useful is balancing data in multimodal learnings?
ICLR, 2024
Ibrahim Alabdulmohsin, Xiao Wang, Andreas Peter Steiner, Priya Goyal, Alexander D'Amour, Xiaohua Zhai
[arXiv] [bib]

Vision Models Are More Robust And Fair When Pretrained On Uncurated Images Without Supervision
arXiv, 2022
Priya Goyal, Quentin Duval, Isaac Seessel, Mathilde Caron, Ishan Misra, Levent Sagun, Armand Joulin, Piotr Bojanowski
[arXiv] [blogpost] [code] [bib]

Fairness Indicators for Systematic Assessments of Visual Feature Extractors
FAccT, 2022
Priya Goyal, Adriana Romero Soriano, Caner Hazirbas, Levent Sagun, Nicolas Usunier
[arXiv] [blogpost] [code] [bib]

A Self-Supervised Descriptor for Image Copy Detection
arXiv, 2022
Ed Pizzi, Sreya Dutta Roy, Sugosh Nagavara Ravindra, Priya Goyal, Matthijs Douze
[arXiv] [bib]

Fully Sharded Data Parallel: faster AI training with fewer GPUs
Facebook Engineering blog, 2021
Myle Ott, Sam Shleifer, Min Xu, Priya Goyal, Quentin Duval, Vittorio Caggiano
[blog] [docs]

Self-supervised pretraining of visual features in the wild
arXiv, 2021
Priya Goyal, Mathilde Caron, Benjamin Lefaudeux, Min Xu, Pengchao Wang, Vivek Pai, Mannat Singh, Vitaliy Liptchinsky, Ishan Misra, Armand Joulin, Piotr Bojanowski
[arXiv] [blogpost] [code] [bib]

VISSL: A library for state-of-the-art self-supervised learning from images
Released Jan'2021
Priya Goyal, Quentin Duval, Jeremy Reizenstein, Matthew Leavitt, Min Xu, Benjamin Lefaudeux, Mannat Singh, Vinicius Reis, Mathilde Caron, Piotr Bojanowski, Armand Joulin, Ishan Misra
[website] [tutorials] [Github] [Docs] [bib]

Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
NeurIPS 2020
Mathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal, Piotr Bojanowski, Armand Joulin
[arXiv] [blogpost] [code] [bib]

Scaling and Benchmarking Self-Supervised Visual Representation Learning
ICCV 2019
Priya Goyal, Dhruv Mahajan, Abhinav Gupta*, Ishan Misra*
[arXiv] [code] [bib]

Focal Loss for Dense Object Detection
ICCV 2017 (best student paper award)
Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, Piotr Dollár
[arXiv] [bib]

Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
arXiv 2017
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He
[arXiv] [NeurIPS 2017 talk]

Resume

PDF