Yamini Bansal
yamini [at] yaminibansal.com
I work in AI research. I did my PhD at Harvard with Boaz Barak and David Cox. My thesis was on the theory of deep learning, taking an empirical approach to theoretical questions around over-parameterization, generalization and representations. Before that, I was an undergraduate at IIT Bombay.
I also paint, mostly in oil, some of it outdoors. I trained at the Art Students League in 2024 and 2025 with Ricky Mujica, Garin Baker and Sherry Camhy. Paintings.
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. Gemini Team (contributor). arXiv:2403.05530, 2024. [arXiv]
- Gemini: A family of highly capable multimodal models. Gemini Team (contributor). arXiv:2312.11805, 2023. [arXiv]
- Beyond human data: Scaling self-training for problem-solving with language models. A. Singh, J. D. Co-Reyes, R. Agarwal, A. Anand, P. Patil, X. Garcia, P. J. Liu, et al. arXiv:2312.06585, 2023. [arXiv]
- On privileged and convergent bases in neural network representations. D. Brown, N. Vyas, Y. Bansal. arXiv:2307.12941, 2023. [arXiv]
- The unreasonable effectiveness of few-shot learning for machine translation. X. Garcia, Y. Bansal, C. Cherry, G. Foster, M. Krikun, M. Johnson, O. Firat. ICML, 2023. [arXiv]
- Empirical limitations of the NTK for understanding scaling laws in deep learning. N. Vyas, Y. Bansal, P. Nakkiran. TMLR, 2023. [arXiv]
- Data scaling laws in NMT: The effect of noise and architecture. Y. Bansal, B. Ghorbani, A. Garg, B. Zhang, C. Cherry, B. Neyshabur, O. Firat. ICML, 2022. [arXiv]
- Building the theoretical foundations of deep learning: An empirical approach. Y. Bansal. PhD thesis, Harvard University, 2022. [pdf]
- Deep double descent: Where bigger models and more data hurt. P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, I. Sutskever. ICLR 2020; J. Stat. Mech., 2021. [arXiv]
- Revisiting model stitching to compare neural representations. Y. Bansal, P. Nakkiran, B. Barak. NeurIPS, 2021. [arXiv]
- For self-supervised learning, rationality implies generalization, provably. Y. Bansal, G. Kaplun, B. Barak. ICLR, 2021. [arXiv]
- Distributional generalization: A new kind of generalization. P. Nakkiran, Y. Bansal. arXiv:2009.08092, 2020. [arXiv]
- Improving the reconstruction of disentangled representation learners via multi-stage modelling. A. Srivastava, Y. Bansal, Y. Ding, C. Hurwitz, K. Xu, B. Egger, P. Sattigeri, et al. arXiv:2010.13187, 2020. [arXiv]
- On the information bottleneck theory of deep learning. A. M. Saxe, Y. Bansal, J. Dapello, M. Advani, A. Kolchinsky, B. D. Tracey, D. D. Cox. ICLR 2018; J. Stat. Mech., 2019. [link]
- Minnorm training: An algorithm for training over-parameterized deep neural networks. Y. Bansal, M. Advani, D. D. Cox, A. M. Saxe. arXiv:1806.00730, 2018. [arXiv]
In late 2023, after a protracted illness of the soul brought on by thinking for a living, I discovered oil painting. Much like Matisse, I found in it a kind of paradise. This is the work of the two years since, much of it owed to the Art Students League.
Hover or tap a red mark to see the painting.