Computer Science · NYU Courant

Boyang Zheng

I'm currently a first-year CS PhD student at NYU Courant, advised by Saining Xie. I obtained my bachelor degree at Shanghai Jiao Tong University, ACM Honor Class.

My research aims to understand the dynamics of high-dimensional, continuous, and often noisy data spaces. Representations, as a special form of data generated by neural networks, are particularly interesting to me. I'm specifically attracted to how to make predictions (of various types) on these representations and to model them efficiently. Concretely, I mainly work on (visual) representation learning, generative models, and multimodal learning. I also have some experience in low-level vision and adversarial examples.

Portrait of Boyang Zheng
  1. Internship begins! I'm now a summer intern at AMI Labs.

  2. New blog post (with Peter Tong): Lessons from Two Years of TPU Training in Academia — sharing hard-earned TPU debugging wisdom!

  3. Enrolled as a PhD student at NYU Courant, advised by Saining Xie.

Earlier updates
  1. Graduated from Shanghai Jiao Tong University with an Honor degree in Computer Science, ACM Honor Class.

  2. Internship begins! I'm now an intern at NYU VisionX Lab, advised by Saining Xie, doing research on generative models and MLLMs. I'll be on site at July, see you in New York!

  3. Internship begins! I'm now an intern at Shanghai AI Lab, advised by Chao Dong, doing research on MLLM and their possible applications on low-level vision tasks.

Research

Publications

arXiv 2026

Beyond Language Modeling: An Exploration of Multimodal Pretraining

Shengbang Tong*, David Fan*, John Nguyen*, Ellis Brown, Gaoyue Zhou, Shengyi Qian, Boyang Zheng, Théophane Vallaeys, Junlin Han, Rob Fergus, Naila Murray, Marjan Ghazvininejad, Mike Lewis, Nicolas Ballas, Amir Bar, Michael Rabbat, Jakob Verbeek, Luke Zettlemoyer†, Koustuv Sinha†, Yann LeCun†, Saining Xie†

A systematic study of unified multimodal pretraining with representation autoencoders and Mixture-of-Experts, showing how visual data complements language, enables world modeling, and benefits both understanding and generation.

ICLR 2026

Diffusion Transformers with Representation Autoencoders

Boyang Zheng, Nanye Ma, Shengbang Tong, Saining Xie

A class of autoencoders that utilize pretrained, frozen representation encoders as encoders and train ViT decoders on top. Training Diffusion Transformers in the latent space of RAE achieves strong performance and fast convergence on image generation tasks.

Beyond Research

NY Milk Tea Board A personal ranking of milk tea and bottled tea around New York.