Yuta Oshima

I’m a Ph.D. student at The University of Tokyo, mentored by Professor Yutaka Matsuo.

My research goal is to develop algorithmic advances for vision foundation models — simple ideas that hold or even grow as models scale, and become part of how they are built and used.

Toward this goal, I currently work on the alignment and evaluation of image and video generation models to measure and elicit what they can do.

selected publications

  1. CVPR 2026 Main
    multibanana_teaser.png
    MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image Generation
    Yuta Oshima, Daiki Miyake, Kohsei Matsutani, Yusuke Iwasawa, Masahiro Suzuki, Yutaka Matsuo, and Hiroki Furuta
    In the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
  2. NeurIPS 2025
    dlbs_teaser.gif
    Inference-Time Text-to-Video Alignment with Diffusion Latent Beam Search
    Yuta Oshima, Masahiro Suzuki, Yutaka Matsuo, and Hiroki Furuta
    In Neural Information Processing Systems (NeurIPS), 2025
  3. NeurIPS 2024
    adopt.png
    ADOPT: Modified Adam Can Converge with Any \beta_2 with the Optimal Rate
    Shohei Taniguchi, Keno Harada, Gouki Minegishi, Yuta Oshima, Seong Cheol Jeong, Go Nagahara, Tomoshi Iiyama, Masahiro Suzuki, Yusuke Iwasawa, and Yutaka Matsuo
    In Neural Information Processing Systems (NeurIPS), 2024