Yuta Oshima
I’m a Ph.D. student at The University of Tokyo, mentored by Professor Yutaka Matsuo.
My research goal is to develop algorithmic advances for vision foundation models — simple ideas that hold or even grow as models scale, and become part of how they are built and used.
Toward this goal, I currently work on the alignment and evaluation of image and video generation models to measure and elicit what they can do.
selected publications
- CVPR 2026 Main
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationIn the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026