Yuta Oshima
I’m a Ph.D. student at The University of Tokyo, mentored by Professor Yutaka Matsuo.
I conduct research on generative models, image and video generation, and world models. My work aims to develop scalable, general-purpose generative models to create and simulate the visual world.
selected publications
- CVPR 2026 Main
MultiBanana: A Challenging Benchmark for Multi-Reference Text-to-Image GenerationIn the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026 - TMLR 2026
WorldPack: Dynamic Frame Compression for Long-context Video World ModelingIn Transactions on Machine Learning Research, 2026