About Me
I am a master's student at Fudan University, advised by Prof. Tao Chen. My research explores how multimodal systems perceive, reason about, and interact with the physical world.
Earlier, I conducted research with Prof. Chang Tang and Assoc. Prof. Bo Zhao. I value work that identifies meaningful new problems, makes difficult problems tractable, and develops approaches that generalize beyond a single benchmark.
Publications
We introduce the Generative Embedding Benchmark (GEB), which measures how much answer-relevant visual information can be recovered from a frozen embedding through generative readout. Across visual-only and vision-language joint settings, GEB reveals information bottlenecks that separability-based embedding benchmarks do not capture.
STI-Bench evaluates spatial-temporal understanding in MLLMs through object appearance, pose, displacement, and motion tasks. Built on 300+ videos and 2,000+ QA pairs across diverse scenes, it exposes persistent weaknesses in precise distance estimation and motion analysis.
Experience
Kuaishou
Multimodal Understanding and Applications Group, Algorithm
ByteDance
Lark AI, Multimodal Algorithm
Selected Honors
- National First Prize, China Undergraduate Mathematical Contest in Modeling (2024)
- National Third Prize, China Engineering Robot Contest (2023)
- National Endeavor Scholarship (2023 - 2025)