- Vision-language models for video and document understanding
- World models and world action models for robot control
|
Yongjoo Kim Hello! I am an integrated B.S./M.S. student in Computer Science and Engineering at Korea University, co-advised by Prof. Jungbeom Lee and Prof. Jinkyu Kim at the Vision & AI Lab. I received my B.E. in Computer Science and Engineering from Korea University in 2026. My research focuses on vision-language models (VLMs) for video and document understanding, including multimodal retrieval and benchmark design. I am also interested in world models, world action models (WAMs), and vision-language-action (VLA) models. Email / CV / Google Scholar / GitHub / LinkedIn |
|
| Aug. 2026 | Our paper SAGE got accepted to the EMNLP 2026 Main Conference. See you in Budapest! 🇭🇺 |
| Aug. 2026 | Won the Excellence Award (5th place) at the 3rd Wind Power Generation Forecasting AI Competition (BARAM 2026). |
| Mar. 2026 | Started the integrated B.S./M.S. program at Vision & AI Lab @ Korea University. |
| Dec. 2025 | Won the Bronze Prize at the 2025 Inthon Datathon Track. |
| Jun. 2025 | Joined Lingora as an AI Engineer Intern. |
| Apr. 2025 | Joined Vision & AI Lab @ Korea University as an Undergraduate Researcher. |
I work on vision-language models for understanding videos and visually rich documents, with a particular interest in multimodal retrieval and in designing benchmarks that expose what current models fail to capture.
|
SAGE: Semantic Attribute Graphs for Multi-Entity Visual Retrieval
Yongjoo Kim*, Mincheol Kwon, Seonga Choi, Minseung Lee, Kyeong-Jin Oh, Hyunyoung Lee, Yunsu Choi, Jungbeom Lee Conference on Empirical Methods in Natural Language Processing (EMNLP), 2026 We introduce DEAR, a benchmark for visual retrieval over images containing multiple entities, and propose SAGE, a training-free retrieval method that represents each image as a graph of entities and their semantic attributes. |
* Equal contribution