Paper Accepted for the ACCV 2026 International Conference / Research Team Led by Professor Lee SeongWon (Department of Electrical Engineering)
- Development of image search technology that understands complex search intent within long sentences… Proposal of a new benchmark and search framework
- 26.09.30 / 홍유민
A research team led by Professor Lee Seong-won of the Department of Electrical Engineering at Kookmin University has achieved the distinction of having a paper accepted at the 18th Asian Conference on Computer Vision (ACCV 2026), a major international conference in the field of computer vision. This research aims to improve the accuracy of search technologies that utilize both images and natural language, proposing a new method to effectively incorporate users’ detailed requirements—described in multiple sentences—into search results.
ACCV is an international conference sponsored by the Asian Federation of Computer Vision and held every two years. It serves as a forum where researchers from universities and research institutions around the world gather to share the latest research findings in the fields of computer vision, machine learning, and artificial intelligence.
The accepted paper, titled “Beyond Single Sentences: Composed Image Retrieval with Long-Form Modification Texts,” addresses “Composed Image Retrieval,” a technique that finds target images by inputting both a reference image and a description of the desired modifications.
For example, if a user provides a photo along with a description of how they would like to change the subject’s color, background, or position, the technology searches for images that meet those conditions.
Existing Composed Image Retrieval (CIRR) datasets primarily used short, single-sentence descriptions, which had the limitation of being unable to adequately express multiple visual changes or detailed conditions.
To address this, the research team proposed a new benchmark called “L-CIRR,” which enriches the image pairs in the existing CIRR dataset with long, detailed descriptions. This allows for the description of various changes—including not only the subject’s attributes but also spatial relationships and scene context—enabling the system to learn and evaluate complex search requests.
Additionally, the research team developed “PACE,” an image search framework designed to effectively utilize long descriptions. PACE incorporates descriptions sequentially, sentence by sentence, to progressively refine search criteria. Furthermore, by applying “multi-stage hard negative mining”—which identifies images that are easily confused with the ground truth at each stage and uses them for training—the framework enhanced the ability to distinguish subtle differences among visually similar images.
This research is significant in that it highlights not only the structure of image search models but also the importance of training data that explicitly expresses search intent. It is expected to serve as the foundation for image search technologies that reflect users’ specific needs in areas such as product search for online shopping, visual recommendations, and digital content exploration.
Meanwhile, ACCV 2026 will be held from December 14 to 18, 2026, at the Grand Cube Osaka in Osaka, Japan, and this paper is scheduled to be presented during the main conference sessions from December 16 to 18.

|
This content is translated from Korean to English using the AI translation service DeepL and may contain translation errors such as jargon/pronouns. If you find any, please send your feedback to kookminpr@kookmin.ac.kr so we can correct them.
|
|
Paper Accepted for the ACCV 2026 International Conference / Research Team Led by Professor Lee SeongWon (Department of Electrical Engineering) - Development of image search technology that understands complex search intent within long sentences… Proposal of a new benchmark and search framework |
||||
|---|---|---|---|---|
|
A research team led by Professor Lee Seong-won of the Department of Electrical Engineering at Kookmin University has achieved the distinction of having a paper accepted at the 18th Asian Conference on Computer Vision (ACCV 2026), a major international conference in the field of computer vision. This research aims to improve the accuracy of search technologies that utilize both images and natural language, proposing a new method to effectively incorporate users’ detailed requirements—described in multiple sentences—into search results. ACCV is an international conference sponsored by the Asian Federation of Computer Vision and held every two years. It serves as a forum where researchers from universities and research institutions around the world gather to share the latest research findings in the fields of computer vision, machine learning, and artificial intelligence. The accepted paper, titled “Beyond Single Sentences: Composed Image Retrieval with Long-Form Modification Texts,” addresses “Composed Image Retrieval,” a technique that finds target images by inputting both a reference image and a description of the desired modifications. Existing Composed Image Retrieval (CIRR) datasets primarily used short, single-sentence descriptions, which had the limitation of being unable to adequately express multiple visual changes or detailed conditions. Additionally, the research team developed “PACE,” an image search framework designed to effectively utilize long descriptions. PACE incorporates descriptions sequentially, sentence by sentence, to progressively refine search criteria. Furthermore, by applying “multi-stage hard negative mining”—which identifies images that are easily confused with the ground truth at each stage and uses them for training—the framework enhanced the ability to distinguish subtle differences among visually similar images. This research is significant in that it highlights not only the structure of image search models but also the importance of training data that explicitly expresses search intent. It is expected to serve as the foundation for image search technologies that reflect users’ specific needs in areas such as product search for online shopping, visual recommendations, and digital content exploration.
|
||||






