Congratualtions!

Our paper has been accepted to the Integrating Image Processing with Large-Scale Language/Vision Models for Advanced Visual Understanding at the IEEE International Conference on Image Processing (ICIP) 2024 [LINK]

  • Title: Unveiling the Potential of Multimodal Large Language Models for Scene Text Segmentation via Semantic-Enhanced Features

  • Authors: Ho Jun Kim, Hyung Kyu Kim, Sangmin Lee (UIUC), and Hak Gu Kim (*equal contribution)

  • Abstract: Scene text segmentation is to accurately identify text areas within a scene while disregarding non-textual elements like background imagery or graphical elements. However, current text segmentation models often fail to accurately segment text regions due to complex background noises or various font styles and sizes. To address this issue, it is essential to consider not only visual information but also semantic information of text in scene text segmentation. For this purpose, we propose a novel semantic-aware scene text segmentation framework, which incorporates multimodal large language models (MLLMs) to fuse both visual, text and linguistic information. By leveraging semantic-enhanced feature from multimodal LLMs, scene text segmentation model can remove false positives that are visually confusing but not recognized as text. Both qualitative and quantitative evaluations demonstrate that multimodal LLMs improve scene text segmentation performances.