품사 정보를 활용한 이미지 캡션 생성

강필구; 임유빈; 김형주; Philgoo Kang; Yubin Lim; Hyoungjoo Kim

연구문헌

국내 논문지

홈 > 연구문헌 > 국내 논문지 > 한국정보과학회 논문지 > 정보과학회논문지 (Journal of KIISE)

정보과학회논문지 (Journal of KIISE)

Current Result Document :

한글제목(Korean Title)	품사 정보를 활용한 이미지 캡션 생성
영문제목(English Title)	Boosting Image Caption Generation with Parts of Speech
저자(Author)	강필구 임유빈 김형주 Philgoo Kang Yubin Lim Hyoungjoo Kim
원문수록처(Citation)	VOL 48 NO. 03 PP. 0317 ~ 0324 (2021. 03)
한글내용 (Korean Abstract)	일상 생활 속에서의 스마트 기기와 AI에 대한 의존도가 높아지면서, 시각 장애인 보조, 인간 컴퓨터 상호 작용 등 다양한 분야에 접목 가능한 이미지 캡션 생성 기술의 중요성이 높아지고 있다. 본 논 문에서는 캡션 생성 기능의 향상을 위해 명사, 동사와 같은 언어의 품사(POS) 정보를 이미지로부터 추출하여 활용하는 새로운 기법을 제안한다. 제안하는 모델은 복수의 CNN 인코더를 품사 별로 학습하여 품사 별 특징 벡터를 추출한 후, 추출한 품사 벡터를 LSTM에 입력하여 캡션을 생성한다. 제안한 모델은 Flickr30k, MS-COCO 데이터 셋에 대해 실험을 진행하며, 사람을 대상으로 2가지 설문 조사를 진행하여 결과물의 실질적인 유효성을 검증한다.
영문내용 (English Abstract)	With the integration of smart devices and reliance on AI into our daily lives, the ability to generate image caption is becoming increasingly important in various fields such as guidance for visually-impaired individuals, human-computer interaction and so on. In this paper, we propose a novel approach based on parts of speech (POS), such as nouns and verbs extracted from image to enhance the image caption generation. The proposed model exploits multiple CNN encoders, which were specifically trained to identify features related to POS, and feed them into an LSTM decoder to generate image captions. We conducted experiments involving both Flickr30k and MS-COCO datasets using several text metrics and additional human surveys to validate the practical effectiveness of the proposed model.
키워드(Keyword)	이미지 캡션 생성 인코더-디코더 구조 품사 컴퓨터 비전 image caption generation encoder-decoder architec parts of speech computer vision
파일첨부	PDF 다운로드