Pertanika Journal of Science & Technology
Pertanika Journal Home Facebook
Pertanika ยท Universiti Putra Malaysia Press

Pertanika Journal of Science & Technology

Official journal of Universiti Putra Malaysia for scholarly work across science, engineering and related technologies.

e-ISSN 2231-8526 ISSN 0128-7680
Research article

Enhanced Image Captioning Using Sinusoidal Spatial Transform CNN and Fine-tuned BERT

Bhargavi Polepalli, Praveen Kumar Sekharamantry, and Konda Srinivasa Rao

https://doi.org/10.47836/pjst.34.S1.05
KeywordsBright and contrast augmentation, fine-tuned BERT, Flexiframe Filter, Sinusoidal Spatial Transform Convolutional Neural Network (SST-CNN), Stochastic k-sampling DenseGRU Neural Network
Article content

Abstract

Image caption generation is a significant area of study in artificial intelligence and computer vision, focusing on training systems to generate accurate textual descriptions of images. This paper presents an advanced framework for image captioning using a dataset sourced from Kaggle. The process begins with pre-processing techniques, including the Flexiframe Filter, which dynamically adjusts the window size based on local variance to reduce noise in both text and images. Bright- Contrast augmentation is then applied to enhance input images, enriching the dataset and improving feature extraction. The encoder module utilises a Sinusoidal Spatial Transform Convolutional Neural Network (SST-CNN) integrated with a Visual Geometry Group (VGG16) model for feature extraction and encoding. For the decoder, stochastic k-sampling is combined with a Dense Neural Network and a Gated Recurrent Unit (GRU) to generate initial captions. To refine these captions, a fine-tuned Bidirectional Encoder Representations from Transformers (BERT) model is employed for enhanced coherence and accuracy. The model's performance is evaluated using standard metrics, achieving BLEU-1, METEOR, ROUGE-L, and CIDER scores of 1, 0.99, 0.99, and 1, respectively. The proposed approach demonstrates significant improvements in generating descriptive and accurate image captions, making it a robust solution for applications requiring semantic understanding of visual data.