Style-aware contrastive adapters for multi-style image captioning

Authors

  • Riyan Mahmudin Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember, Indonesia
  • Shintami Chusnul Hidayati Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember, Indonesia https://orcid.org/0000-0001-5045-4842
  • Nanik Suciati Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember, Indonesia, Indonesia https://orcid.org/0000-0002-1991-0464

DOI:

https://doi.org/10.22441/sinergi.2026.3.023

Keywords:

Contrastive Learning, Image Captioning, InstructBLIP, Parameter-Efficient Fine-Tuning, Style Adapter, Visual Instruction Tuning, Visual Language Models

Abstract

InstructBLIP demonstrates a strong ability to follow instructions through its instruction-aware Q-Former model. However, for image captioning tasks, this model is trained with instruction templates that are stylistically homogeneous and have similar descriptive intentions. In multi-style food captioning, different styles demand distinct interpretational emphases on identical visual features: ingredient-focused captions require selective attention to texture and compositional patterns, while aesthetic descriptions prioritize color harmony and spatial arrangement. This heterogeneity poses a challenge: whether a single shared Q-Former adaptation mechanism can optimally extract and weight visual features for fundamentally different descriptive objectives. This study proposes an enhanced InstructBLIP architecture integrating parallel Style Adapters for style-specific feature transformations and Style-Conditioned Contrastive Learning to regularize adapter outputs. The finding establishes that effective style-specific adaptation requires architectural complementarity: adapters provide parametric capacity for style-specialized transformations, while contrastive learning enforces geometric constraints that prevent representation collapse across styles. These components exhibit mutual dependency: adapters without regularization drift toward homogeneous representations, while contrastive learning without adaptation capacity creates conflicting optimization pressures between discriminative and generative objectives. Experiments validate this complementarity principle, achieving consistent improvements (CIDEr: 181.4 vs. baseline 179.9, +0.8%; SPICE: +1.9%), with substantial gains for structured styles (ingredient: +3.1% CIDEr). Critically, ablation studies confirm this mutual dependency: contrastive loss alone degrades performance (−2.2% CIDEr) while adapters alone yield minimal gains (−1.4% CIDEr), demonstrating that only their combination effectively stabilizes style-specific adaptations in instruction-aware captioning.


Downloads

Download data is not yet available.

Author Biographies

Riyan Mahmudin, Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember


Shintami Chusnul Hidayati, Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember


Nanik Suciati, Department of Informatics, Faculty of Electrical Technology and Intelligent Informatics, Institut Teknologi Sepuluh Nopember, Indonesia


Published

2026-09-30

How to Cite

[1]
R. Mahmudin, S. C. Hidayati, and N. Suciati, “Style-aware contrastive adapters for multi-style image captioning”, Sinergi, vol. 30, no. 3, pp. 995–1010, Sep. 2026.

Issue

Section

Articles

Similar Articles

<< < > >> 

You may also start an advanced similarity search for this article.