Style-aware contrastive adapters for multi-style image captioning
DOI:
https://doi.org/10.22441/sinergi.2026.3.023Keywords:
Contrastive Learning, Image Captioning, InstructBLIP, Parameter-Efficient Fine-Tuning, Style Adapter, Visual Instruction Tuning, Visual Language ModelsAbstract
InstructBLIP demonstrates a strong ability to follow instructions through its instruction-aware Q-Former model. However, for image captioning tasks, this model is trained with instruction templates that are stylistically homogeneous and have similar descriptive intentions. In multi-style food captioning, different styles demand distinct interpretational emphases on identical visual features: ingredient-focused captions require selective attention to texture and compositional patterns, while aesthetic descriptions prioritize color harmony and spatial arrangement. This heterogeneity poses a challenge: whether a single shared Q-Former adaptation mechanism can optimally extract and weight visual features for fundamentally different descriptive objectives. This study proposes an enhanced InstructBLIP architecture integrating parallel Style Adapters for style-specific feature transformations and Style-Conditioned Contrastive Learning to regularize adapter outputs. The finding establishes that effective style-specific adaptation requires architectural complementarity: adapters provide parametric capacity for style-specialized transformations, while contrastive learning enforces geometric constraints that prevent representation collapse across styles. These components exhibit mutual dependency: adapters without regularization drift toward homogeneous representations, while contrastive learning without adaptation capacity creates conflicting optimization pressures between discriminative and generative objectives. Experiments validate this complementarity principle, achieving consistent improvements (CIDEr: 181.4 vs. baseline 179.9, +0.8%; SPICE: +1.9%), with substantial gains for structured styles (ingredient: +3.1% CIDEr). Critically, ablation studies confirm this mutual dependency: contrastive loss alone degrades performance (−2.2% CIDEr) while adapters alone yield minimal gains (−1.4% CIDEr), demonstrating that only their combination effectively stabilizes style-specific adaptations in instruction-aware captioning.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 SINERGI

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.











