ELIMINATING THE QUALITY PARADOX IN VIZWIZ CAPTIONING VIA DATA-CENTRIC CURATION AND PEFT LORA
DOI:
https://doi.org/10.33480/jitk.v12i1.8060Keywords:
Accessibility, BLIP, Data Curation, Image Captioning, LoRAAbstract
Vision-language models could be very useful for assistive technology, but they struggle significantly with processing low-quality, noisy photos that visually impaired individuals often take. The primary contribution of this research is a data-centric pipeline that combines Fast Fourier Transform (FFT) blur filtering with Parameter-Efficient Fine-Tuning (PEFT) utilizing Low-Rank Adaptation (LoRA) to directly confront and eradicate the “Quality Paradox” bias in accessibility datasets like VizWiz. We use the AdamW optimizer for memory-efficient training and perform comprehensive statistical ablation studies with FFT-based tertile splits to see how image quality affects the results. The optimized model achieves a CIDEr score of 0.4481, a strong BERTScore F1 of 0.8815, and a BLEU-4 of 0.0511. Statistical analysis (p = 0.622, Cohen’s d = -0.038) shows that this bias has been successfully removed, which means that performance is balanced on both severely blurry and crisp images. Further qualitative testing shows that the model can recognize text and objects even in very low light without needing an extra OCR module, demonstrating its viability as a stand-alone navigational aid.
Downloads
References
[1] M. D. Messaoudi, B.-A. J. Menelas, and H. Mcheick, “Review of Navigation Assistive Tools and Technologies for the Visually Impaired,” Sensors, vol. 22, no. 20, 2022, doi: 10.3390/s22207888.
[2] G. I. Okolo, T. Althobaiti, and N. Ramzan, “Assistive Systems for Visually Impaired Persons: Challenges and Opportunities for Navigation Assistance,” Sensors, vol. 24, no. 11, 2024, doi: 10.3390/s24113572.
[3] S. A. Doore, D. Istrati, C. Xu, Y. Qiu, A. Sarrazin, and N. A. Giudice, “Images, Words, and Imagination: Accessible Descriptions to Support Blind and Low Vision Art Exploration and Engagement,” J. Imaging, vol. 10, no. 1, 2024, doi: 10.3390/jimaging10010026.
[4] J. Li, D. Li, C. Xiong, and S. Hoi, “BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation,” in Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato, Eds., in Proceedings of Machine Learning Research, vol. 162. PMLR, Aug. 2022, pp. 12888–12900. [Online]. Available: https://proceedings.mlr.press/v162/li22n.html
[5] V. Deshpande, G. Shelke, and B. Kadam, “Empowering Vision: A Survey on Image Captioning Assistive Technologies for the Visually Impaired,” in Smart Trends in Computing and Communications, T. Senjyu, C. So-In, and A. Joshi, Eds., Singapore: Springer Nature Singapore, 2026, pp. 357–368. doi: 10.1007/978-981-96-7505-0_28.
[6] F. Y. Mekonnen et al., “Development of a Fully Autonomous Offline Assistive System for Visually Impaired Individuals: A Privacy-First Approach,” Sensors, vol. 25, no. 19, 2025, doi: 10.3390/s25196006.
[7] D. Gurari, Y. Zhao, M. Zhang, and N. Bhattacharya, “Captioning Images Taken by People Who Are Blind,” in Computer Vision – ECCV 2020, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds., Cham: Springer International Publishing, 2020, pp. 417–434.
[8] A. Karamolegkou, P. Rust, R. Cui, Y. Cao, A. Søgaard, and D. Hershcovich, “Vision-Language Models under Cultural and Inclusive Considerations,” in Proceedings of the 1st Human-Centered Large Language Modeling Workshop, N. Soni, L. Flek, A. Sharma, D. Yang, S. Hooker, and H. A. Schwartz, Eds., ACL, Aug. 2024, pp. 53–66. doi: 10.18653/v1/2024.hucllm-1.5.
[9] P. Patel, S. Pampaniya, A. Ghosh, R. Raj, D. Karuppaih, and S. Kandasamy, “Enhancing Accessibility Through Machine Learning: A Review on Visual and Hearing Impairment Technologies,” IEEE Access, vol. 13, pp. 33286–33307, 2025, doi: 10.1109/ACCESS.2025.3539081.
[10] J. Guo, J. Ma, Á. F. García-Fernández, Y. Zhang, and H. Liang, “A survey on image enhancement for Low-light images,” Heliyon, vol. 9, no. 4, p. e14558, 2023, doi: 10.1016/j.heliyon.2023.e14558.
[11] A. S. Waghmare and S. S. Satonkar, “A Comparative Study on Blur Detection and Image Restoration Techniques: Traditional Methods vs. Fuzzy Logic,” Journal of Harbin Engineering University, vol. 46, no. 7, 2025.
[12] E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations (ICLR), 2022. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
[13] X. Wang, L. Bai, C. Gu, and Y. Wu, “Universal Approximated Real-Valued Fast Fourier Transform for Image Blur Detection,” in 2025 35th Irish Signals and Systems Conference (ISSC), 2025, pp. 1–5. doi: 10.1109/ISSC67739.2025.11291523.
[14] S. Panjeh, A. Nordahl-Hansen, and H. Cogo-Moreira, “Establishing New Cutoffs for Cohen’s d: An Application Using Known Effect Sizes from Trials for Improving Sleep Quality on Composite Mental Health,” Int. J. Methods Psychiatr. Res., vol. 32, no. 3, p. e1969, 2023, doi: 10.1002/mpr.1969.
[15] L. V Hedges, “Interpretation of the Standardized Mean Difference Effect Size When Distributions Are Not Normal or Homoscedastic,” Educ. Psychol. Meas., vol. 85, no. 2, pp. 245–257, 2025, doi: 10.1177/00131644241278928.
[16] M. Mandal, D. Ghadiyaram, D. Gurari, and A. C. Bovik, “Helping Visually Impaired People Take Better Quality Pictures,” IEEE Transactions on Image Processing, vol. 32, pp. 3873–3884, 2023, doi: 10.1109/TIP.2023.3282067.
[17] R. A. Bafghi and D. Gurari, “A New Dataset Based on Images Taken by Blind People for Testing the Robustness of Image Classification Models Trained for ImageNet Categories,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 16261–16270. doi: 10.1109/CVPR52729.2023.01560.
[18] Y. Wang, Y. Deng, Y. Zheng, P. Chattopadhyay, and L. Wang, “Vision Transformers for Image Classification: A Comparative Survey,” Technologies (Basel)., vol. 13, no. 1, 2025, doi: 10.3390/technologies13010032.
[19] T. Wolf et al., “Transformers: State-of-the-Art Natural Language Processing,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Q. Liu and D. Schlangen, Eds., Online: Association for Computational Linguistics, Oct. 2020, pp. 38–45. doi: 10.18653/v1/2020.emnlp-demos.6.
[20] X. Liu, H. Qi, S. Jia, Y. Guo, and Y. Liu, “Recent Advances in Optimization Methods for Machine Learning: A Systematic Review,” Mathematics, vol. 13, no. 13, 2025, doi: 10.3390/math13132210.
[21] Y. Ming, N. Hu, C. Fan, F. Feng, J. Zhou, and H. Yu, “Visuals to Text: A Comprehensive Review on Automatic Image Captioning,” IEEE/CAA Journal of Automatica Sinica, vol. 9, no. 8, pp. 1339–1365, 2022, doi: 10.1109/JAS.2022.105734.
[22] S. Sarto, M. Cornia, and R. Cucchiara, “Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives,” in Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI), 2025, pp. 10632–10640. doi: 10.24963/ijcai.2025/1180.
[23] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating Text Generation with BERT,” in International Conference on Learning Representations (ICLR), 2020. [Online]. Available: https://openreview.net/forum?id=SkeHuCVFDr
[24] J. Hessel, A. Holtzman, M. Forbes, R. Le Bras, and Y. Choi, “CLIPScore: A Reference-free Evaluation Metric for Image Captioning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, M.-F. Moens, X. Huang, L. Specia, and S. W. Yih, Eds., Online and Punta Cana, Dominican Republic: Association for Computational Linguistics, Nov. 2021, pp. 7514–7528. doi: 10.18653/v1/2021.emnlp-main.595.
[25] D. Santi, A. A. Ilham, Syafaruddin, and I. Nurtanio, “Attention-Driven Image Captioning for Mobile Accessibility of the Visually Impaired,” International Journal of Interactive Mobile Technologies (iJIM), vol. 19, no. 9, pp. 4–18, May 2025, doi: 10.3991/ijim.v19i09.53441.
[26] M. Rifki, A. Bastian, and A. Mardiana, “Development of CNN-LSTM-Based Image Captioning Dataset to Enhance Visual Accessibility for Disabilities,” JITK, vol. 10, no. 4, pp. 980–992, May 2025, doi: 10.33480/jitk.v10i4.6657.
[27] A. Khan and J. Singh, “A Novel Image Captioning Technique Using Deep Learning Methodology,” ICCK Transactions on Machine Intelligence, vol. 1, no. 2, pp. 52–68, 2025, doi: 10.62762/TMI.2025.886122.
[28] C.-S. He, N.-K. Lo, Y.-H. Chien, and S.-S. Lin, “Image Descriptions for Visually Impaired Individuals to Locate Restroom Facilities,” Engineering Proceedings, vol. 92, no. 1, 2025, doi: 10.3390/engproc2025092013.
[29] L. Orynbay, B. Razakhova, P. Peer, B. Meden, and Ž. Emeršič, “Recent Advances in Synthesis and Interaction of Speech, Text, and Vision,” Electronics (Basel)., vol. 13, no. 9, 2024, doi: 10.3390/electronics13091726.
[30] N. Joshika, S. Srikanth, S. Suraj, S. Pradeep, and A. N. Sree, “A Multimodel GPT Based Application to Help Blind Users Visual Picture Extensions,” Journal of Science Engineering Technology and Management Science, vol. 2, no. 8, pp. 437–443, Aug. 2025, doi: 10.63590/jsetms.2025.v02.i08.pp437-443.
[31] A. M. Hilal, F. Alrowais, F. N. Al-Wesabi, and R. Marzouk, “Red Deer Optimization with Artificial Intelligence Enabled Image Captioning System for Visually Impaired People,” Computer Systems Science and Engineering, vol. 46, no. 2, pp. 1929–1945, 2023, doi: 10.32604/csse.2023.035529.
[32] H. Kim, C. Park, J. Jang, J. Lee, J. Yoon, and J. Paik, “Visual Question Answering with Multimodal Learning for VizWiz-VQA,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, Seattle, WA, USA, Jun. 2024.
[33] H. Fadhilah and N. Utama, “Systematic Literature Review on Medical Image Captioning Using CNN-LSTM and Transformer-Based Models,” Jurnal Masyarakat Informatika, vol. 16, no. 1, pp. 32–53, 2025, doi: 10.14710/jmasif.16.1.73127.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Rhama Rangga Dhiputra; Anna Baita

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.






-a.jpg)
-b.jpg)











