REAL-TIME MULTIMODAL LARGE LANGUAGE MODEL FOR ENDOSCOPIC POLYP DETECTION WITH ADAPTIVE FALSE POSITIVE SUPPRESSION VIA NEGATIVE VECTOR DATABASE

Authors

  • Son Do Huu https://scholar.google.com/citations?user=7a9PxlQAAAAJ&hl=vi
  • TRAN HUU PHUC
  • PHAN VAN NAM

DOI:

https://doi.org/10.51453/3093-3706/2026/1473

Keywords:

Multimodal Large Language Model, YOLO, Polyp Detection, Negative Vector Database, Endoscopy, Computer-Aided Detection, FAISS

Abstract

Colorectal cancer (CRC) remains one of the leading causes of cancer-related mortality worldwide, with early detection of polyps during colonoscopy being critical for prevention. Existing computer-aided detection (CADe) systems based on deep learning achieve promising accuracy but suffer from high false positive rates and lack the ability to generate interpretable clinical reports. In this paper, we propose a novel Multimodal Large Language Model (MLLM) framework that integrates YOLOv8n-seg for real-time instance segmentation with Qwen2.5-1.5B for medical report generation via LLaVA-style prefix injection. Our key contribution is the Negative Vector Database (NegDB), an adaptive false positive suppression mechanism that leverages FAISS cosine similarity search on masked-ROI embeddings to learn from doctor feedback in real-time. Unlike conventional approaches that store whole-frame embeddings, our method encodes only the segmentation-masked region of interest, achieving more precise false positive identification. The system further incorporates a voice-controlled clinical workflow using Web Speech API, enabling hands-free operation during endoscopic procedures. Experimental results on the Kvasir-SEG and CVC-ClinicDB datasets demonstrate that our proposed pipeline achieves a precision of 91.2%, recall of 86.4%, and F1-score of 88.7%, outperforming the YOLOv8n-seg baseline by 9.2%, 8.4%, and 8.7% respectively. The Negative Vector Database reduces the false positive rate from 18.3% to 4.1% after 50 doctor feedback iterations, demonstrating effective online learning capability. The complete system operates at 28 FPS on a single NVIDIA RTX 3060 (12GB), meeting real-time clinical requirements.

Downloads

Download data is not yet available.

References

1. H. Sung, J. Ferlay, R. L. Siegel, et al., “Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 71, no. 3, pp. 209–249, 2021.

2. A. G. Zauber, S. J. Winawer, M. J. O'Brien, et al., “Colonoscopic polypectomy and long-term prevention of colorectal-cancer deaths,” New England Journal of Medicine, vol. 366, no. 8, pp. 687–696, 2012.

3. J. C. van Rijn, J. B. Reitsma, J. Stoker, P. M. Bossuyt, S. J. van Deventer, and E. Dekker, “Polyp miss rate determined by tandem colonoscopy: A systematic review,” American Journal of Gastroenterology, vol. 101, no. 2, pp. 343–350, 2006.

4. G. Jocher, A. Chaurasia, and J. Qiu, “YOLO by Ultralytics,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics

5. G. Jocher, A. Chaurasia, and J. Qiu, “YOLOv8: A new state of the art for object detection,” Ultralytics, 2023. [Online]. Available: https://docs.ultralytics.com

6. D. Fan, G. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao, “PraNet: Parallel reverse attention network for polyp segmentation,” in Proc. MICCAI, 2020, pp. 263–273.

7. J. Wei, Y. Hu, R. Zhang, Z. Li, S. K. Zhou, and S. Cui, “Shallow attention network for polyp segmentation,” in Proc. MICCAI, 2021, pp. 699–708.

8. B. Dong, W. Wang, D. Fan, J. Li, H. Fu, and L. Shao, “Polyp-PVT: Polyp segmentation with pyramid vision transformers,” CAAI Artificial Intelligence Research, vol. 2, pp. 1–11, 2023.

9. OpenAI, “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023.

10. H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. NeurIPS, 2023.

11. C. Li, C. Wong, S. Zhang, et al., “LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day,” in Proc. NeurIPS, 2023.

12. Qwen Team, “Qwen2.5: A party of foundation models,” arXiv preprint arXiv:2412.15115, 2024.

13. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. ICLR, 2022.

14. J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021.

15. O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. MICCAI, 2015, pp. 234–241.

16. T. Tu, S. Azizi, D. Driess, et al., “Towards generalist biomedical AI,” NEJM AI, vol. 1, no. 3, 2024.

17. X. Zhang, C. Wu, Z. Zhao, et al., “PMC-VQA: Visual instruction tuning for medical visual question answering,” arXiv preprint arXiv:2305.10415, 2023.

18. F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and clustering,” in Proc. CVPR, 2015, pp. 815–823.

19. S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Trans. PAMI, vol. 39, no. 6, pp. 1137–1149, 2017.

20. A. Hermans, L. Beyer, and B. Leibe, “In defense of the triplet loss for person re-identification,” arXiv preprint arXiv:1703.07737, 2017.

21. T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. ICCV, 2017, pp. 2980–2988.

22. A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in Proc. CVPR, 2016, pp. 761–769.

23. H. R. Roth, L. Lu, A. Seff, et al., “A new 2.5D representation for lymph node detection using random sets of deep convolutional neural network observations,” in Proc. MICCAI, 2014, pp. 520–527.

24. J. Guo, T. He, Z. Lin, et al., “Voice-controlled surgical robots: Current status and future directions,” Surgical Innovation, vol. 27, no. 4, pp. 440–449, 2020.

25. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML, 2023, pp. 28492–28518.

26. T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Proc. NeurIPS, 2023.

27. D. Jha, P. H. Smedsrud, M. A. Riegler, et al., “Kvasir-SEG: A segmented polyp dataset,” in Proc. MMM, 2020, pp. 451–462.

28. J. Bernal, F. J. Sánchez, G. Fernández-Esparrach, D. Gil, C. Rodríguez, and F. Vilariño, “WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,” Computerized Medical Imaging and Graphics, vol. 43, pp. 99–111, 2015.

29. J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proc. CVPR, 2016, pp. 779–788.

30. A. Vaswani, N. Shazeer, N. Parmar, et al., “Attention is all you need,” in Proc. NeurIPS, 2017, pp. 5998–6008.

31. A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.

32. J. Chen, Y. Lu, Q. Yu, et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv preprint arXiv:2102.04306, 2021.

33. A. Kirillov, E. Mintun, N. Ravi, et al., “Segment Anything,” in Proc. ICCV, 2023, pp. 4015–4026.

34. H. Touvron, T. Lavril, G. Izacard, et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.

35. W. X. Zhao, K. Zhou, J. Li, et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023.

Downloads

Published

2026-07-06

How to Cite

Do Huu, S., TRAN HUU PHUC, & PHAN VAN NAM. (2026). REAL-TIME MULTIMODAL LARGE LANGUAGE MODEL FOR ENDOSCOPIC POLYP DETECTION WITH ADAPTIVE FALSE POSITIVE SUPPRESSION VIA NEGATIVE VECTOR DATABASE. SCIENTIFIC JOURNAL OF TAN TRAO UNIVERSITY, 12(2). https://doi.org/10.51453/3093-3706/2026/1473