REAL-TIME MULTIMODAL LARGE LANGUAGE MODEL FOR ENDOSCOPIC POLYP DETECTION WITH ADAPTIVE FALSE POSITIVE SUPPRESSION VIA NEGATIVE VECTOR DATABASE
DOI:
https://doi.org/10.51453/3093-3706/2026/1473Keywords:
Multimodal Large Language Model, YOLO, Polyp Detection, Negative Vector Database, Endoscopy, Computer-Aided Detection, FAISSAbstract
Colorectal cancer (CRC) remains one of the leading causes of cancer-related mortality worldwide, with early detection of polyps during colonoscopy being critical for prevention. Existing computer-aided detection (CADe) systems based on deep learning achieve promising accuracy but suffer from high false positive rates and lack the ability to generate interpretable clinical reports. In this paper, we propose a novel Multimodal Large Language Model (MLLM) framework that integrates YOLOv8n-seg for real-time instance segmentation with Qwen2.5-1.5B for medical report generation via LLaVA-style prefix injection. Our key contribution is the Negative Vector Database (NegDB), an adaptive false positive suppression mechanism that leverages FAISS cosine similarity search on masked-ROI embeddings to learn from doctor feedback in real-time. Unlike conventional approaches that store whole-frame embeddings, our method encodes only the segmentation-masked region of interest, achieving more precise false positive identification. The system further incorporates a voice-controlled clinical workflow using Web Speech API, enabling hands-free operation during endoscopic procedures. Experimental results on the Kvasir-SEG and CVC-ClinicDB datasets demonstrate that our proposed pipeline achieves a precision of 91.2%, recall of 86.4%, and F1-score of 88.7%, outperforming the YOLOv8n-seg baseline by 9.2%, 8.4%, and 8.7% respectively. The Negative Vector Database reduces the false positive rate from 18.3% to 4.1% after 50 doctor feedback iterations, demonstrating effective online learning capability. The complete system operates at 28 FPS on a single NVIDIA RTX 3060 (12GB), meeting real-time clinical requirements.
Downloads
References
1. H. Sung, J. Ferlay, R. L. Siegel, et al., “Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA: A Cancer Journal for Clinicians, vol. 71, no. 3, pp. 209–249, 2021.
2. A. G. Zauber, S. J. Winawer, M. J. O'Brien, et al., “Colonoscopic polypectomy and long-term prevention of colorectal-cancer deaths,” New England Journal of Medicine, vol. 366, no. 8, pp. 687–696, 2012.
3. J. C. van Rijn, J. B. Reitsma, J. Stoker, P. M. Bossuyt, S. J. van Deventer, and E. Dekker, “Polyp miss rate determined by tandem colonoscopy: A systematic review,” American Journal of Gastroenterology, vol. 101, no. 2, pp. 343–350, 2006.
4. G. Jocher, A. Chaurasia, and J. Qiu, “YOLO by Ultralytics,” 2023. [Online]. Available: https://github.com/ultralytics/ultralytics
5. G. Jocher, A. Chaurasia, and J. Qiu, “YOLOv8: A new state of the art for object detection,” Ultralytics, 2023. [Online]. Available: https://docs.ultralytics.com
6. D. Fan, G. Ji, T. Zhou, G. Chen, H. Fu, J. Shen, and L. Shao, “PraNet: Parallel reverse attention network for polyp segmentation,” in Proc. MICCAI, 2020, pp. 263–273.
7. J. Wei, Y. Hu, R. Zhang, Z. Li, S. K. Zhou, and S. Cui, “Shallow attention network for polyp segmentation,” in Proc. MICCAI, 2021, pp. 699–708.
8. B. Dong, W. Wang, D. Fan, J. Li, H. Fu, and L. Shao, “Polyp-PVT: Polyp segmentation with pyramid vision transformers,” CAAI Artificial Intelligence Research, vol. 2, pp. 1–11, 2023.
9. OpenAI, “GPT-4 technical report,” arXiv preprint arXiv:2303.08774, 2023.
10. H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. NeurIPS, 2023.
11. C. Li, C. Wong, S. Zhang, et al., “LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day,” in Proc. NeurIPS, 2023.
12. Qwen Team, “Qwen2.5: A party of foundation models,” arXiv preprint arXiv:2412.15115, 2024.
13. E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-rank adaptation of large language models,” in Proc. ICLR, 2022.
14. J. Johnson, M. Douze, and H. Jégou, “Billion-scale similarity search with GPUs,” IEEE Trans. Big Data, vol. 7, no. 3, pp. 535–547, 2021.
15. O. Ronneberger, P. Fischer, and T. Brox, “U-Net: Convolutional networks for biomedical image segmentation,” in Proc. MICCAI, 2015, pp. 234–241.
16. T. Tu, S. Azizi, D. Driess, et al., “Towards generalist biomedical AI,” NEJM AI, vol. 1, no. 3, 2024.
17. X. Zhang, C. Wu, Z. Zhao, et al., “PMC-VQA: Visual instruction tuning for medical visual question answering,” arXiv preprint arXiv:2305.10415, 2023.
18. F. Schroff, D. Kalenichenko, and J. Philbin, “FaceNet: A unified embedding for face recognition and clustering,” in Proc. CVPR, 2015, pp. 815–823.
19. S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” IEEE Trans. PAMI, vol. 39, no. 6, pp. 1137–1149, 2017.
20. A. Hermans, L. Beyer, and B. Leibe, “In defense of the triplet loss for person re-identification,” arXiv preprint arXiv:1703.07737, 2017.
21. T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proc. ICCV, 2017, pp. 2980–2988.
22. A. Shrivastava, A. Gupta, and R. Girshick, “Training region-based object detectors with online hard example mining,” in Proc. CVPR, 2016, pp. 761–769.
23. H. R. Roth, L. Lu, A. Seff, et al., “A new 2.5D representation for lymph node detection using random sets of deep convolutional neural network observations,” in Proc. MICCAI, 2014, pp. 520–527.
24. J. Guo, T. He, Z. Lin, et al., “Voice-controlled surgical robots: Current status and future directions,” Surgical Innovation, vol. 27, no. 4, pp. 440–449, 2020.
25. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak supervision,” in Proc. ICML, 2023, pp. 28492–28518.
26. T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient finetuning of quantized LLMs,” in Proc. NeurIPS, 2023.
27. D. Jha, P. H. Smedsrud, M. A. Riegler, et al., “Kvasir-SEG: A segmented polyp dataset,” in Proc. MMM, 2020, pp. 451–462.
28. J. Bernal, F. J. Sánchez, G. Fernández-Esparrach, D. Gil, C. Rodríguez, and F. Vilariño, “WM-DOVA maps for accurate polyp highlighting in colonoscopy: Validation vs. saliency maps from physicians,” Computerized Medical Imaging and Graphics, vol. 43, pp. 99–111, 2015.
29. J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,” in Proc. CVPR, 2016, pp. 779–788.
30. A. Vaswani, N. Shazeer, N. Parmar, et al., “Attention is all you need,” in Proc. NeurIPS, 2017, pp. 5998–6008.
31. A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. ICLR, 2021.
32. J. Chen, Y. Lu, Q. Yu, et al., “TransUNet: Transformers make strong encoders for medical image segmentation,” arXiv preprint arXiv:2102.04306, 2021.
33. A. Kirillov, E. Mintun, N. Ravi, et al., “Segment Anything,” in Proc. ICCV, 2023, pp. 4015–4026.
34. H. Touvron, T. Lavril, G. Izacard, et al., “LLaMA: Open and efficient foundation language models,” arXiv preprint arXiv:2302.13971, 2023.
35. W. X. Zhao, K. Zhou, J. Li, et al., “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All articles published in SJTTU are licensed under a Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA) license. This means anyone is free to copy, transform, or redistribute articles for any lawful purpose in any medium, provided they give appropriate attribution to the original author(s) and SJTTU, link to the license, indicate if changes were made, and redistribute any derivative work under the same license.
Copyright on articles is retained by the respective author(s), without restrictions. A non-exclusive license is granted to SJTTU to publish the article and identify itself as its original publisher, along with the commercial right to include the article in a hardcopy issue for sale to libraries and individuals.
Although the conditions of the CC BY-SA license don't apply to authors (as the copyright holder of your article, you have no restrictions on your rights), by submitting to SJTTU, authors recognize the rights of readers, and must grant any third party the right to use their article to the extent provided by the license.