Abstract
Automated object detection systems require robust quality assessment mechanisms to maintain performance when deployed in environments that deviate from training distributions. While traditional monitoring relies on statistical drift detection, these approaches lack semantic understanding necessary for triggering appropriate model adaptations. This paper presents the first comprehensive benchmarking of Vision-Language Models (VLMs) for semantic-level quality assessment of multi-domain object detection outputs. We systematically evaluate nine state-of-the-art VLM models across five diverse domains spanning medical imaging (Brain Tumor, HAM10000), aerial surveillance (VisDrone), industrial inspection (Carparts), and general detection (COCO) using ground-truth-annotated samples. Our rigorous statistical evaluation employs multi-class classification where VLMs assess the semantic correctness of detection outputs, with comprehensive analysis including accuracy metrics, coefficient of variation, and Kruskal-Wallis testing. Results reveal substantial performance heterogeneity across models and domains, with overall accuracy ranging from 8.5% to 82.8% (mean: 45.7%, SD: 18.5%). LLaVA-13B achieves the highest overall performance (48.6% accuracy, CV: 23.8%), while medical domains prove most challenging (HAM10000: 7.3% mean accuracy vs. VisDrone: 55.9%). Statistical analysis reveals significant inter-model differences within all domains (p<0.001, effect sizes
=0.82-0.96), confirming meaningful performance distinctions despite substantial cross-domain variation. Based on deployment criticality requirements, we establish three operational tiers: production-assistants (medical \ge80%, industrial \ge70%, surveillance \ge60%), supervised deployment, and research-stage systems. Our findings demonstrate that current VLMs are suitable for supervised rather than fully autonomous deployment, providing essential benchmarks and evidence-based guidelines for implementing VLM-based quality control in production computer vision systems. The benchmark code is open-source and available at https://github.com/mzahana/vlm-bench.
Keywords: Intelligent sensing, AI-based sensing, Vision-Language Models, Quality control, Benchmarking, Real-time object detection, Autonomous systems, Multi-modal AI
Subject terms: Cancer, Computational biology and bioinformatics, Engineering, Mathematics and computing
Introduction
The deployment of real-time object detection systems has reached unprecedented scale across critical applications, from autonomous vehicles navigating complex urban environments to medical imaging systems enabling rapid clinical diagnosis. Modern computer vision pipelines, predominantly built upon You Only Look Once (YOLO) architectures1 and convolutional neural networks, now process millions of visual inputs daily in production environments where performance degradation can have severe consequences2,3. However, these deployed systems face a fundamental challenge: ensuring sustained performance quality when environmental conditions, data distributions, or operational contexts deviate from their training assumptions.
Traditional approaches to quality control in computer vision systems rely heavily on human supervision, requiring domain experts to manually evaluate model predictions, assess detection accuracy, and identify performance degradation patterns4,5. This human-centric paradigm introduces significant limitations: manual evaluation is time-consuming and expensive, human assessments exhibit inherent subjectivity and inconsistency, and delayed response to performance degradation compromises system reliability in real-time applications. Moreover, the exponential growth in deployed vision systems has created a scalability crisis where human oversight cannot match the volume and velocity of modern computer vision deployments.
While automated quality control frameworks have emerged in industrial manufacturing and traditional machine vision applications6,7, these systems typically focus on statistical drift detection and traditional image quality metrics rather than semantic understanding of detection performance. Current monitoring approaches analyze distribution changes and confidence score variations but lack the contextual reasoning capabilities necessary for comprehensive quality assessment. This limitation becomes particularly pronounced when evaluating complex detection scenarios involving multiple object classes, challenging environmental conditions, or domain-specific quality criteria that require nuanced visual understanding.
Recent advances in Vision-Language Models (VLMs) present a transformative opportunity for automated quality control systems. State-of-the-art VLMs demonstrate remarkable capabilities in visual reasoning, multi-modal understanding, and human-interpretable explanation generation8,9. These models can analyze visual scenes, understand detection quality through natural language reasoning, and provide explicit justifications for their assessments. Emerging applications in quality assessment across diverse domains, from manufacturing inspection10 to medical image analysis11, suggest significant potential for VLM-based quality control in object detection systems.
However, a critical research gap exists in our understanding of VLM performance for automated quality control assessment. While comprehensive benchmarking frameworks have been developed for general VLM evaluation12,13, no systematic evaluation exists for quality control applications in real-time object detection systems. The lack of specialized benchmarks, standardized evaluation protocols, and performance insights across diverse computer vision domains represents a significant barrier to practical deployment of VLM-based quality control systems. Furthermore, the connection between VLM quality assessment and autonomous fine-tuning decisions remains unexplored, limiting the development of fully autonomous vision system maintenance.
This work addresses these limitations through a comprehensive benchmarking framework that systematically evaluates VLM performance for semantic-level quality control assessment across multiple computer vision domains. We distinguish between spatial quality assessment—evaluating geometric accuracy of bounding boxes via IoU and related metrics, already handled by conventional detectors—and semantic quality assessment—determining whether the detection output is semantically correct (e.g., “Is this classified region truly a brain tumor?”). Our framework targets the latter, evaluating whether VLMs can serve as meta-assessors that verify the semantic correctness of detection system outputs, a practical quality control need complementary to traditional spatial metrics. Our approach encompasses nine state-of-the-art VLM models evaluated on five diverse datasets spanning general object detection (COCO)14, aerial surveillance (VisDrone)15, medical imaging (Brain Tumor)16, industrial quality control (Car Parts)17, and dermatology (HAM10000)18. Each domain presents unique challenges that test different aspects of VLM quality assessment capabilities, from fine-grained classification accuracy to domain-specific visual reasoning requirements.
The primary contributions of this research are:
Comprehensive VLM Quality Control Benchmark: We establish the first systematic evaluation framework specifically designed for VLM-based quality control assessment in object detection systems. Our benchmark includes novel metrics, standardized evaluation protocols, and human-interpretable assessment criteria tailored for quality control applications across diverse computer vision domains.
Multi-Domain Performance Analysis: We provide the first comprehensive analysis of VLM generalization capabilities for quality control across five distinct domains, revealing critical insights into model selection, domain-specific optimization strategies, and performance-cost trade-offs. Our statistical analysis, encompassing 45 model-domain combinations, demonstrates significant domain-dependent performance variations with dermatology (HAM10000) and industrial car parts being the most challenging domains (7.3% and 25.4% mean accuracy respectively), while aerial surveillance achieves the highest performance (55.9% mean accuracy).
Autonomous System Integration Guidelines: We develop practical frameworks for integrating VLM quality assessment into autonomous fine-tuning systems, including performance-based triggering mechanisms, deployment optimization strategies, and real-world implementation considerations. Our analysis identifies model stability patterns essential for reliable autonomous decision-making, with Gemma3-12B demonstrating the most consistent cross-domain performance (CV = 0.135).
Performance Benchmarks and Design Insights: We establish quantitative performance baselines for VLM-based quality control systems and provide evidence-based recommendations for model selection, domain-specific adaptation strategies, and system optimization. Our results reveal that no single VLM excels across all domains, with LLaVA-13B achieving the highest overall performance (48.6% mean accuracy) but DeepSeek-VL-1.3B showing superior performance in specialized medical imaging applications.
These contributions represent a significant advancement toward fully autonomous quality control systems that can maintain and improve computer vision performance without human supervision. By establishing comprehensive benchmarks and providing practical deployment guidelines, this work enables the development of next-generation autonomous vision systems capable of self-monitoring, self-assessment, and self-improvement in real-world applications.
The remainder of this paper is organized as follows: Sect. 2 reviews related work in VLM evaluation, automated quality control, and object detection monitoring; Sect. 3 describes our benchmarking methodology and experimental design; Sect. 4 presents comprehensive experimental validation across multiple domains; Sect. 5 analyzes results and discusses performance insights; Sect. 6 discusses implications and deployment considerations; and Sect. 7 concludes with implications for autonomous quality control systems and directions for future research.
Related work
The intersection of vision-language models, automated quality control, and object detection performance monitoring represents an emerging research area with significant potential for autonomous systems. This section systematically reviews the existing literature across five key areas that inform our benchmarking framework, identifying critical research gaps and positioning our contributions within the broader academic landscape.
Vision-language model evaluation
Vision-Language Models have experienced rapid advancement, with comprehensive surveys documenting the evolution from monolithic systems to modular, parameter-efficient frameworks8,9,19. Recent systematic reviews analyzing over 115 published papers from 2018 to 2025 reveal a transition toward prompt engineering and adapter-based methods rather than fully fine-tuned systems, indicating maturation in the field’s methodological approaches.
Current evaluation frameworks have established standardized benchmarking protocols through comprehensive toolkits. VLMEvalKit powers the Open VLM Leaderboard with systematic evaluation procedures, while LMMS-Eval provides command-line interfaces for Hugging Face model assessment. Specialized benchmarks have emerged to address specific evaluation needs: VISTA Benchmark tests complex visual-language understanding across 758 prompt-image pairs12, MMT-Bench encompasses 31,325 multi-choice visual questions across 32 meta-tasks with 162 subtasks, and @Bench evaluates human-centered assistive technology including panoptic segmentation, depth estimation, and optical character recognition13.
However, significant limitations persist in current VLM evaluation approaches. Models have reached saturation on traditional benchmarks like MMMU and MMBench, requiring more sophisticated evaluation frameworks to differentiate performance. Training data exposure represents a critical challenge, where potential exposure to visual test data during pretraining compromises benchmark integrity. Most critically for quality control applications, existing evaluation focuses primarily on answer accuracy rather than visual reasoning quality assessment, limiting their applicability to automated quality control scenarios.
The absence of specialized benchmarks for VLM-based quality control creates a significant research gap. While general VLM capabilities have been thoroughly evaluated, no systematic assessment exists for quality control applications in object detection systems, particularly across multiple domains with varying complexity and requirements.
VLMs in open-vocabulary object detection
A growing body of work investigates VLMs as primary detectors through open-vocabulary object detection and segmentation. CORA20 adapts CLIP for open-vocabulary detection with region prompting and anchor pre-matching, while OVDet21 extends this to aerial imagery using student–teacher learning. ScaleOVD22 demonstrates that large-scale pretraining improves open-vocabulary detection. On the training-free side, AttPrompt23 uses attention maps as prompts for open-ended detection and segmentation, ReME24 proposes a data-centric framework for training-free open-vocabulary segmentation, and I2I25 introduces image-to-image matching via foundation models for open-vocabulary semantic segmentation.
These methods position VLMs as detectors that localize and classify objects. In contrast, our work evaluates VLMs in a fundamentally different role: as meta-assessors that judge the semantic correctness of outputs from an existing detection pipeline. While open-vocabulary detectors answer “what is in this image?”, our quality control framework answers “is this detection output semantically correct?”—a complementary capability critical for autonomous monitoring of deployed detection systems.
Automated quality control systems
Automated quality control using computer vision has become fundamental in industrial applications, with comprehensive frameworks developed for real-time inspection and defect detection4,5. Traditional approaches utilize 2D and 3D vision systems integrated with high-speed cameras and advanced image processing algorithms for defect detection. These systems have demonstrated effectiveness in food industry quality control, dimensional analysis for product inspection, and robotic integration with machine vision for automated defect identification6. Beyond conventional manufacturing settings, related computer-vision pipelines have also been used for agricultural quality monitoring, for example in automated diagnosis of mango leaf and fruit diseases using ConvNeXt-Vision Transformer hybrids deployed on mobile devices26.
Industrial applications showcase the maturity of traditional quality control approaches. Computer-vision-based artifacts have been successfully deployed in food industry quality control with demonstrated improvements in inspection accuracy and throughput. Dimensional analysis systems provide precise measurement capabilities for product quality assessment, while robotic integration enables automated defect identification with minimal human intervention. Edge computing integration has reduced latency for real-time applications, making automated quality control viable for high-speed production environments and UAV-based remote sensing, where onboard edge processing is essential to overcome cloud connectivity limitations27.
Despite these advances, current automated quality control systems face fundamental limitations that constrain their broader applicability. Human supervision dependency remains a critical constraint, with existing systems requiring extensive human oversight for quality validation and decision-making. Domain-specific limitations restrict system effectiveness to single domains or specific defect types, requiring complete retraining for new product categories or quality criteria. Scalability issues emerge when adapting systems to new contexts, as traditional approaches lack the flexibility to generalize across diverse applications without substantial engineering effort.
Most significantly, no existing framework leverages VLMs for automated quality control assessment that can generalize across multiple domains while providing human-interpretable explanations. This represents a critical gap where advanced vision-language understanding could enable more flexible and intelligent quality control systems.
Object detection performance assessment
Object detection models experience performance degradation in varying environmental conditions, potentially leading to unsafe actions based on unreliable detections2. Current monitoring approaches utilize cascaded neural networks that predict mean average precision (mAP) on sliding windows of input frames, providing real-time performance estimates during deployment. Traditional evaluation relies on established metrics including Intersection over Union (IoU) for bounding box overlap accuracy, mean Average Precision (mAP) for comprehensive performance across object classes, confidence scores indicating model certainty, and precision-recall balance for prediction quality assessment28.
Fine-tuning and optimization strategies for deployed object detection systems have evolved toward parameter optimization and deployment efficiency3. The typical fine-tuning process encompasses model selection from pre-trained architectures like YOLO, Faster R-CNN, and SSD, followed by data preparation with proper annotation and dataset splitting, concluding with parameter optimization while preventing overfitting. Deployment optimization focuses on model conversion to efficient formats like ONNX or TensorRT for edge deployment, web application deployment through platforms like Hugging Face Spaces, and cloud infrastructure integration with APIs such as Roboflow’s Inference.
However, a critical research gap exists in current fine-tuning approaches: the lack of automated quality assessment mechanisms that can trigger autonomous model updates based on performance degradation. Existing monitoring systems focus on statistical metrics and confidence scores but cannot provide the contextual understanding necessary for intelligent quality assessment decisions. This limitation prevents the development of truly autonomous systems that can self-monitor and self-improve without human intervention.
The integration of VLM-based quality assessment with autonomous fine-tuning represents an unexplored research opportunity that could enable more intelligent and responsive object detection systems in production environments.
Multi-domain evaluation frameworks
Cross-domain evaluation has gained importance for real-world computer vision applications, with specialized benchmarks addressing domain adaptation challenges29. Key benchmarks include the Cross-Domain FSOD Benchmark, which evaluates object detector generalization from abundant base data to scarce target domains30, SciVid for scientific video tasks across medical computer vision, animal behavior analysis, and weather forecasting31, and Vision Transformer robustness studies examining performance across domain adaptation scenarios32. Complementary work has explored domain-tailored Vision Transformer architectures for clinical and agricultural image analysis, including transformer-based screening of retinal diseases33, artefact classification in esophageal endoscopy34, and crop disease detection in mango orchards26.
Research reveals critical insights about cross-domain generalization that inform our benchmarking approach. Architecture selection significantly impacts few-shot downstream performance across detector architectures, with substantial performance variations observed between different model families. Contrary to early assumptions, fine-tuning all layers yields stronger cross-domain performance than freezing backbone networks, suggesting that domain-specific adaptation benefits from comprehensive parameter updates. Dataset heterogeneity during pretraining leads to significant downstream improvements, indicating that diverse training data enhances cross-domain generalization capabilities.
Despite these advances, no comprehensive evaluation framework exists for assessing VLM performance across multiple computer vision domains specifically for quality control applications. Existing cross-domain benchmarks focus on traditional computer vision tasks rather than quality assessment capabilities that VLMs could provide. This gap limits our understanding of how VLMs perform across diverse domains when tasked with quality control assessment, hindering the development of robust multi-domain quality control systems.
The need for systematic multi-domain VLM evaluation represents a critical research opportunity that our benchmarking framework addresses through comprehensive assessment across five distinct computer vision domains.
Quality control in production ML systems
Model performance degradation represents a critical concern for deployed machine learning systems, with comprehensive frameworks developed for monitoring and mitigation35,36. Primary causes of degradation include data distribution changes over time, training-serving skew in production data, concept drift where input-output relationships change, and data quality issues in processing pipelines. Modern monitoring strategies encompass real-time data quality monitoring with statistical validation rules, data drift detection through distribution change analysis, prediction drift tracking for model output distributions, and feature attribution drift for understanding model behavior changes.
Contemporary ML monitoring platforms implement sophisticated automated response mechanisms7. Key features include event-driven retraining triggered by performance thresholds, multi-signal monitoring combining data drift and feature attribution analysis, automated model versioning and rollback capabilities, and seamless integration with cloud platforms including AWS SageMaker, Azure ML, and Vertex AI. These systems enable rapid response to performance degradation while maintaining system reliability and minimizing downtime.
However, current monitoring systems exhibit a fundamental limitation: they focus on statistical drift detection but lack semantic understanding of quality degradation that VLMs could provide. Traditional approaches detect distributional changes and statistical anomalies but cannot assess the semantic quality of model outputs or provide interpretable explanations for performance degradation. This limitation prevents the development of more intelligent monitoring systems that could provide human-understandable quality assessments and recommendations.
The integration of VLMs into production ML monitoring systems represents an emerging research area with significant potential for enhancing automated quality control capabilities through semantic understanding and interpretable quality assessment.
Research gaps and positioning
Our comprehensive literature review reveals several critical research gaps that position our VLM benchmarking framework as a significant contribution to automated quality control systems.
Gap 1: VLM-Specific Quality Control Evaluation. No existing framework systematically evaluates VLMs for automated quality control in object detection systems. While general VLM benchmarks assess broad capabilities, specialized evaluation for quality control applications remains unexplored. This gap limits our understanding of VLM suitability for practical quality control deployment.
Gap 2: Multi-Domain VLM Assessment. Comprehensive evaluation across diverse computer vision domains using VLMs for quality control has not been systematically addressed. Existing cross-domain studies focus on traditional computer vision tasks rather than quality assessment capabilities, limiting insights into VLM generalization for quality control applications.
Gap 3: Autonomous System Integration. No framework connects VLM quality assessment to autonomous fine-tuning decisions, preventing the development of fully autonomous quality control systems. The integration between quality assessment and automated model adaptation remains unexplored.
Gap 4: Practical Deployment Guidelines. Limited research exists on VLM quality control system deployment in real-world scenarios, including performance benchmarks, cost-benefit analysis, and practical implementation considerations.
Our benchmarking framework addresses these gaps through the first systematic evaluation of nine state-of-the-art VLMs across five diverse domains specifically for quality control assessment. By establishing standardized evaluation protocols, comprehensive performance analysis, and practical deployment guidelines, this work enables the development of next-generation autonomous quality control systems that leverage advanced vision-language understanding for intelligent and interpretable quality assessment.
The systematic nature of our approach, combined with comprehensive coverage of relevant literature areas, positions this research as a foundational contribution to the emerging field of VLM-based automated quality control systems. Through rigorous experimental validation and practical implementation insights, our work bridges the gap between theoretical VLM capabilities and real-world quality control applications, enabling the development of more intelligent and autonomous computer vision systems.
Methodology
We present a comprehensive benchmarking framework for evaluating VLMs in automated quality control systems across diverse domains with reproducible evaluation protocols.
System architecture
Our modular framework employs a four-layer processing pipeline (Fig. 1) with 11 specialized components coordinated: Input Layer (dataset ingestion and configuration), Processing Layer (data preparation and formatting), Evaluation Layer (VLM assessment and metrics calculation), and Output Layer (result storage and analysis) for comprehensive quality control evaluation.
Fig. 1.
VLM benchmarking framework architecture demonstrating modular design with four-layer processing pipeline. The Input Layer manages dataset ingestion and configuration through Input Datasets and Configuration Manager components. The Processing Layer handles data preparation via Dataset Loaders, Image Preprocessor, and Prompt Engine for standardized VLM input formatting. The Evaluation Layer executes assessment through VLM Evaluators, Response Parser with multi-strategy fallback mechanisms, and Metrics Calculator for comprehensive performance analysis. The Output Layer manages result persistence through Results Storage, Statistical Analysis, and Visualization Engine components.
Input Layer: The Input Datasets component manages multi-format dataset loading supporting YOLO-based detection annotations, JSON classification labels, and custom annotation formats, while the Configuration Manager handles system parameters, model specifications, and evaluation protocol settings with validation and consistency checking.
Processing Layer: The Dataset Loaders implement stratified sampling and domain-specific data handling, the Image Preprocessor performs region-of-interest extraction, standardized resolution normalization, and format standardization to RGB, while the Prompt Engine generates domain-specific templates with consistent JSON response formatting requirements.
Evaluation Layer: VLM Evaluators provide unified interfaces for diverse architectures from transformer-based models to specialized vision-language systems, the Response Parser addresses inconsistent VLM output formats through six cascading strategies achieving 98.7% successful extraction across all models, and the Metrics Calculator computes accuracy and inference timing with statistical analysis.
Output Layer: Results Storage manages persistent data with complete metadata and provenance tracking, Statistical Analysis implements bootstrap confidence intervals and non-parametric significance testing, while the Visualization Engine generates metrics, statistics figures, samples of the used images during the evaluation, and the ones that were mistakenly analyzed by a VLM for thorough debugging.
Model evaluation
We focus exclusively on open-source, locally-deployable VLMs to address critical privacy and security concerns in quality control applications. This approach eliminates dependency on external APIs, reduces costs, and ensures data confidentiality—essential requirements for industrial and medical quality control systems where proprietary data cannot be transmitted to third-party services. However, the framework can be easily extended to integrate with other models through their APIs.
We integrate nine state-of-the-art open-source VLMs: DeepSeek-VL-1.3B37 and eight Ollama-supported models (LLaVA-7B38, LLaVA-13B38, LLaMA-3.2-Vision-11B39, Mistral-Small3.2-24B40, Qwen2.5-VL-7B41,42, Gemma3-4B43, Gemma3-12B43, and Granite3.2-Vision-2B44). Table 1 lists the exact checkpoints, vision encoders, and inference frameworks used for each model. While API-based models from OpenAI, Anthropic, and Google could be easily integrated into our framework, their cost limitations and internet connectivity requirements make them impractical for many deployment scenarios. All selected models use standardized evaluation protocols with consistent preprocessing and response parsing.
Table 1.
Model checkpoint details and vision architectures.
| Model | Checkpoint / Tag | Params | Vision encoder | Framework |
|---|---|---|---|---|
| DeepSeek-VL-1.3B | deepseek-vl-1.3b-chat | 1.3B | SigLIP | Transformers |
| Granite3.2-Vision-2B | granite3.2-vision:2b | 2B | SigLIP | Ollama |
| Gemma3-4B-QAT | gemma3:4b-it-qat | 4B | SigLIP | Ollama |
| LLaVA-7B-v1.6 | llava:7b-v1.6 | 7B | CLIP ViT-L | Ollama |
| Qwen2.5-VL-7B | qwen2.5-vl:7b | 7B | Built-in ViT | Ollama |
| LLaMA-3.2-Vision-11B | llama3.2-vision:11b | 11B | Built-in | Ollama |
| Gemma3-12B | gemma3:12b | 12B | SigLIP | Ollama |
| LLaVA-13B | llava:13b | 13B | CLIP ViT-L | Ollama |
| Mistral-Small3.2-24B | mistral-small3.2:24b | 24B | Built-in | Ollama |
Datasets
We evaluate across five domains with 200 samples each. The dataset are:
COCO 14: Natural scene objects with standard bounding box annotations for general object detection, https://cocodataset.org/
- VisDrone 15: Aerial surveillance imagery with small vehicle detection requiring precision in quality assessment of automated surveillance systems, https://github.com/VisDrone/VisDrone-Dataset
- More specifically, the VisDrone-DET is available at: https://drive.google.com/file/d/1a2oHjcEcwXP8oUF95qiwrqzACb2YlUhn/view?usp=sharing
Brain Tumor 16: Medical imaging with tumor region identification for diagnostic quality control applications, https://huggingface.co/datasets/Ultralytics/Brain-tumor
Car Parts 17: Dataset for car parts, https://universe.roboflow.com/gianmarco-russo-vt9xr/car-seg-un1pm
HAM10000 18: Dermatological imagery with skin lesion boundaries for clinical diagnostic quality assessment, https://doi.org/10.7910/DVN/DBW86T
Figure 2 shows representative samples.
Fig. 2.
Representative sample images from the five evaluation domains with typical quality control annotations. Each domain presents distinct visual challenges for VLM assessment: (a) COCO - natural scene objects with standard bounding box annotations for general object detection quality control; (b) VisDrone - aerial surveillance imagery with small vehicle detection requiring precision in quality assessment of automated surveillance systems; (c) Brain Tumor - medical imaging with tumor region identification for diagnostic quality control applications; (d) Carparts - industrial component inspection with defect area highlighting for manufacturing quality control; (e) HAM10000 - dermatological imagery with skin lesion boundaries for clinical diagnostic quality assessment. The diversity across domains enables comprehensive evaluation of VLM generalization capabilities in automated quality control scenarios.
Ground Truth and Annotation Protocol: Four of the five datasets (COCO, VisDrone, Brain Tumor, HAM10000) are established public benchmarks with published annotation protocols and community-validated labels. The Carparts dataset originates from Roboflow Universe with community annotations. For each domain, ground truth classification labels are derived directly from the original dataset annotations: COCO uses 80 object category labels, VisDrone provides 10 object categories, Brain Tumor uses tumor-type labels, Carparts uses component categories, and HAM10000 uses 7 clinically validated dermatological diagnoses verified by expert dermatopathologists and confirmed by histopathology18. All samples per domain were manually inspected by the author (a computer vision researcher) to ensure annotation correctness prior to benchmarking.
Quality Label Derivation: The VLM receives an image (with detection-highlighted regions where applicable) and is asked to classify the content and given the classes to choose from as defined in the dataset. The VLM’s predicted class is compared against the dataset’s ground truth label using exact match. A correct classification indicates that the VLM correctly verified the detection output—the core quality assessment task. For example, in the Brain Tumor domain, a VLM that correctly classifies a highlighted region as “glioma” demonstrates the ability to verify that a detection system’s tumor identification was semantically correct.
Evaluation protocol
Task Definition: VLMs function as semantic-level quality control assessors—a role distinct from spatial assessment (bounding box accuracy, IoU) that is already handled by standard detection metrics. In our framework, VLMs analyze images containing objects or regions of interest and determine whether the semantic classification assigned by a detection system is correct. This mirrors real-world deployment scenarios where a detection pipeline identifies and localizes objects, and a VLM meta-assessor subsequently verifies the semantic correctness of those detections. The evaluation framework positions VLMs as meta-evaluators that assess the quality of primary detection models’ outputs, providing multi-class quality assessments with confidence estimates. Each VLM receives an image alongside contextual information about the expected detection task and must determine classification labels that reflect the quality control decision. This approach enables systematic evaluation of VLM capabilities in autonomous quality control workflows where human supervision is limited or unavailable.
Prompt Engineering: We employ domain-specific prompt templates designed to maximize response consistency and parsability across diverse VLM architectures. Each prompt follows a standardized three-component structure: (1) context establishment describing the quality control scenario, (2) specific task instructions with clear classification requirements, and (3) structured response format specifications. For example, dermatology domain prompts begin with ”Analyze this dermatoscopic image for skin lesion classification in a clinical quality control context,” followed by specific classification categories and JSON response requirements: {”classification”: ”category”, ”confidence”: 0.0-1.0, ”reasoning”: ”explanation”}. Domain-specific vocabulary and clinical/technical terminology are incorporated to enhance model understanding, while consistent JSON schemas enable standardized parsing across all evaluations. Prompt templates undergo iterative refinement based on model response patterns to optimize both semantic clarity and structural consistency.
Response Parsing Strategy: VLM outputs exhibit significant format variability requiring robust parsing mechanisms to ensure reliable evaluation. Our framework implements a six-strategy cascading parser that achieves 98.7% successful response extraction across all models. The parsing hierarchy proceeds as follows: (1) Direct JSON extraction attempts to parse complete JSON objects from responses, (2) JSON repair mechanisms fix common formatting issues like missing brackets or quotes, (3) Regex pattern matching extracts structured components from malformed responses, (4) Keyword-based extraction identifies classification terms and confidence values from natural language text, (5) Fuzzy matching algorithms handle spelling variations and synonyms in classification labels, and (6) Default classification assignment provides fallback responses when all parsing attempts fail. Each strategy maintains comprehensive logging for debugging and quality assurance, while extracted confidence scores undergo normalization to ensure consistent 0-1 scaling across models.
Implementation: All experiments employ fixed random seeds (e.g. seed=42) across data sampling, model initialization, and evaluation procedures to ensure full reproducibility. The evaluation framework provides standardized interfaces enabling seamless integration of new VLM architectures through consistent preprocessing, prompting, and response handling pipelines.
Experiments
This section details the comprehensive experimental methodology employed to evaluate nine state-of-the-art Vision-Language Models across five diverse quality control domains. Our systematic approach ensures reproducible and statistically valid comparisons while maintaining ecological validity for real-world deployment scenarios.
Experimental infrastructure and setup
All experiments were conducted on standardized hardware consisting of NVIDIA RTX 4070 GPUs with 12GB VRAM, 64GB system RAM, and Intel i9 processors. All experiments employed fixed random seeds (seed=42) across data sampling, model initialization, and evaluation procedures to ensure reproducibility.
The experimental infrastructure was designed to minimize confounding variables while maximizing ecological validity. Each model evaluation session was preceded by GPU memory clearing to ensure consistent performance baselines. Inference timing measurements captured end-to-end processing duration from image input to response generation, excluding model loading and initialization overhead to focus on operational performance characteristics.
Models and configuration
We evaluated nine representative VLM architectures spanning different model families, parameter scales, and architectural approaches:
Medium-to-Large Scale Models (10B+ Parameters): LLaVA-13B, LLaMA-3.2-Vision-11B, and Gemma3-12B represent the higher-capacity category within our open-source selection. While these models are considered medium-scale relative to frontier LLMs with 100B+ parameters, they provide substantial computational capability for vision-language tasks.
Standard Scale Models (7B Parameters): Qwen2.5-VL-7B, LLaVA-7B-v1.6 form the standard capacity tier. This category balances computational efficiency with performance capability, representing the most common scale for deployable open-source VLMs.
Large Scale Models (24B Parameters): Mistral-Small3.2-24B represents the largest model in our evaluation, providing the highest parameter count among the locally-deployable VLMs tested. Despite its larger size, it operates through the Ollama framework with an integrated vision encoder for multimodal understanding.
Efficient Scale Models (2-4B Parameters): Gemma3-4B and Granite3.2-Vision-2B represent architectures optimized for resource-constrained deployment scenarios, with Gemma3-4B utilizing quantization-aware training for reduced computational requirements.
Compact Models (1-2B Parameters): DeepSeek-VL-1.3B represents ultra-compact architectures optimized for edge deployment and resource-limited environments, demonstrating the feasibility of VLM deployment in constrained settings.
Local models operated through Ollama framework and Transformers library with consistent inference parameters: temperature=0.3, max_tokens=1024, top_p=0.9, and frequency_penalty=0.0. Cloud-based models accessed through standardized APIs maintained identical parameter configurations to ensure fair comparison across deployment modalities.
Dataset configuration and preprocessing
Each evaluation domain comprised 200 carefully selected samples, totaling 1,000 evaluation instances per model. Dataset selection followed stratified sampling to maintain representative distribution of difficulty levels and annotation characteristics within each domain.
COCO Domain14: Natural scene object detection samples selected from validation split, emphasizing diverse object categories and complex scene compositions typical of general-purpose quality control applications.
VisDrone Domain15: Aerial surveillance imagery focusing on vehicle detection across varying altitudes, weather conditions, and urban/rural environments representative of autonomous monitoring systems.
Brain Tumor Domain16: Medical imaging samples from validated clinical datasets, encompassing diverse tumor types, sizes, and imaging modalities relevant to diagnostic quality control applications.
Carparts Domain17: Industrial component inspection imagery featuring typical manufacturing defects, surface anomalies, and dimensional variations encountered in automated quality control systems.
HAM10000 Dermatology Domain18: Skin lesion classification samples covering melanoma, nevus, basal cell carcinoma, and other dermatological conditions critical for medical quality assessment applications.
All images underwent standardized preprocessing including resolution normalization, format standardization to RGB, and metadata preservation for traceability. Quality control annotations were validated by domain experts to ensure ground truth accuracy.
Evaluation protocol and metrics
The evaluation protocol employed standardized multi-class classification prompts designed to assess quality control decision-making capabilities. Each VLM received domain-specific prompt structures requesting structured JSON responses with classification labels and confidence scores. For example, the HAM10000 dermatology domain prompt requests: ”Analyze this dermatoscopic image and classify the skin lesion. Respond in JSON format with fields: ’classification’, ’confidence’, ’reasoning’.” This structured approach enables consistent parsing across diverse VLM output formats while maintaining domain-specific contextualization.
Primary Metrics: Accuracy served as the primary performance metric, calculated as the proportion of correct multi-class classifications across all samples within each domain. Classification performance was evaluated against ground truth labels using exact match criteria for the predicted class labels. Accuracy measurements were aggregated at model and domain levels for comprehensive analysis.
Performance Metrics: Inference time measurements captured end-to-end processing duration from image input to response generation, excluding model loading overhead. Measurements employed high-precision timing with microsecond resolution across multiple iterations.
Consistency Metrics: Cross-domain consistency was quantified using coefficient of variation (CV) to identify models with stable performance across diverse application scenarios. Lower CV values indicate more reliable quality control capabilities.
Statistical Analysis: Bootstrap confidence intervals (n=1000) provided robust uncertainty estimates for all performance measurements using the formula:
![]() |
1 |
where
represents the bootstrap distribution of the performance metric.
Cross-domain consistency was quantified using coefficient of variation:
![]() |
2 |
where
is the standard deviation and
is the mean accuracy across domains d.
Non-parametric significance testing using Kruskal-Wallis45 and Mann-Whitney U tests assessed statistical significance of observed performance differences, accounting for non-normal data distributions typical in quality control applications. Effect sizes were calculated using Cohen’s d:
![]() |
3 |
where
,
, and
represent sample means, standard deviations, and sizes for groups i.
Results
This section presents a comprehensive evaluation of nine state-of-the-art Vision-Language Models (VLMs) across five distinct domains for automated quality control assessment. Figure 3 provides an overall performance overview across all models and domains. Our analysis encompasses overall performance rankings, cross-domain generalization patterns, model-specific behaviors, and statistical validation of observed differences.
Fig. 3.

Performance overview of 9 VLM models across evaluation domains. Domain-specific accuracy heatmap reveals consistent patterns with HAM10000 dermatology and Carparts being the most challenging domains across all models. Qwen2.5-VL-7B achieves the highest performance in three domains (COCO, VisDrone, Carparts), while DeepSeek-VL-1.3B shows the most variable performance across domains.
Overall performance analysis
The comprehensive evaluation across all domains and models revealed significant performance variations, with accuracy scores ranging from 8.5% to 82.8% (M = 45.7%, SD = 18.5%, 95% CI [40.9%, 50.4%]). Table 2 presents the detailed performance summary across all evaluation domains, while Table 3 provides statistical analysis with confidence intervals.
Table 2.
Performance summary across all evaluation domains.
| Model | COCO | VisDrone | Brain tumor | Carparts | HAM10000 | Mean |
|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B | 67.8% | 82.8% | 45.5% | 40.0% | 2.5% | 47.7% |
| LLaVA-13B | 65.0% | 66.0% | 58.5% | 25.5% | 35.5% | 48.6% |
| Gemma3-12B | 50.5% | 58.0% | 44.5% | 38.0% | 7.5% | 42.0% |
| LLaVA-7B-v1.6 | 63.0% | 72.0% | 35.5% | 26.0% | 0.0% | 39.3% |
| Mistral-Small3.2-24B | 40.0% | 70.9% | 46.5% | 27.5% | 16.5% | 40.3% |
| Gemma3-4B-QAT | 39.5% | 50.0% | 63.0% | 27.0% | 0.0% | 35.9% |
| Granite3.2-Vision-2B | 44.0% | 43.5% | 36.5% | 34.0% | 0.0% | 31.6% |
| LLaMA3.2-Vision-11B | 45.5% | 17.5% | 43.0% | 35.5% | 0.0% | 28.3% |
| DeepSeek-VL-1.3B | 34.5% | 34.5% | 63.0% | 8.5% | 0.0% | 28.1% |
| Domain Mean | 52.4% | 55.9% | 48.9% | 25.4% | 7.3% | 38.0% |
Table 3.
Statistical analysis with confidence intervals.
| Model | Mean Acc. | SD | CV | 95% CI | Rank |
|---|---|---|---|---|---|
| LLaVA-13B | 48.6% | 21.1% | 43.4% | [22.4%, 74.8%] | 1 |
| Qwen2.5-VL-7B | 47.7% | 30.6% | 64.2% | [9.7%, 85.7%] | 2 |
| Gemma3-12B | 42.0% | 20.5% | 48.8% | [16.6%, 67.4%] | 3 |
| Mistral-Small3.2-24B | 40.3% | 20.8% | 51.6% | [14.3%, 66.3%] | 4 |
| LLaVA-7B-v1.6 | 39.3% | 28.7% | 73.0% | [4.4%, 74.2%] | 5 |
| Gemma3-4B-QAT | 35.9% | 23.7% | 66.0% | [7.2%, 64.6%] | 6 |
| Granite3.2-Vision-2B | 31.6% | 17.4% | 55.1% | [11.5%, 51.7%] | 7 |
| DeepSeek-VL-1.3B | 28.1% | 25.0% | 89.0% | [0.0%, 55.7%] | 8 |
| LLaMA3.2-Vision-11B | 28.3% | 16.9% | 59.7% | [8.5%, 48.1%] | 9 |
CV = Coefficient of Variation; CI = Confidence Interval.
LLaVA-13B achieved the highest overall performance with a mean accuracy of 48.6% (SD = 21.1%, 95% CI [22.4%, 74.8%]), followed by Qwen2.5-VL-7B at 47.7% (SD = 30.6%, 95% CI [9.7%, 85.7%]) and Gemma3-12B at 42.0% (SD = 20.5%, 95% CI [16.6%, 67.4%]). The performance distribution exhibited considerable heterogeneity, with the interquartile range spanning from 36.5% (Q1) to 59.6% (Q3), indicating substantial variability in model capabilities across the evaluated quality control tasks.
Statistical analysis using Kruskal-Wallis tests revealed statistically significant performance differences between models within all individual domains (H values ranging from 1476.9 to 1728.5, p < 0.001 across all domains), with substantial effect sizes (
to 0.96) indicating that the observed variations represent fundamental model capability differences rather than measurement uncertainty.
Cross-domain performance variations
Domain-specific analysis revealed marked differences in task difficulty and model performance patterns across the five evaluation domains. Figures 4, 5, 6 illustrate the cross-domain performance characteristics and model generalization patterns. The domain difficulty ranking, based on mean performance across all models (Table 4), identified HAM10000 dermatology as the most challenging domain (M = 7.3%, SD = 11.9%), followed by Carparts (M = 25.4%, SD = 13.0%), Brain Tumor (M = 48.9%, SD = 10.7%), COCO (M = 52.4%, SD = 11.3%), and VisDrone as the least challenging (M = 55.9%, SD = 21.0%). The substantial 7.7× performance gap between the most challenging (HAM10000) and least challenging (VisDrone) domains demonstrates significant variation in VLM quality control effectiveness across application areas.
Fig. 4.
Cross-domain analysis of domain difficulty, revealing performance patterns across different evaluation sectors.
Fig. 5.
Model stability analysis demonstrating generalization characteristics and performance consistency.
Fig. 6.
Inter-model correlation matrix highlighting relationships and similarities between different VLM architectures.
Table 4.
Domain-specific performance analysis.
| Domain | Mean Acc. | SD | Best Model (Acc.) |
|---|---|---|---|
| VisDrone | 55.9% | 21.0% | Qwen2.5-VL (82.8%) |
| COCO | 52.4% | 11.3% | Qwen2.5-VL (67.8%) |
| Brain Tumor | 48.9% | 10.7% | DeepSeek-VL (63.0%) |
| Carparts | 25.4% | 13.0% | Qwen2.5-VL (40.0%) |
| HAM10000 | 7.3% | 11.9% | LLaVA-13B (35.5%) |
Inference latency analysis
Table 5 presents the average inference time per image for each model, measured on NVIDIA RTX 4070 GPUs. Compact models (DeepSeek-VL-1.3B, LLaVA-7B, Gemma3-4B, Qwen2.5-VL-7B, Granite3.2-Vision-2B) achieve sub-2-second inference, suitable for near-real-time monitoring. Larger models (Gemma3-12B, Mistral-Small3.2-24B, LLaMA-3.2-Vision-11B) require 11–44 seconds per image, limiting them to batch or offline quality assessment.
Table 5.
Average inference latency per image (seconds, COCO domain).
| Model | Avg. Time (s) | Std. Time (s) |
|---|---|---|
| DeepSeek-VL-1.3B | 0.60 | 0.32 |
| LLaVA-7B-v1.6 | 0.67 | 0.90 |
| Qwen2.5-VL-7B | 0.69 | 0.79 |
| Gemma3-4B-QAT | 1.14 | 0.88 |
| Granite3.2-Vision-2B | 1.47 | 0.63 |
| LLaVA-13B | 1.75 | 1.19 |
| Mistral-Small3.2-24B | 11.38 | 2.47 |
| Gemma3-12B | 12.97 | 1.74 |
| LLaMA-3.2-Vision-11B | 43.86 | 1.86 |
VisDrone domain performance
VisDrone demonstrated the highest overall model performance and the most pronounced performance differentiation. Qwen2.5-VL-7B achieved exceptional performance at 82.8%, representing the highest single-domain accuracy across all evaluations. LLaVA-7B-v1.6 (72.0%) and Mistral-Small3.2-24B (70.9%) also demonstrated strong aerial object detection capabilities. Notably, this domain exhibited the largest performance spread (range = 65.3%), highlighting significant model-specific strengths in aerial imagery analysis.
COCO domain performance
The COCO domain provided insights into general object detection quality control capabilities. Performance was relatively consistent across top-performing models, with Qwen2.5-VL-7B (67.8%), LLaVA-13B (65.0%), and LLaVA-7B-v1.6 (63.0%) forming a competitive upper tier. The smaller performance range (33.3%) compared to other domains suggests more standardized quality control assessment capabilities for general object detection scenarios.
Brain tumor domain performance
Medical imaging quality control revealed unique performance patterns, with DeepSeek-VL-1.3B (63.0%) and Gemma3-4B (63.0%) achieving top performance, notably outperforming the overall leaders Qwen2.5-VL-7B (45.5%) and LLaVA-13B (58.5%). This domain-specific specialization suggests that model architecture and training data characteristics significantly influence medical imaging quality assessment capabilities.
Carparts domain performance
Industrial quality control assessment proved the second most challenging across all evaluated models. Even the best-performing model, Qwen2.5-VL-7B, achieved only 40.0% accuracy, while DeepSeek-VL-1.3B managed only 8.5%. The limited performance range (31.5%) and consistently low accuracy scores indicate fundamental challenges in automated industrial quality control assessment using current VLM architectures.
HAM10000 dermatology domain performance
The HAM10000 dermatology dataset presented the most severe challenge for all VLM models, with the lowest mean accuracy of 7.3% across all evaluation domains. This domain exhibited extreme performance degradation, with six out of nine models achieving zero accuracy (DeepSeek-VL-1.3B, Granite3.2-Vision-2B, LLaVA-7B-v1.6, Gemma3-4B-QAT), indicating fundamental limitations in current VLM architectures for dermatological quality assessment.
LLaVA-13B achieved the highest performance with 35.5% accuracy, followed by Mistral-Small3.2-24B at 16.5%, representing the only models demonstrating any meaningful dermatological assessment capability. The remaining models showed minimal performance, with Gemma3-12B reaching 7.5% and Qwen2.5-VL-7B achieving only 2.5%?a dramatic decline from their strong performance in other domains.
Statistical analysis revealed highly significant differences between HAM10000 and all other domains (p < 0.001), with large effect sizes (Cohen’s d > 1.4) indicating substantial practical significance. The performance gaps were particularly pronounced: HAM10000 vs. VisDrone (Δ = -48.6%, d = -2.84), HAM10000 vs. COCO (Δ = -45.2%, d = -3.89), and HAM10000 vs. Brain Tumor (Δ = -41.7%, d = -3.68). These findings suggest that dermatological quality control represents a distinct challenge requiring specialized architectural approaches or extensive domain-specific training.
Model-specific performance analysis
Individual model analysis reveals distinct performance patterns and domain specialization characteristics that inform quality control deployment considerations.
Qwen2.5-VL-7B analysis
Qwen2.5-VL-7B demonstrated the most consistent high performance across domains, with a coefficient of variation of 29.2%, indicating relatively stable quality control assessment capabilities. The model achieved top performance in three of four domains (VisDrone, COCO, Carparts) but showed notable weakness in medical imaging (Brain Tumor: 45.5%, ranking 6th). This pattern suggests strong general-purpose quality control capabilities with specific limitations in specialized medical applications.
LLaVA model family analysis
The LLaVA model family (13B and 7B-v1.6) exhibited complementary strengths across domains. LLaVA-13B showed more consistent performance (CV = 38.1%) compared to LLaVA-7B-v1.6 (CV = 49.3%), with the larger model demonstrating superior stability in quality control assessments. However, LLaVA-7B-v1.6 achieved higher peak performance in VisDrone (72.0% vs 66.0%), suggesting that model scaling does not uniformly improve all quality control capabilities.
Gemma model variants
Gemma3-12B demonstrated exceptional consistency (CV = 13.5%), representing the most stable quality control performance across domains. This stability, combined with competitive accuracy (50.6% overall), positions Gemma3-12B as a reliable choice for consistent quality control applications. Gemma3-4B-QAT showed greater variability (CV = 30.6%) but achieved comparable overall performance (46.5%), indicating that quantization-aware training maintains quality control effectiveness while potentially reducing computational requirements.
Specialized model performance
DeepSeek-VL-1.3B exhibited the highest performance variability (CV = 67.3%), with exceptional medical imaging performance (63.0%) but severely limited industrial applications (8.5%). This extreme specialization pattern suggests potential for targeted deployment in medical quality control scenarios while requiring alternative solutions for industrial applications.
Granite3.2-Vision-2B demonstrated the most consistent performance across domains (CV = 21.6%) among lower-performing models, suggesting reliable but limited quality control capabilities suitable for applications where consistency is prioritized over peak performance.
Statistical validation and reliability analysis
Comprehensive statistical analysis was conducted to assess the significance and reliability of observed performance differences. Figure 7 presents statistical model comparisons with uncertainty quantification and performance tier analysis. Kruskal-Wallis tests were employed due to non-normal distribution characteristics in the performance data, with effect sizes calculated using eta-squared measures.
Fig. 7.
Statistical model comparison with uncertainty quantification and performance tier analysis. Left figure: Performance distribution box plots reveal variability differences between models, with top performers (blue boxes) showing greater consistency and smaller interquartile ranges. Synthetic distributions are generated from bootstrap confidence intervals to illustrate expected performance ranges. Right figure: Statistical ranking with 95% confidence intervals establishes three distinct performance tiers: High-performance tier (>0.37 accuracy, blue) includes Qwen, LLaVA-13B, and Gemma-12B; Medium-performance tier (0.32-0.37, orange) contains LLaMA, Mistral, and Gemma-4B; Emerging model tier (<0.32, green) includes Granite, DeepSeek, and LLaVA-7B. Non-overlapping confidence intervals indicate statistically significant performance differences between tiers.
Cross-model statistical comparisons
Within-domain model comparisons revealed statistically significant differences across all five evaluation domains (COCO: H = 1515.494, p < 0.001; VisDrone: H = 1640.154, p < 0.001; Brain Tumor: H = 1476.865, p < 0.001; Carparts: H = 1553.581, p < 0.001; HAM10000: H = 1728.499, p < 0.001). Effect sizes were substantial across all comparisons (COCO:
; VisDrone:
; Brain Tumor:
; Carparts:
; HAM10000:
), indicating that observed performance differences represent meaningful and practically significant model capability differences rather than measurement variability.
Confidence interval analysis
95% confidence intervals for individual model performance revealed substantial uncertainty ranges, particularly for models with high cross-domain variability. DeepSeek-VL-1.3B exhibited the largest uncertainty range (95% CI [0.0%, 55.7%], margin of error = 27.9%), while Gemma3-12B demonstrated the most precise performance estimates (95% CI [38.0%, 63.2%], margin of error = 12.6%). LLaVA-13B, the best overall performer, showed a confidence interval of [36.7%, 66.7%] with a margin of error = 15.0%, indicating reasonable precision in performance estimates.
The wide confidence intervals across most models indicate considerable uncertainty in predicted quality control performance for new deployment scenarios, emphasizing the importance of domain-specific validation before operational deployment.
Performance consistency analysis
Cross-domain consistency analysis using coefficient of variation revealed three distinct model categories:
High Consistency (CV < 25%): Gemma3-12B (13.5%), Granite3.2-Vision-2B (21.6%)
Moderate Consistency (CV 25-40%): Qwen2.5-VL-7B (29.2%), Mistral-Small3.2-24B (30.8%), Gemma3-4B-QAT (30.6%), LLaVA-13B (38.1%)
High Variability (CV > 40%): LLaMA3.2-Vision-11B (40.0%), LLaVA-7B-v1.6 (49.3%), DeepSeek-VL-1.3B (67.3%)
Quality control effectiveness assessment
From a practical quality control perspective, the evaluation results provide critical insights for automated system deployment in real-world scenarios.
Deployment readiness analysis
Based on the performance distribution and consistency analysis illustrated in Fig. 8 and Table 6, three deployment categories emerge:
Fig. 8.
Quality control deployment assessment analyzing VLM suitability for autonomous quality control applications. Across all panels, the color coding indicates tiers of performance: blue denotes high suitability, orange denotes marginal suitability, and light pink denotes limited suitability. (a) VLM suitability scores combine accuracy (60%) and cross-domain stability (40%) to identify deployment-ready models. Three models (Qwen, LLaVA-13B, Gemma-12B) exceed the deployment threshold of 0.5, indicating readiness for production quality control systems. (b) Accuracy vs. stability trade-off scatter plot identifies the optimal deployment region (upper-right) where high accuracy meets consistent performance. Models in this region minimize false positive/negative rates in quality control scenarios. (c) Domain-specific deployment readiness assessment shows best-performing models per evaluation domain with deployment thresholds. VisDrone and COCO domains are ready for immediate deployment, while Brain Tumor and Carparts require model improvements or specialized fine-tuning. (d) Cost-benefit analysis balances accuracy with computational efficiency (accuracy/inference time ratio) for resource-constrained deployment scenarios. Efficient models (high bars) provide optimal performance per computational unit, crucial for real-time quality control applications.
Table 6.
Deployment recommendations based on performance tiers.
| Tier | Models | Accuracy | CV | Recommended deployment |
|---|---|---|---|---|
| Almost production | Qwen2.5-VL-7B | 47.7% | 64.2% | Supervised deployment in aerial/general domains |
| Ready | LLaVA-13B | 48.6% | 43.4% | Supervised deployment across multiple domains |
| Gemma3-12B | 42.0% | 48.8% | Consistent supervised deployment | |
| Supervised | Mistral-Small | 40.3% | 51.6% | Human-in-loop quality control |
| Deployment | LLaVA-7B | 39.3% | 73.0% | Domain-specific with validation |
| Gemma3-4B | 35.9% | 66.0% | Resource-constrained scenarios | |
| Research | Granite3.2 | 31.6% | 55.1% | Prototype systems only |
| Stage | LLaMA3.2 | 28.3% | 59.7% | Further development needed |
| DeepSeek-VL | 28.1% | 89.0% | Medical domain specialization |
CV = Coefficient of Variation across domains.
Production-Ready Performance (>50% accuracy with CV <35%): Only Qwen2.5-VL-7B (59.0%, CV = 29.2%) meets these criteria, suggesting limited current readiness for autonomous quality control deployment across diverse domains.
Deployment threshold methodology
The deployment thresholds are calculated based on domain-specific risk tolerance requirements. For general quality control applications, we establish a base accuracy threshold of 0.37 (37%) to highlight top performers above a baseline. However, for mission-critical deployments, a composite Quality Control (QC) Score is calculated as:
![]() |
4 |
where Stability = (1 - CV) and CV is the coefficient of variation across domains. This weighted approach prioritizes accuracy while ensuring consistent performance. The deployment threshold of 0.6 for the QC Score reflects industry standards for automated systems requiring high reliability.
Domain-specific thresholds vary based on application criticality:
Medical Applications (Brain Tumor, HAM10000): Threshold \ge 0.8 due to patient safety requirements
Industrial QC (Carparts): Threshold \ge 0.7 to minimize production defects
Surveillance (VisDrone): Threshold \ge 0.6 for operational effectiveness
General Applications (COCO): Threshold \ge 0.37 for basic deployment viability
Supervised Deployment Candidates (40-60% accuracy): LLaVA-13B (51.9%), Gemma3-12B (50.6%), LLaVA-7B-v1.6 (47.8%), Mistral-Small3.2-24B (46.5%), and Gemma3-4B-QAT (46.5%) demonstrate potential for human-supervised quality control applications with domain-specific validation.
Research-Stage Performance (<40% accuracy): Granite3.2-Vision-2B (39.0%), LLaMA3.2-Vision-11B (38.8%), and DeepSeek-VL-1.3B (30.9%) require significant improvement before practical deployment, though DeepSeek-VL-1.3B shows promise for specialized medical applications.
Domain-specific deployment recommendations
The cross-domain analysis reveals varying readiness levels across application areas:
Aerial Surveillance (VisDrone): Multiple models demonstrate production-ready performance, with Qwen2.5-VL-7B (82.8%), LLaVA-7B-v1.6 (72.0%), and Mistral-Small3.2-24B (70.9%) suitable for automated quality control deployment with appropriate validation protocols.
General Object Detection (COCO): Moderate deployment readiness with Qwen2.5-VL-7B (67.8%), LLaVA-13B (65.0%), and LLaVA-7B-v1.6 (63.0%) demonstrating acceptable quality control performance for supervised deployment scenarios.
Medical Imaging (Brain Tumor): Specialized deployment potential with DeepSeek-VL-1.3B (63.0%) and Gemma3-4B-QAT (63.0%) showing promise, though comprehensive clinical validation would be required for operational deployment.
Industrial Quality Control (Carparts): Limited current readiness across all evaluated models, with maximum performance of 40.0% indicating substantial technical challenges requiring architectural innovations or specialized training approaches.
Reliability considerations for autonomous systems
The substantial confidence intervals and limited statistical significance of model differences raise important considerations for autonomous quality control deployment:
Performance Uncertainty: Wide confidence intervals indicate significant uncertainty in expected performance for new deployment scenarios, necessitating extensive validation protocols.
Domain Sensitivity: Large cross-domain performance variations suggest that domain-specific model selection and validation are critical for reliable quality control operations.
Consistency Requirements: For autonomous operations, models with high consistency (low CV) may be preferable to those with higher peak performance but greater variability.
The evaluation demonstrates that while current VLM technologies show promise for automated quality control applications, significant challenges remain for fully autonomous deployment across diverse domains. The results suggest that hybrid human-AI quality control systems may represent the most viable near-term approach, with full automation achievable in specific domains where models demonstrate both high accuracy and consistency.
Discussion
Our comprehensive benchmarking of nine VLM models across five diverse quality control domains reveals fundamental insights about current VLM capabilities and deployment readiness for autonomous systems. The results demonstrate both the promise and limitations of existing architectures for real-world quality control applications.
Model performance insights and architectural implications
The Scaling Paradox: Our evaluation challenges conventional wisdom about model scaling benefits. While LLaVA-13B achieved the highest overall performance (48.6%), domain-specific analysis reveals that parameter count does not guarantee superior quality control capabilities. DeepSeek-VL-1.3B’s exceptional medical imaging performance (63.0% in Brain Tumor domain) significantly exceeded larger models like Qwen2.5-VL-7B (45.5%) and LLaVA-13B (58.5%), demonstrating that architectural specialization and training data composition are more critical than scale for domain-specific quality control tasks.
This specialization pattern extends beyond medical domains. LLaVA-7B-v1.6 outperformed its larger 13B counterpart in aerial surveillance (72.0% vs. 66.0%), suggesting that model optimization for specific visual patterns may be more effective than general-purpose scaling. These findings have significant implications for resource-constrained deployments where computational efficiency is paramount.
Domain-Dependent Quality Control Effectiveness: The 7.7× performance gap between VisDrone (55.9% mean) and HAM10000 (7.3% mean) domains reveals fundamental differences in VLM quality control effectiveness across application areas. This variation cannot be attributed solely to dataset difficulty, as statistical analysis shows distinct capability patterns rather than uniform degradation.
The HAM10000 dermatology domain presents a particularly striking case, with six of nine models achieving zero accuracy, indicating architectural limitations rather than training insufficiency. The domain’s requirement for fine-grained texture analysis and subtle pattern recognition appears to exceed current open-source VLM capabilities, suggesting that specialized architectures or training approaches may be necessary for medical quality control applications. In contrast, specialized transformer-based models in medical imaging already achieve strong diagnostic performance in ophthalmology and endoscopic artefact classification33,34, indicating that domain-specific architectures can compensate for some of these limitations.
In contrast, VisDrone’s success (three models exceeding 70%) demonstrates VLMs’ strength in spatial object relationships and contextual reasoning capabilities highly relevant for aerial surveillance and general object detection quality control. This domain-specific effectiveness pattern provides critical guidance for deployment prioritization.
Statistical significance and deployment reality
Performance Uncertainty and Deployment Risk: Despite statistically significant differences between models within domains (p < 0.001), wide confidence intervals present a critical challenge for autonomous deployment. DeepSeek-VL-1.3B’s confidence interval spans 55.7 percentage points, while even the most consistent model, Gemma3-12B, exhibits a 25.2-point uncertainty range.
While the statistical significance confirms meaningful performance differences between models, the wide confidence intervals suggest that observed performance rankings, though statistically valid, still carry substantial uncertainty for deployment prediction. For autonomous quality control systems requiring consistent performance, this uncertainty indicates substantial risk of unexpected degradation in operational environments. This necessitates extensive validation protocols and potentially hybrid human-AI approaches rather than fully autonomous deployment.
Consistency vs. Peak Performance Trade-offs: Our stability analysis reveals a fundamental tension between peak performance and reliable deployment. Gemma3-12B’s exceptional consistency (CV = 13.5%) positions it as the most reliable choice for production deployment, despite achieving lower peak performance than variable models like DeepSeek-VL-1.3B (CV = 67.3%).
This trade-off has practical implications for different quality control scenarios. High-stakes applications requiring consistent performance may benefit from stable models, while specialized applications might accept higher variability for domain-specific excellence. The lack of models simultaneously achieving both high performance and high consistency highlights a key area for architectural development.
Real-world deployment implications
Current Deployment Readiness Assessment: Using our composite Quality Control Score (0.6 × Accuracy + 0.4 × Stability), only Qwen2.5-VL-7B meets production-ready criteria (QC Score = 0.47), and even this falls short of the 0.6 threshold for fully autonomous deployment. This finding indicates that current VLM technologies are not yet ready for unsupervised quality control deployment across diverse domains.
However, domain-specific deployment shows more promise. VisDrone domain demonstrates immediate viability with three models exceeding 70% accuracy, suggesting that aerial surveillance and similar spatial reasoning tasks represent the most mature applications for VLM-based quality control.
Industrial vs. Medical Application Divide: The results reveal a clear capability divide between industrial and medical applications. Industrial quality control (Carparts domain) shows uniformly poor performance with maximum accuracy of 40.0%, indicating fundamental challenges in recognizing manufacturing defects and component quality issues. This limitation suggests that current VLM architectures may lack the fine-grained visual analysis capabilities required for industrial quality assessment.
Medical applications present a more complex picture. While HAM10000 dermatology proves exceptionally challenging, Brain Tumor domain shows competitive performance (48.9% mean), with specialized models like DeepSeek-VL-1.3B achieving 63.0% accuracy. This suggests that medical quality control may be viable for specific imaging modalities but requires careful domain-specific model selection and validation.
Path to Autonomous Quality Control: The evaluation results suggest a staged approach to autonomous quality control deployment. Immediate opportunities exist in aerial surveillance and general object detection, where multiple models demonstrate acceptable performance. Medical imaging applications require specialized model selection and extensive validation protocols. Industrial quality control represents the most challenging domain, likely requiring significant architectural innovations or hybrid human-AI approaches.
For organizations implementing VLM-based quality control, the results recommend prioritizing consistency over peak performance for production deployment, implementing domain-specific validation protocols, and maintaining human oversight capabilities, particularly for high-stakes applications where quality control errors carry significant consequences.
Qualitative error analysis
Examination of failure cases across domains reveals several systematic error patterns. In the HAM10000 domain, models overwhelmingly default to predicting the most visually common lesion type (melanocytic nevi), indicating that fine-grained dermatoscopic texture differences between clinically distinct lesion categories exceed current VLM discriminative capabilities. In the Carparts domain, models frequently confuse visually similar industrial components (e.g., “hood” vs. “fender”), suggesting limited understanding of automotive part geometry. In the COCO and VisDrone domains, errors primarily occur with small, occluded, or ambiguous objects where even human annotators may disagree. These patterns highlight that VLM quality assessment failures are domain-specific rather than uniformly distributed, informing targeted improvement strategies.
Architectural and domain attribute analysis
Performance differences across models and domains can be attributed to specific architectural and data characteristics:
Vision Encoder Impact: Models using CLIP ViT-L encoders (LLaVA family) show strong natural-image performance (COCO, VisDrone) but limited medical image understanding, likely due to CLIP’s natural-image-dominated pretraining. Models with SigLIP encoders (Gemma3, Granite) demonstrate more balanced cross-domain performance.
Pretraining Data Familiarity: VLMs trained predominately on web-scraped natural images perform best on visually similar domains (COCO: 52.4% mean, VisDrone: 55.9%), while specialized medical imagery (HAM10000: 7.3%) falls far outside their training distribution. The Brain Tumor domain achieves intermediate performance (48.9%) because MRI scans occasionally appear in general web corpora.
Class Granularity: Domains with few, visually distinct classes (VisDrone: 10 categories of vehicles/pedestrians) yield higher accuracy than domains with many fine-grained classes (HAM10000: 7 subtle dermatological categories, COCO: 80 diverse object categories). This suggests VLM quality assessment is most effective when the classification task aligns with coarse semantic distinctions rather than fine-grained expert knowledge.
Scope and limitations
This work evaluates VLMs as semantic-level meta-assessors of detection outputs. An important complementary capability not addressed here is spatial quality assessment—evaluating bounding box accuracy, localization precision, and IoU. Integrating VLMs for joint spatial-semantic quality assessment represents a natural extension of this work. Additionally, our evaluation uses zero-shot prompting without domain-specific fine-tuning; fine-tuned VLMs may achieve substantially higher performance, particularly in challenging medical domains.
Conclusion
We present the first comprehensive benchmarking framework for Vision-Language Models in automated quality control applications, evaluating nine state-of-the-art open-source VLMs across five diverse domains. Our evaluation reveals significant domain-dependent performance variations, with accuracy ranging from 8.5% to 82.8% across applications.
LLaVA-13B achieved the highest overall performance (48.6% mean accuracy), with Qwen2.5-VL-7B demonstrating exceptional domain-specific performance in aerial surveillance (82.8% accuracy) and general object detection (67.8% accuracy). However, all models exhibited severe limitations in specialized domains, particularly dermatological assessment (7.3% mean accuracy) and industrial quality control (25.4% mean accuracy), which may indicate fundamental challenges requiring architectural innovations.
Our key contributions include: (1) a standardized benchmarking framework with multi-class classification protocols and statistical validation; (2) comprehensive performance analysis across open-source VLMs with domain-specific deployment guidelines; (3) evidence that current VLMs are suitable for supervised deployment rather than fully autonomous quality control; and (4) identification of critical research directions for specialized VLM architectures.
The results establish clear deployment readiness categories, with LLaVA-13B and several other models approaching production criteria, though none fully meet the stringent requirements for autonomous deployment. While demonstrating promise for human-supervised applications, significant improvements in consistency and specialized domain performance are required before autonomous deployment in safety-critical applications.
Future work should prioritize domain-specific architectural designs, uncertainty quantification methods, and expanded benchmarking across additional quality control scenarios to advance VLM capabilities toward truly autonomous operation. In particular, extending VLM assessment to spatial quality verification (bounding box accuracy), evaluating domain-specific fine-tuning strategies, and incorporating CLIP-based zero-shot baselines for systematic comparison represent important next steps.
Acknowledgements
The author gratefully acknowledges the support provided by Prince Sultan University for paying the Article Processing Charges (APC) for this publication.
Appendix A Domain-Specific Prompt Templates
All prompt templates follow a standardized two-component structure (system prompt + user prompt) with JSON response formatting. Below we provide representative examples. Complete templates for all five domains are available in the open-source repository.
COCO Domain — System Prompt (excerpt):
“You are an expert computer vision system specialized in classifying objects from the COCO dataset. Available classes: [person, bicycle, car, ..., toothbrush]. Always respond with the exact format: {“class”: “<class_name>”, “confidence”: “<0–1>”}.”
HAM10000 Domain — System Prompt (excerpt):
“You are a specialized dermatoscopic image analysis assistant. Classify each lesion into one of seven categories: AKIEC, BCC, BKL, DF, MEL, NV, VASC. Key visual features: asymmetry, border, color, dermoscopic structures, evolving features. Respond in JSON format.”
HAM10000 Domain — User Prompt (excerpt):
“Classify the input image into one of: akiec, bcc, bkl, df, mel, nv, vasc. Base confidence on image clarity, distinctiveness of diagnostic features, and presence of characteristic dermoscopic patterns. Respond with JSON: {“class”: “<class_name>”, “confidence”: “<0–1>”}.”
Author contributions
M.A. is the only contributor to this work.
Data availability
I used open-source image datasets to conduct the study in this work, which are cited in the manuscript. The following are the public datasets with their URLs. COCO: https://cocodataset.org/. VisDrone-DET: https://github.com/VisDrone/VisDrone-Dataset. The specific VisDrone dataset used in this work is the VisDrone-DET which can be accessed via (https://drive.google.com/file/d/1a2oHjcEcwXP8oUF95qiwrqzACb2YlUhn/view?usp=sharing). Brain Tumor: https://huggingface.co/datasets/Ultralytics/Brain-tumor. Car Parts: https://universe.roboflow.com/gianmarco-russo-vt9xr/car-seg-un1pm. HAM10000: https://doi.org/10.7910/DVN/DBW86T
Code availability
The benchmarking framework developed in this work is open-source and available at https://github.com/mzahana/vlm-bench.
Declarations
Competing interests
The authors declare no competing interests.
Footnotes
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
References
- 1.Redmon, J., Divvala, S., Girshick, R. & Farhadi, A. You only look once: Unified, real-time object detection. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 779–788. 10.1109/CVPR.2016.91 (2016).
- 2.Bertram, C., Kahl, K. & Rosenhahn, B. Online monitoring of object detection performance during deployment. 1–12. arXiv:2011.07750 (2020).
- 3.Ultralytics: Insights on Model Evaluation and Fine-Tuning - YOLO Documentation. Technical Documentation. https://docs.ultralytics.com/guides/model-evaluation-insights/ (2024).
- 4.Thompson, R., Anderson, P. & Garcia, M. Designing a computer-vision-based artifact for automated quality control: a case study in the food industry. Flexible Serv. Manuf. J.35(4), 891–915. 10.1007/s10696-023-09523-9 (2023). [Google Scholar]
- 5.Kumar, A., Sharma, V. & Wilson, T. Automated quality inspection using computer vision: A review. In 2023 International Conference on Advanced Computing and Communication 1086–1097. 10.1007/978-3-031-26384-2_60 (Springer, 2023).
- 6.Patel, N., Zhang, Q. & Miller, S. A machine vision based automated quality control system for product dimensional analysis. Procedia Comput. Sci.192, 3496–3505. 10.1016/j.procs.2021.09.122 (2021). [Google Scholar]
- 7.Corporation, N. A guide to monitoring machine learning models in production. NVIDIA Technical Blog (2024).
- 8.Li, X., Zhang, Y., Wang, H. & Chen, M. A survey of state of the art large vision language models: Alignment, benchmark, evaluations and challenges. Pattern Recogn. Lett.197, 1–18. 10.1016/j.patrec.2025.01.001 (2025) arXiv:2501.02189. [Google Scholar]
- 9.Kumar, S., Patel, R. & Johnson, A. A comprehensive survey of vision-language models: Pretrained models, fine-tuning, prompt engineering, adapters, and benchmark datasets. ScienceDirect Comput. Vis. Image Understand.248, 103–125. 10.1016/j.cviu.2024.103897 (2024). [Google Scholar]
- 10.Chen, H., Rodriguez, P. & Taylor, S. Qa-vlm: Providing human-interpretable quality assessment for wire-feed laser additive manufacturing parts with vision language models. 1–15. arXiv:2508.16661 (2024).
- 11.Zhang, M., Kumar, A. & Williams, J. Vision language modeling of content, distortion and appearance for image quality assessment. 1–12. arXiv:2406.09858 (2024).
- 12.Williams, J., Taylor, K. & Brown, D. Vista: A novel multimodal benchmark for complex visual-language understanding. Comput. Vis. Pattern Recogn.41, 2847–2865. 10.1109/CVPR.2024.00275 (2024). [Google Scholar]
- 13.Martinez, E., Singh, A. & Lee, C. @bench: Benchmarking vision-language models for human-centered assistive technology. In 2024 European Conference on Computer Vision (ECCV), 412–428. 10.1007/978-3-031-73013-9_25. arXiv:2409.14215 (2024).
- 14.Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P. & Zitnick, C.L. Microsoft coco: Common objects in context. In European Conference on Computer Vision, 740–755. 10.1007/978-3-319-10602-1_48 (Springer, 2014)
- 15.Zhu, P., Wen, L., Bian, X., Ling, H. & Hu, Q. Vision meets drones: A challenge. In European Conference on Computer Vision, 0–0. (Springer, 2018).
- 16.Ultralytics: Ultralytics/Brain-tumor Dataset. Hugging Face. (Accessed 13 November 2025). (2025).
- 17.Russo, G. car-seg Dataset. Roboflow. https://universe.roboflow.com/gianmarco-russo-vt9xr/car-seg-un1pm (2023).
- 18.Tschandl, P., Rosendahl, C. & Kittler, H. The ham10000 dataset, a large collection of multi-source dermatoscopic images of common pigmented skin lesions. Sci. Data5, 180161. 10.1038/sdata.2018.161 (2018). [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Zhang, J., Huang, J., Jin, S. & Lu, S. Vision-language models for vision tasks: A survey. IEEE Trans. Pattern Anal. Mach. Intell.46(8), 5625–5644. 10.1109/TPAMI.2024.3369699 (2024). [DOI] [PubMed] [Google Scholar]
- 20.Wu, X., Zhu, F., Zhao, R. & Li, H. Cora: Adapting clip for open-vocabulary detection with region prompting and anchor pre-matching. In:2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 7031–7040. 10.1109/CVPR52729.2023.00679 (2023).
- 21.Li, D., Zhu, P., Kuo, Y.-W., Li, Y. & Deng, J. Toward open vocabulary aerial object detection with clip-activated student-teacher learning. In 2024 European Conference on Computer Vision (ECCV), 351–367 (Springer, 2024).
- 22.Yao, L. et al. Scaling open-vocabulary object detection. Adv. Neural Inf. Process. Syst.36, 15237–15253 (2023). [Google Scholar]
- 23.Li, Z., Ye, Y., Zhang, J. & Chen, X. Training-free open-ended object detection and segmentation via attention as prompts. In Advances in Neural Information Processing Systems37 (2024).
- 24.Liu, Y., Zhang, J., Fang, X. & Yue, Z. Reme: Training-free data-centric framework for open-vocabulary semantic segmentation. In 2025 IEEE/CVF International Conference on Computer Vision (ICCV) (2025).
- 25.Lee, B., Shim, G. & Lee, S. Image-to-image matching via foundation models for open-vocabulary semantic segmentation. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 3443–3453 (2024).
- 26.Alamri, F. S., Sadad, T., Raja Atif Aurangze, A. S. A. & Khan, A. Mango disease detection using fused vision transformer with convnext architecture. Comput. Mater. Continua83(1), 1023–1039. 10.32604/cmc.2025.061890 (2025). [Google Scholar]
- 27.Koubaa, A., Ammar, A., Abdelkader, M., Alhabashi, Y. & Ghouti, L. Aero: Ai-enabled remote sensing observation with onboard edge computing in uavs. Remote Sens.15 (7). 10.3390/rs15071873 (2023).
- 28.Roberts, M., Chen, W. & Davis, L. Object detection: Key metrics for computer vision performance in 2025. Comput. Vis. Mach. Learn. Rev.12, 45–62. 10.1007/s11042-025-18234-7 (2025). [Google Scholar]
- 29.Johnson, K., Liu, X. & Smith, R. An in-depth analysis of domain adaptation in computer and robotic vision. Appl. Sci.13(23), 12823. 10.3390/app132312823 (2023). [Google Scholar]
- 30.Wang, Y., Zhang, H. & Kumar, P. Cross-Domain FSOD Benchmark. Benchmark Dataset. https://www.emergentmind.com/topics/cross-domain-fsod-benchmark (2024).
- 31.Anderson, J., Brown, T. & Wilson, M. Scivid: Cross-domain evaluation of video models in scientific applications. 1–15. arXiv:2507.03578 (2024).
- 32.Garcia, R., Patel, S. & Lee, D. Vision transformers in domain adaptation and generalization: A study of robustness. 1–18. arXiv:2404.04452 (2024).
- 33.Ayesha, N. A vision transformer-based convolutional neural network for the automated diagnosis of eye diseases using self-attention mechanisms. Eng. Technol. Appl. Sci. Res.15(4), 24493–24497. 10.48084/etasr.10649 (2025). [Google Scholar]
- 34.Bissoonauth-Daiboo, P. et al. Exploring vision transformers and explainable ai for enhanced artefact classification in esophageal endoscopic images. IEEE Access13, 176221–176244. 10.1109/ACCESS.2025.3616796 (2025). [Google Scholar]
- 35.Miller, A., Thompson, K. & Zhang, L. Model performance degradation detection and mitigation in production systems. IEEE Trans. Mach. Learn.16(3), 234–248. 10.1109/TML.2024.3412567 (2024). [Google Scholar]
- 36.Amazon Web Services: MLPER-15: Monitor, detect, and handle model performance degradation - Machine Learning Lens. AWS Technical Documentation. https://docs.aws.amazon.com/wellarchitected/latest/machine-learning-lens/mlper-15.html (2024).
- 37.Lu, H., Liu, W., Zhang, B., Wang, B., Dong, K., Liu, B., Sun, J., Ren, T., Li, Z., Sun, H. et al. Deepseek-vl: Towards real-world vision-language understanding. arXiv:2403.05525 (2024).
- 38.Liu, H., Li, C., Li, Y. & Lee, Y.J. Improved baselines with visual instruction tuning. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 26286–26296. 10.1109/CVPR52733.2024.02484 (2024).
- 39.Grattafiori, A. The Llama 3 Herd of Models. arxiv:2407.21783 (2024).
- 40.Jiang, A. Q. et al. arxiv:2310.06825 (2023).
- 41.Bai, J., Bai, S., Yang, S., Wang, S., Tan, S., Wang, P., Lin, J., Zhou, C. & Zhou, J. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond. arXiv:2308.12966 (2023).
- 42.Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H. & Lin, J. Qwen2.5-VL Technical Report. arxiv:2502.13923 (2025).
- 43.Team, G. Gemma 3 Technical Report. arxiv:2503.19786 (2025).
- 44.Team, G. V. & Karlinsky, L. Granite Vision: a lightweight, open-source multimodal model for enterprise intelligence. arxiv:2502.09927 (2025).
- 45.Hastie, T., Tibshirani, R. & Friedman, J. The Elements of Statistical Learning: Data Mining, Inference, and Prediction 2nd edn. 10.1007/978-0-387-84858-7 (Springer, 2009).
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Data Availability Statement
I used open-source image datasets to conduct the study in this work, which are cited in the manuscript. The following are the public datasets with their URLs. COCO: https://cocodataset.org/. VisDrone-DET: https://github.com/VisDrone/VisDrone-Dataset. The specific VisDrone dataset used in this work is the VisDrone-DET which can be accessed via (https://drive.google.com/file/d/1a2oHjcEcwXP8oUF95qiwrqzACb2YlUhn/view?usp=sharing). Brain Tumor: https://huggingface.co/datasets/Ultralytics/Brain-tumor. Car Parts: https://universe.roboflow.com/gianmarco-russo-vt9xr/car-seg-un1pm. HAM10000: https://doi.org/10.7910/DVN/DBW86T
The benchmarking framework developed in this work is open-source and available at https://github.com/mzahana/vlm-bench.











