Explorerβ€ΊBiomedical Engineeringβ€ΊEngineering
Research PaperResearchia:202609.10034

Cost-Aware Vision--Language Model Arbitration for Fabric Structure Recognition A Deployable Multi-Agent System

Chenwei Wang

Abstract

Recognizing a fabric's structure is a prerequisite for translating textile-specific material information into structured digital form for downstream supply-chain systems. Pure CNN classifiers are cost-efficient but fail on visually ambiguous categories; vision--language models (VLMs) generalize more broadly but cost much more per image and are unstable on specialist domains. We present a multi-agent system in which a CNN cascade handles the easy majority and a VLM is invoked only as a selective ...

Submitted: September 10, 2026Subjects: Engineering; Biomedical Engineering

Description / Details

Recognizing a fabric's structure is a prerequisite for translating textile-specific material information into structured digital form for downstream supply-chain systems. Pure CNN classifiers are cost-efficient but fail on visually ambiguous categories; vision--language models (VLMs) generalize more broadly but cost much more per image and are unstable on specialist domains. We present a multi-agent system in which a CNN cascade handles the easy majority and a VLM is invoked only as a selective arbiter, constrained to a top-3 taxonomy-consistent choice. The fabric taxonomy performs as a constraint for the whole recognition process to increase the accuracy and reduce the VLM calls. Meanwhile, the CNN cascade is distilled to a small parameter size to reduce the inference time and meet the needs of practical deployment. On a newly curated 14-class benchmark, a flat ConvNeXt-Tiny baseline reaches 90.45 %90.45\,\% top-1 and 76.9 %76.9\,\% on the four hardest classes; \method's hierarchical cascade reaches 93.94 %93.94\,\% top-1 and 94.50 %94.50\,\% hard (+17.6+17.6,pp). Tightening the VLM trigger from 60 %60\,\% to < ⁣10 %<\!10\,\% cuts API cost by ∼90 %{\sim}90\,\% with no measurable accuracy loss. CPU inference is ≀ ⁣93\le\!93,ms without a VLM call (9.39.3,ms distilled). Each prediction carries a machine-readable reasoning record, offered as an entry point for future supply-chain documentation.


Source: arXiv:2609.10065v1 - http://arxiv.org/abs/2609.10065v1 PDF: https://arxiv.org/pdf/2609.10065v1 Original Link: http://arxiv.org/abs/2609.10065v1

Please sign in to join the discussion.

No comments yet. Be the first to share your thoughts!

Access Paper
View Source PDF
Submission Info
Date:
Sep 10, 2026
Topic:
Biomedical Engineering
Area:
Engineering
Comments:
0
Bookmark