Skip navigation

NEXT: Multi-grained mixture of experts via text-modulation for multi-modal object Re-Identification

NEXT: Multi-grained mixture of experts via text-modulation for multi-modal object Re-Identification

Li, Shihao ORCID logoORCID: https://orcid.org/0009-0001-0923-3965, Huang, Huaibo ORCID logoORCID: https://orcid.org/0000-0001-5866-2283, Duan, Junxian ORCID logoORCID: https://orcid.org/0000-0002-0218-6924, Zheng, Aihua ORCID logoORCID: https://orcid.org/0000-0002-9820-4743, Tang, Jin ORCID logoORCID: https://orcid.org/0000-0001-8375-3590 and Ma, Jixin ORCID logoORCID: https://orcid.org/0000-0001-7458-7412 (2026) NEXT: Multi-grained mixture of experts via text-modulation for multi-modal object Re-Identification. IEEE Transactions on Information Forensics and Security. ISSN 1556-6013 (Print), 1556-6021 (Online) (doi:10.1109/TIFS.2026.3729508)

[thumbnail of Author's Accepted Manuscript]
Preview
PDF (Author's Accepted Manuscript)
54353 MA_NEXT_Multi-Grained_Mixture_Of_Experts_Via_Text-Modulation_(IEEE AAM)_2026.pdf - Accepted Version
Available under License Creative Commons Attribution.

Download (7MB) | Preview

Abstract

Multi-modal object Re-IDentification (ReID) aims to obtain complete identity features across heterogeneous modalities. However, most existing methods rely on implicit feature fusion modules, making it difficult to model fine-grained recognition patterns under various challenges in real world. Benefiting from the powerful Multi-modal Large Language Models (MLLMs), the object appearances are effectively translated into descriptive captions. In this paper, we propose a reliable caption generation pipeline based on attribute confidence, which significantly reduces the unknown recognition rate of MLLMs and improves the quality of generated text. Additionally, to model diverse identity patterns, we propose a novel ReID framework, named NEXT, the Multi-grained Mixture of Experts via Text-Modulation for Multi-modal Object Re-Identification. Specifically, we decouple the recognition problem into semantic and structural branches to separately capture fine-grained appearance features and coarse-grained structure features. For semantic recognition, we first propose the Text-Modulated Semantic Experts (TMSE) module, which randomly samples high-quality captions to modulate experts capturing semantic features and mining inter-modality complementary cues. Second, to recognize structure features, we propose the Context-Shared Structure Experts (CSSE) module, which focuses on the holistic object structure and maintains identity structural consistency via a soft routing mechanism. Finally, we propose the Multi-Grained Features Aggregation (MGFA) module, which adopts a unified fusion strategy to effectively integrate multi-grained expert features into the final identity representations. Extensive experiments on two public person datasets and three vehicle datasets demonstrate the effectiveness of our method, showing that it significantly outperforms existing state-of-the-art methods.

Item Type: Article
Uncontrolled Keywords: modeling, conferences, visualization, ranking (statistics), routing, learning (artificial intelligence), vehicles, lighting, identification of persons, modulation
Subjects: H Social Sciences > HD Industries. Land use. Labor > HD61 Risk Management
Q Science > Q Science (General)
Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Faculty / School / Research Centre / Research Group: Faculty of Engineering & Science
Faculty of Engineering & Science > School of Computing & Mathematical Sciences (CMS)
Last Modified: 07 Sep 2026 11:55
URI: https://gala.gre.ac.uk/id/eprint/54353

Actions (login required)

View Item View Item

Downloads

Downloads per month over past year

View more statistics