Journals & Magazines >IEEE Transactions on Multimedia >Volume: 26

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations betwee...Show More

Metadata

Abstract:

Text-to-image person re-identification (ReID) aims to retrieve images of a person based on a given textual description. The key challenge is to learn the relations between detailed information from visual and textual modalities. Existing work focuses on learning a latent space to narrow the modality gap and further build local correspondences between two modalities. However, these methods assume that image-to-text and text-to-image associations are modality-agnostic, resulting in suboptimal associations. In this work, we demonstrate the discrepancy between image-to-text association and text-to-image association and proposecross-modal adaptive dual association (CADA) to build fine bidirectional image-text detailed associations. Our approach features a decoder-based adaptive dual association module that enables full interaction between visual and textual modalities, enabling bidirectional and adaptive cross-modal correspondence associations. Specifically, this paper proposes a bidirectional association mechanism: Association of text Tokens to image Patches (ATP) and Association of image Regions to text Attributes (ARA). We adaptively model the ATP based on the fact that aggregating cross-modal features based on mistaken associations will lead to feature distortion. For modeling the ARA, since attributes are typically the first distinguishing cues of a person, we explore attribute-level associations by predicting the masked text phrase using the related image region. Finally, we learn the dual associations between texts and images, and the experimental results demonstrate the superiority of our dual formulation.

Published in: IEEE Transactions on Multimedia ( Volume: 26)

Page(s): 6609 - 6620

Date of Publication: 18 January 2024

ISSN Information:

DOI: 10.1109/TMM.2024.3355644

Funding Agency:

Contents

I. Introduction

Text-to-image person retrieval is a task that involves retrieving a person of interest from a large image gallery that best matches a given textual description query [1]. Textual descriptions provide a natural and comprehensive way to describe a person's attributes and are more easily accessible than images. As a result, text-to-image person retrieval has received increasing attention in recent years, benefiting a variety of applications from personal photo album searches to public security.

References is not available for this document.

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

References

IEEE Account

Purchase Details

Profile Information

Need Help?

Cross-Modal Adaptive Dual Association for Text-to-Image Person Retrieval

Alerts

Abstract:

Metadata

Abstract:

ISSN Information:

Funding Agency:

I. Introduction

Authors

Figures

References

Citations

Keywords

Metrics

Footnotes

References

IEEE Account

Purchase Details

Profile Information

Need Help?