Abstract

Compositional Zero-Shot Learning (CZSL) is a particular Zero-Shot Learning (ZSL) task that aims to utilize known concepts (e.g., states and objects) to identify novel state-object compositions for Image Classification. Previous works have primarily focused on disentangling concept compositions or exploring the complex interactions between the states and objects while neglecting the critical fact that the inference of many states and compositions is related to different frequency components, which should be analyzed from the global perspective. Therefore, we propose a Spatial-frequency Feature Fusion Network (SFFNet) to introduce a new branch that utilizes a frequency-domain filtering encoder to enhance key frequency components and capture non-local interactions adaptively. Besides, we also find that the widely used backbone in conventional CZSL settings behaves superior in perceiving local features. Thus, we construct a fusion block to combine both strengths to capture the local and non-local information. In addition, the traditional one-hot ground-truth distribution in the training phase does not reflect the accurate relationships between compositions, so we propose a composition-relation based label distribution regularization to encourage the model to actively learn the inner relationships between compositions, and extend this method to construct unseen composition pseudo distribution to further enhance the model’s generalization ability to unseen compositions. Extensive experiments and detailed analysis are conducted on three popular datasets, and the results show that our method can achieve state-of-the-art performance, which reveals its superiority in identifying novel compositions. Code is available at https://github.com/lisuyi/SFFNet_czsl.

Full Text
Paper version not known

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call

Disclaimer: All third-party content on this website/platform is and will remain the property of their respective owners and is provided on "as is" basis without any warranties, express or implied. Use of third-party content does not indicate any affiliation, sponsorship with or endorsement by them. Any references to third-party content is to identify the corresponding services and shall be considered fair use under The CopyrightLaw.