Abstract

Recently, video object segmentation (VOS) is a new challenging research direction from DAVIS competition. Carrying on with these researches, we propose a visual attention guided framework in video object segmentation, which includes four main components: segmentation network, visual encoder, spatial encoder and guide. The segmentation network predicts the object mask in the current video frame, and the visual guide force segmentation network to focus on the annotated object by visual information from visual encoder, and the spatial guide provide spatial location by spatial encoder from previous frame. Visual attention mechanism plays an important role in the model on capturing annotated object without online fine-tuning as previous models. This approach has an advantage over previous methods on accuracy and efficiency, especially avoid the online fine-tuning in those one-shot learning approaches.

Full Text
Published version (Free)

Talk to us

Join us for a 30 min session where you can share your feedback and ask us any queries you have

Schedule a call