Review: Dual Attention Network for Scene Segmentation

https://cdn-images-1.medium.com/max/1557/0*9KqZX7G1EC2Rx8mk

Original Source Here

Review: Dual Attention Network for Scene Segmentation

The aim of this article is to provide a brief overview of this paper Dual Attention Network for Scene Segmentation.

Paper:
The paper can be found online here.

Publication:
CVPR 2019

Institution:
National Laboratory of Pattern Recognition, Institute of Automation, Chinese Academy of Sciences

Datasets:
The paper uses Cityscapes Dataset, Pascal VOC Dataset, Pascal Context Dataset and COCO-Stuff Dataset.

Overview:

Fig.1. Dual Attention Network

The paper asserts that although encoder-decoder architecture is a standard method for semantic segmentation and has achieved a lot of traction in recent years, it heavily relies on local information which may bring some bias as global information is not seen. The paper addresses this problem by capturing rich contextual dependencies based on the self-attention mechanism.

The paper proposes Dual Attention Network (DANet) which appends two attention modules (position attention module and channel attention module) on top of dilated FCN, which model the semantic interdependencies in spatial and channel dimensions respectively. The position attention module selectively aggregates the features at all positions. Similar features would be related to each other regardless of their distances. Meanwhile, the channel attention module selectively emphasizes interdependent channel maps. The network takes the sum of the outputs of the two attention modules to further improve feature representation which yields improvement in accuracy of semantic segmentation.

The succeeding paragraphs provide an overview of the following:

  • Dual Attention Network
  • Position Attention Module
  • Channel Attention Module
  • Experimental Results

The Fig.1. illustrates the overall network that takes the input image and passes it through a dilated residual block. Then the local features from the dilated ResNet are fed to two parallel convolution layers and then passed to position and channel attention blocks respectively. The outputs of the two attention modules are transformed by another convolutional layer and element-wise sum is computed to accomplish feature fusion. At last a convolution layer is followed to generate the final prediction map.

Fig.2. Position Attention Module

In order to model rich contextual relationships over local features, a position attention module is introduced. The position attention module encodes a wider range of contextual information into local features, thus enhancing their representation capability.

In the position attention module, given a local feature A ∈ R (C×H×W), it is first fed into a convolution layer to generate two new feature maps B and C, respectively, where {B, C} ∈ R (C×H×W). Then these feature maps are reshaped to R (C×N), where N = H×W is the number of pixels. After that matrix multiplication is performed between the transpose of C and B, and a softmax layer is applied to calculate the spatial attention map S ∈ R(N×N). The more similar feature representations of the two positions contributes to greater correlation between them. Meanwhile, feature A is fed into a convolution layer to generate a new feature map D ∈ R (C×H×W) and reshaped to R(C×N). Then a matrix multiplication is performed between D and the transpose of S and reshape the result to R (C×H×W). Finally, it is multiplied by a scale parameter α and performs an element-wise sum operation with the features A to obtain the final output E ∈ R (C×H×W). The scale parameter α is initialized to 0 and gradually learns the weight.

Fig.3. Channel Attention Module

Each channel map of high level features can be regarded as a class-specific response, and different semantic responses are associated with each other. By exploiting the interdependencies between channel maps, interdependent feature maps could be emphasized and the feature representation of specific semantics could be improved. Therefore, a channel attention module to explicitly model interdependencies between channels is proposed.

In the channel attention module, the channel attention map X ∈ R (C×C) is directly calculated from the original features A ∈ R (C×H×W). Specifically, A is reshaped to R(C×N), and then a matrix multiplication is performed between A and the transpose of A. Finally, a softmax layer is applied to obtain the channel attention map X ∈ R (C×C). In addition, a matrix multiplication between the transpose of X and A is performed and their result is reshaped to R (C×H×W). Then the result is multiplied by a scale parameter β and perform an element-wise sum operation with A to obtain the final output E ∈ R (C×H×W). The parameter β gradually learns a weight from 0.

Note that the attention modules are simple and can be directly inserted in the existing FCN pipeline.

The paper provides a detailed comparison of the results obtained with and without the attention modules respectively. The paper conducts a comprehensive ablation study to compare the results obtained with other state-of-the-art networks for semantic segmentation.

Fig.4. Visualization Results of Position Attention Module on Cityscapes Dataset
Fig.5. Visualization Results of Channel Attention Module on Cityscapes Dataset

The ablation experiments show that dual attention modules capture long-range contextual information effectively and give more precise segmentation results. The attention network achieves outstanding performance consistently on four scene segmentation datasets.

AI/ML

Trending AI/ML Article Identified & Digested via Granola by Ramsey Elbasheer; a Machine-Driven RSS Bot



via WordPress https://ramseyelbasheer.wordpress.com/2021/01/31/review-dual-attention-network-for-scene-segmentation/

Popular posts from this blog

Fully Explained DBScan Clustering Algorithm with Python

The 2021 machine learning, AI, and data landscape

Hierarchical clustering explained