{
  "id": 70650,
  "title": "6th place solution: U-net-like segmentation",
  "url": "/competitions/rsna-pneumonia-detection-challenge/writeups/pfneumonia-6th-place-solution-u-net-like-segmentat",
  "author_name": "",
  "post_date": "2018-11-07T12:31:28.117Z",
  "votes": 26,
  "comment_count": 5,
  "views": 0,
  "content": "<p>We are software engineers at PFN -- Preferred Networks Inc., Japan. Our solution is based on semantic segmentation, using U-net-like network architecture. One key insight of ours is that it is crucial to detect the ‘edges’ of bounding-boxes. We tried to predict top, bottom, left and right edges of bounding-boxes separately.</p>\n\n<h2>U-net-like backbone</h2>\n\n<p>We adopted Imagenet-pretrained ResNet152 as the feature extractor. Input image is resized to 512x512, processed by the extractor to the size 16x16, then unpooled four times to get 256x256 output. For each step of unpooling, the feature vector is concatenated with the corresponding ResNet layer, then processed 3x3 convolution and ReLU twice. The shape of the final output is 64x256x256.</p>\n\n<h2>‘seg’ and ‘edge’ predictions</h2>\n\n<p>A 1-channel 1x1 convolution is applied to this output to get ‘seg’ layer. This layer indicates the confidence of each pixel being inside a bounding-box. Sigmoid of this layer is copied and concatenated to the 64-channel output, applied 3x3 convolution and ReLU twice, and finally applied a 1x1 convolution to produce 4-channel ‘edge’ layer. The first channel of this layer indicates the confidence of each pixel being one of the ‘topmost’ pixels of a bounding-box. The other 3 channels are for bottommost, leftmost and rightmost.</p>\n\n<p>The loss for ‘seg’ layer is <code>1 - f1</code>, where f1 is the differentiable F1 score of segmentation. The loss for ‘edge’ layer is cross-entropy loss. The whole loss of the network is the sum of these two losses.\nAs most of pixels of ground-truth for ‘edge’ layers are negative (even for images that have bounding-boxes), such negative pixels are appropriately undersampled.</p>\n\n<h2>Training details</h2>\n\n<p>We used x-flip, 90-degree rotation, zoom-in/out and random contrast changes (let <code>k</code> be a random int between [-10, 10], add <code>k</code> over all pixels) as data augmentation. Positive samples were three times oversampled. The batch size was 10.</p>\n\n<p>We used Adam optimizer with weight decay 1e-4, and trained for 30 epochs. Lr was divided by 10 after 20 and 27 epochs finished. We freezed the weight of extractor for the first 200 iterations, then fine-tuned over the rest of training.</p>\n\n<p>All trainings were done with eight Tesla P100s. It took about 6 hours to complete a training.</p>\n\n<h2>Inference</h2>\n\n<p>At inference time, first we split each image into two pieces (right and left) by the line <code>x = c</code>, where c is the weighted average of x-coordinates of all pixels (weight = pixel values). For each piece of image, we examine every possible rectangle with at least 40 px height and width, and find one that maximizes <code>p = p_top * p_bottom * p_left * p_right</code> where <code>p_top</code> is the geometric mean of sigmoid of the first channel of ‘edge’ layer, over the topmost pixels of the rectangle being examined (same as <code>p_bottom</code>, <code>p_left</code> and <code>p_right</code>). Finally, if p &gt;= 0.3, it is considered as a prediction.</p>\n\n<h2>Test-time augmentation and ensembles</h2>\n\n<p>At test time, we perform x-flip augmentation in which ‘seg’ layers and ‘edge’ layers are both averaged over flipped and non-flipped images.\nWe perform 10-fold CV and ensemble them only by averaging ‘edge’ layers, as averaging ‘seg’ layers over different models turned out to score worse. We divide 10 models evenly into 2 groups, and predict with them independently to get 2 submission files. Finally, we merge those two submissions: if two boxes overlap with an IoU more than 0.5, we adopt the intersection as the final prediction. Otherwise, we preserve the whole bbox as the prediction.</p>\n\n<h2>Things that didn’t work</h2>\n\n<ul>\n<li>Predicting the class labels of images as well as ‘seg’ and ‘edge' layers made the training unstable, and scored worse than the solution above even at the best of times.</li>\n<li>Using deconvolutions instead of unpoolings scored about the same or slightly worse, probably because the final layers should be simple rectangles or lines.</li>\n<li>Using cross-entropy loss instead of f1-loss for ’seg’ layer made the training longer to converge. F1-loss allows higher variance of confidence values, which might help ‘edge’ layer to guess. </li>\n</ul>",
  "messages": [
    {
      "id": "416137",
      "postDate": "11/06/2018 08:47:35",
      "content": "<p>We are software engineers at PFN -- Preferred Networks Inc., Japan. Our solution is based on semantic segmentation, using U-net-like network architecture. One key insight of ours is that it is crucial to detect the ‘edges’ of bounding-boxes. We tried to predict top, bottom, left and right edges of bounding-boxes separately.</p>\n\n<h2>U-net-like backbone</h2>\n\n<p>We adopted Imagenet-pretrained ResNet152 as the feature extractor. Input image is resized to 512x512, processed by the extractor to the size 16x16, then unpooled four times to get 256x256 output. For each step of unpooling, the feature vector is concatenated with the corresponding ResNet layer, then processed 3x3 convolution and ReLU twice. The shape of the final output is 64x256x256.</p>\n\n<h2>‘seg’ and ‘edge’ predictions</h2>\n\n<p>A 1-channel 1x1 convolution is applied to this output to get ‘seg’ layer. This layer indicates the confidence of each pixel being inside a bounding-box. Sigmoid of this layer is copied and concatenated to the 64-channel output, applied 3x3 convolution and ReLU twice, and finally applied a 1x1 convolution to produce 4-channel ‘edge’ layer. The first channel of this layer indicates the confidence of each pixel being one of the ‘topmost’ pixels of a bounding-box. The other 3 channels are for bottommost, leftmost and rightmost.</p>\n\n<p>The loss for ‘seg’ layer is <code>1 - f1</code>, where f1 is the differentiable F1 score of segmentation. The loss for ‘edge’ layer is cross-entropy loss. The whole loss of the network is the sum of these two losses.\nAs most of pixels of ground-truth for ‘edge’ layers are negative (even for images that have bounding-boxes), such negative pixels are appropriately undersampled.</p>\n\n<h2>Training details</h2>\n\n<p>We used x-flip, 90-degree rotation, zoom-in/out and random contrast changes (let <code>k</code> be a random int between [-10, 10], add <code>k</code> over all pixels) as data augmentation. Positive samples were three times oversampled. The batch size was 10.</p>\n\n<p>We used Adam optimizer with weight decay 1e-4, and trained for 30 epochs. Lr was divided by 10 after 20 and 27 epochs finished. We freezed the weight of extractor for the first 200 iterations, then fine-tuned over the rest of training.</p>\n\n<p>All trainings were done with eight Tesla P100s. It took about 6 hours to complete a training.</p>\n\n<h2>Inference</h2>\n\n<p>At inference time, first we split each image into two pieces (right and left) by the line <code>x = c</code>, where c is the weighted average of x-coordinates of all pixels (weight = pixel values). For each piece of image, we examine every possible rectangle with at least 40 px height and width, and find one that maximizes <code>p = p_top * p_bottom * p_left * p_right</code> where <code>p_top</code> is the geometric mean of sigmoid of the first channel of ‘edge’ layer, over the topmost pixels of the rectangle being examined (same as <code>p_bottom</code>, <code>p_left</code> and <code>p_right</code>). Finally, if p &gt;= 0.3, it is considered as a prediction.</p>\n\n<h2>Test-time augmentation and ensembles</h2>\n\n<p>At test time, we perform x-flip augmentation in which ‘seg’ layers and ‘edge’ layers are both averaged over flipped and non-flipped images.\nWe perform 10-fold CV and ensemble them only by averaging ‘edge’ layers, as averaging ‘seg’ layers over different models turned out to score worse. We divide 10 models evenly into 2 groups, and predict with them independently to get 2 submission files. Finally, we merge those two submissions: if two boxes overlap with an IoU more than 0.5, we adopt the intersection as the final prediction. Otherwise, we preserve the whole bbox as the prediction.</p>\n\n<h2>Things that didn’t work</h2>\n\n<ul>\n<li>Predicting the class labels of images as well as ‘seg’ and ‘edge' layers made the training unstable, and scored worse than the solution above even at the best of times.</li>\n<li>Using deconvolutions instead of unpoolings scored about the same or slightly worse, probably because the final layers should be simple rectangles or lines.</li>\n<li>Using cross-entropy loss instead of f1-loss for ’seg’ layer made the training longer to converge. F1-loss allows higher variance of confidence values, which might help ‘edge’ layer to guess. </li>\n</ul>",
      "rawMarkdown": "We are software engineers at PFN -- Preferred Networks Inc., Japan. Our solution is based on semantic segmentation, using U-net-like network architecture. One key insight of ours is that it is crucial to detect the ‘edges’ of bounding-boxes. We tried to predict top, bottom, left and right edges of bounding-boxes separately.\n\n## U-net-like backbone\nWe adopted Imagenet-pretrained ResNet152 as the feature extractor. Input image is resized to 512x512, processed by the extractor to the size 16x16, then unpooled four times to get 256x256 output. For each step of unpooling, the feature vector is concatenated with the corresponding ResNet layer, then processed 3x3 convolution and ReLU twice. The shape of the final output is 64x256x256.\n\n## ‘seg’ and ‘edge’ predictions\nA 1-channel 1x1 convolution is applied to this output to get ‘seg’ layer. This layer indicates the confidence of each pixel being inside a bounding-box. Sigmoid of this layer is copied and concatenated to the 64-channel output, applied 3x3 convolution and ReLU twice, and finally applied a 1x1 convolution to produce 4-channel ‘edge’ layer. The first channel of this layer indicates the confidence of each pixel being one of the ‘topmost’ pixels of a bounding-box. The other 3 channels are for bottommost, leftmost and rightmost.\n\nThe loss for ‘seg’ layer is `1 - f1`, where f1 is the differentiable F1 score of segmentation. The loss for ‘edge’ layer is cross-entropy loss. The whole loss of the network is the sum of these two losses.\nAs most of pixels of ground-truth for ‘edge’ layers are negative (even for images that have bounding-boxes), such negative pixels are appropriately undersampled.\n\n## Training details\nWe used x-flip, 90-degree rotation, zoom-in/out and random contrast changes (let `k` be a random int between [-10, 10], add `k` over all pixels) as data augmentation. Positive samples were three times oversampled. The batch size was 10.\n\nWe used Adam optimizer with weight decay 1e-4, and trained for 30 epochs. Lr was divided by 10 after 20 and 27 epochs finished. We freezed the weight of extractor for the first 200 iterations, then fine-tuned over the rest of training.\n\nAll trainings were done with eight Tesla P100s. It took about 6 hours to complete a training.\n\n## Inference\nAt inference time, first we split each image into two pieces (right and left) by the line `x = c`, where c is the weighted average of x-coordinates of all pixels (weight = pixel values). For each piece of image, we examine every possible rectangle with at least 40 px height and width, and find one that maximizes `p = p_top * p_bottom * p_left * p_right` where `p_top` is the geometric mean of sigmoid of the first channel of ‘edge’ layer, over the topmost pixels of the rectangle being examined (same as `p_bottom`, `p_left` and `p_right`). Finally, if p &gt;= 0.3, it is considered as a prediction.\n\n## Test-time augmentation and ensembles\nAt test time, we perform x-flip augmentation in which ‘seg’ layers and ‘edge’ layers are both averaged over flipped and non-flipped images.\nWe perform 10-fold CV and ensemble them only by averaging ‘edge’ layers, as averaging ‘seg’ layers over different models turned out to score worse. We divide 10 models evenly into 2 groups, and predict with them independently to get 2 submission files. Finally, we merge those two submissions: if two boxes overlap with an IoU more than 0.5, we adopt the intersection as the final prediction. Otherwise, we preserve the whole bbox as the prediction.\n\n## Things that didn’t work\n- Predicting the class labels of images as well as ‘seg’ and ‘edge' layers made the training unstable, and scored worse than the solution above even at the best of times.\n- Using deconvolutions instead of unpoolings scored about the same or slightly worse, probably because the final layers should be simple rectangles or lines.\n- Using cross-entropy loss instead of f1-loss for ’seg’ layer made the training longer to converge. F1-loss allows higher variance of confidence values, which might help ‘edge’ layer to guess.",
      "votes": null
    },
    {
      "id": "416160",
      "postDate": "11/06/2018 09:48:30",
      "content": "<p>Congratulation! I wish I were in your team. The approach of your team looks so tricky. Did you select a segmentation approach from first? Did you compare it with other object detection approaches?</p>",
      "rawMarkdown": "Congratulation! I wish I were in your team. The approach of your team looks so tricky. Did you select a segmentation approach from first? Did you compare it with other object detection approaches?",
      "votes": null
    },
    {
      "id": "416193",
      "postDate": "11/06/2018 11:24:20",
      "content": "<p>Thank you. We tried both segmentation based and object-detection (Faster R-CNN) based approaches in parallel. We got the better result with segmentation based approach, so we decided to dedicate ourselves to this solution.</p>",
      "rawMarkdown": "Thank you. We tried both segmentation based and object-detection (Faster R-CNN) based approaches in parallel. We got the better result with segmentation based approach, so we decided to dedicate ourselves to this solution.",
      "votes": null
    },
    {
      "id": "416538",
      "postDate": "11/06/2018 21:02:04",
      "content": "<p>Congrats and thanks for sharing. Interesting solution!</p>",
      "rawMarkdown": "Congrats and thanks for sharing. Interesting solution!",
      "votes": null
    },
    {
      "id": "416795",
      "postDate": "11/07/2018 09:46:11",
      "content": "<p>Congratulations! This is a very different and neat approach, particularly the inference part. Our team thought about segmenting, but we had no idea how to turn a mask into bbox coordinates, mostly due to multiple bboxes in some cases. Well done!</p>",
      "rawMarkdown": "Congratulations! This is a very different and neat approach, particularly the inference part. Our team thought about segmenting, but we had no idea how to turn a mask into bbox coordinates, mostly due to multiple bboxes in some cases. Well done!",
      "votes": null
    },
    {
      "id": "417951",
      "postDate": "11/09/2018 03:38:50",
      "content": "<p>Our source codes are now available at <a href=\"https://github.com/pfnet-research/pfneumonia\">https://github.com/pfnet-research/pfneumonia</a></p>",
      "rawMarkdown": "Our source codes are now available at https://github.com/pfnet-research/pfneumonia",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 416160,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "11/06/2018 09:48:30",
      "content": "<p>Congratulation! I wish I were in your team. The approach of your team looks so tricky. Did you select a segmentation approach from first? Did you compare it with other object detection approaches?</p>",
      "votes": null,
      "replies": [
        {
          "id": 416193,
          "author_name": "yhirano",
          "author_url": "",
          "post_date": "11/06/2018 11:24:20",
          "content": "<p>Thank you. We tried both segmentation based and object-detection (Faster R-CNN) based approaches in parallel. We got the better result with segmentation based approach, so we decided to dedicate ourselves to this solution.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 416538,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "11/06/2018 21:02:04",
      "content": "<p>Congrats and thanks for sharing. Interesting solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 416795,
      "author_name": "felipekitamura",
      "author_url": "",
      "post_date": "11/07/2018 09:46:11",
      "content": "<p>Congratulations! This is a very different and neat approach, particularly the inference part. Our team thought about segmenting, but we had no idea how to turn a mask into bbox coordinates, mostly due to multiple bboxes in some cases. Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417951,
      "author_name": "yhirano",
      "author_url": "",
      "post_date": "11/09/2018 03:38:50",
      "content": "<p>Our source codes are now available at <a href=\"https://github.com/pfnet-research/pfneumonia\">https://github.com/pfnet-research/pfneumonia</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "416137": "We are software engineers at PFN -- Preferred Networks Inc., Japan. Our solution is based on semantic segmentation, using U-net-like network architecture. One key insight of ours is that it is crucial to detect the ‘edges’ of bounding-boxes. We tried to predict top, bottom, left and right edges of bounding-boxes separately.\n\n## U-net-like backbone\nWe adopted Imagenet-pretrained ResNet152 as the feature extractor. Input image is resized to 512x512, processed by the extractor to the size 16x16, then unpooled four times to get 256x256 output. For each step of unpooling, the feature vector is concatenated with the corresponding ResNet layer, then processed 3x3 convolution and ReLU twice. The shape of the final output is 64x256x256.\n\n## ‘seg’ and ‘edge’ predictions\nA 1-channel 1x1 convolution is applied to this output to get ‘seg’ layer. This layer indicates the confidence of each pixel being inside a bounding-box. Sigmoid of this layer is copied and concatenated to the 64-channel output, applied 3x3 convolution and ReLU twice, and finally applied a 1x1 convolution to produce 4-channel ‘edge’ layer. The first channel of this layer indicates the confidence of each pixel being one of the ‘topmost’ pixels of a bounding-box. The other 3 channels are for bottommost, leftmost and rightmost.\n\nThe loss for ‘seg’ layer is `1 - f1`, where f1 is the differentiable F1 score of segmentation. The loss for ‘edge’ layer is cross-entropy loss. The whole loss of the network is the sum of these two losses.\nAs most of pixels of ground-truth for ‘edge’ layers are negative (even for images that have bounding-boxes), such negative pixels are appropriately undersampled.\n\n## Training details\nWe used x-flip, 90-degree rotation, zoom-in/out and random contrast changes (let `k` be a random int between [-10, 10], add `k` over all pixels) as data augmentation. Positive samples were three times oversampled. The batch size was 10.\n\nWe used Adam optimizer with weight decay 1e-4, and trained for 30 epochs. Lr was divided by 10 after 20 and 27 epochs finished. We freezed the weight of extractor for the first 200 iterations, then fine-tuned over the rest of training.\n\nAll trainings were done with eight Tesla P100s. It took about 6 hours to complete a training.\n\n## Inference\nAt inference time, first we split each image into two pieces (right and left) by the line `x = c`, where c is the weighted average of x-coordinates of all pixels (weight = pixel values). For each piece of image, we examine every possible rectangle with at least 40 px height and width, and find one that maximizes `p = p_top * p_bottom * p_left * p_right` where `p_top` is the geometric mean of sigmoid of the first channel of ‘edge’ layer, over the topmost pixels of the rectangle being examined (same as `p_bottom`, `p_left` and `p_right`). Finally, if p &gt;= 0.3, it is considered as a prediction.\n\n## Test-time augmentation and ensembles\nAt test time, we perform x-flip augmentation in which ‘seg’ layers and ‘edge’ layers are both averaged over flipped and non-flipped images.\nWe perform 10-fold CV and ensemble them only by averaging ‘edge’ layers, as averaging ‘seg’ layers over different models turned out to score worse. We divide 10 models evenly into 2 groups, and predict with them independently to get 2 submission files. Finally, we merge those two submissions: if two boxes overlap with an IoU more than 0.5, we adopt the intersection as the final prediction. Otherwise, we preserve the whole bbox as the prediction.\n\n## Things that didn’t work\n- Predicting the class labels of images as well as ‘seg’ and ‘edge' layers made the training unstable, and scored worse than the solution above even at the best of times.\n- Using deconvolutions instead of unpoolings scored about the same or slightly worse, probably because the final layers should be simple rectangles or lines.\n- Using cross-entropy loss instead of f1-loss for ’seg’ layer made the training longer to converge. F1-loss allows higher variance of confidence values, which might help ‘edge’ layer to guess.",
    "416160": "Congratulation! I wish I were in your team. The approach of your team looks so tricky. Did you select a segmentation approach from first? Did you compare it with other object detection approaches?",
    "416193": "Thank you. We tried both segmentation based and object-detection (Faster R-CNN) based approaches in parallel. We got the better result with segmentation based approach, so we decided to dedicate ourselves to this solution.",
    "416538": "Congrats and thanks for sharing. Interesting solution!",
    "416795": "Congratulations! This is a very different and neat approach, particularly the inference part. Our team thought about segmenting, but we had no idea how to turn a mask into bbox coordinates, mostly due to multiple bboxes in some cases. Well done!",
    "417951": "Our source codes are now available at https://github.com/pfnet-research/pfneumonia"
  },
  "source": "meta"
}