{
  "id": 77299,
  "title": "31st place solution (16th in public lb)",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77299",
  "author_name": "zhangboshen",
  "post_date": "2019-01-11T08:28:32.838000",
  "votes": 7,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thanks to Kaggle and HPA team for holding this interesting competition. \nAnd thanks to my teammates: <a href=\"https://www.kaggle.com/xf1994\">xf1994</a>, <a href=\"https://www.kaggle.com/cheninghan\">hanhan chen</a>, <a href=\"https://www.kaggle.com/zhouyanghust\">T-mac</a>, ...</p>\n\n<p>Below is some points of our solution.</p>\n\n<p><strong>Input</strong>: we trained our models using three different input sizes 1024*1024, 786*786, 512*512,  RGBY images. We initialize conv weights of Y channel from red channel in pretrain models. Besides, we tried self-designed dual-path network using two pathways of inputs (one for RGB, one for Y).</p>\n\n<p><strong>Models</strong>: we trained around 20+ models including se-resnet, se-resnext, resnet, resnext, dpn incpetions, inception-resnet and some self-designed networks. Generally these networks show similar performance except dpn. Dpn performs bad in this task.  Further, we reproduce a so-called MLFN network from : <a href=\"https://arxiv.org/abs/1803.09132\">https://arxiv.org/abs/1803.09132</a> , it did gave us some improvements compared with its counterpart ResNext50.</p>\n\n<p>Some good single models include:\nseresnext101: public 0.572, private 0.515\nself-designed dualpath: public 0.574, private 0.502</p>\n\n<p><strong>Data augmentation</strong>: rand horizontal flip and random vertical flip.</p>\n\n<p><strong>Loss</strong>: weighted BCE loss, thanks to <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">Tilii</a>.</p>\n\n<p><strong>Thresholds</strong>: One useful trick on public lb for us is finding thresholds for each class separately. We find the threshold according to the amount of each class. Adjusting the threshold according to the public lb helps a lot. Finally for most classes, we set the thresholds around 0.21. for some scarce classes, thresholds are around 0.1.</p>\n\n<p><strong>Ensemble</strong>: The most useful trick is ensemble. We ensemble our model totally according to the public lb performance. This gives us 16th in public lb. However, it drops to 31st in private lb. We checked our submission on private lb. Our best model on private lb is 0.554, whose public lb score is only 0.608. This model ensembles models more evenly compared with our best model on public lb.</p>\n\n<p><strong>Tried but didn't work</strong>: focal loss, oversampling, some augmentations(scale, rotate, shift)...</p>\n\n<p>We will release our code soon...</p>",
  "messages": [
    {
      "id": 454202,
      "postDate": "2019-01-11T08:28:32.840Z",
      "content": "<p>Thanks to Kaggle and HPA team for holding this interesting competition. \nAnd thanks to my teammates: <a href=\"https://www.kaggle.com/xf1994\">xf1994</a>, <a href=\"https://www.kaggle.com/cheninghan\">hanhan chen</a>, <a href=\"https://www.kaggle.com/zhouyanghust\">T-mac</a>, ...</p>\n\n<p>Below is some points of our solution.</p>\n\n<p><strong>Input</strong>: we trained our models using three different input sizes 1024*1024, 786*786, 512*512,  RGBY images. We initialize conv weights of Y channel from red channel in pretrain models. Besides, we tried self-designed dual-path network using two pathways of inputs (one for RGB, one for Y).</p>\n\n<p><strong>Models</strong>: we trained around 20+ models including se-resnet, se-resnext, resnet, resnext, dpn incpetions, inception-resnet and some self-designed networks. Generally these networks show similar performance except dpn. Dpn performs bad in this task.  Further, we reproduce a so-called MLFN network from : <a href=\"https://arxiv.org/abs/1803.09132\">https://arxiv.org/abs/1803.09132</a> , it did gave us some improvements compared with its counterpart ResNext50.</p>\n\n<p>Some good single models include:\nseresnext101: public 0.572, private 0.515\nself-designed dualpath: public 0.574, private 0.502</p>\n\n<p><strong>Data augmentation</strong>: rand horizontal flip and random vertical flip.</p>\n\n<p><strong>Loss</strong>: weighted BCE loss, thanks to <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065\">Tilii</a>.</p>\n\n<p><strong>Thresholds</strong>: One useful trick on public lb for us is finding thresholds for each class separately. We find the threshold according to the amount of each class. Adjusting the threshold according to the public lb helps a lot. Finally for most classes, we set the thresholds around 0.21. for some scarce classes, thresholds are around 0.1.</p>\n\n<p><strong>Ensemble</strong>: The most useful trick is ensemble. We ensemble our model totally according to the public lb performance. This gives us 16th in public lb. However, it drops to 31st in private lb. We checked our submission on private lb. Our best model on private lb is 0.554, whose public lb score is only 0.608. This model ensembles models more evenly compared with our best model on public lb.</p>\n\n<p><strong>Tried but didn't work</strong>: focal loss, oversampling, some augmentations(scale, rotate, shift)...</p>\n\n<p>We will release our code soon...</p>",
      "rawMarkdown": "Thanks to Kaggle and HPA team for holding this interesting competition. \nAnd thanks to my teammates: [xf1994](https://www.kaggle.com/xf1994), [hanhan chen](https://www.kaggle.com/cheninghan), [T-mac](https://www.kaggle.com/zhouyanghust), ...\n\nBelow is some points of our solution.\n\n**Input**: we trained our models using three different input sizes 1024*1024, 786*786, 512*512,  RGBY images. We initialize conv weights of Y channel from red channel in pretrain models. Besides, we tried self-designed dual-path network using two pathways of inputs (one for RGB, one for Y).\n\n**Models**: we trained around 20+ models including se-resnet, se-resnext, resnet, resnext, dpn incpetions, inception-resnet and some self-designed networks. Generally these networks show similar performance except dpn. Dpn performs bad in this task.  Further, we reproduce a so-called MLFN network from : https://arxiv.org/abs/1803.09132 , it did gave us some improvements compared with its counterpart ResNext50.\n\nSome good single models include:\nseresnext101: public 0.572, private 0.515\nself-designed dualpath: public 0.574, private 0.502\n\n**Data augmentation**: rand horizontal flip and random vertical flip.\n\n**Loss**: weighted BCE loss, thanks to [Tilii](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065).\n\n**Thresholds**: One useful trick on public lb for us is finding thresholds for each class separately. We find the threshold according to the amount of each class. Adjusting the threshold according to the public lb helps a lot. Finally for most classes, we set the thresholds around 0.21. for some scarce classes, thresholds are around 0.1.\n\n**Ensemble**: The most useful trick is ensemble. We ensemble our model totally according to the public lb performance. This gives us 16th in public lb. However, it drops to 31st in private lb. We checked our submission on private lb. Our best model on private lb is 0.554, whose public lb score is only 0.608. This model ensembles models more evenly compared with our best model on public lb.\n\n\n**Tried but didn't work**: focal loss, oversampling, some augmentations(scale, rotate, shift)...\n\nWe will release our code soon...",
      "votes": 7
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "454202": "Thanks to Kaggle and HPA team for holding this interesting competition. \nAnd thanks to my teammates: [xf1994](https://www.kaggle.com/xf1994), [hanhan chen](https://www.kaggle.com/cheninghan), [T-mac](https://www.kaggle.com/zhouyanghust), ...\n\nBelow is some points of our solution.\n\n**Input**: we trained our models using three different input sizes 1024*1024, 786*786, 512*512,  RGBY images. We initialize conv weights of Y channel from red channel in pretrain models. Besides, we tried self-designed dual-path network using two pathways of inputs (one for RGB, one for Y).\n\n**Models**: we trained around 20+ models including se-resnet, se-resnext, resnet, resnext, dpn incpetions, inception-resnet and some self-designed networks. Generally these networks show similar performance except dpn. Dpn performs bad in this task.  Further, we reproduce a so-called MLFN network from : https://arxiv.org/abs/1803.09132 , it did gave us some improvements compared with its counterpart ResNext50.\n\nSome good single models include:\nseresnext101: public 0.572, private 0.515\nself-designed dualpath: public 0.574, private 0.502\n\n**Data augmentation**: rand horizontal flip and random vertical flip.\n\n**Loss**: weighted BCE loss, thanks to [Tilii](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/74065).\n\n**Thresholds**: One useful trick on public lb for us is finding thresholds for each class separately. We find the threshold according to the amount of each class. Adjusting the threshold according to the public lb helps a lot. Finally for most classes, we set the thresholds around 0.21. for some scarce classes, thresholds are around 0.1.\n\n**Ensemble**: The most useful trick is ensemble. We ensemble our model totally according to the public lb performance. This gives us 16th in public lb. However, it drops to 31st in private lb. We checked our submission on private lb. Our best model on private lb is 0.554, whose public lb score is only 0.608. This model ensembles models more evenly compared with our best model on public lb.\n\n\n**Tried but didn't work**: focal loss, oversampling, some augmentations(scale, rotate, shift)...\n\nWe will release our code soon..."
  }
}