{
  "id": 238373,
  "title": "46th place solution - Simple Image Level Multilabel Classifier",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238373",
  "author_name": "Sunghyun Jun",
  "post_date": "2021-05-12T03:17:01.163000",
  "votes": 15,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congrats to all winners! Thanks to the organizers of this competition.</p>\n<p>And specially thanks to <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a>, <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>, <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>, <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>, <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> for sharing knowledge and datasets. I have learned a lot from you.</p>\n<h2>Summary</h2>\n<p>\"Simple Image Level Multilabel Classifier\"</p>\n<p>The label was determined by applying a classifier to the single cell mask obtained by HPA-Cell-Segmentation.</p>\n<p>Classifier was trained using the full dataset.</p>\n<h2>Tools</h2>\n<ul>\n<li>Colab Pro, GCE, Tesla V100 16GB single GPU</li>\n<li>GCS</li>\n<li>Pytorch Lightning</li>\n<li>Neptune</li>\n<li>Kaggle API</li>\n</ul>\n<h2>Dataset</h2>\n<p>I used both the Competitions default dataset and the extra dataset.</p>\n<p><a href=\"https://www.kaggle.com/phalanx/hpa-512512\" target=\"_blank\">HPA 512 PNG Dataset</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></p>\n<p><a href=\"https://www.kaggle.com/phalanx/hpa-768768\" target=\"_blank\">HPA 768 PNG Dataset</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></p>\n<p><a href=\"https://www.kaggle.com/sunghyunjun/hpa-1024-png-dataset\" target=\"_blank\">HPA 1024 PNG Dataset</a></p>\n<p><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/223822\" target=\"_blank\">HPA Public Data 768x768 \"rare classes\" dataset</a> by <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@Alexander Riedel</a></p>\n<p>The extra dataset was downloaded by referring to the public note.<br>\nImages saved to 768px png. The size is approximately 200 GB.</p>\n<p><a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">HPA public data download and HPACellSeg</a></p>\n<h2>Validation</h2>\n<p>MultilabelStratifiedKFold, 5-fold split was used.</p>\n<p>The performance of Multilabel Classifier was verified with Macro-F1, Micro-F1 Score.</p>\n<p><a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a></p>\n<h2>Model training</h2>\n<p>3-channel RGB images<br>\nThe image size is 1024px, and trained with the following dataset.</p>\n<ul>\n<li>1024px Competition default dataset + 768px rare classes dataset(1024 resized)</li>\n<li>1024px Competition default dataset + 768px extra dataset(1024 resized)</li>\n<li>AdamW</li>\n<li>CosineAnnealingLR</li>\n<li>epochs = 5 for full, 10 for rare<br>\nbce, focal loss was used.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>dataset</th>\n<th>folds</th>\n<th>loss</th>\n<th>batch_size</th>\n<th>init_lr</th>\n<th>weight_decay</th>\n<th>macro F1</th>\n<th>micro F1</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b0</td>\n<td>full</td>\n<td>2 of 5</td>\n<td>bce</td>\n<td>16</td>\n<td>6.0e-4</td>\n<td>1.0e-5</td>\n<td>0.7663</td>\n<td>0.8171</td>\n<td>0.454</td>\n<td>0.429</td>\n</tr>\n<tr>\n<td>efficientnet_b0</td>\n<td>rare classes</td>\n<td>single</td>\n<td>bce</td>\n<td>16</td>\n<td>6.0e-4</td>\n<td>1.0e-5</td>\n<td>0.8154</td>\n<td>0.8368</td>\n<td>0.394</td>\n<td>0.360</td>\n</tr>\n<tr>\n<td>seresnext26d_32x4d</td>\n<td>full</td>\n<td>single</td>\n<td>alpha=0.75, gamma=0.0</td>\n<td>14</td>\n<td>6.5e-5</td>\n<td>1.0e-5</td>\n<td>0.7317</td>\n<td>0.7956</td>\n<td>0.381</td>\n<td>0.335</td>\n</tr>\n<tr>\n<td><strong>final ensemble</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td><strong>0.471</strong></td>\n<td><strong>0.433</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Segmentation</h2>\n<p>HPA-Cell-Segmentation was used, and the speed was improved by referring to <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a> 's notebook.</p>\n<p>The input image was resized by 1/4, and the CellSegmentator scale_factor=1.0.</p>\n<p>The related values of the label_cell function have been adjusted to 1/4.</p>\n<p><a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPA-Cell-Segmentation</a></p>\n<p><a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">Faster HPA Cell Segmentation</a><br>\nby <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a></p>\n<h2>Augmentation</h2>\n<pre><code>A.Compose(\n    [\n        A.Resize(height=resize_height, width=resize_width),\n        A.RandomScale(scale_limit=(-0.2, 0.2), p=1.0),\n        A.PadIfNeeded(\n            min_height=resize_height,\n            min_width=resize_width,\n            border_mode=cv2.BORDER_CONSTANT,\n            value=0,\n            p=1.0,\n        ),\n        A.RandomCrop(height=resize_height, width=resize_width, p=1.0),\n        A.RandomBrightnessContrast(p=0.8),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(border_mode=cv2.BORDER_CONSTANT, value=0, p=0.5),\n        A.Normalize(mean=norm_mean, std=norm_std),\n        ToTensorV2(),\n    ]\n)\n</code></pre>\n<h2>TTA 4x</h2>\n<p>HorizontalFlip, VerticalFlip, Resize 0.8, Resize 1.2</p>\n<h2>What did not work</h2>\n<ul>\n<li><p>Label Smoothing</p></li>\n<li><p>pos/neg balanced weighted loss</p>\n<p>X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers.<br>\n(Dec. 2017). \"ChestX-ray8: Hospital-scale chest X-ray database and<br>\nbenchmarks on weakly-supervised classification and localization of common thorax diseases.\", (p. 5) <a href=\"https://arxiv.org/abs/1705.02315\" target=\"_blank\">https://arxiv.org/abs/1705.02315</a></p></li>\n</ul>\n<h2>Using GCS and errors</h2>\n<p>The dataset size of this competition is really huge. I had some difficult to download extra public data. It took a lot of time. Multiprocessing was not helped because more than two files couldn't be downloaded at the same time.</p>\n<p>I use colab-pro.I usually downloaded dataset to colab VM for convinience. But to train huge extra dataset I loaded dataset directly from my GCS bucket.</p>\n<p>I refer the article <a href=\"https://medium.com/pytorch/training-faster-with-large-datasets-using-scale-and-pytorch-946dfe774d8c\" target=\"_blank\">Training Faster With Large Datasets using Scale and PyTorch</a> And I didn't implement Asynchronous dataload. In my case, multiprocessing of torch.utils.data.DataLoader is enough for latency hiding.</p>\n<p>But training from GCS had got some rare errors. (504 GatewayTimeout, 104 Connection reset by peer)</p>\n<p>I don't know exact reason but it seems relate belows.</p>\n<ul>\n<li><p>opencv multithreading deadlock with pytorch DataLoader (num_workers&gt;0, pin_memory=True)<br>\n<a href=\"https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method\" target=\"_blank\">https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method</a><br>\n<a href=\"https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985</a><br>\nsolution: cv2.setNumThreads(0)</p></li>\n<li><p>CPU memory leaks of copy on write<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812</a><br>\nsolution<br>\n<a href=\"https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57\" target=\"_blank\">https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57</a></p></li>\n</ul>\n<h2>Source Code</h2>\n<p>Source code is available at <a href=\"https://github.com/sunghyunjun/kaggle-hpa\" target=\"_blank\">https://github.com/sunghyunjun/kaggle-hpa</a></p>\n<p>Submission notebook is <a href=\"https://www.kaggle.com/sunghyunjun/hpa-faster-final-ensemble-w-o-rot-exp-2\" target=\"_blank\">HPA faster final ensemble w/o rot exp 2</a></p>",
  "messages": [
    {
      "id": 1303356,
      "postDate": "2021-05-12T03:17:01.163Z",
      "content": "<p>Congrats to all winners! Thanks to the organizers of this competition.</p>\n<p>And specially thanks to <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a>, <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>, <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@alexanderriedel</a>, <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a>, <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a>, <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, <a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> for sharing knowledge and datasets. I have learned a lot from you.</p>\n<h2>Summary</h2>\n<p>\"Simple Image Level Multilabel Classifier\"</p>\n<p>The label was determined by applying a classifier to the single cell mask obtained by HPA-Cell-Segmentation.</p>\n<p>Classifier was trained using the full dataset.</p>\n<h2>Tools</h2>\n<ul>\n<li>Colab Pro, GCE, Tesla V100 16GB single GPU</li>\n<li>GCS</li>\n<li>Pytorch Lightning</li>\n<li>Neptune</li>\n<li>Kaggle API</li>\n</ul>\n<h2>Dataset</h2>\n<p>I used both the Competitions default dataset and the extra dataset.</p>\n<p><a href=\"https://www.kaggle.com/phalanx/hpa-512512\" target=\"_blank\">HPA 512 PNG Dataset</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></p>\n<p><a href=\"https://www.kaggle.com/phalanx/hpa-768768\" target=\"_blank\">HPA 768 PNG Dataset</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a></p>\n<p><a href=\"https://www.kaggle.com/sunghyunjun/hpa-1024-png-dataset\" target=\"_blank\">HPA 1024 PNG Dataset</a></p>\n<p><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/223822\" target=\"_blank\">HPA Public Data 768x768 \"rare classes\" dataset</a> by <a href=\"https://www.kaggle.com/alexanderriedel\" target=\"_blank\">@Alexander Riedel</a></p>\n<p>The extra dataset was downloaded by referring to the public note.<br>\nImages saved to 768px png. The size is approximately 200 GB.</p>\n<p><a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">HPA public data download and HPACellSeg</a></p>\n<h2>Validation</h2>\n<p>MultilabelStratifiedKFold, 5-fold split was used.</p>\n<p>The performance of Multilabel Classifier was verified with Macro-F1, Micro-F1 Score.</p>\n<p><a href=\"https://github.com/trent-b/iterative-stratification\" target=\"_blank\">iterative-stratification</a></p>\n<h2>Model training</h2>\n<p>3-channel RGB images<br>\nThe image size is 1024px, and trained with the following dataset.</p>\n<ul>\n<li>1024px Competition default dataset + 768px rare classes dataset(1024 resized)</li>\n<li>1024px Competition default dataset + 768px extra dataset(1024 resized)</li>\n<li>AdamW</li>\n<li>CosineAnnealingLR</li>\n<li>epochs = 5 for full, 10 for rare<br>\nbce, focal loss was used.</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>dataset</th>\n<th>folds</th>\n<th>loss</th>\n<th>batch_size</th>\n<th>init_lr</th>\n<th>weight_decay</th>\n<th>macro F1</th>\n<th>micro F1</th>\n<th>public LB</th>\n<th>private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b0</td>\n<td>full</td>\n<td>2 of 5</td>\n<td>bce</td>\n<td>16</td>\n<td>6.0e-4</td>\n<td>1.0e-5</td>\n<td>0.7663</td>\n<td>0.8171</td>\n<td>0.454</td>\n<td>0.429</td>\n</tr>\n<tr>\n<td>efficientnet_b0</td>\n<td>rare classes</td>\n<td>single</td>\n<td>bce</td>\n<td>16</td>\n<td>6.0e-4</td>\n<td>1.0e-5</td>\n<td>0.8154</td>\n<td>0.8368</td>\n<td>0.394</td>\n<td>0.360</td>\n</tr>\n<tr>\n<td>seresnext26d_32x4d</td>\n<td>full</td>\n<td>single</td>\n<td>alpha=0.75, gamma=0.0</td>\n<td>14</td>\n<td>6.5e-5</td>\n<td>1.0e-5</td>\n<td>0.7317</td>\n<td>0.7956</td>\n<td>0.381</td>\n<td>0.335</td>\n</tr>\n<tr>\n<td><strong>final ensemble</strong></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td></td>\n<td><strong>0.471</strong></td>\n<td><strong>0.433</strong></td>\n</tr>\n</tbody>\n</table>\n<h2>Segmentation</h2>\n<p>HPA-Cell-Segmentation was used, and the speed was improved by referring to <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a> 's notebook.</p>\n<p>The input image was resized by 1/4, and the CellSegmentator scale_factor=1.0.</p>\n<p>The related values of the label_cell function have been adjusted to 1/4.</p>\n<p><a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPA-Cell-Segmentation</a></p>\n<p><a href=\"https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation\" target=\"_blank\">Faster HPA Cell Segmentation</a><br>\nby <a href=\"https://www.kaggle.com/linshokaku\" target=\"_blank\">@linshokaku</a></p>\n<h2>Augmentation</h2>\n<pre><code>A.Compose(\n    [\n        A.Resize(height=resize_height, width=resize_width),\n        A.RandomScale(scale_limit=(-0.2, 0.2), p=1.0),\n        A.PadIfNeeded(\n            min_height=resize_height,\n            min_width=resize_width,\n            border_mode=cv2.BORDER_CONSTANT,\n            value=0,\n            p=1.0,\n        ),\n        A.RandomCrop(height=resize_height, width=resize_width, p=1.0),\n        A.RandomBrightnessContrast(p=0.8),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(border_mode=cv2.BORDER_CONSTANT, value=0, p=0.5),\n        A.Normalize(mean=norm_mean, std=norm_std),\n        ToTensorV2(),\n    ]\n)\n</code></pre>\n<h2>TTA 4x</h2>\n<p>HorizontalFlip, VerticalFlip, Resize 0.8, Resize 1.2</p>\n<h2>What did not work</h2>\n<ul>\n<li><p>Label Smoothing</p></li>\n<li><p>pos/neg balanced weighted loss</p>\n<p>X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers.<br>\n(Dec. 2017). \"ChestX-ray8: Hospital-scale chest X-ray database and<br>\nbenchmarks on weakly-supervised classification and localization of common thorax diseases.\", (p. 5) <a href=\"https://arxiv.org/abs/1705.02315\" target=\"_blank\">https://arxiv.org/abs/1705.02315</a></p></li>\n</ul>\n<h2>Using GCS and errors</h2>\n<p>The dataset size of this competition is really huge. I had some difficult to download extra public data. It took a lot of time. Multiprocessing was not helped because more than two files couldn't be downloaded at the same time.</p>\n<p>I use colab-pro.I usually downloaded dataset to colab VM for convinience. But to train huge extra dataset I loaded dataset directly from my GCS bucket.</p>\n<p>I refer the article <a href=\"https://medium.com/pytorch/training-faster-with-large-datasets-using-scale-and-pytorch-946dfe774d8c\" target=\"_blank\">Training Faster With Large Datasets using Scale and PyTorch</a> And I didn't implement Asynchronous dataload. In my case, multiprocessing of torch.utils.data.DataLoader is enough for latency hiding.</p>\n<p>But training from GCS had got some rare errors. (504 GatewayTimeout, 104 Connection reset by peer)</p>\n<p>I don't know exact reason but it seems relate belows.</p>\n<ul>\n<li><p>opencv multithreading deadlock with pytorch DataLoader (num_workers&gt;0, pin_memory=True)<br>\n<a href=\"https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method\" target=\"_blank\">https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method</a><br>\n<a href=\"https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985</a><br>\nsolution: cv2.setNumThreads(0)</p></li>\n<li><p>CPU memory leaks of copy on write<br>\n<a href=\"https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812</a><br>\nsolution<br>\n<a href=\"https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57\" target=\"_blank\">https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57</a></p></li>\n</ul>\n<h2>Source Code</h2>\n<p>Source code is available at <a href=\"https://github.com/sunghyunjun/kaggle-hpa\" target=\"_blank\">https://github.com/sunghyunjun/kaggle-hpa</a></p>\n<p>Submission notebook is <a href=\"https://www.kaggle.com/sunghyunjun/hpa-faster-final-ensemble-w-o-rot-exp-2\" target=\"_blank\">HPA faster final ensemble w/o rot exp 2</a></p>",
      "rawMarkdown": "Congrats to all winners! Thanks to the organizers of this competition.\n\nAnd specially thanks to @rdizzl3, @phalanx, @alexanderriedel, @thedrcat, @linshokaku, @its7171, @dschettler8845, @samusram for sharing knowledge and datasets. I have learned a lot from you.\n\n## Summary\n\n\"Simple Image Level Multilabel Classifier\"\n\nThe label was determined by applying a classifier to the single cell mask obtained by HPA-Cell-Segmentation.\n\nClassifier was trained using the full dataset.\n\n## Tools\n\n- Colab Pro, GCE, Tesla V100 16GB single GPU\n- GCS\n- Pytorch Lightning\n- Neptune\n- Kaggle API\n\n## Dataset\n\nI used both the Competitions default dataset and the extra dataset.\n\n[HPA 512 PNG Dataset](https://www.kaggle.com/phalanx/hpa-512512) by [@phalanx](https://www.kaggle.com/phalanx)\n\n[HPA 768 PNG Dataset](https://www.kaggle.com/phalanx/hpa-768768) by [@phalanx](https://www.kaggle.com/phalanx)\n\n[HPA 1024 PNG Dataset](https://www.kaggle.com/sunghyunjun/hpa-1024-png-dataset)\n\n[HPA Public Data 768x768 \"rare classes\" dataset](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/223822) by [@Alexander Riedel](https://www.kaggle.com/alexanderriedel)\n\nThe extra dataset was downloaded by referring to the public note.\nImages saved to 768px png. The size is approximately 200 GB.\n\n[HPA public data download and HPACellSeg](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg)\n\n## Validation\n\nMultilabelStratifiedKFold, 5-fold split was used.\n\nThe performance of Multilabel Classifier was verified with Macro-F1, Micro-F1 Score.\n\n[iterative-stratification](https://github.com/trent-b/iterative-stratification)\n\n## Model training\n\n3-channel RGB images\nThe image size is 1024px, and trained with the following dataset.\n\n- 1024px Competition default dataset + 768px rare classes dataset(1024 resized)\n- 1024px Competition default dataset + 768px extra dataset(1024 resized)\n- AdamW\n- CosineAnnealingLR\n- epochs = 5 for full, 10 for rare\nbce, focal loss was used.\n\n|model|dataset|folds|loss|batch_size|init_lr|weight_decay|macro F1|micro F1|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|---|---|\n|efficientnet_b0|full|2 of 5|bce|16|6.0e-4|1.0e-5|0.7663|0.8171|0.454|0.429|\n|efficientnet_b0|rare classes|single|bce|16|6.0e-4|1.0e-5|0.8154|0.8368|0.394|0.360|\n|seresnext26d_32x4d|full|single|alpha=0.75, gamma=0.0|14|6.5e-5|1.0e-5|0.7317|0.7956|0.381|0.335|\n|**final ensemble**|||||||||**0.471**|**0.433**|\n\n## Segmentation\n\nHPA-Cell-Segmentation was used, and the speed was improved by referring to @linshokaku 's notebook.\n\nThe input image was resized by 1/4, and the CellSegmentator scale_factor=1.0.\n\nThe related values of the label_cell function have been adjusted to 1/4.\n\n[HPA-Cell-Segmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)\n\n[Faster HPA Cell Segmentation](https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation)\nby @linshokaku\n\n## Augmentation\n\n```python\nA.Compose(\n    [\n        A.Resize(height=resize_height, width=resize_width),\n        A.RandomScale(scale_limit=(-0.2, 0.2), p=1.0),\n        A.PadIfNeeded(\n            min_height=resize_height,\n            min_width=resize_width,\n            border_mode=cv2.BORDER_CONSTANT,\n            value=0,\n            p=1.0,\n        ),\n        A.RandomCrop(height=resize_height, width=resize_width, p=1.0),\n        A.RandomBrightnessContrast(p=0.8),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(border_mode=cv2.BORDER_CONSTANT, value=0, p=0.5),\n        A.Normalize(mean=norm_mean, std=norm_std),\n        ToTensorV2(),\n    ]\n)\n```\n\n## TTA 4x\n\nHorizontalFlip, VerticalFlip, Resize 0.8, Resize 1.2\n\n## What did not work\n\n- Label Smoothing\n\n- pos/neg balanced weighted loss\n\n    X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers.\n(Dec. 2017). \"ChestX-ray8: Hospital-scale chest X-ray database and\nbenchmarks on weakly-supervised classification and localization of common thorax diseases.\", (p. 5) [https://arxiv.org/abs/1705.02315](https://arxiv.org/abs/1705.02315)\n\n## Using GCS and errors\n\nThe dataset size of this competition is really huge. I had some difficult to download extra public data. It took a lot of time. Multiprocessing was not helped because more than two files couldn't be downloaded at the same time.\n\nI use colab-pro.I usually downloaded dataset to colab VM for convinience. But to train huge extra dataset I loaded dataset directly from my GCS bucket.\n\nI refer the article [Training Faster With Large Datasets using Scale and PyTorch](https://medium.com/pytorch/training-faster-with-large-datasets-using-scale-and-pytorch-946dfe774d8c) And I didn't implement Asynchronous dataload. In my case, multiprocessing of torch.utils.data.DataLoader is enough for latency hiding.\n\nBut training from GCS had got some rare errors. (504 GatewayTimeout, 104 Connection reset by peer)\n\nI don't know exact reason but it seems relate belows.\n\n- opencv multithreading deadlock with pytorch DataLoader (num_workers>0, pin_memory=True)\n[https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method](https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method)\n[https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985](https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985)\nsolution: cv2.setNumThreads(0)\n\n- CPU memory leaks of copy on write\n[https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812](https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812)\nsolution\n[https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57](https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57)\n\n## Source Code\nSource code is available at [https://github.com/sunghyunjun/kaggle-hpa](https://github.com/sunghyunjun/kaggle-hpa)\n\nSubmission notebook is [HPA faster final ensemble w/o rot exp 2](https://www.kaggle.com/sunghyunjun/hpa-faster-final-ensemble-w-o-rot-exp-2)",
      "votes": 15
    },
    {
      "id": 1303711,
      "postDate": "2021-05-12T07:49:57.743Z",
      "content": "<p>Thanks for the explanation, and congratulations for your achievement </p>",
      "rawMarkdown": "Thanks for the explanation, and congratulations for your achievement ",
      "votes": 1
    },
    {
      "id": 1303446,
      "postDate": "2021-05-12T04:32:44.627Z",
      "content": "<p>Congratulations on your result! :) Thank you for the great write-up and for your kind words!</p>",
      "rawMarkdown": "Congratulations on your result! :) Thank you for the great write-up and for your kind words!",
      "votes": 1
    },
    {
      "id": 1303396,
      "postDate": "2021-05-12T03:54:25.617Z",
      "content": "<p><a href=\"https://www.kaggle.com/sunghyunjun\" target=\"_blank\">@sunghyunjun</a> Congratulations and Thanks for sharing the approach</p>",
      "rawMarkdown": "@sunghyunjun Congratulations and Thanks for sharing the approach",
      "votes": 1
    },
    {
      "id": 1303359,
      "postDate": "2021-05-12T03:19:46.830Z",
      "content": "<p>Awesome write up! Congratulations !</p>",
      "rawMarkdown": "Awesome write up! Congratulations !",
      "votes": 1,
      "replies": [
        {
          "id": 1303363,
          "postDate": "2021-05-12T03:24:58.220Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1303711,
      "author_name": "Salim Khazem",
      "author_url": "",
      "post_date": "2021-05-12T07:49:57.743000",
      "content": "<p>Thanks for the explanation, and congratulations for your achievement </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1303446,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-12T04:32:44.627000",
      "content": "<p>Congratulations on your result! :) Thank you for the great write-up and for your kind words!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1303396,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-12T03:54:25.617000",
      "content": "<p><a href=\"https://www.kaggle.com/sunghyunjun\" target=\"_blank\">@sunghyunjun</a> Congratulations and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1303359,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-12T03:19:46.830000",
      "content": "<p>Awesome write up! Congratulations !</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1303363,
          "author_name": "Sunghyun Jun",
          "author_url": "",
          "post_date": "2021-05-12T03:24:58.220000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1303356": "Congrats to all winners! Thanks to the organizers of this competition.\n\nAnd specially thanks to @rdizzl3, @phalanx, @alexanderriedel, @thedrcat, @linshokaku, @its7171, @dschettler8845, @samusram for sharing knowledge and datasets. I have learned a lot from you.\n\n## Summary\n\n\"Simple Image Level Multilabel Classifier\"\n\nThe label was determined by applying a classifier to the single cell mask obtained by HPA-Cell-Segmentation.\n\nClassifier was trained using the full dataset.\n\n## Tools\n\n- Colab Pro, GCE, Tesla V100 16GB single GPU\n- GCS\n- Pytorch Lightning\n- Neptune\n- Kaggle API\n\n## Dataset\n\nI used both the Competitions default dataset and the extra dataset.\n\n[HPA 512 PNG Dataset](https://www.kaggle.com/phalanx/hpa-512512) by [@phalanx](https://www.kaggle.com/phalanx)\n\n[HPA 768 PNG Dataset](https://www.kaggle.com/phalanx/hpa-768768) by [@phalanx](https://www.kaggle.com/phalanx)\n\n[HPA 1024 PNG Dataset](https://www.kaggle.com/sunghyunjun/hpa-1024-png-dataset)\n\n[HPA Public Data 768x768 \"rare classes\" dataset](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/223822) by [@Alexander Riedel](https://www.kaggle.com/alexanderriedel)\n\nThe extra dataset was downloaded by referring to the public note.\nImages saved to 768px png. The size is approximately 200 GB.\n\n[HPA public data download and HPACellSeg](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg)\n\n## Validation\n\nMultilabelStratifiedKFold, 5-fold split was used.\n\nThe performance of Multilabel Classifier was verified with Macro-F1, Micro-F1 Score.\n\n[iterative-stratification](https://github.com/trent-b/iterative-stratification)\n\n## Model training\n\n3-channel RGB images\nThe image size is 1024px, and trained with the following dataset.\n\n- 1024px Competition default dataset + 768px rare classes dataset(1024 resized)\n- 1024px Competition default dataset + 768px extra dataset(1024 resized)\n- AdamW\n- CosineAnnealingLR\n- epochs = 5 for full, 10 for rare\nbce, focal loss was used.\n\n|model|dataset|folds|loss|batch_size|init_lr|weight_decay|macro F1|micro F1|public LB|private LB|\n|---|---|---|---|---|---|---|---|---|---|---|\n|efficientnet_b0|full|2 of 5|bce|16|6.0e-4|1.0e-5|0.7663|0.8171|0.454|0.429|\n|efficientnet_b0|rare classes|single|bce|16|6.0e-4|1.0e-5|0.8154|0.8368|0.394|0.360|\n|seresnext26d_32x4d|full|single|alpha=0.75, gamma=0.0|14|6.5e-5|1.0e-5|0.7317|0.7956|0.381|0.335|\n|**final ensemble**|||||||||**0.471**|**0.433**|\n\n## Segmentation\n\nHPA-Cell-Segmentation was used, and the speed was improved by referring to @linshokaku 's notebook.\n\nThe input image was resized by 1/4, and the CellSegmentator scale_factor=1.0.\n\nThe related values of the label_cell function have been adjusted to 1/4.\n\n[HPA-Cell-Segmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)\n\n[Faster HPA Cell Segmentation](https://www.kaggle.com/linshokaku/faster-hpa-cell-segmentation)\nby @linshokaku\n\n## Augmentation\n\n```python\nA.Compose(\n    [\n        A.Resize(height=resize_height, width=resize_width),\n        A.RandomScale(scale_limit=(-0.2, 0.2), p=1.0),\n        A.PadIfNeeded(\n            min_height=resize_height,\n            min_width=resize_width,\n            border_mode=cv2.BORDER_CONSTANT,\n            value=0,\n            p=1.0,\n        ),\n        A.RandomCrop(height=resize_height, width=resize_width, p=1.0),\n        A.RandomBrightnessContrast(p=0.8),\n        A.HorizontalFlip(p=0.5),\n        A.VerticalFlip(p=0.5),\n        A.Rotate(border_mode=cv2.BORDER_CONSTANT, value=0, p=0.5),\n        A.Normalize(mean=norm_mean, std=norm_std),\n        ToTensorV2(),\n    ]\n)\n```\n\n## TTA 4x\n\nHorizontalFlip, VerticalFlip, Resize 0.8, Resize 1.2\n\n## What did not work\n\n- Label Smoothing\n\n- pos/neg balanced weighted loss\n\n    X. Wang, Y. Peng, L. Lu, Z. Lu, M. Bagheri, and R. M. Summers.\n(Dec. 2017). \"ChestX-ray8: Hospital-scale chest X-ray database and\nbenchmarks on weakly-supervised classification and localization of common thorax diseases.\", (p. 5) [https://arxiv.org/abs/1705.02315](https://arxiv.org/abs/1705.02315)\n\n## Using GCS and errors\n\nThe dataset size of this competition is really huge. I had some difficult to download extra public data. It took a lot of time. Multiprocessing was not helped because more than two files couldn't be downloaded at the same time.\n\nI use colab-pro.I usually downloaded dataset to colab VM for convinience. But to train huge extra dataset I loaded dataset directly from my GCS bucket.\n\nI refer the article [Training Faster With Large Datasets using Scale and PyTorch](https://medium.com/pytorch/training-faster-with-large-datasets-using-scale-and-pytorch-946dfe774d8c) And I didn't implement Asynchronous dataload. In my case, multiprocessing of torch.utils.data.DataLoader is enough for latency hiding.\n\nBut training from GCS had got some rare errors. (504 GatewayTimeout, 104 Connection reset by peer)\n\nI don't know exact reason but it seems relate belows.\n\n- opencv multithreading deadlock with pytorch DataLoader (num_workers>0, pin_memory=True)\n[https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method](https://stackoverflow.com/questions/54013846/pytorch-dataloader-stucked-if-using-opencv-resize-method)\n[https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985](https://github.com/pytorch/pytorch/issues/1355#issuecomment-675018985)\nsolution: cv2.setNumThreads(0)\n\n- CPU memory leaks of copy on write\n[https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812](https://github.com/pytorch/pytorch/issues/13246#issuecomment-737442812)\nsolution\n[https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57](https://gist.github.com/vadimkantorov/86c3a46bf25bed3ad45d043ae86fff57)\n\n## Source Code\nSource code is available at [https://github.com/sunghyunjun/kaggle-hpa](https://github.com/sunghyunjun/kaggle-hpa)\n\nSubmission notebook is [HPA faster final ensemble w/o rot exp 2](https://www.kaggle.com/sunghyunjun/hpa-faster-final-ensemble-w-o-rot-exp-2)",
    "1303711": "Thanks for the explanation, and congratulations for your achievement ",
    "1303446": "Congratulations on your result! :) Thank you for the great write-up and for your kind words!",
    "1303396": "@sunghyunjun Congratulations and Thanks for sharing the approach",
    "1303359": "Awesome write up! Congratulations !"
  }
}