{
  "id": 238487,
  "title": "5th place solution - my part",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238487",
  "author_name": "narsil (jobs-in-data.com)",
  "post_date": "2021-05-12T10:49:18.622000",
  "votes": 35,
  "comment_count": 14,
  "views": 0,
  "content": "<p>First of all, I would like to thank my fantastic teammates <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.</p>\n<p>Big congrats to <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.</p>\n<p>Congrats to all other teams - it was great to compete with you. </p>\n<p>Finally, I would like to thank the hosts for creating such an interesting problem for us to tackle.</p>\n<p>I will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:</p>\n<p><strong>1. Model on cell-level and progressive pseudo-labeling</strong> </p>\n<p>I started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:</p>\n<pre><code>threshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           &amp; ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] &gt; threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] &gt; threshold_pred)  ) ).astype(int)\n</code></pre>\n<p>The logic behind the above formula is the following:<br>\nI set the label for a single-cell image to 1 for a given class only if:</p>\n<ul>\n<li>The whole image has label 1 for this class</li>\n<li>This particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image</li>\n</ul>\n<p>Those parameters were tuned using feedback from LB. I did 3 iterations -&gt; models -&gt; preds -&gt; labels. This was the single biggest source of boost for my models. </p>\n<p><strong>2. Filtering our cells detected by segmentation model, but invisible to humans</strong> </p>\n<p>When the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):</p>\n<pre><code>cell_img[2,:,:][cell_img[2,:,:]&gt;5] &lt; 25 \n</code></pre>\n<p>Such cells were removed from the predictions</p>\n<p><strong>3. Manual review of mitotic spindle</strong></p>\n<p>I could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.</p>",
  "messages": [
    {
      "id": 1303941,
      "postDate": "2021-05-12T10:49:18.623Z",
      "content": "<p>First of all, I would like to thank my fantastic teammates <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.</p>\n<p>Big congrats to <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.</p>\n<p>Congrats to all other teams - it was great to compete with you. </p>\n<p>Finally, I would like to thank the hosts for creating such an interesting problem for us to tackle.</p>\n<p>I will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:</p>\n<p><strong>1. Model on cell-level and progressive pseudo-labeling</strong> </p>\n<p>I started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:</p>\n<pre><code>threshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           &amp; ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] &gt; threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] &gt; threshold_pred)  ) ).astype(int)\n</code></pre>\n<p>The logic behind the above formula is the following:<br>\nI set the label for a single-cell image to 1 for a given class only if:</p>\n<ul>\n<li>The whole image has label 1 for this class</li>\n<li>This particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image</li>\n</ul>\n<p>Those parameters were tuned using feedback from LB. I did 3 iterations -&gt; models -&gt; preds -&gt; labels. This was the single biggest source of boost for my models. </p>\n<p><strong>2. Filtering our cells detected by segmentation model, but invisible to humans</strong> </p>\n<p>When the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):</p>\n<pre><code>cell_img[2,:,:][cell_img[2,:,:]&gt;5] &lt; 25 \n</code></pre>\n<p>Such cells were removed from the predictions</p>\n<p><strong>3. Manual review of mitotic spindle</strong></p>\n<p>I could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.</p>",
      "rawMarkdown": "First of all, I would like to thank my fantastic teammates @its7171 and @tivfrvqhs5 . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.\n\nBig congrats to @bestfitting who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.\n\nCongrats to all other teams - it was great to compete with you. \n\nFinally, I would like to thank the hosts for creating such an interesting problem for us to tackle.\n\nI will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:\n\n**1. Model on cell-level and progressive pseudo-labeling** \n\nI started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:\n\n```\nthreshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           & ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] > threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] > threshold_pred)  ) ).astype(int)\n```\n\nThe logic behind the above formula is the following:\nI set the label for a single-cell image to 1 for a given class only if:\n- The whole image has label 1 for this class\n- This particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image\n\nThose parameters were tuned using feedback from LB. I did 3 iterations -> models -> preds -> labels. This was the single biggest source of boost for my models. \n\n**2. Filtering our cells detected by segmentation model, but invisible to humans** \n\nWhen the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):\n\n```\ncell_img[2,:,:][cell_img[2,:,:]>5] < 25 \n```\nSuch cells were removed from the predictions\n\n**3. Manual review of mitotic spindle**\n\nI could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.",
      "votes": 35
    },
    {
      "id": 1610759,
      "postDate": "2021-12-07T13:49:58.043Z",
      "content": "<p>日本語訳</p>\n<p>First of all, I would like to thank my fantastic teammates <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.</p>\n<p>Big congrats to <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.</p>\n<p>Congrats to all other teams - it was great to compete with you.</p>\n<p>Finally, I would like to thank the hosts for creating such an interesting problem for us to tackle.</p>\n<p>I will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:</p>\n<ol>\n<li>Model on cell-level and progressive pseudo-labeling</li>\n</ol>\n<p>I started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:<br>\n最初は画像レベル全体でトレーニングされたモデルから始め、次に単一セルレベルでトレーニングされたモデルに移行しました。 シングルセル画像にラベルを割り当てるとき、私は次のアプローチを使用しました。<br>\nthreshold_std_above_mean = 0.5<br>\nthreshold_pred = 0.9</p>\n<p>for i in range(num_classes):<br>\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) <br>\n                           &amp; ( (cell_level_df[f'img_pred_rank_{i}'] == 1) <br>\n                                   | (cell_level_df[f'std_from_mean_{i}'] &gt; threshold_std_above_mean)<br>\n                                   | (cell_level_df[f'pred_class_{i}'] &gt; threshold_pred)  ) ).astype(int)<br>\nThe logic behind the above formula is the following:<br>\nI set the label for a single-cell image to 1 for a given class only if:</p>\n<p>The whole image has label 1 for this class<br>\nThis particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image<br>\nThose parameters were tuned using feedback from LB. I did 3 iterations -&gt; models -&gt; preds -&gt; labels. This was the single biggest source of boost for my models.</p>\n<p>上記の式の背後にあるロジックは次のとおりです。<br>\n次の場合にのみ、特定のクラスの単一セル画像のラベルを1に設定します。</p>\n<p>このクラスの画像全体にラベル1が付いています<br>\nこの特定のセルは、画像内のすべてのセルの中でこのクラスの予測が最も高いか、0.9を超えるか、この画像のこのクラスの平均予測よりも0.5標準偏差高くなっています。<br>\nこれらのパラメーターは、LBからのフィードバックを使用して調整されました。 3回の反復-&gt;モデル-&gt;プレデズ-&gt;ラベルを実行しました。 これは私のモデルにとって唯一最大のブーストの源でした。</p>\n<ol>\n<li>Filtering our cells detected by segmentation model, but invisible to humans</li>\n</ol>\n<p>When the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):</p>\n<p>cell_img[2,:,:][cell_img[2,:,:]&gt;5] &lt; 25 <br>\nSuch cells were removed from the predictions</p>\n<ol>\n<li>Manual review of mitotic spindle</li>\n</ol>\n<p>I could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.</p>",
      "rawMarkdown": "日本語訳\n\nFirst of all, I would like to thank my fantastic teammates @its7171 and @tivfrvqhs5 . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.\n\nBig congrats to @bestfitting who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.\n\nCongrats to all other teams - it was great to compete with you.\n\nFinally, I would like to thank the hosts for creating such an interesting problem for us to tackle.\n\nI will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:\n\n1. Model on cell-level and progressive pseudo-labeling\n\nI started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:\n最初は画像レベル全体でトレーニングされたモデルから始め、次に単一セルレベルでトレーニングされたモデルに移行しました。 シングルセル画像にラベルを割り当てるとき、私は次のアプローチを使用しました。\nthreshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           & ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] > threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] > threshold_pred)  ) ).astype(int)\nThe logic behind the above formula is the following:\nI set the label for a single-cell image to 1 for a given class only if:\n\nThe whole image has label 1 for this class\nThis particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image\nThose parameters were tuned using feedback from LB. I did 3 iterations -> models -> preds -> labels. This was the single biggest source of boost for my models.\n\n上記の式の背後にあるロジックは次のとおりです。\n次の場合にのみ、特定のクラスの単一セル画像のラベルを1に設定します。\n\nこのクラスの画像全体にラベル1が付いています\nこの特定のセルは、画像内のすべてのセルの中でこのクラスの予測が最も高いか、0.9を超えるか、この画像のこのクラスの平均予測よりも0.5標準偏差高くなっています。\nこれらのパラメーターは、LBからのフィードバックを使用して調整されました。 3回の反復->モデル->プレデズ->ラベルを実行しました。 これは私のモデルにとって唯一最大のブーストの源でした。\n\n2. Filtering our cells detected by segmentation model, but invisible to humans\n\nWhen the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):\n\ncell_img[2,:,:][cell_img[2,:,:]>5] < 25 \nSuch cells were removed from the predictions\n\n3. Manual review of mitotic spindle\n\nI could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.",
      "votes": 1
    },
    {
      "id": 1305461,
      "postDate": "2021-05-13T09:41:47.903Z",
      "content": "<blockquote>\n  <p>I can release this dataset if anyone is interested.</p>\n</blockquote>\n<p>That will be very interesting to compare our labels. I released the same dataset <a href=\"https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "> I can release this dataset if anyone is interested.\n\nThat will be very interesting to compare our labels. I released the same dataset [here](https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup).\n",
      "votes": 1
    },
    {
      "id": 1305029,
      "postDate": "2021-05-13T04:27:27.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@narsil Congratulations  and Thanks for sharing the approach",
      "votes": 1,
      "replies": [
        {
          "id": 1305240,
          "postDate": "2021-05-13T07:08:12.740Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "rawMarkdown": "Thanks @usharengaraju !"
        }
      ]
    },
    {
      "id": 1304271,
      "postDate": "2021-05-12T14:19:41.073Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! We were both inspired by the data-centric approach, but you were able to make it work, this is amazing :) </p>",
      "rawMarkdown": "Congratulations @narsil! We were both inspired by the data-centric approach, but you were able to make it work, this is amazing :) ",
      "votes": 1,
      "replies": [
        {
          "id": 1305232,
          "postDate": "2021-05-13T07:04:52.247Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> ! Congrats on your amazing result too!</p>",
          "rawMarkdown": "Thanks @thedrcat ! Congrats on your amazing result too!"
        }
      ]
    },
    {
      "id": 1304209,
      "postDate": "2021-05-12T13:37:21.550Z",
      "content": "<p>Congrats, Paweł! Thank you for sharing the details, impressive combination of cell-level and image-level predictions in your progressive pseudo-labeling part!</p>\n<p>p.s. And thanks a lot for your activity in the discussions! To me, it made the competition even more awesome 😊</p>",
      "rawMarkdown": "Congrats, Paweł! Thank you for sharing the details, impressive combination of cell-level and image-level predictions in your progressive pseudo-labeling part!\n\np.s. And thanks a lot for your activity in the discussions! To me, it made the competition even more awesome 😊",
      "votes": 1,
      "replies": [
        {
          "id": 1305234,
          "postDate": "2021-05-13T07:06:00.187Z",
          "content": "<p>Thanks Raman! I enjoy discussions on Kaggle - the community here is simply amazing .</p>",
          "rawMarkdown": "Thanks Raman! I enjoy discussions on Kaggle - the community here is simply amazing ."
        }
      ]
    },
    {
      "id": 1304103,
      "postDate": "2021-05-12T12:29:05.830Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>. progressive pseudo-labeling approach is interesting.</p>",
      "rawMarkdown": "Congrats @narsil, @its7171 and @tivfrvqhs5. progressive pseudo-labeling approach is interesting.",
      "votes": 1,
      "replies": [
        {
          "id": 1305236,
          "postDate": "2021-05-13T07:06:43.640Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> Congrats on your result too! </p>",
          "rawMarkdown": "Thanks @corochann Congrats on your result too! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304060,
      "postDate": "2021-05-12T11:57:48.107Z",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> and team! Congratulations and thanks for your solution summary!</p>",
      "rawMarkdown": "Great job @narsil and team! Congratulations and thanks for your solution summary!",
      "votes": 1,
      "replies": [
        {
          "id": 1305237,
          "postDate": "2021-05-13T07:07:04.340Z",
          "content": "<p>Thanks Sasza! You know where the next stop is :)</p>",
          "rawMarkdown": "Thanks Sasza! You know where the next stop is :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1304037,
      "postDate": "2021-05-12T11:46:11.163Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> and team on 5th place and thanks for sharing details approach </p>",
      "rawMarkdown": "Congratulations @narsil and team on 5th place and thanks for sharing details approach ",
      "votes": 1,
      "replies": [
        {
          "id": 1305238,
          "postDate": "2021-05-13T07:07:44.390Z",
          "content": "<p>Thanks! it was a super challenging and interesting problem to work on</p>",
          "rawMarkdown": "Thanks! it was a super challenging and interesting problem to work on"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1610759,
      "author_name": "pixyz0130",
      "author_url": "",
      "post_date": "2021-12-07T13:49:58.043000",
      "content": "<p>日本語訳</p>\n<p>First of all, I would like to thank my fantastic teammates <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a> . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.</p>\n<p>Big congrats to <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.</p>\n<p>Congrats to all other teams - it was great to compete with you.</p>\n<p>Finally, I would like to thank the hosts for creating such an interesting problem for us to tackle.</p>\n<p>I will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:</p>\n<ol>\n<li>Model on cell-level and progressive pseudo-labeling</li>\n</ol>\n<p>I started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:<br>\n最初は画像レベル全体でトレーニングされたモデルから始め、次に単一セルレベルでトレーニングされたモデルに移行しました。 シングルセル画像にラベルを割り当てるとき、私は次のアプローチを使用しました。<br>\nthreshold_std_above_mean = 0.5<br>\nthreshold_pred = 0.9</p>\n<p>for i in range(num_classes):<br>\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) <br>\n                           &amp; ( (cell_level_df[f'img_pred_rank_{i}'] == 1) <br>\n                                   | (cell_level_df[f'std_from_mean_{i}'] &gt; threshold_std_above_mean)<br>\n                                   | (cell_level_df[f'pred_class_{i}'] &gt; threshold_pred)  ) ).astype(int)<br>\nThe logic behind the above formula is the following:<br>\nI set the label for a single-cell image to 1 for a given class only if:</p>\n<p>The whole image has label 1 for this class<br>\nThis particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image<br>\nThose parameters were tuned using feedback from LB. I did 3 iterations -&gt; models -&gt; preds -&gt; labels. This was the single biggest source of boost for my models.</p>\n<p>上記の式の背後にあるロジックは次のとおりです。<br>\n次の場合にのみ、特定のクラスの単一セル画像のラベルを1に設定します。</p>\n<p>このクラスの画像全体にラベル1が付いています<br>\nこの特定のセルは、画像内のすべてのセルの中でこのクラスの予測が最も高いか、0.9を超えるか、この画像のこのクラスの平均予測よりも0.5標準偏差高くなっています。<br>\nこれらのパラメーターは、LBからのフィードバックを使用して調整されました。 3回の反復-&gt;モデル-&gt;プレデズ-&gt;ラベルを実行しました。 これは私のモデルにとって唯一最大のブーストの源でした。</p>\n<ol>\n<li>Filtering our cells detected by segmentation model, but invisible to humans</li>\n</ol>\n<p>When the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):</p>\n<p>cell_img[2,:,:][cell_img[2,:,:]&gt;5] &lt; 25 <br>\nSuch cells were removed from the predictions</p>\n<ol>\n<li>Manual review of mitotic spindle</li>\n</ol>\n<p>I could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1305461,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2021-05-13T09:41:47.903000",
      "content": "<blockquote>\n  <p>I can release this dataset if anyone is interested.</p>\n</blockquote>\n<p>That will be very interesting to compare our labels. I released the same dataset <a href=\"https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup\" target=\"_blank\">here</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1305029,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:27:27.847000",
      "content": "<p><a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305240,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:08:12.740000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304271,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-05-12T14:19:41.073000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>! We were both inspired by the data-centric approach, but you were able to make it work, this is amazing :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305232,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:04:52.247000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> ! Congrats on your amazing result too!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304209,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-12T13:37:21.550000",
      "content": "<p>Congrats, Paweł! Thank you for sharing the details, impressive combination of cell-level and image-level predictions in your progressive pseudo-labeling part!</p>\n<p>p.s. And thanks a lot for your activity in the discussions! To me, it made the competition even more awesome 😊</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305234,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:06:00.187000",
          "content": "<p>Thanks Raman! I enjoy discussions on Kaggle - the community here is simply amazing .</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304103,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-12T12:29:05.830000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a>, <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a> and <a href=\"https://www.kaggle.com/tivfrvqhs5\" target=\"_blank\">@tivfrvqhs5</a>. progressive pseudo-labeling approach is interesting.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305236,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:06:43.640000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/corochann\" target=\"_blank\">@corochann</a> Congrats on your result too! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304060,
      "author_name": "Alvor",
      "author_url": "",
      "post_date": "2021-05-12T11:57:48.107000",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> and team! Congratulations and thanks for your solution summary!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305237,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:07:04.340000",
          "content": "<p>Thanks Sasza! You know where the next stop is :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1304037,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2021-05-12T11:46:11.163000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> and team on 5th place and thanks for sharing details approach </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305238,
          "author_name": "narsil (jobs-in-data.com)",
          "author_url": "",
          "post_date": "2021-05-13T07:07:44.390000",
          "content": "<p>Thanks! it was a super challenging and interesting problem to work on</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1303941": "First of all, I would like to thank my fantastic teammates @its7171 and @tivfrvqhs5 . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.\n\nBig congrats to @bestfitting who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.\n\nCongrats to all other teams - it was great to compete with you. \n\nFinally, I would like to thank the hosts for creating such an interesting problem for us to tackle.\n\nI will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:\n\n**1. Model on cell-level and progressive pseudo-labeling** \n\nI started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:\n\n```\nthreshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           & ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] > threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] > threshold_pred)  ) ).astype(int)\n```\n\nThe logic behind the above formula is the following:\nI set the label for a single-cell image to 1 for a given class only if:\n- The whole image has label 1 for this class\n- This particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image\n\nThose parameters were tuned using feedback from LB. I did 3 iterations -> models -> preds -> labels. This was the single biggest source of boost for my models. \n\n**2. Filtering our cells detected by segmentation model, but invisible to humans** \n\nWhen the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):\n\n```\ncell_img[2,:,:][cell_img[2,:,:]>5] < 25 \n```\nSuch cells were removed from the predictions\n\n**3. Manual review of mitotic spindle**\n\nI could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.",
    "1610759": "日本語訳\n\nFirst of all, I would like to thank my fantastic teammates @its7171 and @tivfrvqhs5 . We realized 5 days before the deadline that we have to recalculate everything, and we managed to do so, selecting our submissions 1 hour before the deadline. The fact that it worked is a miracle.\n\nBig congrats to @bestfitting who once again has shown his greatness, to the surprise of nobody :) We were suspecting you were #1 already for a long time, even when we were higher on public LB.\n\nCongrats to all other teams - it was great to compete with you.\n\nFinally, I would like to thank the hosts for creating such an interesting problem for us to tackle.\n\nI will share key parts of my solution, which brought the largest score boost. Other components of my models are fairly standard:\n\n1. Model on cell-level and progressive pseudo-labeling\n\nI started with models trained on a whole image level, then I moved to models trained on a single-cell level. When assigning labels to single-cell images, I used the following approach:\n最初は画像レベル全体でトレーニングされたモデルから始め、次に単一セルレベルでトレーニングされたモデルに移行しました。 シングルセル画像にラベルを割り当てるとき、私は次のアプローチを使用しました。\nthreshold_std_above_mean = 0.5\nthreshold_pred = 0.9\n\nfor i in range(num_classes):\n    cell_level_df[f'cell_label_class{i}'] = ((cell_level_df[f'gt_class_{i}'] == 1) \n                           & ( (cell_level_df[f'img_pred_rank_{i}'] == 1) \n                                   | (cell_level_df[f'std_from_mean_{i}'] > threshold_std_above_mean)\n                                   | (cell_level_df[f'pred_class_{i}'] > threshold_pred)  ) ).astype(int)\nThe logic behind the above formula is the following:\nI set the label for a single-cell image to 1 for a given class only if:\n\nThe whole image has label 1 for this class\nThis particular cell has the highest prediction for this class among all cells in the image, or is above 0.9 or is 0.5 standard deviations higher than the mean prediction for this class on this image\nThose parameters were tuned using feedback from LB. I did 3 iterations -> models -> preds -> labels. This was the single biggest source of boost for my models.\n\n上記の式の背後にあるロジックは次のとおりです。\n次の場合にのみ、特定のクラスの単一セル画像のラベルを1に設定します。\n\nこのクラスの画像全体にラベル1が付いています\nこの特定のセルは、画像内のすべてのセルの中でこのクラスの予測が最も高いか、0.9を超えるか、この画像のこのクラスの平均予測よりも0.5標準偏差高くなっています。\nこれらのパラメーターは、LBからのフィードバックを使用して調整されました。 3回の反復->モデル->プレデズ->ラベルを実行しました。 これは私のモデルにとって唯一最大のブーストの源でした。\n\n2. Filtering our cells detected by segmentation model, but invisible to humans\n\nWhen the blue channel is very weak, sometimes the official segmentation model provided by the hosts detects a cell, even when it is nearly invisible to the human eye, and could surely be removed by manual labelers. This is a simple condition I used, which brought like 0.04 improvement on the LB ( I assume blue is the 2nd channel):\n\ncell_img[2,:,:][cell_img[2,:,:]>5] < 25 \nSuch cells were removed from the predictions\n\n3. Manual review of mitotic spindle\n\nI could not resist :) I spent a couple of evenings manually reviewing all images with mitotic spindle (label for class_11 == 1) and some high-predictions for class 0. This improved score of the mitotic spindle from 0.024 to 0.032. I can release this dataset if anyone is interested.",
    "1305461": "> I can release this dataset if anyone is interested.\n\nThat will be very interesting to compare our labels. I released the same dataset [here](https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup).\n",
    "1305029": "@narsil Congratulations  and Thanks for sharing the approach",
    "1304271": "Congratulations @narsil! We were both inspired by the data-centric approach, but you were able to make it work, this is amazing :) ",
    "1304209": "Congrats, Paweł! Thank you for sharing the details, impressive combination of cell-level and image-level predictions in your progressive pseudo-labeling part!\n\np.s. And thanks a lot for your activity in the discussions! To me, it made the competition even more awesome 😊",
    "1304103": "Congrats @narsil, @its7171 and @tivfrvqhs5. progressive pseudo-labeling approach is interesting.",
    "1304060": "Great job @narsil and team! Congratulations and thanks for your solution summary!",
    "1304037": "Congratulations @narsil and team on 5th place and thanks for sharing details approach "
  }
}