{
  "id": 238512,
  "title": "23rd Place Solution: You don't need heuristics or expensive GPU",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238512",
  "author_name": "Raman",
  "post_date": "2021-05-12T12:11:44.332000",
  "votes": 35,
  "comment_count": 17,
  "views": 0,
  "content": "<p>[<strong>UPDATE</strong>]:<br>\nFor the latest version of the solution write-up, please check <a href=\"https://github.com/SamusRam/hpa-single-cell/blob/main/README.md\" target=\"_blank\"><strong>the GitHub repo README</strong></a>. There, all the graphics should be present, but I am not maintaining external hosting links for graphics in the current post. Please, feel free to drop me a personal email via Kaggle or raise a GitHub issue in case of any questions. Thanks!</p>\n<hr>\n<p>Hello, everyone!</p>\n<p>First of all, I'd like to congratulate&nbsp;the winners! It's a privilege to learn from your solutions!&nbsp;</p>\n<p>Secondly, I'd like to thank the organizers of the competition! Thank you for your kind, attentive and responsive approach on forums!<br>\nAnd thank you, all my fellow HPA competitors! 😊Thanks to all of you it felt like an awesome Team of amazing like-minded enthusiastic colleagues. As Darek <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> has put it nicely, I wish we would meet and celebrate. Hopefully, it'd happen during some offline KaggleDays hackathon in the future 😉</p>\n<h1>Solution</h1>\n<h2>Intro</h2>\n<p>Analogously to <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/232396\" target=\"_blank\">the post</a> by Paweł <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> , my approach was <strong>very much data-centric</strong> and <strong>theoretically rigorous</strong>.</p>\n<p>The core of my solution was <strong>de-noising</strong>, including de-noising based on graph Laplacian regularization theory. De-noising itself allowed me to achieve a competitive position <strong>without compute-intensive large models</strong>, <strong>without heuristics</strong> like ranking, multiplication of predictions, etc.</p>\n<h2>De-noising</h2>\n<p><a href=\"https://drive.google.com/file/d/1RDrcZ9_boh4u6gTX6O5ObETK7TpSl3ns/view?usp=sharing\" target=\"_blank\">IMAGE</a>:<br>\n<img src=\"https://svgshare.com/i/XB5.svg\" alt=\"\"></p>\n<p>I managed to implement the computationally efficient graph signal denoising from <a href=\"https://ieeexplore.ieee.org/document/7450177\" target=\"_blank\">this paper</a>. My Python implementation can be found <a href=\"https://github.com/SamusRam/hpa-single-cell/blob/main/src/denoising/graph_denoising.py\" target=\"_blank\">here</a>. The algorithm from the paper enforces label sparsity so that each object would have a single confident label after the de-noising. Therefore, outputs of the de-noising were not used directly. I focused on the group of cells having the highest mode in the de-noised soft labels. E.g., I've added around 500 cells from the highest de-noised labels:<br>\nIMAGE:<br>\n<img src=\"https://i.imgur.com/3GWvpWn.png\" alt=\"\"></p>\n<h2>Negative label</h2>\n<p><br>\n$$ P_{neg} = \\prod_{c\\ in\\ patterns} (1 - P_{c}) $$.</p>\n<p></p>\n<p><em>Update</em>: I noticed that by mistake I did not include the independence-based estimation into my final submission. In the submission corresponding to my private LB position I estimated negative label probability as <br>\n$$ P_{neg} = \\min_{c\\ in\\ patterns} (1 - P_{c}) $$. </p>\n<p>Such a way to estimate P_{neg} was worse compared to the independence-based estimation by 0.0007 on public LB. But it turned out to give a result by 0.0003 better on the private LB, luckily.</p>\n<h2>Segmentation</h2>\n<p>I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527.</p>\n<h2>Simple model</h2>\n<p>5 folds, DenseNet121 by&nbsp;Shubin <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, referenced in <a href=\"https://www.nature.com/articles/s41592-019-0658-6\" target=\"_blank\">the Nature paper</a> and <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214925\" target=\"_blank\">on forums</a>. Thank you for your awesome work, Shubin <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> !</p>\n<p>Data-related challenged seemed to be of higher priority, I focused on those and had no time/hardware to switch from DenseNet121.&nbsp;</p>\n<p><em>For reference I checked that on ImageNet DenseNet121 is currently ranked 325th with top-1 accuracy of 75%, while state-of-the-art is 90% (<a href=\"https://paperswithcode.com/sota/image-classification-on-imagenet\" target=\"_blank\">source</a>). Other successful teams used EfficientNets, Swin transformers which achieve&nbsp;up to 87% on ImageNet.</em></p>\n<h2>Strict avoidance of heuristics</h2>\n<p>The hosts mentioned that the hidden test set was designed to have&nbsp;high single-cell variation. For this reason, I fought the temptation of heuristical re-ranking based on image-level labels, etc., even though the community reported that it helped on the public test set. Regarding the private LB placement, it seems that I've been unnecessarily cautious.</p>\n<h2>Splitting cells</h2>\n<p>I've invested quite some time into modifications of the HPA-Cell-Segmentation to split crowded&nbsp;cells, e.g. in tissues, but it didn't provide a public LB boost. The work-in-progress splitting algorithm can be found&nbsp;<a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation/tree/closer_to_orig\" target=\"_blank\">in this branch on GitHub</a></p>\n<h2>Masking cells</h2>\n<p>When predicting a single cell, I masked all surrounding cells out, so that in the high SCV scenario predictions would not be influenced by neighboring cells in the bounding box. I later realized that it might make it harder for the model to spot later phases of mitosis. Yesterday I quickly tried to create a separate binary classifier for the mitotic spindle, where I included the surrounding cells in the extended bounding box. But the quick experiment with including the mitotic spindle classifier didn't provide a score improvement.</p>\n<h2>Source code</h2>\n<p>.. is in <a href=\"https://github.com/SamusRam/hpa-single-cell\" target=\"_blank\">this GitHub repo</a>. The high-level entry point would be <code>orchestration scripts</code>. </p>",
  "messages": [
    {
      "id": 1304080,
      "postDate": "2021-05-12T12:11:44.333Z",
      "content": "<p>[<strong>UPDATE</strong>]:<br>\nFor the latest version of the solution write-up, please check <a href=\"https://github.com/SamusRam/hpa-single-cell/blob/main/README.md\" target=\"_blank\"><strong>the GitHub repo README</strong></a>. There, all the graphics should be present, but I am not maintaining external hosting links for graphics in the current post. Please, feel free to drop me a personal email via Kaggle or raise a GitHub issue in case of any questions. Thanks!</p>\n<hr>\n<p>Hello, everyone!</p>\n<p>First of all, I'd like to congratulate&nbsp;the winners! It's a privilege to learn from your solutions!&nbsp;</p>\n<p>Secondly, I'd like to thank the organizers of the competition! Thank you for your kind, attentive and responsive approach on forums!<br>\nAnd thank you, all my fellow HPA competitors! 😊Thanks to all of you it felt like an awesome Team of amazing like-minded enthusiastic colleagues. As Darek <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> has put it nicely, I wish we would meet and celebrate. Hopefully, it'd happen during some offline KaggleDays hackathon in the future 😉</p>\n<h1>Solution</h1>\n<h2>Intro</h2>\n<p>Analogously to <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/232396\" target=\"_blank\">the post</a> by Paweł <a href=\"https://www.kaggle.com/narsil\" target=\"_blank\">@narsil</a> , my approach was <strong>very much data-centric</strong> and <strong>theoretically rigorous</strong>.</p>\n<p>The core of my solution was <strong>de-noising</strong>, including de-noising based on graph Laplacian regularization theory. De-noising itself allowed me to achieve a competitive position <strong>without compute-intensive large models</strong>, <strong>without heuristics</strong> like ranking, multiplication of predictions, etc.</p>\n<h2>De-noising</h2>\n<p><a href=\"https://drive.google.com/file/d/1RDrcZ9_boh4u6gTX6O5ObETK7TpSl3ns/view?usp=sharing\" target=\"_blank\">IMAGE</a>:<br>\n<img src=\"https://svgshare.com/i/XB5.svg\" alt=\"\"></p>\n<p>I managed to implement the computationally efficient graph signal denoising from <a href=\"https://ieeexplore.ieee.org/document/7450177\" target=\"_blank\">this paper</a>. My Python implementation can be found <a href=\"https://github.com/SamusRam/hpa-single-cell/blob/main/src/denoising/graph_denoising.py\" target=\"_blank\">here</a>. The algorithm from the paper enforces label sparsity so that each object would have a single confident label after the de-noising. Therefore, outputs of the de-noising were not used directly. I focused on the group of cells having the highest mode in the de-noised soft labels. E.g., I've added around 500 cells from the highest de-noised labels:<br>\nIMAGE:<br>\n<img src=\"https://i.imgur.com/3GWvpWn.png\" alt=\"\"></p>\n<h2>Negative label</h2>\n<p><br>\n$$ P_{neg} = \\prod_{c\\ in\\ patterns} (1 - P_{c}) $$.</p>\n<p></p>\n<p><em>Update</em>: I noticed that by mistake I did not include the independence-based estimation into my final submission. In the submission corresponding to my private LB position I estimated negative label probability as <br>\n$$ P_{neg} = \\min_{c\\ in\\ patterns} (1 - P_{c}) $$. </p>\n<p>Such a way to estimate P_{neg} was worse compared to the independence-based estimation by 0.0007 on public LB. But it turned out to give a result by 0.0003 better on the private LB, luckily.</p>\n<h2>Segmentation</h2>\n<p>I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527.</p>\n<h2>Simple model</h2>\n<p>5 folds, DenseNet121 by&nbsp;Shubin <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a>, referenced in <a href=\"https://www.nature.com/articles/s41592-019-0658-6\" target=\"_blank\">the Nature paper</a> and <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214925\" target=\"_blank\">on forums</a>. Thank you for your awesome work, Shubin <a href=\"https://www.kaggle.com/bestfitting\" target=\"_blank\">@bestfitting</a> !</p>\n<p>Data-related challenged seemed to be of higher priority, I focused on those and had no time/hardware to switch from DenseNet121.&nbsp;</p>\n<p><em>For reference I checked that on ImageNet DenseNet121 is currently ranked 325th with top-1 accuracy of 75%, while state-of-the-art is 90% (<a href=\"https://paperswithcode.com/sota/image-classification-on-imagenet\" target=\"_blank\">source</a>). Other successful teams used EfficientNets, Swin transformers which achieve&nbsp;up to 87% on ImageNet.</em></p>\n<h2>Strict avoidance of heuristics</h2>\n<p>The hosts mentioned that the hidden test set was designed to have&nbsp;high single-cell variation. For this reason, I fought the temptation of heuristical re-ranking based on image-level labels, etc., even though the community reported that it helped on the public test set. Regarding the private LB placement, it seems that I've been unnecessarily cautious.</p>\n<h2>Splitting cells</h2>\n<p>I've invested quite some time into modifications of the HPA-Cell-Segmentation to split crowded&nbsp;cells, e.g. in tissues, but it didn't provide a public LB boost. The work-in-progress splitting algorithm can be found&nbsp;<a href=\"https://github.com/SamusRam/HPA-Cell-Segmentation/tree/closer_to_orig\" target=\"_blank\">in this branch on GitHub</a></p>\n<h2>Masking cells</h2>\n<p>When predicting a single cell, I masked all surrounding cells out, so that in the high SCV scenario predictions would not be influenced by neighboring cells in the bounding box. I later realized that it might make it harder for the model to spot later phases of mitosis. Yesterday I quickly tried to create a separate binary classifier for the mitotic spindle, where I included the surrounding cells in the extended bounding box. But the quick experiment with including the mitotic spindle classifier didn't provide a score improvement.</p>\n<h2>Source code</h2>\n<p>.. is in <a href=\"https://github.com/SamusRam/hpa-single-cell\" target=\"_blank\">this GitHub repo</a>. The high-level entry point would be <code>orchestration scripts</code>. </p>",
      "rawMarkdown": "[**UPDATE**]:\nFor the latest version of the solution write-up, please check [**the GitHub repo README**](https://github.com/SamusRam/hpa-single-cell/blob/main/README.md). There, all the graphics should be present, but I am not maintaining external hosting links for graphics in the current post. Please, feel free to drop me a personal email via Kaggle or raise a GitHub issue in case of any questions. Thanks!\n\n---------------------------------------------------------------------\nHello, everyone!\n\nFirst of all, I'd like to congratulate the winners! It's a privilege to learn from your solutions! \n\nSecondly, I'd like to thank the organizers of the competition! Thank you for your kind, attentive and responsive approach on forums!\nAnd thank you, all my fellow HPA competitors! 😊Thanks to all of you it felt like an awesome Team of amazing like-minded enthusiastic colleagues. As Darek @thedrcat has put it nicely, I wish we would meet and celebrate. Hopefully, it'd happen during some offline KaggleDays hackathon in the future 😉\n\n# Solution\n## Intro\nAnalogously to [the post](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/232396) by Paweł @narsil , my approach was **very much data-centric** and **theoretically rigorous**.\n \nThe core of my solution was **de-noising**, including de-noising based on graph Laplacian regularization theory. De-noising itself allowed me to achieve a competitive position **without compute-intensive large models**, **without heuristics** like ranking, multiplication of predictions, etc.\n\n## De-noising\n[IMAGE](https://drive.google.com/file/d/1RDrcZ9_boh4u6gTX6O5ObETK7TpSl3ns/view?usp=sharing):\n![](https://svgshare.com/i/XB5.svg)\n\nI managed to implement the computationally efficient graph signal denoising from [this paper](https://ieeexplore.ieee.org/document/7450177). My Python implementation can be found [here](https://github.com/SamusRam/hpa-single-cell/blob/main/src/denoising/graph_denoising.py). The algorithm from the paper enforces label sparsity so that each object would have a single confident label after the de-noising. Therefore, outputs of the de-noising were not used directly. I focused on the group of cells having the highest mode in the de-noised soft labels. E.g., I've added around 500 cells from the highest de-noised labels:\nIMAGE:\n![](https://i.imgur.com/3GWvpWn.png)\n\n## Negative label\n~~I estimated negative label probability under the assumption of labels independence:~~\n$$ P_{neg} = \\prod_{c\\ in\\ patterns} (1 - P_{c}) $$.\n\n~~Estimating P_{neg} this way instead of trying to predict the negative label boosted my public LB score from 0.517 to 0.523.~~\n\n*Update*: I noticed that by mistake I did not include the independence-based estimation into my final submission. In the submission corresponding to my private LB position I estimated negative label probability as \n$$ P_{neg} = \\min_{c\\ in\\ patterns} (1 - P_{c}) $$. \n\nSuch a way to estimate P_{neg} was worse compared to the independence-based estimation by 0.0007 on public LB. But it turned out to give a result by 0.0003 better on the private LB, luckily.\n\n## Segmentation\nI implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527.\n\n## Simple model\n \n 5 folds, DenseNet121 by Shubin @bestfitting, referenced in [the Nature paper](https://www.nature.com/articles/s41592-019-0658-6) and [on forums](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214925). Thank you for your awesome work, Shubin @bestfitting !\n\nData-related challenged seemed to be of higher priority, I focused on those and had no time/hardware to switch from DenseNet121. \n\n*For reference I checked that on ImageNet DenseNet121 is currently ranked 325th with top-1 accuracy of 75%, while state-of-the-art is 90% ([source](https://paperswithcode.com/sota/image-classification-on-imagenet)). Other successful teams used EfficientNets, Swin transformers which achieve up to 87% on ImageNet.*\n\n## Strict avoidance of heuristics\nThe hosts mentioned that the hidden test set was designed to have high single-cell variation. For this reason, I fought the temptation of heuristical re-ranking based on image-level labels, etc., even though the community reported that it helped on the public test set. Regarding the private LB placement, it seems that I've been unnecessarily cautious.\n\n## Splitting cells\nI've invested quite some time into modifications of the HPA-Cell-Segmentation to split crowded cells, e.g. in tissues, but it didn't provide a public LB boost. The work-in-progress splitting algorithm can be found [in this branch on GitHub](https://github.com/SamusRam/HPA-Cell-Segmentation/tree/closer_to_orig)\n\n## Masking cells\nWhen predicting a single cell, I masked all surrounding cells out, so that in the high SCV scenario predictions would not be influenced by neighboring cells in the bounding box. I later realized that it might make it harder for the model to spot later phases of mitosis. Yesterday I quickly tried to create a separate binary classifier for the mitotic spindle, where I included the surrounding cells in the extended bounding box. But the quick experiment with including the mitotic spindle classifier didn't provide a score improvement.\n\n## Source code\n.. is in [this GitHub repo](https://github.com/SamusRam/hpa-single-cell). The high-level entry point would be `orchestration scripts`. ",
      "votes": 35
    },
    {
      "id": 1308381,
      "postDate": "2021-05-15T06:48:30.773Z",
      "content": "<p>Great work Raman!! Very proud of your achievement mate!!</p>",
      "rawMarkdown": "Great work Raman!! Very proud of your achievement mate!!\n",
      "votes": 1,
      "replies": [
        {
          "id": 1308421,
          "postDate": "2021-05-15T07:45:40.990Z",
          "content": "<p>Thank you, Dave!! 😊</p>",
          "rawMarkdown": "Thank you, Dave!! 😊",
          "votes": 1
        }
      ]
    },
    {
      "id": 1306153,
      "postDate": "2021-05-13T16:28:47.610Z",
      "content": "<p>Jeez! This is an incredible write-up. Thank you for this incredible information. I am learning so much!</p>",
      "rawMarkdown": "Jeez! This is an incredible write-up. Thank you for this incredible information. I am learning so much!",
      "votes": 1,
      "replies": [
        {
          "id": 1306169,
          "postDate": "2021-05-13T16:40:22.133Z",
          "content": "<p>Darien, thank you so much for the kind words! 😊 I enjoyed our discussions in notebooks and on the forum enormously!  </p>",
          "rawMarkdown": "Darien, thank you so much for the kind words! 😊 I enjoyed our discussions in notebooks and on the forum enormously!  ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305769,
      "postDate": "2021-05-13T13:19:33.637Z",
      "content": "<p>Thanks! I like your denoising approach.</p>\n<blockquote>\n  <p>I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527</p>\n</blockquote>\n<p>Interesting, I've started implemeting the same post-processing then I stopped to concentrate to models training. Now I know that it was worth. Is the the same boost on private LB?</p>",
      "rawMarkdown": "Thanks! I like your denoising approach.\n\n> I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527\n\nInteresting, I've started implemeting the same post-processing then I stopped to concentrate to models training. Now I know that it was worth. Is the the same boost on private LB?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1305826,
          "postDate": "2021-05-13T13:49:14.690Z",
          "content": "<p>Thank you for the kind words, <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> ! It means a lot to me!</p>\n<blockquote>\n  <p>Is there the same boost on private LB?</p>\n</blockquote>\n<p>unfortunately, I can't say for sure as I didn't run every submission on the private test. <br>\nI happen to have a private LB score for the submission without the border cell removal and with a different small cell removal threshold (the threshold was already median-based; there were no major differences between this scored submission and the final submission). Based on the scored submission, fine-tuning of the median-based threshold and removal of the border cells improved my scores by 0.003 on the public LB and by 0.002 on the private LB. </p>",
          "rawMarkdown": "Thank you for the kind words, @mpware ! It means a lot to me!\n\n> Is there the same boost on private LB?\n\nunfortunately, I can't say for sure as I didn't run every submission on the private test. \nI happen to have a private LB score for the submission without the border cell removal and with a different small cell removal threshold (the threshold was already median-based; there were no major differences between this scored submission and the final submission). Based on the scored submission, fine-tuning of the median-based threshold and removal of the border cells improved my scores by 0.003 on the public LB and by 0.002 on the private LB. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1305026,
      "postDate": "2021-05-13T04:26:24.257Z",
      "content": "<p><a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Congratulations  and Thanks for sharing the approach</p>",
      "rawMarkdown": "@samusram Congratulations  and Thanks for sharing the approach",
      "votes": 1,
      "replies": [
        {
          "id": 1305160,
          "postDate": "2021-05-13T05:59:31.013Z",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "rawMarkdown": "Thank you, @usharengaraju !"
        }
      ]
    },
    {
      "id": 1304248,
      "postDate": "2021-05-12T14:08:24.693Z",
      "content": "<p>Congratulations Raman! I'll take some time to study your approach, looks very interesting! There were quite a few competitors from Central Europe in this challenge, maybe we can meet up indeed :) </p>",
      "rawMarkdown": "Congratulations Raman! I'll take some time to study your approach, looks very interesting! There were quite a few competitors from Central Europe in this challenge, maybe we can meet up indeed :) ",
      "votes": 1,
      "replies": [
        {
          "id": 1304494,
          "postDate": "2021-05-12T16:46:25.940Z",
          "content": "<p>Thanks for the kind words, Darek! :) That'd be awesome to meet :)</p>",
          "rawMarkdown": "Thanks for the kind words, Darek! :) That'd be awesome to meet :)"
        }
      ]
    },
    {
      "id": 1304153,
      "postDate": "2021-05-12T13:03:21.730Z",
      "content": "<p>Hi. Thanks for the post.<br>\nYour denoising approach is quite interesting.<br>\nBTW, at the last part, you wrote you did all surrounding cells out. <br>\nDoes this mean all pixel values other than the targeted cell convert to 0?<br>\nI also tried this, but LB dropped. Have you compared the LB score with or without it??</p>",
      "rawMarkdown": "Hi. Thanks for the post.\nYour denoising approach is quite interesting.\nBTW, at the last part, you wrote you did all surrounding cells out. \nDoes this mean all pixel values other than the targeted cell convert to 0?\nI also tried this, but LB dropped. Have you compared the LB score with or without it??",
      "votes": 1,
      "replies": [
        {
          "id": 1304176,
          "postDate": "2021-05-12T13:21:20.367Z",
          "content": "<blockquote>\n  <p>BTW, at the last part, you wrote you did all surrounding cells out.<br>\n  Does this mean all pixel values other than the targeted cell convert to 0?</p>\n</blockquote>\n<p>Yes, exactly.</p>\n<blockquote>\n  <p>I also tried this, but LB dropped. Have you compared the LB score with or without it??</p>\n</blockquote>\n<p>I did not, on purpose. </p>\n<p>I thought that the private test set would have significantly higher SCV compared to the public test set.</p>\n<p>So, if a prediction based on the cells with its surrounding led to a better public LB score, I'd explain it with the fact that surrounding cells supported correct predictions (because all the cells are often similar in a sample with low SCV). But for the higher SCV samples, I expected surrounding cells to actually poison the single-cell prediction. Again, it seems I was overcautious and the private test SCV is not as severe as I expected.</p>",
          "rawMarkdown": "> BTW, at the last part, you wrote you did all surrounding cells out.\nDoes this mean all pixel values other than the targeted cell convert to 0?\n\nYes, exactly.\n\n> I also tried this, but LB dropped. Have you compared the LB score with or without it??\n\nI did not, on purpose. \n\nI thought that the private test set would have significantly higher SCV compared to the public test set.\n\nSo, if a prediction based on the cells with its surrounding led to a better public LB score, I'd explain it with the fact that surrounding cells supported correct predictions (because all the cells are often similar in a sample with low SCV). But for the higher SCV samples, I expected surrounding cells to actually poison the single-cell prediction. Again, it seems I was overcautious and the private test SCV is not as severe as I expected.",
          "votes": 1
        },
        {
          "id": 1304187,
          "postDate": "2021-05-12T13:27:32.700Z",
          "content": "<p>I see. I also selected 2 final subs because of the same expectation as you; one with only cell level label model, the other with cell level label + image level label model.<br>\nAnd the result was private score higher with combined model, so private SCV might be close to public, I guess.</p>",
          "rawMarkdown": "I see. I also selected 2 final subs because of the same expectation as you; one with only cell level label model, the other with cell level label + image level label model.\nAnd the result was private score higher with combined model, so private SCV might be close to public, I guess.",
          "votes": 1
        },
        {
          "id": 1304216,
          "postDate": "2021-05-12T13:39:44.400Z",
          "content": "<p>seems like… thanks for sharing those additional details!</p>",
          "rawMarkdown": "seems like... thanks for sharing those additional details!"
        }
      ]
    },
    {
      "id": 1312398,
      "postDate": "2021-05-18T03:16:41.533Z",
      "content": "<p><em>Post-deadline experiments</em>:<br>\nI checked that the label de-noising could have been used to remove labels as well. As described in the write-up, during the competition I used de-noising only to get additional mitotic spindle labels.</p>\n<p>I computed candidates for label removal as follows (without strong candidates for removal though, please see the normalized de-noised labels in the Fig.):</p>\n<pre><code># Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\npositive_init_label_threshold = 0.95\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\nremoved_bool_idx = ((Y_init[:, class_i] &gt;= positive_init_label_threshold) &amp; \n              (Y_opt[:, class_i] &lt; positive_denoised_threshold))\n</code></pre>\n<p>The computed post-deadline candidates for removal are depicted in the following figure:<br>\n<img src=\"https://i.imgur.com/ZWlROr8.png\" alt=\"\"></p>\n<p>I believe that the outcomes should be viewed as candidates for data cleaning which help an annotator to focus on the most important samples for potential label correction. <br>\nThe outcomes of label de-noising depend on:</p>\n<ul>\n<li>the quality of appearance embedding (how well the extracted embeddings encode cell appearance)</li>\n<li>noise level of the initial labels (how well the labels cover various mitotic spindle patterns)</li>\n</ul>\n<p><strong>Credits, once again</strong> <br>\nThe label de-noising procedure, which I took the effort to implement, is from the paper <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\">Learning from Weak and Noisy Labels for Semantic Segmentation</a></p>",
      "rawMarkdown": "*Post-deadline experiments*:\nI checked that the label de-noising could have been used to remove labels as well. As described in the write-up, during the competition I used de-noising only to get additional mitotic spindle labels.\n\nI computed candidates for label removal as follows (without strong candidates for removal though, please see the normalized de-noised labels in the Fig.):\n```\n\n# Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\npositive_init_label_threshold = 0.95\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\nremoved_bool_idx = ((Y_init[:, class_i] >= positive_init_label_threshold) & \n              (Y_opt[:, class_i] < positive_denoised_threshold))\n```\n\n\nThe computed post-deadline candidates for removal are depicted in the following figure:\n![](https://i.imgur.com/ZWlROr8.png)\n\nI believe that the outcomes should be viewed as candidates for data cleaning which help an annotator to focus on the most important samples for potential label correction. \nThe outcomes of label de-noising depend on:\n- the quality of appearance embedding (how well the extracted embeddings encode cell appearance)\n- noise level of the initial labels (how well the labels cover various mitotic spindle patterns)\n\n**Credits, once again** \nThe label de-noising procedure, which I took the effort to implement, is from the paper [Learning from Weak and Noisy Labels for Semantic Segmentation](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y)",
      "replies": [
        {
          "id": 1313519,
          "postDate": "2021-05-18T15:40:45.620Z",
          "content": "<p>The visualization above suggests that removing labels via de-noising is not effective for the problem at hand, as false-positive mitotic spindle labels do not seem to be the main issue. False-negative mitotic spindle cells seem to be a more significant data issue, based on the cells observed among results of de-noising via</p>\n<pre><code># Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\nnegative_init_label_threshold = 0.4\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\naddition_bool_idx = ((Y_init[:, class_i]  &lt; negative_init_label_threshold) &amp; \n              (Y_opt[:, class_i] &gt;= positive_denoised_threshold))\n</code></pre>\n<p><img src=\"https://i.imgur.com/AxKJ2t2.png\" alt=\"\"></p>",
          "rawMarkdown": "The visualization above suggests that removing labels via de-noising is not effective for the problem at hand, as false-positive mitotic spindle labels do not seem to be the main issue. False-negative mitotic spindle cells seem to be a more significant data issue, based on the cells observed among results of de-noising via\n\n```\n\n# Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\nnegative_init_label_threshold = 0.4\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\naddition_bool_idx = ((Y_init[:, class_i]  < negative_init_label_threshold) & \n              (Y_opt[:, class_i] >= positive_denoised_threshold))\n```\n\n![](https://i.imgur.com/AxKJ2t2.png)"
        }
      ]
    },
    {
      "id": 1304140,
      "postDate": "2021-05-12T12:55:03.017Z",
      "content": "<p>Thanks for sharing detailed post.</p>",
      "rawMarkdown": "Thanks for sharing detailed post.",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1308381,
      "author_name": "Dave Rattan",
      "author_url": "",
      "post_date": "2021-05-15T06:48:30.773000",
      "content": "<p>Great work Raman!! Very proud of your achievement mate!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1308421,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-15T07:45:40.990000",
          "content": "<p>Thank you, Dave!! 😊</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1306153,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-05-13T16:28:47.610000",
      "content": "<p>Jeez! This is an incredible write-up. Thank you for this incredible information. I am learning so much!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1306169,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-13T16:40:22.133000",
          "content": "<p>Darien, thank you so much for the kind words! 😊 I enjoyed our discussions in notebooks and on the forum enormously!  </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305769,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-05-13T13:19:33.637000",
      "content": "<p>Thanks! I like your denoising approach.</p>\n<blockquote>\n  <p>I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527</p>\n</blockquote>\n<p>Interesting, I've started implemeting the same post-processing then I stopped to concentrate to models training. Now I know that it was worth. Is the the same boost on private LB?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305826,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-13T13:49:14.690000",
          "content": "<p>Thank you for the kind words, <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> ! It means a lot to me!</p>\n<blockquote>\n  <p>Is there the same boost on private LB?</p>\n</blockquote>\n<p>unfortunately, I can't say for sure as I didn't run every submission on the private test. <br>\nI happen to have a private LB score for the submission without the border cell removal and with a different small cell removal threshold (the threshold was already median-based; there were no major differences between this scored submission and the final submission). Based on the scored submission, fine-tuning of the median-based threshold and removal of the border cells improved my scores by 0.003 on the public LB and by 0.002 on the private LB. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1305026,
      "author_name": "Tensor Girl",
      "author_url": "",
      "post_date": "2021-05-13T04:26:24.257000",
      "content": "<p><a href=\"https://www.kaggle.com/samusram\" target=\"_blank\">@samusram</a> Congratulations  and Thanks for sharing the approach</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1305160,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-13T05:59:31.013000",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/usharengaraju\" target=\"_blank\">@usharengaraju</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304248,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-05-12T14:08:24.693000",
      "content": "<p>Congratulations Raman! I'll take some time to study your approach, looks very interesting! There were quite a few competitors from Central Europe in this challenge, maybe we can meet up indeed :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304494,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-12T16:46:25.940000",
          "content": "<p>Thanks for the kind words, Darek! :) That'd be awesome to meet :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304153,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-05-12T13:03:21.730000",
      "content": "<p>Hi. Thanks for the post.<br>\nYour denoising approach is quite interesting.<br>\nBTW, at the last part, you wrote you did all surrounding cells out. <br>\nDoes this mean all pixel values other than the targeted cell convert to 0?<br>\nI also tried this, but LB dropped. Have you compared the LB score with or without it??</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1304176,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-12T13:21:20.367000",
          "content": "<blockquote>\n  <p>BTW, at the last part, you wrote you did all surrounding cells out.<br>\n  Does this mean all pixel values other than the targeted cell convert to 0?</p>\n</blockquote>\n<p>Yes, exactly.</p>\n<blockquote>\n  <p>I also tried this, but LB dropped. Have you compared the LB score with or without it??</p>\n</blockquote>\n<p>I did not, on purpose. </p>\n<p>I thought that the private test set would have significantly higher SCV compared to the public test set.</p>\n<p>So, if a prediction based on the cells with its surrounding led to a better public LB score, I'd explain it with the fact that surrounding cells supported correct predictions (because all the cells are often similar in a sample with low SCV). But for the higher SCV samples, I expected surrounding cells to actually poison the single-cell prediction. Again, it seems I was overcautious and the private test SCV is not as severe as I expected.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1304187,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-05-12T13:27:32.700000",
          "content": "<p>I see. I also selected 2 final subs because of the same expectation as you; one with only cell level label model, the other with cell level label + image level label model.<br>\nAnd the result was private score higher with combined model, so private SCV might be close to public, I guess.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1304216,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-12T13:39:44.400000",
          "content": "<p>seems like… thanks for sharing those additional details!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1312398,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-05-18T03:16:41.533000",
      "content": "<p><em>Post-deadline experiments</em>:<br>\nI checked that the label de-noising could have been used to remove labels as well. As described in the write-up, during the competition I used de-noising only to get additional mitotic spindle labels.</p>\n<p>I computed candidates for label removal as follows (without strong candidates for removal though, please see the normalized de-noised labels in the Fig.):</p>\n<pre><code># Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\npositive_init_label_threshold = 0.95\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\nremoved_bool_idx = ((Y_init[:, class_i] &gt;= positive_init_label_threshold) &amp; \n              (Y_opt[:, class_i] &lt; positive_denoised_threshold))\n</code></pre>\n<p>The computed post-deadline candidates for removal are depicted in the following figure:<br>\n<img src=\"https://i.imgur.com/ZWlROr8.png\" alt=\"\"></p>\n<p>I believe that the outcomes should be viewed as candidates for data cleaning which help an annotator to focus on the most important samples for potential label correction. <br>\nThe outcomes of label de-noising depend on:</p>\n<ul>\n<li>the quality of appearance embedding (how well the extracted embeddings encode cell appearance)</li>\n<li>noise level of the initial labels (how well the labels cover various mitotic spindle patterns)</li>\n</ul>\n<p><strong>Credits, once again</strong> <br>\nThe label de-noising procedure, which I took the effort to implement, is from the paper <a href=\"https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y\" target=\"_blank\">Learning from Weak and Noisy Labels for Semantic Segmentation</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1313519,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-18T15:40:45.620000",
          "content": "<p>The visualization above suggests that removing labels via de-noising is not effective for the problem at hand, as false-positive mitotic spindle labels do not seem to be the main issue. False-negative mitotic spindle cells seem to be a more significant data issue, based on the cells observed among results of de-noising via</p>\n<pre><code># Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\nnegative_init_label_threshold = 0.4\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\naddition_bool_idx = ((Y_init[:, class_i]  &lt; negative_init_label_threshold) &amp; \n              (Y_opt[:, class_i] &gt;= positive_denoised_threshold))\n</code></pre>\n<p><img src=\"https://i.imgur.com/AxKJ2t2.png\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1304140,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-12T12:55:03.017000",
      "content": "<p>Thanks for sharing detailed post.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304080": "[**UPDATE**]:\nFor the latest version of the solution write-up, please check [**the GitHub repo README**](https://github.com/SamusRam/hpa-single-cell/blob/main/README.md). There, all the graphics should be present, but I am not maintaining external hosting links for graphics in the current post. Please, feel free to drop me a personal email via Kaggle or raise a GitHub issue in case of any questions. Thanks!\n\n---------------------------------------------------------------------\nHello, everyone!\n\nFirst of all, I'd like to congratulate the winners! It's a privilege to learn from your solutions! \n\nSecondly, I'd like to thank the organizers of the competition! Thank you for your kind, attentive and responsive approach on forums!\nAnd thank you, all my fellow HPA competitors! 😊Thanks to all of you it felt like an awesome Team of amazing like-minded enthusiastic colleagues. As Darek @thedrcat has put it nicely, I wish we would meet and celebrate. Hopefully, it'd happen during some offline KaggleDays hackathon in the future 😉\n\n# Solution\n## Intro\nAnalogously to [the post](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/232396) by Paweł @narsil , my approach was **very much data-centric** and **theoretically rigorous**.\n \nThe core of my solution was **de-noising**, including de-noising based on graph Laplacian regularization theory. De-noising itself allowed me to achieve a competitive position **without compute-intensive large models**, **without heuristics** like ranking, multiplication of predictions, etc.\n\n## De-noising\n[IMAGE](https://drive.google.com/file/d/1RDrcZ9_boh4u6gTX6O5ObETK7TpSl3ns/view?usp=sharing):\n![](https://svgshare.com/i/XB5.svg)\n\nI managed to implement the computationally efficient graph signal denoising from [this paper](https://ieeexplore.ieee.org/document/7450177). My Python implementation can be found [here](https://github.com/SamusRam/hpa-single-cell/blob/main/src/denoising/graph_denoising.py). The algorithm from the paper enforces label sparsity so that each object would have a single confident label after the de-noising. Therefore, outputs of the de-noising were not used directly. I focused on the group of cells having the highest mode in the de-noised soft labels. E.g., I've added around 500 cells from the highest de-noised labels:\nIMAGE:\n![](https://i.imgur.com/3GWvpWn.png)\n\n## Negative label\n~~I estimated negative label probability under the assumption of labels independence:~~\n$$ P_{neg} = \\prod_{c\\ in\\ patterns} (1 - P_{c}) $$.\n\n~~Estimating P_{neg} this way instead of trying to predict the negative label boosted my public LB score from 0.517 to 0.523.~~\n\n*Update*: I noticed that by mistake I did not include the independence-based estimation into my final submission. In the submission corresponding to my private LB position I estimated negative label probability as \n$$ P_{neg} = \\min_{c\\ in\\ patterns} (1 - P_{c}) $$. \n\nSuch a way to estimate P_{neg} was worse compared to the independence-based estimation by 0.0007 on public LB. But it turned out to give a result by 0.0003 better on the private LB, luckily.\n\n## Segmentation\nI implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527.\n\n## Simple model\n \n 5 folds, DenseNet121 by Shubin @bestfitting, referenced in [the Nature paper](https://www.nature.com/articles/s41592-019-0658-6) and [on forums](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/214925). Thank you for your awesome work, Shubin @bestfitting !\n\nData-related challenged seemed to be of higher priority, I focused on those and had no time/hardware to switch from DenseNet121. \n\n*For reference I checked that on ImageNet DenseNet121 is currently ranked 325th with top-1 accuracy of 75%, while state-of-the-art is 90% ([source](https://paperswithcode.com/sota/image-classification-on-imagenet)). Other successful teams used EfficientNets, Swin transformers which achieve up to 87% on ImageNet.*\n\n## Strict avoidance of heuristics\nThe hosts mentioned that the hidden test set was designed to have high single-cell variation. For this reason, I fought the temptation of heuristical re-ranking based on image-level labels, etc., even though the community reported that it helped on the public test set. Regarding the private LB placement, it seems that I've been unnecessarily cautious.\n\n## Splitting cells\nI've invested quite some time into modifications of the HPA-Cell-Segmentation to split crowded cells, e.g. in tissues, but it didn't provide a public LB boost. The work-in-progress splitting algorithm can be found [in this branch on GitHub](https://github.com/SamusRam/HPA-Cell-Segmentation/tree/closer_to_orig)\n\n## Masking cells\nWhen predicting a single cell, I masked all surrounding cells out, so that in the high SCV scenario predictions would not be influenced by neighboring cells in the bounding box. I later realized that it might make it harder for the model to spot later phases of mitosis. Yesterday I quickly tried to create a separate binary classifier for the mitotic spindle, where I included the surrounding cells in the extended bounding box. But the quick experiment with including the mitotic spindle classifier didn't provide a score improvement.\n\n## Source code\n.. is in [this GitHub repo](https://github.com/SamusRam/hpa-single-cell). The high-level entry point would be `orchestration scripts`. ",
    "1308381": "Great work Raman!! Very proud of your achievement mate!!\n",
    "1306153": "Jeez! This is an incredible write-up. Thank you for this incredible information. I am learning so much!",
    "1305769": "Thanks! I like your denoising approach.\n\n> I implemented the removal of border cells. I also modified the removal of small nuclei, doing it based on medium nucleus size instead of the hardcoded threshold. These modifications boosted my public LB score from 0.523 to 0.527\n\nInteresting, I've started implemeting the same post-processing then I stopped to concentrate to models training. Now I know that it was worth. Is the the same boost on private LB?\n",
    "1305026": "@samusram Congratulations  and Thanks for sharing the approach",
    "1304248": "Congratulations Raman! I'll take some time to study your approach, looks very interesting! There were quite a few competitors from Central Europe in this challenge, maybe we can meet up indeed :) ",
    "1304153": "Hi. Thanks for the post.\nYour denoising approach is quite interesting.\nBTW, at the last part, you wrote you did all surrounding cells out. \nDoes this mean all pixel values other than the targeted cell convert to 0?\nI also tried this, but LB dropped. Have you compared the LB score with or without it??",
    "1312398": "*Post-deadline experiments*:\nI checked that the label de-noising could have been used to remove labels as well. As described in the write-up, during the competition I used de-noising only to get additional mitotic spindle labels.\n\nI computed candidates for label removal as follows (without strong candidates for removal though, please see the normalized de-noised labels in the Fig.):\n```\n\n# Y_init contains labels before application of de-noising\n# Y_opt contains labels after de-noising\n\nclass_i = class_names.index('Mitotic spindle')\npositive_init_label_threshold = 0.95\npositive_denoised_threshold = Y_opt[:, class_i].max()*0.92\n\nremoved_bool_idx = ((Y_init[:, class_i] >= positive_init_label_threshold) & \n              (Y_opt[:, class_i] < positive_denoised_threshold))\n```\n\n\nThe computed post-deadline candidates for removal are depicted in the following figure:\n![](https://i.imgur.com/ZWlROr8.png)\n\nI believe that the outcomes should be viewed as candidates for data cleaning which help an annotator to focus on the most important samples for potential label correction. \nThe outcomes of label de-noising depend on:\n- the quality of appearance embedding (how well the extracted embeddings encode cell appearance)\n- noise level of the initial labels (how well the labels cover various mitotic spindle patterns)\n\n**Credits, once again** \nThe label de-noising procedure, which I took the effort to implement, is from the paper [Learning from Weak and Noisy Labels for Semantic Segmentation](https://qmro.qmul.ac.uk/xmlui/bitstream/handle/123456789/12661/imparsing_final.pdf?sequence=1&amp;isAllowed=y)",
    "1304140": "Thanks for sharing detailed post."
  }
}