{
  "id": 238910,
  "title": "3rd place solution - ZFTurbo part",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/238910",
  "author_name": "ZFTurbo",
  "post_date": "2021-05-13T22:03:14.370000",
  "votes": 28,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Before further reading it's better to read solution of my teammates <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> first:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238862\" target=\"_blank\">MPWARE part</a></li>\n<li><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238898\" target=\"_blank\">Dieter part</a></li>\n</ul>\n<h2>Introduction</h2>\n<p>My initial solution was made on cell level only. E.g. I extracted all cells from all large images (including all available external data) and then used them independently. We merge with MPWARE early, and he used image level models. So I didn’t start to create my own pipeline for image level and concentrate on cell level only. We investigated early that ensembling of our approaches gives good boost. ~0.550 public.</p>\n<h2>ZFTurbo Part 1. Cell level models.</h2>\n<p>The training of cell level models was pretty straight-forward. I used image level labels for each independent cell. I used KFold split which were created using CellLine variable for external data. Actually training on all data or only external data gave me similar results.</p>\n<p>Models were different types of EffNet: EffNetB0, EffNetB3 and EffNetB5. The quality for all of them was similar. In first ensembles we usually took smallest ones. I used sigmoid on final layer with BCE loss and later switch to Focal Loss which improved score a little. I also used soft labels with (0.01*num_label)s koeff. I added large Dropout (0.5) to prevent overfitting.</p>\n<p>For validation I used several metrics: Avg AUC per class, Avg Accuracy per class and LogLoss. I mostly used AUC for early stopping. And ReduceLROnPlateu for the AUC metric as well. It was actually pretty hard to find where to stop. After some point model tends to overfit a little.<br>\nI trained on 6 channel images [red, green, blue, yellow, mask, nuclei]. I used many different augmentations, including random crops, rotations etc. On inference I used TTA2 – original and mirror image.</p>\n<p>My simple cell level models had typical score from 0.460 on public LB up to 0.480 (around 0.465 on private). While this approach worked I had no real improvements after some point. It’s understandable because of weak labels. For some classes there was only small number of cells which actually of given class, while others are from other classes etc. Model start to give large probability for incorrect cells e.g. increasing number of false positives. I tried to train on single class images but it didn’t give additional boost.</p>\n<h2>ZFTurbo Part 2. OOF models.</h2>\n<p>I created OOF predictions for all cells on all images for some of my models.  Eric (MPWARE) also created OOF for his models. We mixed OOF predictions together in the same way we ensemble our models on kaggle LB. So in some way we created mark up which is closer to reality e.g. “decrease” weakness of labels.<br>\nI had some problems with creating OOF because we extract cells with slightly different algorithms. I wasn’t always able to find the same cell in Eric predictions. So only around 90% of cells were included in OOF with some small noise.<br>\nSo I adapt my models from previous chapter and used OOF predictions as new target. I used mean square error loss function here and tried to minimize it. </p>\n<p>These models gives around 0.528 on public LB, and after competition ends I find out that private is very similar to public score for these models. I also trained such model using green channel only.<br>\nThese two models, trained on OOF, were included in our final ensemble instead of standard classification models. </p>\n<h2>ZFTurbo Part 3. Single class models.</h2>\n<p>At some point we checked which classes had low score for our solution. And as we expected class 11 was the worst. This class has low amount of images available. So I spend one evening on the following: I created small labeler tool. It shows you single cell from full image and you must click on one button if cell of target class or other button otherwise (see screenshot). It stores your answer in cache, so you can continue from last point.</p>\n<p><img src=\"//soloro.ru/images/2021/hpa-kaggle.png\" alt=\"img\"></p>\n<p>I’m not biologist so I just googled “Mitotic spindle” and check images how it looks like. And actually it was pretty simple to find these cells on images. Also I probably found out why our models were bad for this class. The main reason is that Mitotic spindle cells are rare on images. It can have many cells and only single class 11. So when I trained my cell level models I show them mostly incorrect cells in training process. Using labeler tool I easily made good labels for all images in 3-4 hours.<br>\nAfter all images were processed I just create single CSV with markup for class 11. And you can find it in this <a href=\"https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup/\" target=\"_blank\">open dataset</a>.<br>\nSo I trained binary classifier model on new labels and our score increased on public +0.008 (and as I find out recently it probably even more around 0.014 on private).<br>\nSo actually in my opinion handlabeling is key to win in this competition. ) But it was hard to do mark up for other classes because of large number of images and not so obvious differences in cells between some classes. Also I was too lazy to do additional markup. So class 11 was the only one where we made labels.</p>\n<p>I also created some other single models using default training labels, but they mostly wasn’t very good comparing to multiclass models, so we didn’t use them.</p>",
  "messages": [
    {
      "id": 1306568,
      "postDate": "2021-05-13T22:03:14.370Z",
      "content": "<p>Before further reading it's better to read solution of my teammates <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> first:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238862\" target=\"_blank\">MPWARE part</a></li>\n<li><a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238898\" target=\"_blank\">Dieter part</a></li>\n</ul>\n<h2>Introduction</h2>\n<p>My initial solution was made on cell level only. E.g. I extracted all cells from all large images (including all available external data) and then used them independently. We merge with MPWARE early, and he used image level models. So I didn’t start to create my own pipeline for image level and concentrate on cell level only. We investigated early that ensembling of our approaches gives good boost. ~0.550 public.</p>\n<h2>ZFTurbo Part 1. Cell level models.</h2>\n<p>The training of cell level models was pretty straight-forward. I used image level labels for each independent cell. I used KFold split which were created using CellLine variable for external data. Actually training on all data or only external data gave me similar results.</p>\n<p>Models were different types of EffNet: EffNetB0, EffNetB3 and EffNetB5. The quality for all of them was similar. In first ensembles we usually took smallest ones. I used sigmoid on final layer with BCE loss and later switch to Focal Loss which improved score a little. I also used soft labels with (0.01*num_label)s koeff. I added large Dropout (0.5) to prevent overfitting.</p>\n<p>For validation I used several metrics: Avg AUC per class, Avg Accuracy per class and LogLoss. I mostly used AUC for early stopping. And ReduceLROnPlateu for the AUC metric as well. It was actually pretty hard to find where to stop. After some point model tends to overfit a little.<br>\nI trained on 6 channel images [red, green, blue, yellow, mask, nuclei]. I used many different augmentations, including random crops, rotations etc. On inference I used TTA2 – original and mirror image.</p>\n<p>My simple cell level models had typical score from 0.460 on public LB up to 0.480 (around 0.465 on private). While this approach worked I had no real improvements after some point. It’s understandable because of weak labels. For some classes there was only small number of cells which actually of given class, while others are from other classes etc. Model start to give large probability for incorrect cells e.g. increasing number of false positives. I tried to train on single class images but it didn’t give additional boost.</p>\n<h2>ZFTurbo Part 2. OOF models.</h2>\n<p>I created OOF predictions for all cells on all images for some of my models.  Eric (MPWARE) also created OOF for his models. We mixed OOF predictions together in the same way we ensemble our models on kaggle LB. So in some way we created mark up which is closer to reality e.g. “decrease” weakness of labels.<br>\nI had some problems with creating OOF because we extract cells with slightly different algorithms. I wasn’t always able to find the same cell in Eric predictions. So only around 90% of cells were included in OOF with some small noise.<br>\nSo I adapt my models from previous chapter and used OOF predictions as new target. I used mean square error loss function here and tried to minimize it. </p>\n<p>These models gives around 0.528 on public LB, and after competition ends I find out that private is very similar to public score for these models. I also trained such model using green channel only.<br>\nThese two models, trained on OOF, were included in our final ensemble instead of standard classification models. </p>\n<h2>ZFTurbo Part 3. Single class models.</h2>\n<p>At some point we checked which classes had low score for our solution. And as we expected class 11 was the worst. This class has low amount of images available. So I spend one evening on the following: I created small labeler tool. It shows you single cell from full image and you must click on one button if cell of target class or other button otherwise (see screenshot). It stores your answer in cache, so you can continue from last point.</p>\n<p><img src=\"//soloro.ru/images/2021/hpa-kaggle.png\" alt=\"img\"></p>\n<p>I’m not biologist so I just googled “Mitotic spindle” and check images how it looks like. And actually it was pretty simple to find these cells on images. Also I probably found out why our models were bad for this class. The main reason is that Mitotic spindle cells are rare on images. It can have many cells and only single class 11. So when I trained my cell level models I show them mostly incorrect cells in training process. Using labeler tool I easily made good labels for all images in 3-4 hours.<br>\nAfter all images were processed I just create single CSV with markup for class 11. And you can find it in this <a href=\"https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup/\" target=\"_blank\">open dataset</a>.<br>\nSo I trained binary classifier model on new labels and our score increased on public +0.008 (and as I find out recently it probably even more around 0.014 on private).<br>\nSo actually in my opinion handlabeling is key to win in this competition. ) But it was hard to do mark up for other classes because of large number of images and not so obvious differences in cells between some classes. Also I was too lazy to do additional markup. So class 11 was the only one where we made labels.</p>\n<p>I also created some other single models using default training labels, but they mostly wasn’t very good comparing to multiclass models, so we didn’t use them.</p>",
      "rawMarkdown": "Before further reading it's better to read solution of my teammates @mpware and @christofhenkel first:\n* [MPWARE part](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238862)\n* [Dieter part](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238898)\n\n## Introduction\n\nMy initial solution was made on cell level only. E.g. I extracted all cells from all large images (including all available external data) and then used them independently. We merge with MPWARE early, and he used image level models. So I didn’t start to create my own pipeline for image level and concentrate on cell level only. We investigated early that ensembling of our approaches gives good boost. ~0.550 public.\n\n## ZFTurbo Part 1. Cell level models.\n\nThe training of cell level models was pretty straight-forward. I used image level labels for each independent cell. I used KFold split which were created using CellLine variable for external data. Actually training on all data or only external data gave me similar results.\n\nModels were different types of EffNet: EffNetB0, EffNetB3 and EffNetB5. The quality for all of them was similar. In first ensembles we usually took smallest ones. I used sigmoid on final layer with BCE loss and later switch to Focal Loss which improved score a little. I also used soft labels with (0.01*num_label)s koeff. I added large Dropout (0.5) to prevent overfitting.\n\nFor validation I used several metrics: Avg AUC per class, Avg Accuracy per class and LogLoss. I mostly used AUC for early stopping. And ReduceLROnPlateu for the AUC metric as well. It was actually pretty hard to find where to stop. After some point model tends to overfit a little.\nI trained on 6 channel images [red, green, blue, yellow, mask, nuclei]. I used many different augmentations, including random crops, rotations etc. On inference I used TTA2 – original and mirror image.\n\nMy simple cell level models had typical score from 0.460 on public LB up to 0.480 (around 0.465 on private). While this approach worked I had no real improvements after some point. It’s understandable because of weak labels. For some classes there was only small number of cells which actually of given class, while others are from other classes etc. Model start to give large probability for incorrect cells e.g. increasing number of false positives. I tried to train on single class images but it didn’t give additional boost.\n\n## ZFTurbo Part 2. OOF models.\n\nI created OOF predictions for all cells on all images for some of my models.  Eric (MPWARE) also created OOF for his models. We mixed OOF predictions together in the same way we ensemble our models on kaggle LB. So in some way we created mark up which is closer to reality e.g. “decrease” weakness of labels.\nI had some problems with creating OOF because we extract cells with slightly different algorithms. I wasn’t always able to find the same cell in Eric predictions. So only around 90% of cells were included in OOF with some small noise.\nSo I adapt my models from previous chapter and used OOF predictions as new target. I used mean square error loss function here and tried to minimize it. \n\nThese models gives around 0.528 on public LB, and after competition ends I find out that private is very similar to public score for these models. I also trained such model using green channel only.\nThese two models, trained on OOF, were included in our final ensemble instead of standard classification models. \n\n## ZFTurbo Part 3. Single class models.\n\nAt some point we checked which classes had low score for our solution. And as we expected class 11 was the worst. This class has low amount of images available. So I spend one evening on the following: I created small labeler tool. It shows you single cell from full image and you must click on one button if cell of target class or other button otherwise (see screenshot). It stores your answer in cache, so you can continue from last point.\n\n![img](//soloro.ru/images/2021/hpa-kaggle.png)\n\nI’m not biologist so I just googled “Mitotic spindle” and check images how it looks like. And actually it was pretty simple to find these cells on images. Also I probably found out why our models were bad for this class. The main reason is that Mitotic spindle cells are rare on images. It can have many cells and only single class 11. So when I trained my cell level models I show them mostly incorrect cells in training process. Using labeler tool I easily made good labels for all images in 3-4 hours.\nAfter all images were processed I just create single CSV with markup for class 11. And you can find it in this [open dataset](https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup/).\nSo I trained binary classifier model on new labels and our score increased on public +0.008 (and as I find out recently it probably even more around 0.014 on private).\nSo actually in my opinion handlabeling is key to win in this competition. ) But it was hard to do mark up for other classes because of large number of images and not so obvious differences in cells between some classes. Also I was too lazy to do additional markup. So class 11 was the only one where we made labels.\n\nI also created some other single models using default training labels, but they mostly wasn’t very good comparing to multiclass models, so we didn’t use them.\n",
      "votes": 28
    },
    {
      "id": 1307011,
      "postDate": "2021-05-14T07:47:46.413Z",
      "content": "<p>I released the dataset here <a href=\"https://www.kaggle.com/narsil/hpa2021-class11-manuallabels\" target=\"_blank\">https://www.kaggle.com/narsil/hpa2021-class11-manuallabels</a></p>",
      "rawMarkdown": "I released the dataset here https://www.kaggle.com/narsil/hpa2021-class11-manuallabels",
      "votes": 1
    },
    {
      "id": 1306965,
      "postDate": "2021-05-14T07:09:26.707Z",
      "content": "<p>This is a snapshot if images in my dataset<br>\n<img src=\"https://i.ibb.co/g6mfxZ5/mitotic-spindle.png\" alt=\"mitotic_spindle\"></p>",
      "rawMarkdown": "This is a snapshot if images in my dataset\n![mitotic_spindle](https://i.ibb.co/g6mfxZ5/mitotic-spindle.png)",
      "votes": 2,
      "replies": [
        {
          "id": 1306978,
          "postDate": "2021-05-14T07:21:23.690Z",
          "content": "<p>Looks very similar to what I marked up as positive for class 11 )</p>",
          "rawMarkdown": "Looks very similar to what I marked up as positive for class 11 )",
          "votes": 2
        }
      ]
    },
    {
      "id": 1306895,
      "postDate": "2021-05-14T06:11:12.417Z",
      "content": "<p>Congrats on the very impressive result <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> !</p>\n<p>Regarding manual re-labeling of class 11, I had exactly the same intuition and followed a very similar procedure. I am working to make the dataset public.</p>",
      "rawMarkdown": "Congrats on the very impressive result @zfturbo !\n\nRegarding manual re-labeling of class 11, I had exactly the same intuition and followed a very similar procedure. I am working to make the dataset public.",
      "votes": 2
    },
    {
      "id": 1306961,
      "postDate": "2021-05-14T07:04:40.183Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1307360,
      "postDate": "2021-05-14T12:01:49.597Z",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "rawMarkdown": "Congrats! Thank you for sharing."
    }
  ],
  "comments": [
    {
      "id": 1307011,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2021-05-14T07:47:46.413000",
      "content": "<p>I released the dataset here <a href=\"https://www.kaggle.com/narsil/hpa2021-class11-manuallabels\" target=\"_blank\">https://www.kaggle.com/narsil/hpa2021-class11-manuallabels</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1306965,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2021-05-14T07:09:26.707000",
      "content": "<p>This is a snapshot if images in my dataset<br>\n<img src=\"https://i.ibb.co/g6mfxZ5/mitotic-spindle.png\" alt=\"mitotic_spindle\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 1306978,
          "author_name": "ZFTurbo",
          "author_url": "",
          "post_date": "2021-05-14T07:21:23.690000",
          "content": "<p>Looks very similar to what I marked up as positive for class 11 )</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1306895,
      "author_name": "narsil (jobs-in-data.com)",
      "author_url": "",
      "post_date": "2021-05-14T06:11:12.417000",
      "content": "<p>Congrats on the very impressive result <a href=\"https://www.kaggle.com/zfturbo\" target=\"_blank\">@zfturbo</a> !</p>\n<p>Regarding manual re-labeling of class 11, I had exactly the same intuition and followed a very similar procedure. I am working to make the dataset public.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1306961,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-14T07:04:40.183000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1307360,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-05-14T12:01:49.597000",
      "content": "<p>Congrats! Thank you for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1306568": "Before further reading it's better to read solution of my teammates @mpware and @christofhenkel first:\n* [MPWARE part](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238862)\n* [Dieter part](https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/238898)\n\n## Introduction\n\nMy initial solution was made on cell level only. E.g. I extracted all cells from all large images (including all available external data) and then used them independently. We merge with MPWARE early, and he used image level models. So I didn’t start to create my own pipeline for image level and concentrate on cell level only. We investigated early that ensembling of our approaches gives good boost. ~0.550 public.\n\n## ZFTurbo Part 1. Cell level models.\n\nThe training of cell level models was pretty straight-forward. I used image level labels for each independent cell. I used KFold split which were created using CellLine variable for external data. Actually training on all data or only external data gave me similar results.\n\nModels were different types of EffNet: EffNetB0, EffNetB3 and EffNetB5. The quality for all of them was similar. In first ensembles we usually took smallest ones. I used sigmoid on final layer with BCE loss and later switch to Focal Loss which improved score a little. I also used soft labels with (0.01*num_label)s koeff. I added large Dropout (0.5) to prevent overfitting.\n\nFor validation I used several metrics: Avg AUC per class, Avg Accuracy per class and LogLoss. I mostly used AUC for early stopping. And ReduceLROnPlateu for the AUC metric as well. It was actually pretty hard to find where to stop. After some point model tends to overfit a little.\nI trained on 6 channel images [red, green, blue, yellow, mask, nuclei]. I used many different augmentations, including random crops, rotations etc. On inference I used TTA2 – original and mirror image.\n\nMy simple cell level models had typical score from 0.460 on public LB up to 0.480 (around 0.465 on private). While this approach worked I had no real improvements after some point. It’s understandable because of weak labels. For some classes there was only small number of cells which actually of given class, while others are from other classes etc. Model start to give large probability for incorrect cells e.g. increasing number of false positives. I tried to train on single class images but it didn’t give additional boost.\n\n## ZFTurbo Part 2. OOF models.\n\nI created OOF predictions for all cells on all images for some of my models.  Eric (MPWARE) also created OOF for his models. We mixed OOF predictions together in the same way we ensemble our models on kaggle LB. So in some way we created mark up which is closer to reality e.g. “decrease” weakness of labels.\nI had some problems with creating OOF because we extract cells with slightly different algorithms. I wasn’t always able to find the same cell in Eric predictions. So only around 90% of cells were included in OOF with some small noise.\nSo I adapt my models from previous chapter and used OOF predictions as new target. I used mean square error loss function here and tried to minimize it. \n\nThese models gives around 0.528 on public LB, and after competition ends I find out that private is very similar to public score for these models. I also trained such model using green channel only.\nThese two models, trained on OOF, were included in our final ensemble instead of standard classification models. \n\n## ZFTurbo Part 3. Single class models.\n\nAt some point we checked which classes had low score for our solution. And as we expected class 11 was the worst. This class has low amount of images available. So I spend one evening on the following: I created small labeler tool. It shows you single cell from full image and you must click on one button if cell of target class or other button otherwise (see screenshot). It stores your answer in cache, so you can continue from last point.\n\n![img](//soloro.ru/images/2021/hpa-kaggle.png)\n\nI’m not biologist so I just googled “Mitotic spindle” and check images how it looks like. And actually it was pretty simple to find these cells on images. Also I probably found out why our models were bad for this class. The main reason is that Mitotic spindle cells are rare on images. It can have many cells and only single class 11. So when I trained my cell level models I show them mostly incorrect cells in training process. Using labeler tool I easily made good labels for all images in 3-4 hours.\nAfter all images were processed I just create single CSV with markup for class 11. And you can find it in this [open dataset](https://www.kaggle.com/zfturbo/hpa-single-cell-classification-class-11-markup/).\nSo I trained binary classifier model on new labels and our score increased on public +0.008 (and as I find out recently it probably even more around 0.014 on private).\nSo actually in my opinion handlabeling is key to win in this competition. ) But it was hard to do mark up for other classes because of large number of images and not so obvious differences in cells between some classes. Also I was too lazy to do additional markup. So class 11 was the only one where we made labels.\n\nI also created some other single models using default training labels, but they mostly wasn’t very good comparing to multiclass models, so we didn’t use them.\n",
    "1307011": "I released the dataset here https://www.kaggle.com/narsil/hpa2021-class11-manuallabels",
    "1306965": "This is a snapshot if images in my dataset\n![mitotic_spindle](https://i.ibb.co/g6mfxZ5/mitotic-spindle.png)",
    "1306895": "Congrats on the very impressive result @zfturbo !\n\nRegarding manual re-labeling of class 11, I had exactly the same intuition and followed a very similar procedure. I am working to make the dataset public.",
    "1306961": "",
    "1307360": "Congrats! Thank you for sharing."
  }
}