{
  "id": 221550,
  "title": "Rapid experimentation with fastai [0.342 lb]",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/221550",
  "author_name": "Darek Kłeczek",
  "post_date": "2021-02-23T06:27:52.090000",
  "votes": 64,
  "comment_count": 28,
  "views": 0,
  "content": "<p>I'd like to share my current approach for this competition, inspired by <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114214\" target=\"_blank\">solution for RSNA Intracranial Hemorrhage Detection</a>. I will release my code and datasets as well.</p>\n<p>I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that (similar to the solution shared by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>). </p>\n<h1>1. Create a prototyping dataset</h1>\n<p>What I need as an input to the classification model are images of individual cells. For experimentation I don't need all the images, instead I create a sample from the train set. The additional benefit is that my sample is more balanced than train. I use RGB channels only, which has proven to work well in the previous HPA challenge. I save the extracted cells as RGB jpg images so that I can feed them easily into my classifier. <br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced\" target=\"_blank\">Notebook for creating the training dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset/\" target=\"_blank\">Training Dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/\" target=\"_blank\">Notebook for creating the public test dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc-dataset\" target=\"_blank\">Public Test Dataset</a></p>\n<h1>2. Use fastai training loop with the data-block API</h1>\n<p>fastai is a great tool to create a strong baseline quickly. I use pretty much out of the box approach for multilabel classification, with resnet50 backbone, one cycle training, lr finder etc. The data block API is a great way to prepare the data, and comes with a default set of augmentations that I use as well. <br>\n<a href=\"https://www.kaggle.com/thedrcat/fastai-cell-tile-prototyping-training/\" target=\"_blank\">Training notebook</a></p>\n<h1>3. Quick submission template</h1>\n<p>I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. These submissions will get zero score on private, but there is still lots of time in the competition and better approaches will be developed.<br>\n<a href=\"https://www.kaggle.com/thedrcat/fastai-quick-submission-template\" target=\"_blank\">Submission template</a> </p>\n<p>PS. After a couple of years getting serious about ML, I've finally decided to get my own DL rig, so that I can train bigger models for longer at a later stage of the competition. </p>",
  "messages": [
    {
      "id": 1214816,
      "postDate": "2021-02-23T06:27:52.090Z",
      "content": "<p>I'd like to share my current approach for this competition, inspired by <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> <a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114214\" target=\"_blank\">solution for RSNA Intracranial Hemorrhage Detection</a>. I will release my code and datasets as well.</p>\n<p>I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that (similar to the solution shared by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>). </p>\n<h1>1. Create a prototyping dataset</h1>\n<p>What I need as an input to the classification model are images of individual cells. For experimentation I don't need all the images, instead I create a sample from the train set. The additional benefit is that my sample is more balanced than train. I use RGB channels only, which has proven to work well in the previous HPA challenge. I save the extracted cells as RGB jpg images so that I can feed them easily into my classifier. <br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced\" target=\"_blank\">Notebook for creating the training dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset/\" target=\"_blank\">Training Dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/\" target=\"_blank\">Notebook for creating the public test dataset</a><br>\n<a href=\"https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc-dataset\" target=\"_blank\">Public Test Dataset</a></p>\n<h1>2. Use fastai training loop with the data-block API</h1>\n<p>fastai is a great tool to create a strong baseline quickly. I use pretty much out of the box approach for multilabel classification, with resnet50 backbone, one cycle training, lr finder etc. The data block API is a great way to prepare the data, and comes with a default set of augmentations that I use as well. <br>\n<a href=\"https://www.kaggle.com/thedrcat/fastai-cell-tile-prototyping-training/\" target=\"_blank\">Training notebook</a></p>\n<h1>3. Quick submission template</h1>\n<p>I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. These submissions will get zero score on private, but there is still lots of time in the competition and better approaches will be developed.<br>\n<a href=\"https://www.kaggle.com/thedrcat/fastai-quick-submission-template\" target=\"_blank\">Submission template</a> </p>\n<p>PS. After a couple of years getting serious about ML, I've finally decided to get my own DL rig, so that I can train bigger models for longer at a later stage of the competition. </p>",
      "rawMarkdown": "I'd like to share my current approach for this competition, inspired by @jhoward [solution for RSNA Intracranial Hemorrhage Detection](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114214). I will release my code and datasets as well.\n\nI am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that (similar to the solution shared by @dschettler8845). \n\n# 1. Create a prototyping dataset\nWhat I need as an input to the classification model are images of individual cells. For experimentation I don't need all the images, instead I create a sample from the train set. The additional benefit is that my sample is more balanced than train. I use RGB channels only, which has proven to work well in the previous HPA challenge. I save the extracted cells as RGB jpg images so that I can feed them easily into my classifier. \n[Notebook for creating the training dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced)\n[Training Dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset/)\n[Notebook for creating the public test dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/)\n[Public Test Dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc-dataset)\n\n# 2. Use fastai training loop with the data-block API\nfastai is a great tool to create a strong baseline quickly. I use pretty much out of the box approach for multilabel classification, with resnet50 backbone, one cycle training, lr finder etc. The data block API is a great way to prepare the data, and comes with a default set of augmentations that I use as well. \n[Training notebook](https://www.kaggle.com/thedrcat/fastai-cell-tile-prototyping-training/)\n\n# 3. Quick submission template\nI want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. These submissions will get zero score on private, but there is still lots of time in the competition and better approaches will be developed.\n[Submission template](https://www.kaggle.com/thedrcat/fastai-quick-submission-template) \n\nPS. After a couple of years getting serious about ML, I've finally decided to get my own DL rig, so that I can train bigger models for longer at a later stage of the competition. ",
      "votes": 62
    },
    {
      "id": 1239547,
      "postDate": "2021-03-15T19:48:10.887Z",
      "content": "<p>Hi Darek, thanks for sharing. It really provided me with some insights. I just started this competition, may I ask when you say <br>\n<code>I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that</code><br>\nDoes it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label? Thanks a lot for clarifying. Additionally, how should we make predictions with the trained classifier?</p>",
      "rawMarkdown": "Hi Darek, thanks for sharing. It really provided me with some insights. I just started this competition, may I ask when you say \n```I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that ```\nDoes it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label? Thanks a lot for clarifying. Additionally, how should we make predictions with the trained classifier?",
      "votes": 3,
      "replies": [
        {
          "id": 1240125,
          "postDate": "2021-03-16T08:39:11.903Z",
          "content": "<blockquote>\n  <p>Does it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label?</p>\n</blockquote>\n<p>Yes, correct. </p>\n<blockquote>\n  <p>Additionally, how should we make predictions with the trained classifier?</p>\n</blockquote>\n<p>You also need to follow the 2 steps approach - extract the cells with HPA cell segmentator, then use the classifier to predict each cell. You can see my notebooks shared in the post for how I am doing that. </p>",
          "rawMarkdown": "> Does it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label?\n\nYes, correct. \n\n> Additionally, how should we make predictions with the trained classifier?\n\nYou also need to follow the 2 steps approach - extract the cells with HPA cell segmentator, then use the classifier to predict each cell. You can see my notebooks shared in the post for how I am doing that. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1226922,
      "postDate": "2021-03-05T03:01:48.667Z",
      "content": "<p>Hi I am finding little correlation between the mAP obtained on a single validation fold using variations of this pipeline and leaderboard score have you observed the same?</p>",
      "rawMarkdown": "Hi I am finding little correlation between the mAP obtained on a single validation fold using variations of this pipeline and leaderboard score have you observed the same?",
      "votes": 1,
      "replies": [
        {
          "id": 1226952,
          "postDate": "2021-03-05T03:50:34.510Z",
          "content": "<p>Indeed - figuring out how to do CV is a key challenge in this competition. The cell-level labels are super noisy because they are inherited from image-level labels so may be often inaccurate.</p>\n<p>Right now, I'm doing 5-fold experiments and using public LB to validate. </p>",
          "rawMarkdown": "Indeed - figuring out how to do CV is a key challenge in this competition. The cell-level labels are super noisy because they are inherited from image-level labels so may be often inaccurate.\n\nRight now, I'm doing 5-fold experiments and using public LB to validate. "
        },
        {
          "id": 1248160,
          "postDate": "2021-03-22T11:48:21.690Z",
          "content": "<p>Hi, if you're using tf2, could you please share good mAP metric implementation pls?</p>",
          "rawMarkdown": "Hi, if you're using tf2, could you please share good mAP metric implementation pls?"
        }
      ]
    },
    {
      "id": 1269687,
      "postDate": "2021-04-10T19:31:21.353Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Thank you for sharing. I'm a huge fan of fastai. Love when I see someone doing something really great with it on Kaggle.</p>",
      "rawMarkdown": "@thedrcat Thank you for sharing. I'm a huge fan of fastai. Love when I see someone doing something really great with it on Kaggle.",
      "votes": 2
    },
    {
      "id": 1215654,
      "postDate": "2021-02-23T20:41:21.663Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">Darek</a></p>\n<p>If you have built the machine - share your rig design ?</p>\n<p>If you have not yet built the machine - send me an email thru kaggle - I have built 4 so far for kaggle use and can share some decent advice and experience.</p>",
      "rawMarkdown": "[Darek](https://www.kaggle.com/thedrcat)\n\nIf you have built the machine - share your rig design ?\n\nIf you have not yet built the machine - send me an email thru kaggle - I have built 4 so far for kaggle use and can share some decent advice and experience.",
      "votes": 2,
      "replies": [
        {
          "id": 1215914,
          "postDate": "2021-02-24T03:37:44.663Z",
          "content": "<p>I'm not really a hardware person, so I bought a pre-configured gaming desktop with Intel i9-10900K, 64GB RAM and an RTX3090. Too bad I didn't know you before, maybe I'd be more adventurous! Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> !</p>",
          "rawMarkdown": "I'm not really a hardware person, so I bought a pre-configured gaming desktop with Intel i9-10900K, 64GB RAM and an RTX3090. Too bad I didn't know you before, maybe I'd be more adventurous! Thanks @pcjimmmy !"
        },
        {
          "id": 1216005,
          "postDate": "2021-02-24T04:48:56.203Z",
          "content": "<p>Sounds like it will be fun!  </p>\n<p>Things that made mine run better for ML</p>\n<ol>\n<li>Have your linux on a 2TB SSD - use 200GB for swap.  Pretty much never run out of CPU RAM that way.  Let me know if you need the terminal commands to get that size swap drive - when the swap is on a SSD you notice the speed bump when swap starts being used, but pretty hard IMO to do ML with the standard 2GB swap that Ubuntu sets up.</li>\n<li>Put your spare change in a large coffee can and when its full buy a second graphics card.</li>\n<li>I have had very good success with Ubuntu 20.04.</li>\n<li>Getting the right combination of tensorflow version, graphics card driver, cudatoolkit and cudnn is always a real joy, same joy if your Pytorching it.  I have had lots and lots of practice doing a full clean Ubuntu install when I get things so mixed up they cannot recover.   Plan like your going to reinstall Linux every few months.  This means none of your personal stuff resides on the Linux SSD.  </li>\n<li>I built a cheap freenas server with 10TB hard drives to keep as much off the PC and on the server - since I have 4 machines using a server for central file storage a must.  For this competition I still put all the images on the PC SSD but tabular competitions can almost always keep data on the server.</li>\n<li>Pretty much always do mixed precision - near the end of a competition when the model is pretty well done I will finish off with a training session not using mixed.</li>\n<li>Have a battery backup - I only plug the PC into mine - it's really enough to make you cry to have the power interrupt near the end of a three or four day training run.  When you get that second card your going to be able to keep the house warm with the power going to the dual GPU - make sure you buy a big backup.</li>\n</ol>",
          "rawMarkdown": "Sounds like it will be fun!  \n\nThings that made mine run better for ML\n\n1.  Have your linux on a 2TB SSD - use 200GB for swap.  Pretty much never run out of CPU RAM that way.  Let me know if you need the terminal commands to get that size swap drive - when the swap is on a SSD you notice the speed bump when swap starts being used, but pretty hard IMO to do ML with the standard 2GB swap that Ubuntu sets up.\n2.  Put your spare change in a large coffee can and when its full buy a second graphics card.\n3.  I have had very good success with Ubuntu 20.04.\n4.  Getting the right combination of tensorflow version, graphics card driver, cudatoolkit and cudnn is always a real joy, same joy if your Pytorching it.  I have had lots and lots of practice doing a full clean Ubuntu install when I get things so mixed up they cannot recover.   Plan like your going to reinstall Linux every few months.  This means none of your personal stuff resides on the Linux SSD.  \n5.  I built a cheap freenas server with 10TB hard drives to keep as much off the PC and on the server - since I have 4 machines using a server for central file storage a must.  For this competition I still put all the images on the PC SSD but tabular competitions can almost always keep data on the server.\n6.  Pretty much always do mixed precision - near the end of a competition when the model is pretty well done I will finish off with a training session not using mixed.\n7.  Have a battery backup - I only plug the PC into mine - it's really enough to make you cry to have the power interrupt near the end of a three or four day training run.  When you get that second card your going to be able to keep the house warm with the power going to the dual GPU - make sure you buy a big backup.",
          "votes": 9
        },
        {
          "id": 1219011,
          "postDate": "2021-02-26T11:18:39.677Z",
          "content": "<p>Nice tips <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>. Thanks for sharing. </p>",
          "rawMarkdown": "Nice tips @pcjimmmy. Thanks for sharing. "
        }
      ]
    },
    {
      "id": 1215645,
      "postDate": "2021-02-23T20:30:07.687Z",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -- Outstanding work!!! Congratulations on the excellent score.</p>\n<hr>\n<p>This gives me hope that I can continue to improve my models past where they currently are! Not being able to do any CV is very frustrating (run a submission and check 9 hours later is very tedious).</p>\n<hr>\n<p>Can you provide more clarity on this piece? </p>\n<blockquote>\n  <p><em>\"I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. <strong>These submissions will get zero scores on private</strong>, but there is still lots of time in the competition and better approaches will be developed.\"</em></p>\n</blockquote>\n<p>Are you saying that if we only submit predictions for the test dataset we will still see the full score for the public LB? If this is true you have legitimately made my week and blown my mind.</p>\n<hr>\n<p>Happy to hear you're getting a DL rig! That's wonderful. I'm still leveraging Colab and Kaggle only. Hopefully, I will be able to get a rig soon too.</p>",
      "rawMarkdown": "@thedrcat -- Outstanding work!!! Congratulations on the excellent score.\n\n---\n\nThis gives me hope that I can continue to improve my models past where they currently are! Not being able to do any CV is very frustrating (run a submission and check 9 hours later is very tedious).\n\n---\n\nCan you provide more clarity on this piece? \n\n> *\"I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. **These submissions will get zero scores on private**, but there is still lots of time in the competition and better approaches will be developed.\"*\n\nAre you saying that if we only submit predictions for the test dataset we will still see the full score for the public LB? If this is true you have legitimately made my week and blown my mind.\n\n---\n\nHappy to hear you're getting a DL rig! That's wonderful. I'm still leveraging Colab and Kaggle only. Hopefully, I will be able to get a rig soon too.",
      "votes": 2,
      "replies": [
        {
          "id": 1215917,
          "postDate": "2021-02-24T03:42:54.543Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, yes indeed - you can submit predictions for the test set only and see the public LB score. I run inference on the test set in my training kernel and save the predictions, then load them in a separate submission notebook so that I can test it quickly using different thresholds etc. You can check my submission template for this. </p>\n<p>I've been working through colab and kaggle so far as well, so looking forward to the new experience :) </p>",
          "rawMarkdown": "Hey @dschettler8845, yes indeed - you can submit predictions for the test set only and see the public LB score. I run inference on the test set in my training kernel and save the predictions, then load them in a separate submission notebook so that I can test it quickly using different thresholds etc. You can check my submission template for this. \n\nI've been working through colab and kaggle so far as well, so looking forward to the new experience :) ",
          "votes": 2
        },
        {
          "id": 1216662,
          "postDate": "2021-02-24T11:58:01.890Z",
          "content": "<p>Gah, this is life changing. Haha.</p>\n<p>Thanks again!! </p>",
          "rawMarkdown": "Gah, this is life changing. Haha.\n\nThanks again!! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1265748,
      "postDate": "2021-04-07T07:22:53.630Z",
      "content": "<p>Hello,if I donnot want to fast submit,what should I do ?</p>",
      "rawMarkdown": "Hello,if I donnot want to fast submit,what should I do ?",
      "replies": [
        {
          "id": 1265762,
          "postDate": "2021-04-07T07:34:57.363Z",
          "content": "<p>Run segmentation first, then inference on the segmented cells. Some other notebooks are showing this, I haven't implemented this yet :) </p>",
          "rawMarkdown": "Run segmentation first, then inference on the segmented cells. Some other notebooks are showing this, I haven't implemented this yet :) "
        },
        {
          "id": 1265774,
          "postDate": "2021-04-07T07:42:47.050Z",
          "content": "<p>Oh , it will cost much time …</p>",
          "rawMarkdown": "Oh , it will cost much time ..."
        }
      ]
    },
    {
      "id": 1262719,
      "postDate": "2021-04-04T16:48:23.847Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> :) </p>\n<p>Thnak you for showing me that Fastai is more powerful than I thought. </p>\n<p>I am reading now This book \"Deep Learning for Coders with fastai and PyTorch\" by Jeremy Howard</p>\n<p>Would you say with Fastai you can do well in most kaggle competitions? </p>\n<p>all the best :)</p>",
      "rawMarkdown": "Hi @thedrcat :) \n\nThnak you for showing me that Fastai is more powerful than I thought. \n\nI am reading now This book \"Deep Learning for Coders with fastai and PyTorch\" by Jeremy Howard\n\nWould you say with Fastai you can do well in most kaggle competitions? \n \nall the best :)",
      "replies": [
        {
          "id": 1263097,
          "postDate": "2021-04-05T04:58:36.680Z",
          "content": "<p>It's a very good book, I enjoyed it! Anything you can do with Pytorch, you can do with fastai, so I don't see a limit for what you can achieve :) </p>",
          "rawMarkdown": "It's a very good book, I enjoyed it! Anything you can do with Pytorch, you can do with fastai, so I don't see a limit for what you can achieve :) "
        },
        {
          "id": 1263324,
          "postDate": "2021-04-05T10:15:03.943Z",
          "content": "<p>as Tensorflow users, fastaiv2 is perfect for me. </p>\n<p>oh, I see.. so fastai v2 is a complete rewrite of fastai. </p>\n<blockquote>\n  <p>fastai v2: A complete rewrite of fastai which is faster, easier, and more flexible, implementing new approaches to deep learning framework design, as discussed in the peer-reviewed academic paper<br>\n  <a href=\"https://www.mdpi.com/2078-2489/11/2/108\" target=\"_blank\">https://www.mdpi.com/2078-2489/11/2/108</a></p>\n</blockquote>\n<p>Thnak you, all the best :)</p>",
          "rawMarkdown": "as Tensorflow users, fastaiv2 is perfect for me. \n\noh, I see.. so fastai v2 is a complete rewrite of fastai. \n\n> fastai v2: A complete rewrite of fastai which is faster, easier, and more flexible, implementing new approaches to deep learning framework design, as discussed in the peer-reviewed academic paper\nhttps://www.mdpi.com/2078-2489/11/2/108\n\nThnak you, all the best :)"
        }
      ]
    },
    {
      "id": 1258974,
      "postDate": "2021-04-01T04:13:20.187Z",
      "content": "<p>Hi Darek, thanks for your sharing! You say that your submission will get zero on private score. But I am still confused about it. Is there any difference between \"public test images\" for pre-processing and \"public test images\" during submission? </p>",
      "rawMarkdown": "Hi Darek, thanks for your sharing! You say that your submission will get zero on private score. But I am still confused about it. Is there any difference between \"public test images\" for pre-processing and \"public test images\" during submission? ",
      "replies": [
        {
          "id": 1258981,
          "postDate": "2021-04-01T04:22:50.027Z",
          "content": "<p>When you submit, the submission csv is swapped with a new version that has appended private test set images. Public test images are counted towards public LB score, and the appended private test is counted towards private LB, which is what really matters :) </p>",
          "rawMarkdown": "When you submit, the submission csv is swapped with a new version that has appended private test set images. Public test images are counted towards public LB score, and the appended private test is counted towards private LB, which is what really matters :) ",
          "votes": 1
        },
        {
          "id": 1259010,
          "postDate": "2021-04-01T05:02:28.520Z",
          "content": "<p>I see, thanks!</p>",
          "rawMarkdown": "I see, thanks!"
        }
      ]
    },
    {
      "id": 1226296,
      "postDate": "2021-03-04T12:23:37.500Z",
      "content": "<p>Hi. I have a question. For the test dataset sometimes we have to make predictions where more than one cell exists in the image. So, in the end do you combine all the possible labels for the cells in an image together. I got this question as I can see that the test dataset has individual cell images.</p>",
      "rawMarkdown": "Hi. I have a question. For the test dataset sometimes we have to make predictions where more than one cell exists in the image. So, in the end do you combine all the possible labels for the cells in an image together. I got this question as I can see that the test dataset has individual cell images.",
      "replies": [
        {
          "id": 1226532,
          "postDate": "2021-03-04T16:01:42.420Z",
          "content": "<p>In this challenge we need to first segment the cells and then predict labels for each cell. In my approach, I use HPAsegmentator to segment the cells in a separate notebook, and then classify the individual cells. For the final submission, your notebook needs to do both segmentation and classification. </p>",
          "rawMarkdown": "In this challenge we need to first segment the cells and then predict labels for each cell. In my approach, I use HPAsegmentator to segment the cells in a separate notebook, and then classify the individual cells. For the final submission, your notebook needs to do both segmentation and classification. ",
          "votes": 1
        },
        {
          "id": 1226728,
          "postDate": "2021-03-04T19:36:26.507Z",
          "content": "<p>Hi. I have another question. For doing segmentation we need a ground truth right? How do you use ground truth here?</p>",
          "rawMarkdown": "Hi. I have another question. For doing segmentation we need a ground truth right? How do you use ground truth here?",
          "votes": 1
        },
        {
          "id": 1226954,
          "postDate": "2021-03-05T03:52:38.103Z",
          "content": "<p>Normally yes, but this is a \"weakly supervised instance segmentation\" challenge so we don't have the ground truth for segmentation. But, there is a pretrained segmentation model shared by the host which they say is at least 90% accurate. I am using this pretrained segmentation model in my approach. </p>",
          "rawMarkdown": "Normally yes, but this is a \"weakly supervised instance segmentation\" challenge so we don't have the ground truth for segmentation. But, there is a pretrained segmentation model shared by the host which they say is at least 90% accurate. I am using this pretrained segmentation model in my approach. "
        },
        {
          "id": 1248919,
          "postDate": "2021-03-22T23:07:50.567Z",
          "content": "<p>Could you share the topic where that is from? I'm trying to join the competition a bit late and your topic was one of the firsts i came up with! Thanks for your insights btw!</p>",
          "rawMarkdown": "Could you share the topic where that is from? I'm trying to join the competition a bit late and your topic was one of the firsts i came up with! Thanks for your insights btw!"
        },
        {
          "id": 1249445,
          "postDate": "2021-03-23T09:59:18.927Z",
          "content": "<p><a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> you should be able to find this topic and others with links to the original post in my EDA notebook: <a href=\"https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda</a></p>",
          "rawMarkdown": "@capiru you should be able to find this topic and others with links to the original post in my EDA notebook: https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1239547,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-03-15T19:48:10.887000",
      "content": "<p>Hi Darek, thanks for sharing. It really provided me with some insights. I just started this competition, may I ask when you say <br>\n<code>I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that</code><br>\nDoes it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label? Thanks a lot for clarifying. Additionally, how should we make predictions with the trained classifier?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1240125,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-16T08:39:11.903000",
          "content": "<blockquote>\n  <p>Does it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label?</p>\n</blockquote>\n<p>Yes, correct. </p>\n<blockquote>\n  <p>Additionally, how should we make predictions with the trained classifier?</p>\n</blockquote>\n<p>You also need to follow the 2 steps approach - extract the cells with HPA cell segmentator, then use the classifier to predict each cell. You can see my notebooks shared in the post for how I am doing that. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1226922,
      "author_name": "Felipe Bivort Haiek",
      "author_url": "",
      "post_date": "2021-03-05T03:01:48.667000",
      "content": "<p>Hi I am finding little correlation between the mAP obtained on a single validation fold using variations of this pipeline and leaderboard score have you observed the same?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1226952,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-05T03:50:34.510000",
          "content": "<p>Indeed - figuring out how to do CV is a key challenge in this competition. The cell-level labels are super noisy because they are inherited from image-level labels so may be often inaccurate.</p>\n<p>Right now, I'm doing 5-fold experiments and using public LB to validate. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248160,
          "author_name": "Vladislav Ostankovich",
          "author_url": "",
          "post_date": "2021-03-22T11:48:21.690000",
          "content": "<p>Hi, if you're using tf2, could you please share good mAP metric implementation pls?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1269687,
      "author_name": "Charlie Craine",
      "author_url": "",
      "post_date": "2021-04-10T19:31:21.353000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> Thank you for sharing. I'm a huge fan of fastai. Love when I see someone doing something really great with it on Kaggle.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1215654,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-23T20:41:21.663000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">Darek</a></p>\n<p>If you have built the machine - share your rig design ?</p>\n<p>If you have not yet built the machine - send me an email thru kaggle - I have built 4 so far for kaggle use and can share some decent advice and experience.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1215914,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-24T03:37:44.663000",
          "content": "<p>I'm not really a hardware person, so I bought a pre-configured gaming desktop with Intel i9-10900K, 64GB RAM and an RTX3090. Too bad I didn't know you before, maybe I'd be more adventurous! Thanks <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1216005,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-02-24T04:48:56.203000",
          "content": "<p>Sounds like it will be fun!  </p>\n<p>Things that made mine run better for ML</p>\n<ol>\n<li>Have your linux on a 2TB SSD - use 200GB for swap.  Pretty much never run out of CPU RAM that way.  Let me know if you need the terminal commands to get that size swap drive - when the swap is on a SSD you notice the speed bump when swap starts being used, but pretty hard IMO to do ML with the standard 2GB swap that Ubuntu sets up.</li>\n<li>Put your spare change in a large coffee can and when its full buy a second graphics card.</li>\n<li>I have had very good success with Ubuntu 20.04.</li>\n<li>Getting the right combination of tensorflow version, graphics card driver, cudatoolkit and cudnn is always a real joy, same joy if your Pytorching it.  I have had lots and lots of practice doing a full clean Ubuntu install when I get things so mixed up they cannot recover.   Plan like your going to reinstall Linux every few months.  This means none of your personal stuff resides on the Linux SSD.  </li>\n<li>I built a cheap freenas server with 10TB hard drives to keep as much off the PC and on the server - since I have 4 machines using a server for central file storage a must.  For this competition I still put all the images on the PC SSD but tabular competitions can almost always keep data on the server.</li>\n<li>Pretty much always do mixed precision - near the end of a competition when the model is pretty well done I will finish off with a training session not using mixed.</li>\n<li>Have a battery backup - I only plug the PC into mine - it's really enough to make you cry to have the power interrupt near the end of a three or four day training run.  When you get that second card your going to be able to keep the house warm with the power going to the dual GPU - make sure you buy a big backup.</li>\n</ol>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 1219011,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-26T11:18:39.677000",
          "content": "<p>Nice tips <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>. Thanks for sharing. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1215645,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-02-23T20:30:07.687000",
      "content": "<p><a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> -- Outstanding work!!! Congratulations on the excellent score.</p>\n<hr>\n<p>This gives me hope that I can continue to improve my models past where they currently are! Not being able to do any CV is very frustrating (run a submission and check 9 hours later is very tedious).</p>\n<hr>\n<p>Can you provide more clarity on this piece? </p>\n<blockquote>\n  <p><em>\"I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. <strong>These submissions will get zero scores on private</strong>, but there is still lots of time in the competition and better approaches will be developed.\"</em></p>\n</blockquote>\n<p>Are you saying that if we only submit predictions for the test dataset we will still see the full score for the public LB? If this is true you have legitimately made my week and blown my mind.</p>\n<hr>\n<p>Happy to hear you're getting a DL rig! That's wonderful. I'm still leveraging Colab and Kaggle only. Hopefully, I will be able to get a rig soon too.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1215917,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-02-24T03:42:54.543000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, yes indeed - you can submit predictions for the test set only and see the public LB score. I run inference on the test set in my training kernel and save the predictions, then load them in a separate submission notebook so that I can test it quickly using different thresholds etc. You can check my submission template for this. </p>\n<p>I've been working through colab and kaggle so far as well, so looking forward to the new experience :) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1216662,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-24T11:58:01.890000",
          "content": "<p>Gah, this is life changing. Haha.</p>\n<p>Thanks again!! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1265748,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-04-07T07:22:53.630000",
      "content": "<p>Hello,if I donnot want to fast submit,what should I do ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1265762,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-04-07T07:34:57.363000",
          "content": "<p>Run segmentation first, then inference on the segmented cells. Some other notebooks are showing this, I haven't implemented this yet :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1265774,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-04-07T07:42:47.050000",
          "content": "<p>Oh , it will cost much time …</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1262719,
      "author_name": "Faisal Alsrheed",
      "author_url": "",
      "post_date": "2021-04-04T16:48:23.847000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> :) </p>\n<p>Thnak you for showing me that Fastai is more powerful than I thought. </p>\n<p>I am reading now This book \"Deep Learning for Coders with fastai and PyTorch\" by Jeremy Howard</p>\n<p>Would you say with Fastai you can do well in most kaggle competitions? </p>\n<p>all the best :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1263097,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-04-05T04:58:36.680000",
          "content": "<p>It's a very good book, I enjoyed it! Anything you can do with Pytorch, you can do with fastai, so I don't see a limit for what you can achieve :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1263324,
          "author_name": "Faisal Alsrheed",
          "author_url": "",
          "post_date": "2021-04-05T10:15:03.943000",
          "content": "<p>as Tensorflow users, fastaiv2 is perfect for me. </p>\n<p>oh, I see.. so fastai v2 is a complete rewrite of fastai. </p>\n<blockquote>\n  <p>fastai v2: A complete rewrite of fastai which is faster, easier, and more flexible, implementing new approaches to deep learning framework design, as discussed in the peer-reviewed academic paper<br>\n  <a href=\"https://www.mdpi.com/2078-2489/11/2/108\" target=\"_blank\">https://www.mdpi.com/2078-2489/11/2/108</a></p>\n</blockquote>\n<p>Thnak you, all the best :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1258974,
      "author_name": "Nin7a1",
      "author_url": "",
      "post_date": "2021-04-01T04:13:20.187000",
      "content": "<p>Hi Darek, thanks for your sharing! You say that your submission will get zero on private score. But I am still confused about it. Is there any difference between \"public test images\" for pre-processing and \"public test images\" during submission? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1258981,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-04-01T04:22:50.027000",
          "content": "<p>When you submit, the submission csv is swapped with a new version that has appended private test set images. Public test images are counted towards public LB score, and the appended private test is counted towards private LB, which is what really matters :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1259010,
          "author_name": "Nin7a1",
          "author_url": "",
          "post_date": "2021-04-01T05:02:28.520000",
          "content": "<p>I see, thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1226296,
      "author_name": "RSASHWIN",
      "author_url": "",
      "post_date": "2021-03-04T12:23:37.500000",
      "content": "<p>Hi. I have a question. For the test dataset sometimes we have to make predictions where more than one cell exists in the image. So, in the end do you combine all the possible labels for the cells in an image together. I got this question as I can see that the test dataset has individual cell images.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1226532,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-04T16:01:42.420000",
          "content": "<p>In this challenge we need to first segment the cells and then predict labels for each cell. In my approach, I use HPAsegmentator to segment the cells in a separate notebook, and then classify the individual cells. For the final submission, your notebook needs to do both segmentation and classification. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1226728,
          "author_name": "RSASHWIN",
          "author_url": "",
          "post_date": "2021-03-04T19:36:26.507000",
          "content": "<p>Hi. I have another question. For doing segmentation we need a ground truth right? How do you use ground truth here?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1226954,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-05T03:52:38.103000",
          "content": "<p>Normally yes, but this is a \"weakly supervised instance segmentation\" challenge so we don't have the ground truth for segmentation. But, there is a pretrained segmentation model shared by the host which they say is at least 90% accurate. I am using this pretrained segmentation model in my approach. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1248919,
          "author_name": "Gabriel Prado",
          "author_url": "",
          "post_date": "2021-03-22T23:07:50.567000",
          "content": "<p>Could you share the topic where that is from? I'm trying to join the competition a bit late and your topic was one of the firsts i came up with! Thanks for your insights btw!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1249445,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-03-23T09:59:18.927000",
          "content": "<p><a href=\"https://www.kaggle.com/capiru\" target=\"_blank\">@capiru</a> you should be able to find this topic and others with links to the original post in my EDA notebook: <a href=\"https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda\" target=\"_blank\">https://www.kaggle.com/thedrcat/hpa-single-cell-classification-eda</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1214816": "I'd like to share my current approach for this competition, inspired by @jhoward [solution for RSNA Intracranial Hemorrhage Detection](https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/114214). I will release my code and datasets as well.\n\nI am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that (similar to the solution shared by @dschettler8845). \n\n# 1. Create a prototyping dataset\nWhat I need as an input to the classification model are images of individual cells. For experimentation I don't need all the images, instead I create a sample from the train set. The additional benefit is that my sample is more balanced than train. I use RGB channels only, which has proven to work well in the previous HPA challenge. I save the extracted cells as RGB jpg images so that I can feed them easily into my classifier. \n[Notebook for creating the training dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced)\n[Training Dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-sample-balanced-dataset/)\n[Notebook for creating the public test dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc/)\n[Public Test Dataset](https://www.kaggle.com/thedrcat/hpa-cell-tiles-test-with-enc-dataset)\n\n# 2. Use fastai training loop with the data-block API\nfastai is a great tool to create a strong baseline quickly. I use pretty much out of the box approach for multilabel classification, with resnet50 backbone, one cycle training, lr finder etc. The data block API is a great way to prepare the data, and comes with a default set of augmentations that I use as well. \n[Training notebook](https://www.kaggle.com/thedrcat/fastai-cell-tile-prototyping-training/)\n\n# 3. Quick submission template\nI want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. These submissions will get zero score on private, but there is still lots of time in the competition and better approaches will be developed.\n[Submission template](https://www.kaggle.com/thedrcat/fastai-quick-submission-template) \n\nPS. After a couple of years getting serious about ML, I've finally decided to get my own DL rig, so that I can train bigger models for longer at a later stage of the competition. ",
    "1239547": "Hi Darek, thanks for sharing. It really provided me with some insights. I just started this competition, may I ask when you say \n```I am extracting individual cells with HPA cell segmentator, labeling them with image-level labels and training a classifier based on that ```\nDoes it mean if an image has 8|5|0 as image level labels, each cell of this image would have 8|5|0 as its label? Thanks a lot for clarifying. Additionally, how should we make predictions with the trained classifier?",
    "1226922": "Hi I am finding little correlation between the mAP obtained on a single validation fold using variations of this pipeline and leaderboard score have you observed the same?",
    "1269687": "@thedrcat Thank you for sharing. I'm a huge fan of fastai. Love when I see someone doing something really great with it on Kaggle.",
    "1215654": "[Darek](https://www.kaggle.com/thedrcat)\n\nIf you have built the machine - share your rig design ?\n\nIf you have not yet built the machine - send me an email thru kaggle - I have built 4 so far for kaggle use and can share some decent advice and experience.",
    "1215645": "@thedrcat -- Outstanding work!!! Congratulations on the excellent score.\n\n---\n\nThis gives me hope that I can continue to improve my models past where they currently are! Not being able to do any CV is very frustrating (run a submission and check 9 hours later is very tedious).\n\n---\n\nCan you provide more clarity on this piece? \n\n> *\"I want to experiment quickly and can't wait for the lb score, especially that with weak labels I haven't found a reasonable way to do CV, and depend on the public lb score. I have pre-processed the public test images in the same way as my prototyping dataset and submit my preds only for this piece. **These submissions will get zero scores on private**, but there is still lots of time in the competition and better approaches will be developed.\"*\n\nAre you saying that if we only submit predictions for the test dataset we will still see the full score for the public LB? If this is true you have legitimately made my week and blown my mind.\n\n---\n\nHappy to hear you're getting a DL rig! That's wonderful. I'm still leveraging Colab and Kaggle only. Hopefully, I will be able to get a rig soon too.",
    "1265748": "Hello,if I donnot want to fast submit,what should I do ?",
    "1262719": "Hi @thedrcat :) \n\nThnak you for showing me that Fastai is more powerful than I thought. \n\nI am reading now This book \"Deep Learning for Coders with fastai and PyTorch\" by Jeremy Howard\n\nWould you say with Fastai you can do well in most kaggle competitions? \n \nall the best :)",
    "1258974": "Hi Darek, thanks for your sharing! You say that your submission will get zero on private score. But I am still confused about it. Is there any difference between \"public test images\" for pre-processing and \"public test images\" during submission? ",
    "1226296": "Hi. I have a question. For the test dataset sometimes we have to make predictions where more than one cell exists in the image. So, in the end do you combine all the possible labels for the cells in an image together. I got this question as I can see that the test dataset has individual cell images."
  }
}