{
  "id": 214616,
  "title": "Welcome to the 2nd HPA challenge!",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/214616",
  "author_name": "Emma Lundberg",
  "post_date": "2021-01-27T07:53:51.368000",
  "votes": 79,
  "comment_count": 50,
  "views": 0,
  "content": "<p>Welcome to the Human Protein Atlas single cell classification challenge!</p>\n<p>In this competition, we ask you to help us classify subcellular protein localization patterns of single cells in microscope images. Solving the single-cell image classification challenge will help us determine precise locations for all human proteins in each individual cell in our large collection of open access images. This will contribute to a better understanding of functional differences between otherwise seemingly identical cells.</p>\n<p>You’re given a set of training images, where each image contains multiple cells that may vary in their protein localization patterns. Importantly, we provide rough labels for the protein pattern or patterns for each image, which may not be correct for every cell in the image. Your model will need to figure out how to train with our weakly labeled data and predict precise labels for each cell in the test image. In addition to these data, you’re given access to the publicly available images within the Human Protein Atlas, should you want more training data. In order to perform well in this challenge, you may need a good segmentation model. There are many options of pretrained models that perform quite well on our images (eg. <a href=\"https://www.cellpose.org/\" target=\"_blank\">Cellpose</a> or <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a>). Please check out <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">this notebook</a> to see how you can make use of public HPA data and HPACellSegmentation model.</p>\n<p>Please also note that this is a code competition where a large part of test images are hidden. To gain a better understanding of the task, check out our <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">guide explaining the patterns</a> and cells in our microscope images.</p>\n<p>Just like in our <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">previous competition</a>, we present here a multi-label problem. The main difference is that we are asking you to classify the label(s) of each cell in every image with only weak image-level training data, compared to the classification of image-level labels before. With this unique competition setup, we hope to see novel solutions that are capable of utilizing large amounts of unlabelled or weakly labelled data to generate robust, discriminative features. We envision that these solutions may help us understand protein localization patterns across a multitude of human cells. We have demonstrated the promise of both citizen science and artificial intelligence approaches in describing the location of human proteins in microscope images. However, none map the protein distribution at a single-cell level or match human level performance (<a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">Sullivan et al, 2018;</a> <a href=\"https://science.sciencemag.org/content/356/6340/eaal3321.full\" target=\"_blank\">Thul et al, 2017</a>; <a href=\"https://www.nature.com/articles/s41592-019-0658-6/\" target=\"_blank\">Ouyang et al, 2020</a>), something we believe you can help us with!</p>\n<p>We in the <a href=\"https://www.proteinatlas.org/\" target=\"_blank\">Human Protein Atlas</a> team are true believers in open science and the power of crowdsourcing. We are excited to see how the Kaggle community will tackle this problem and what new models and innovative solutions will be developed. </p>\n<p>Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you to develop the best models possible!</p>\n<p>Good luck, and have fun!</p>",
  "messages": [
    {
      "id": 1171981,
      "postDate": "2021-01-27T07:53:51.370Z",
      "content": "<p>Welcome to the Human Protein Atlas single cell classification challenge!</p>\n<p>In this competition, we ask you to help us classify subcellular protein localization patterns of single cells in microscope images. Solving the single-cell image classification challenge will help us determine precise locations for all human proteins in each individual cell in our large collection of open access images. This will contribute to a better understanding of functional differences between otherwise seemingly identical cells.</p>\n<p>You’re given a set of training images, where each image contains multiple cells that may vary in their protein localization patterns. Importantly, we provide rough labels for the protein pattern or patterns for each image, which may not be correct for every cell in the image. Your model will need to figure out how to train with our weakly labeled data and predict precise labels for each cell in the test image. In addition to these data, you’re given access to the publicly available images within the Human Protein Atlas, should you want more training data. In order to perform well in this challenge, you may need a good segmentation model. There are many options of pretrained models that perform quite well on our images (eg. <a href=\"https://www.cellpose.org/\" target=\"_blank\">Cellpose</a> or <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a>). Please check out <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">this notebook</a> to see how you can make use of public HPA data and HPACellSegmentation model.</p>\n<p>Please also note that this is a code competition where a large part of test images are hidden. To gain a better understanding of the task, check out our <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">guide explaining the patterns</a> and cells in our microscope images.</p>\n<p>Just like in our <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">previous competition</a>, we present here a multi-label problem. The main difference is that we are asking you to classify the label(s) of each cell in every image with only weak image-level training data, compared to the classification of image-level labels before. With this unique competition setup, we hope to see novel solutions that are capable of utilizing large amounts of unlabelled or weakly labelled data to generate robust, discriminative features. We envision that these solutions may help us understand protein localization patterns across a multitude of human cells. We have demonstrated the promise of both citizen science and artificial intelligence approaches in describing the location of human proteins in microscope images. However, none map the protein distribution at a single-cell level or match human level performance (<a href=\"https://www.nature.com/articles/nbt.4225\" target=\"_blank\">Sullivan et al, 2018;</a> <a href=\"https://science.sciencemag.org/content/356/6340/eaal3321.full\" target=\"_blank\">Thul et al, 2017</a>; <a href=\"https://www.nature.com/articles/s41592-019-0658-6/\" target=\"_blank\">Ouyang et al, 2020</a>), something we believe you can help us with!</p>\n<p>We in the <a href=\"https://www.proteinatlas.org/\" target=\"_blank\">Human Protein Atlas</a> team are true believers in open science and the power of crowdsourcing. We are excited to see how the Kaggle community will tackle this problem and what new models and innovative solutions will be developed. </p>\n<p>Along with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you to develop the best models possible!</p>\n<p>Good luck, and have fun!</p>",
      "rawMarkdown": "Welcome to the Human Protein Atlas single cell classification challenge!\n \nIn this competition, we ask you to help us classify subcellular protein localization patterns of single cells in microscope images. Solving the single-cell image classification challenge will help us determine precise locations for all human proteins in each individual cell in our large collection of open access images. This will contribute to a better understanding of functional differences between otherwise seemingly identical cells.\n\nYou’re given a set of training images, where each image contains multiple cells that may vary in their protein localization patterns. Importantly, we provide rough labels for the protein pattern or patterns for each image, which may not be correct for every cell in the image. Your model will need to figure out how to train with our weakly labeled data and predict precise labels for each cell in the test image. In addition to these data, you’re given access to the publicly available images within the Human Protein Atlas, should you want more training data. In order to perform well in this challenge, you may need a good segmentation model. There are many options of pretrained models that perform quite well on our images (eg. [Cellpose](https://www.cellpose.org/) or [HPACellSegmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)). Please check out [this notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) to see how you can make use of public HPA data and HPACellSegmentation model.\n\nPlease also note that this is a code competition where a large part of test images are hidden. To gain a better understanding of the task, check out our [guide explaining the patterns](https://www.kaggle.com/lnhtrang/single-cell-patterns) and cells in our microscope images.\n\nJust like in our [previous competition](https://www.kaggle.com/c/human-protein-atlas-image-classification), we present here a multi-label problem. The main difference is that we are asking you to classify the label(s) of each cell in every image with only weak image-level training data, compared to the classification of image-level labels before. With this unique competition setup, we hope to see novel solutions that are capable of utilizing large amounts of unlabelled or weakly labelled data to generate robust, discriminative features. We envision that these solutions may help us understand protein localization patterns across a multitude of human cells. We have demonstrated the promise of both citizen science and artificial intelligence approaches in describing the location of human proteins in microscope images. However, none map the protein distribution at a single-cell level or match human level performance ([Sullivan et al, 2018;](https://www.nature.com/articles/nbt.4225) [Thul et al, 2017](https://science.sciencemag.org/content/356/6340/eaal3321.full); [Ouyang et al, 2020](https://www.nature.com/articles/s41592-019-0658-6/)), something we believe you can help us with!\n\nWe in the [Human Protein Atlas](https://www.proteinatlas.org/) team are true believers in open science and the power of crowdsourcing. We are excited to see how the Kaggle community will tackle this problem and what new models and innovative solutions will be developed. \n\nAlong with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you to develop the best models possible!\n\nGood luck, and have fun!\n",
      "votes": 79
    },
    {
      "id": 1248993,
      "postDate": "2021-03-23T02:13:57.300Z",
      "content": "<p>I compared label with previous one, and found weird labels. <br>\ne.g.</p>\n<ol>\n<li>ID: <strong>6795f08c-bb99-11e8-b2b9-ac1f6b6435d0</strong>'s label is Nucleoplasm and Nuclear membrane, but previous train set is only Nuclear membrane.</li>\n<li>ID: <strong>cdcfdff6-bb9a-11e8-b2b9-ac1f6b6435d0</strong>'s label is Negative, but previous train set is Nucleoplasm.</li>\n</ol>\n<p>Which is correct?</p>\n<p>Please take a look into this.<br>\n<a href=\"https://www.kaggle.com/onodera/compare-with-previous-hpa-competition\" target=\"_blank\">https://www.kaggle.com/onodera/compare-with-previous-hpa-competition</a></p>",
      "rawMarkdown": "I compared label with previous one, and found weird labels. \ne.g.\n1. ID: **6795f08c-bb99-11e8-b2b9-ac1f6b6435d0**'s label is Nucleoplasm and Nuclear membrane, but previous train set is only Nuclear membrane.\n2. ID: **cdcfdff6-bb9a-11e8-b2b9-ac1f6b6435d0**'s label is Negative, but previous train set is Nucleoplasm.\n\nWhich is correct?\n\nPlease take a look into this.\nhttps://www.kaggle.com/onodera/compare-with-previous-hpa-competition",
      "votes": 3,
      "replies": [
        {
          "id": 1249250,
          "postDate": "2021-03-23T07:48:34.327Z",
          "content": "<p>Hi! The current labels are correct. These are updated through different HPA releases. (1) <code>Nucleoplasm</code> is added and (2) is found to be unspecific staining, which belongs to <code>Negative</code> class. </p>",
          "rawMarkdown": "Hi! The current labels are correct. These are updated through different HPA releases. (1) `Nucleoplasm` is added and (2) is found to be unspecific staining, which belongs to `Negative` class. ",
          "votes": 4
        },
        {
          "id": 1249255,
          "postDate": "2021-03-23T07:54:48.457Z",
          "content": "<p>Thank you for quick reply! </p>",
          "rawMarkdown": "Thank you for quick reply! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1201117,
      "postDate": "2021-02-15T07:23:57.597Z",
      "content": "<p>Hello M'am,<br>\n Can you describe what is the predictionstring : \"0 1 eNoLCAgIMAEABJkBdQ==\" include with example?<br>\nIs this string encripted?<br>\nWhat is the \"0 1\" , \"eNoLCAgIMAEABJkBdQ==\"?</p>",
      "rawMarkdown": "Hello M'am,\n Can you describe what is the predictionstring : \"0 1 eNoLCAgIMAEABJkBdQ==\" include with example?\nIs this string encripted?\nWhat is the \"0 1\" , \"eNoLCAgIMAEABJkBdQ==\"?",
      "votes": 1,
      "replies": [
        {
          "id": 1201511,
          "postDate": "2021-02-15T12:57:11.283Z",
          "content": "<p>Hi there. Please see <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\"><strong>this page</strong></a> for more information on the submission format details.</p>\n<hr>\n<p><strong><code>0</code></strong> is the predicted class-label corresponding to <strong><code>Nucleoplasm</code></strong><br><br>\n<strong><code>1</code></strong> is the <strong>confidence</strong> of the predicted class label <strong><code>Nucleoplasm</code></strong><br><br>\n<strong><code>eNoLCAgIMA...EABJkBdQ==</code></strong> is the RLE, ZLIB compressed and Base64 encoded segmentation mask for the cell we are predicting on (that we are saying has the <strong><code>Nucleoplasm</code></strong> class).</p>\n<blockquote>\n  <p><strong><em>FROM THE LINK ABOVE</em></strong></p>\n  <p><em>For each image in the test set, you must predict a list of instance segmentation masks and their associated detection score (Confidence)</em></p>\n  <p><em>The binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format as EncodedMask. Specifically, we use the Coco masks RLE encoding/decoding (see the encode method of COCO’s mask API), the zlib compression/decompression (RFC1950), and vanilla base64 encoding.</em></p>\n</blockquote>\n<hr>",
          "rawMarkdown": "Hi there. Please see [**this page**](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation) for more information on the submission format details.\n\n---\n\n**`0`** is the predicted class-label corresponding to **`Nucleoplasm`**<br>\n**`1`** is the **confidence** of the predicted class label **`Nucleoplasm`**<br>\n**`eNoLCAgIMA...EABJkBdQ==`** is the RLE, ZLIB compressed and Base64 encoded segmentation mask for the cell we are predicting on (that we are saying has the **`Nucleoplasm`** class).\n\n> ***FROM THE LINK ABOVE***\n>\n>*For each image in the test set, you must predict a list of instance segmentation masks and their associated detection score (Confidence)*\n>\n>*The binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format as EncodedMask. Specifically, we use the Coco masks RLE encoding/decoding (see the encode method of COCO’s mask API), the zlib compression/decompression (RFC1950), and vanilla base64 encoding.*\n\n---",
          "votes": 2
        },
        {
          "id": 1205650,
          "postDate": "2021-02-16T21:01:46.213Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1174970,
      "postDate": "2021-01-28T20:09:17.467Z",
      "content": "<blockquote>\n  <p>There are many options of pretrained models that perform quite well on our images (eg. Cellpose or HPACellSegmentation). Please check out this notebook to see how you can make use of public HPA data and HPACellSegmentation model.</p>\n</blockquote>\n<p>That's really helpful, thanks. Was HPACellSegmentation model used to create the ground truth cell segmentation for this competition, with classes assigned by annotators?</p>",
      "rawMarkdown": "> There are many options of pretrained models that perform quite well on our images (eg. Cellpose or HPACellSegmentation). Please check out this notebook to see how you can make use of public HPA data and HPACellSegmentation model.\n\nThat's really helpful, thanks. Was HPACellSegmentation model used to create the ground truth cell segmentation for this competition, with classes assigned by annotators?",
      "votes": 1,
      "replies": [
        {
          "id": 1174981,
          "postDate": "2021-01-28T20:17:22.603Z",
          "content": "<p>HPACellSegmentation was used as a basis followed my manual correction of the masks for the ground truth. Classes per single cell were subsequently assigned by annotators.</p>",
          "rawMarkdown": "HPACellSegmentation was used as a basis followed my manual correction of the masks for the ground truth. Classes per single cell were subsequently assigned by annotators.",
          "votes": 7
        },
        {
          "id": 1177078,
          "postDate": "2021-01-30T03:24:01.597Z",
          "content": "<p>I also created this notebook that allows to use the HPACellSegmentation with the constraints of the competition with no internet - <a href=\"https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet\" target=\"_blank\">https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet</a></p>",
          "rawMarkdown": "I also created this notebook that allows to use the HPACellSegmentation with the constraints of the competition with no internet - https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet",
          "votes": 1
        }
      ]
    },
    {
      "id": 1173418,
      "postDate": "2021-01-27T21:28:11.227Z",
      "content": "<p>Thank you for hosting this competition. It is fascinating and I’m excited to learn!</p>",
      "rawMarkdown": "Thank you for hosting this competition. It is fascinating and I’m excited to learn!\n",
      "votes": 1,
      "replies": [
        {
          "id": 1174983,
          "postDate": "2021-01-28T20:17:56.233Z",
          "content": "<p>Thank you! And best of luck solving the problem.</p>",
          "rawMarkdown": "Thank you! And best of luck solving the problem.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1173396,
      "postDate": "2021-01-27T21:03:07.487Z",
      "content": "<p>It is strange to me that:</p>\n<blockquote>\n  <p>External data, freely &amp; publicly available, is allowed. This includes pre-trained models.<br>\n  <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements</a></p>\n</blockquote>\n<p>Since a \"pre-trained model\" could embed learning from non-freely and non-publicly available data sets… how does that work?</p>",
      "rawMarkdown": "It is strange to me that:\n> External data, freely & publicly available, is allowed. This includes pre-trained models.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements\n\nSince a \"pre-trained model\" could embed learning from non-freely and non-publicly available data sets... how does that work?",
      "votes": 1,
      "replies": [
        {
          "id": 1174106,
          "postDate": "2021-01-28T09:14:53.780Z",
          "content": "<p>The rule is, as far as I know, standard for these kinds of competitions to make sure there is equal access to data-sets and pre-trained models.</p>\n<p>As long as the pre-trained models are freely available (and their licenses allow for usage in open source contexts) it should be fine to use regardless of them having been embedded information from non-freely available datasets.</p>\n<p>However, I will ask members of Kaggle staff to clarify the idea of the rule.</p>",
          "rawMarkdown": "The rule is, as far as I know, standard for these kinds of competitions to make sure there is equal access to data-sets and pre-trained models.\n\nAs long as the pre-trained models are freely available (and their licenses allow for usage in open source contexts) it should be fine to use regardless of them having been embedded information from non-freely available datasets.\n\nHowever, I will ask members of Kaggle staff to clarify the idea of the rule.",
          "votes": 1
        },
        {
          "id": 1176707,
          "postDate": "2021-01-29T18:32:28.500Z",
          "content": "<p>We agree with Casper's response, pre-trained models are allowed but the expectation is that the pre-trained models are freely available to all participants. </p>",
          "rawMarkdown": "We agree with Casper's response, pre-trained models are allowed but the expectation is that the pre-trained models are freely available to all participants. "
        }
      ]
    },
    {
      "id": 1245260,
      "postDate": "2021-03-19T16:14:52.860Z",
      "content": "<p>Hello everyone, </p>\n<p>I'm a novice on Kaggle, so sorry if that's an obvious question, but :)</p>\n<p>Kaggle coding interface allows uploading files. It calls it 'Create a New Dataset', but I see that I can upload any file including weights of a model that I have pre-trained in different environment. In that environment I did not have any limitations like 30 hours of GPU per week or 20 GB of HDD and so on. </p>\n<p>So the question is - am I really allowed to do this move? Is it fair? Or all the training has to be done in Kaggle notebook strictly and I will be penalized for such a submission?</p>",
      "rawMarkdown": "Hello everyone, \n\nI'm a novice on Kaggle, so sorry if that's an obvious question, but :)\n\nKaggle coding interface allows uploading files. It calls it 'Create a New Dataset', but I see that I can upload any file including weights of a model that I have pre-trained in different environment. In that environment I did not have any limitations like 30 hours of GPU per week or 20 GB of HDD and so on. \n\nSo the question is - am I really allowed to do this move? Is it fair? Or all the training has to be done in Kaggle notebook strictly and I will be penalized for such a submission?",
      "votes": 2,
      "replies": [
        {
          "id": 1249693,
          "postDate": "2021-03-23T13:07:32.190Z",
          "content": "<p>Yes, you can certainly train a neural net in a different environment and upload the weights. </p>",
          "rawMarkdown": "Yes, you can certainly train a neural net in a different environment and upload the weights. ",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 1250018,
          "postDate": "2021-03-23T17:56:23.490Z",
          "content": "<p>Thanks for the info!</p>",
          "rawMarkdown": "Thanks for the info!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1196044,
      "postDate": "2021-02-11T08:11:07.537Z",
      "content": "<p>Got a question about desired cell segmentation. Here is the case I have noticed (sample <code>0a83291a-bbc5-11e8-b2bc-ac1f6b6435d0</code>):<br>\n<img src=\"https://i.ibb.co/X8vyfKH/Screenshot-2021-02-11-at-09-56-03.png\" alt=\"adjacent cells\"></p>\n<p>On the picture, there are two nucleis (hence, two cells as I get) are tightly adjacent. What would be a segmentation mask for these cells? And how would that be presented in the submission file? </p>",
      "rawMarkdown": "Got a question about desired cell segmentation. Here is the case I have noticed (sample `0a83291a-bbc5-11e8-b2bc-ac1f6b6435d0`):\n![adjacent cells](https://i.ibb.co/X8vyfKH/Screenshot-2021-02-11-at-09-56-03.png)\n\nOn the picture, there are two nucleis (hence, two cells as I get) are tightly adjacent. What would be a segmentation mask for these cells? And how would that be presented in the submission file? ",
      "votes": 2,
      "replies": [
        {
          "id": 1198647,
          "postDate": "2021-02-13T07:22:00.013Z",
          "content": "<p>Hi! in creating the ground-truth for test set, our annotators used the baseline prediction from <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a> and then manually adjusted every images. So in this case, if the model returns a mask with 2 nuclei that can clearly be separated into 2 cells, an annotator would manually correct the segmentation masks.</p>",
          "rawMarkdown": "Hi! in creating the ground-truth for test set, our annotators used the baseline prediction from [HPACellSegmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation) and then manually adjusted every images. So in this case, if the model returns a mask with 2 nuclei that can clearly be separated into 2 cells, an annotator would manually correct the segmentation masks.",
          "votes": 1
        },
        {
          "id": 1199526,
          "postDate": "2021-02-13T21:34:29.970Z",
          "content": "<p>I am very much enjoying this competition and learning a lot.</p>\n<p>BUT</p>\n<p>I very confused on the metric and how this example would be scored.  Since we don't get to manually correct the segmentation masks using baseline HPACellSegmentation my submission would show this as a single cell.  Just using my limited eyeball it appears that both cells would report containing the same features if segmented into two cells.</p>\n<p>If my submission is a single cell with the correct features will it hurt my score?</p>\n<p>I ask because it almost appears that in addition to creating a model that gets the feature predictions correct I also need to revise the segmentation so that it matches manually annotated cell output from HPACellSegmentation.  </p>\n<p>So I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set.</p>",
          "rawMarkdown": "I am very much enjoying this competition and learning a lot.\n\nBUT\n\nI very confused on the metric and how this example would be scored.  Since we don't get to manually correct the segmentation masks using baseline HPACellSegmentation my submission would show this as a single cell.  Just using my limited eyeball it appears that both cells would report containing the same features if segmented into two cells.\n\nIf my submission is a single cell with the correct features will it hurt my score?\n\nI ask because it almost appears that in addition to creating a model that gets the feature predictions correct I also need to revise the segmentation so that it matches manually annotated cell output from HPACellSegmentation.  \n\nSo I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set."
        },
        {
          "id": 1201097,
          "postDate": "2021-02-15T06:53:31.693Z",
          "content": "<p>The main purpose of the segmentation task is to identify the cell that you are predicting. And the threshold for matching is only 0.6. By using the HPACellSegmentation, you would already be matching ~90% of the cells in test set. You can choose to spend time improving the rest 10% of segmentation (which could lead to a better generalized model), or spend time developing an efficient model for single cell classification.   </p>\n<p><code>So I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set.</code></p>\n<p>If you didn't change the submitted masks, then your LB improvement will be based on your improved model.</p>\n<p>Hope this helps!</p>",
          "rawMarkdown": "The main purpose of the segmentation task is to identify the cell that you are predicting. And the threshold for matching is only 0.6. By using the HPACellSegmentation, you would already be matching ~90% of the cells in test set. You can choose to spend time improving the rest 10% of segmentation (which could lead to a better generalized model), or spend time developing an efficient model for single cell classification.   \n\n`So I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set.`\n\nIf you didn't change the submitted masks, then your LB improvement will be based on your improved model.\n\nHope this helps!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1263843,
      "postDate": "2021-04-05T17:55:47.737Z",
      "content": "<p>Interesting and important initiative! KTH, sounds like an initiative wholly or partly initiated from Sweden? Extra fun if that’s the case, and first time I come across a Swedish host at Kaggle, on the other hand, I've only been at Kaggle roughly a year. I will do my best to put the extra pressure aside as a swede 😉</p>\n<p>Vi får se om det blir en intern tävling om vem blir den bästa svensken/svenska teamet, finns så många tävlingsmänniskor därute som ska tävla om allt 😉 Samarbete inom Kaggle communityn är såklart också viktigt, är så vi blir bättre tillsammans och finner den bästa lösningen. Trevlig fortsättning på påsken! 😊</p>",
      "rawMarkdown": "Interesting and important initiative! KTH, sounds like an initiative wholly or partly initiated from Sweden? Extra fun if that’s the case, and first time I come across a Swedish host at Kaggle, on the other hand, I've only been at Kaggle roughly a year. I will do my best to put the extra pressure aside as a swede 😉\n\nVi får se om det blir en intern tävling om vem blir den bästa svensken/svenska teamet, finns så många tävlingsmänniskor därute som ska tävla om allt 😉 Samarbete inom Kaggle communityn är såklart också viktigt, är så vi blir bättre tillsammans och finner den bästa lösningen. Trevlig fortsättning på påsken! 😊",
      "replies": [
        {
          "id": 1263876,
          "postDate": "2021-04-05T18:31:45.900Z",
          "content": "<p>Yes this is a Kaggle challenge hosted from Sweden! Great to see some local participation :-) <br>\nAlso, this is not the first one hosted from Sweden. We've hosted one before and also the PANDA challenge was in part hosted by researchers at the Karolinska Institute.</p>",
          "rawMarkdown": "Yes this is a Kaggle challenge hosted from Sweden! Great to see some local participation :-) \nAlso, this is not the first one hosted from Sweden. We've hosted one before and also the PANDA challenge was in part hosted by researchers at the Karolinska Institute.",
          "votes": 1
        },
        {
          "id": 1263947,
          "postDate": "2021-04-05T19:19:28.587Z",
          "content": "<p>So it was a swe. hosted comp., that’s great, glad to hear! Alright, PANDA also, that competition I remember well, learned a lot from it, the other HPA was before I joined Kaggle 😊 Good choice to hosting at Kaggle, many brilliant minds here and SOTA AI developments 😊</p>",
          "rawMarkdown": "So it was a swe. hosted comp., that’s great, glad to hear! Alright, PANDA also, that competition I remember well, learned a lot from it, the other HPA was before I joined Kaggle 😊 Good choice to hosting at Kaggle, many brilliant minds here and SOTA AI developments 😊"
        }
      ]
    },
    {
      "id": 1263365,
      "postDate": "2021-04-05T10:56:42.447Z",
      "content": "<p>Dear Organization Team.<br>\nPlease direct us how to link submission notebook to my submission</p>",
      "rawMarkdown": "Dear Organization Team.\nPlease direct us how to link submission notebook to my submission",
      "replies": [
        {
          "id": 1264849,
          "postDate": "2021-04-06T13:16:56.027Z",
          "content": "<p>Hi! I am not sure what you mean.<br>\nThrough my understanding, you can submit solutions with <code>Submit Predictions</code> button, which triggers a pop-up where you can choose which notebook, version and file that you want to submit. This means your submission is already link to the notebook.<br>\nTo make sure, you can first test by copying and submitting a public notebook that works, such as <a href=\"https://www.kaggle.com/samusram/hpa-rgb-model-rgby-cell-level-classification\" target=\"_blank\">this</a><br>\nHope it helps!</p>",
          "rawMarkdown": "Hi! I am not sure what you mean.\nThrough my understanding, you can submit solutions with `Submit Predictions` button, which triggers a pop-up where you can choose which notebook, version and file that you want to submit. This means your submission is already link to the notebook.\nTo make sure, you can first test by copying and submitting a public notebook that works, such as [this](https://www.kaggle.com/samusram/hpa-rgb-model-rgby-cell-level-classification)\nHope it helps!"
        }
      ]
    },
    {
      "id": 1205651,
      "postDate": "2021-02-16T21:02:32.363Z",
      "content": "<p>Hi Emma,</p>\n<p>Thank you for posting this beautiful competition.</p>\n<p>I have a question regarding the \"Mitotic Spindles\" (=Number 11) - if I look in the \"train data\" (i.e. train.csv) I can find 78 IDs which images contain \"Mitotic Spindles\". But if I take a closer look in your database in order to understand how \"Mitotic Spindles\" look like I can on the 78 IDs only detect 14 Spindles. Is there a possibility to get more images which contain \"Mitotic Spindles\"?  </p>",
      "rawMarkdown": "Hi Emma,\n\nThank you for posting this beautiful competition.\n\nI have a question regarding the \"Mitotic Spindles\" (=Number 11) - if I look in the \"train data\" (i.e. train.csv) I can find 78 IDs which images contain \"Mitotic Spindles\". But if I take a closer look in your database in order to understand how \"Mitotic Spindles\" look like I can on the 78 IDs only detect 14 Spindles. Is there a possibility to get more images which contain \"Mitotic Spindles\"?  \n\n",
      "replies": [
        {
          "id": 1213853,
          "postDate": "2021-02-22T11:33:25.260Z",
          "content": "<p>Hi Markus,</p>\n<p>Yes you can find more images with mitotic spindles in the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg?sort=recent-comments\" target=\"_blank\">HPA public image dataset</a> (external data).</p>\n<p>I would also like to explain why you only see few mitotic spindle patterns in the images with such labels. The image-level labels are what we refer to as weak or noisy. During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label.</p>\n<p>The test set consists of images where each single cell has been annotated independently. Hence the accuracy of these labels is much better, and will be correct for each cell in every image. </p>",
          "rawMarkdown": "Hi Markus,\n\nYes you can find more images with mitotic spindles in the [HPA public image dataset](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg?sort=recent-comments) (external data).\n\nI would also like to explain why you only see few mitotic spindle patterns in the images with such labels. The image-level labels are what we refer to as weak or noisy. During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label.\n\nThe test set consists of images where each single cell has been annotated independently. Hence the accuracy of these labels is much better, and will be correct for each cell in every image. ",
          "votes": 3
        },
        {
          "id": 1214410,
          "postDate": "2021-02-22T20:32:19.927Z",
          "content": "<p>Hi Emma,<br>\nThanks, that is very helpful info. A related question: do you use the same hand-labeling accuracy/method  for both the public (31%) and private test sets (69%) ? Or just for the private test set?<br>\nZoltan</p>",
          "rawMarkdown": "Hi Emma,\nThanks, that is very helpful info. A related question: do you use the same hand-labeling accuracy/method  for both the public (31%) and private test sets (69%) ? Or just for the private test set?\nZoltan",
          "isDeleted": true
        },
        {
          "id": 1214418,
          "postDate": "2021-02-22T20:38:10.690Z",
          "content": "<p>The entire test set was hand-labeled the same way, and later split up into the public and test set. So the labeling accuracy should be similar in the private and public test set.</p>",
          "rawMarkdown": "The entire test set was hand-labeled the same way, and later split up into the public and test set. So the labeling accuracy should be similar in the private and public test set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1198906,
      "postDate": "2021-02-13T11:40:19.650Z",
      "content": "<p>Hi,  </p>\n<p>I am new to the challenge and a beginner in Kaggle. I think a challenge similar to this was hosted 2 years back, What are the points of differences between the two competitions?</p>\n<p>My understanding is that the challenge which was hosted 2 years back had a lot of labels and this challenge may not have that much labels as it is stated weakly labelled. The previous challenge had around 28 labels and in this challenge I can see around 18 labels.  </p>",
      "rawMarkdown": "Hi,  \n\nI am new to the challenge and a beginner in Kaggle. I think a challenge similar to this was hosted 2 years back, What are the points of differences between the two competitions?\n\nMy understanding is that the challenge which was hosted 2 years back had a lot of labels and this challenge may not have that much labels as it is stated weakly labelled. The previous challenge had around 28 labels and in this challenge I can see around 18 labels.  ",
      "replies": [
        {
          "id": 1200159,
          "postDate": "2021-02-14T13:13:32.530Z",
          "content": "<p>Hi,</p>\n<p>Welcome to the challenge! You are correct in that we hosted a similar competition two years ago. The similarity is that the competitions include multi-label classification in similar microscope images. The main difference is that you need to classify the label(s) of each cell in every image with only weak image-level training data in this competition, compared to the classification of image-level labels before. You are also correct in that fewer labels are included in this competition. You can learn more about the single cell classifications we are asking for <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">here</a>. </p>",
          "rawMarkdown": "Hi,\n\nWelcome to the challenge! You are correct in that we hosted a similar competition two years ago. The similarity is that the competitions include multi-label classification in similar microscope images. The main difference is that you need to classify the label(s) of each cell in every image with only weak image-level training data in this competition, compared to the classification of image-level labels before. You are also correct in that fewer labels are included in this competition. You can learn more about the single cell classifications we are asking for [here](https://www.kaggle.com/lnhtrang/single-cell-patterns). "
        }
      ]
    },
    {
      "id": 1196809,
      "postDate": "2021-02-11T17:09:47.420Z",
      "content": "<p>How many test images?</p>\n<p>The public test has 2236/4 images.  All my submissions so far have timed out - so I need to speed up my code.  But I cannot find any reference to the expected number of images in the private test.</p>\n<p>Can you provide a number or a ballpark for the number of images that I will need to process?</p>",
      "rawMarkdown": "How many test images?\n\nThe public test has 2236/4 images.  All my submissions so far have timed out - so I need to speed up my code.  But I cannot find any reference to the expected number of images in the private test.\n\nCan you provide a number or a ballpark for the number of images that I will need to process?",
      "replies": [
        {
          "id": 1196963,
          "postDate": "2021-02-11T19:46:05.907Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, I had the same problem previously. I ran multiple experiments and it appears that the private test dataset takes 3-4 times longer to process than the public test-dataset.</p>\n<ul>\n<li><em>This is just a theory, but the public leaderboard is calculated on 31% of the test data… if we assume this is the test data that is provided to us, we can calculate that scoring all of the test data will take</em> <strong><em>3.33 times longer</em></strong> <em>than scoring the public test data.</em></li>\n</ul>\n<p>I recently refactored my code to be much faster to combat this challenge (9-hour submission down to 6-hour). So, please don't hesitate to reach out or <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-inference?scriptVersionId=54064021\" target=\"_blank\"><strong>check my submission kernel</strong></a> for more information.</p>\n<p>I hope this helps!</p>",
          "rawMarkdown": "Hi @pcjimmmy, I had the same problem previously. I ran multiple experiments and it appears that the private test dataset takes 3-4 times longer to process than the public test-dataset.\n* *This is just a theory, but the public leaderboard is calculated on 31% of the test data... if we assume this is the test data that is provided to us, we can calculate that scoring all of the test data will take* ***3.33 times longer*** *than scoring the public test data.*\n\nI recently refactored my code to be much faster to combat this challenge (9-hour submission down to 6-hour). So, please don't hesitate to reach out or [**check my submission kernel**](https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-inference?scriptVersionId=54064021) for more information.\n\nI hope this helps!",
          "votes": 4
        },
        {
          "id": 1197034,
          "postDate": "2021-02-11T21:37:16.553Z",
          "content": "<p>Thanks - I am likely close to sneaking under the wire based on your 3.33 - was worried I needed to trim lots of hours.  </p>\n<p>Will grab your code - already using your data sets :)    Pretty sure without your shared stuff I would not be still active in this competition.  </p>\n<p>Thanks again</p>",
          "rawMarkdown": "Thanks - I am likely close to sneaking under the wire based on your 3.33 - was worried I needed to trim lots of hours.  \n\nWill grab your code - already using your data sets :)    Pretty sure without your shared stuff I would not be still active in this competition.  \n\nThanks again",
          "votes": 1
        },
        {
          "id": 1197059,
          "postDate": "2021-02-11T22:39:24.540Z",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">Darien</a></p>\n<p>Successfully forked your submission kernel to one of my local machines.  Running my latest model now and it looks like about 1 hr 40 minutes is projected run time for test while it was 2 Hrs 37 minutes on this machine for my previous code.</p>\n<p>In addition to faster yours is so much sweeter - I love the demo images.  Thanks again for the share.</p>",
          "rawMarkdown": "[Darien](https://www.kaggle.com/dschettler8845)\n\nSuccessfully forked your submission kernel to one of my local machines.  Running my latest model now and it looks like about 1 hr 40 minutes is projected run time for test while it was 2 Hrs 37 minutes on this machine for my previous code.\n\nIn addition to faster yours is so much sweeter - I love the demo images.  Thanks again for the share.",
          "votes": 1
        },
        {
          "id": 1197062,
          "postDate": "2021-02-11T22:41:23.060Z",
          "content": "<p>You're very welcome! Glad it helped :)</p>",
          "rawMarkdown": "You're very welcome! Glad it helped :)"
        }
      ]
    },
    {
      "id": 1186403,
      "postDate": "2021-02-04T19:57:13.737Z",
      "content": "<p>Is there any way to get the data available with a wget command? I'd like to run analysis on a server that I don't have root access to install the kaggle api.  </p>",
      "rawMarkdown": "Is there any way to get the data available with a wget command? I'd like to run analysis on a server that I don't have root access to install the kaggle api.  ",
      "replies": [
        {
          "id": 1187033,
          "postDate": "2021-02-05T07:15:53.607Z",
          "content": "<p>You don't need root access to install the kaggle api, but yes you can use wget - you start download in the browser and then copy the download URL.</p>",
          "rawMarkdown": "You don't need root access to install the kaggle api, but yes you can use wget - you start download in the browser and then copy the download URL."
        }
      ]
    },
    {
      "id": 1178250,
      "postDate": "2021-01-30T18:54:22.830Z",
      "content": "<p>I asked this as a separate discussion topic so feel free to reply there as well.</p>\n<p><b></b></p>\n<blockquote>\n  <p>As a point of curiosity, can we find out (or do we know) what the various \"proteins of interest\" are that we are identifying?</p>\n  <p>For instance, I would assume the proteins responsible for Apoptosis might be of interest.</p>\n</blockquote>\n<p></p>",
      "rawMarkdown": "I asked this as a separate discussion topic so feel free to reply there as well.\n\n<b>\n\n>As a point of curiosity, can we find out (or do we know) what the various \"proteins of interest\" are that we are identifying?\n>\n> For instance, I would assume the proteins responsible for Apoptosis might be of interest.\n\n</b>",
      "replies": [
        {
          "id": 1178335,
          "postDate": "2021-01-30T19:49:44.003Z",
          "content": "<p>Good question! We cannot tell you exactly what the proteins are now. All we can say is that in this competition a large part of all human proteins are included in either the train or test set. These proteins are involved in a variety of biological processes.</p>",
          "rawMarkdown": "Good question! We cannot tell you exactly what the proteins are now. All we can say is that in this competition a large part of all human proteins are included in either the train or test set. These proteins are involved in a variety of biological processes.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1177418,
      "postDate": "2021-01-30T09:12:49.043Z",
      "content": "<p>Thank you for hosting this competition.<br>\nI found a tiny bug.<br>\nYou can find the following information on the <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">evaluation page</a>.</p>\n<pre>ImageID,ImageWidth,ImageHeight,PredictionString\n</pre>\n<p>But It will cause <strong>Submission Scoring Error</strong>.<br>\nIt seems that header should be:</p>\n<pre>ID,ImageWidth,ImageHeight,PredictionString\n</pre>",
      "rawMarkdown": "Thank you for hosting this competition.\nI found a tiny bug.\nYou can find the following information on the [evaluation page](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation).\n<pre>\nImageID,ImageWidth,ImageHeight,PredictionString\n</pre>\n\nBut It will cause **Submission Scoring Error**.\nIt seems that header should be:\n<pre>\nID,ImageWidth,ImageHeight,PredictionString\n</pre>",
      "replies": [
        {
          "id": 1177672,
          "postDate": "2021-01-30T12:57:03.993Z",
          "content": "<p>Thanks for noticing that. The format in <code>sample_submission.csv</code> is correct. I have updated Evaluation page.</p>",
          "rawMarkdown": "Thanks for noticing that. The format in `sample_submission.csv` is correct. I have updated Evaluation page.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1174398,
      "postDate": "2021-01-28T12:50:47.260Z",
      "content": "<p>The link from the email sent yesterday 2021-01-27 to join this competition takes you to the wrong page(<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a>)</p>",
      "rawMarkdown": "The link from the email sent yesterday 2021-01-27 to join this competition takes you to the wrong page(https://www.kaggle.com/c/human-protein-atlas-image-classification)",
      "replies": [
        {
          "id": 1174427,
          "postDate": "2021-01-28T13:21:46.047Z",
          "content": "<p>We had not noticed this unfortunate mistake. Thank you for letting us know.</p>",
          "rawMarkdown": "We had not noticed this unfortunate mistake. Thank you for letting us know."
        }
      ]
    },
    {
      "id": 1173711,
      "postDate": "2021-01-28T04:35:07.830Z",
      "content": "<p>Hi Emma, can a participant use Apache Spark in one's Kaggle Notebook for this competition ? Let me know. <br>\nThank you,<br>\nBharat</p>",
      "rawMarkdown": "Hi Emma, can a participant use Apache Spark in one's Kaggle Notebook for this competition ? Let me know. \nThank you,\nBharat",
      "replies": [
        {
          "id": 1174493,
          "postDate": "2021-01-28T14:11:40.157Z",
          "content": "<p>I see no reason to disallow Spark in the competition. I have not used it in the context of Kaggle notebooks, but assuming the Spark cluster manager can be set up to run for a code competition submission it should be fine.</p>",
          "rawMarkdown": "I see no reason to disallow Spark in the competition. I have not used it in the context of Kaggle notebooks, but assuming the Spark cluster manager can be set up to run for a code competition submission it should be fine.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1193657,
      "postDate": "2021-02-09T19:12:24.593Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1197538,
      "postDate": "2021-02-12T08:23:04.970Z",
      "content": "<p>Thank you for hosting this competition</p>",
      "rawMarkdown": "Thank you for hosting this competition"
    },
    {
      "id": 1187246,
      "postDate": "2021-02-05T10:09:22.987Z",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much"
    }
  ],
  "comments": [
    {
      "id": 1248993,
      "author_name": "ONODERA",
      "author_url": "",
      "post_date": "2021-03-23T02:13:57.300000",
      "content": "<p>I compared label with previous one, and found weird labels. <br>\ne.g.</p>\n<ol>\n<li>ID: <strong>6795f08c-bb99-11e8-b2b9-ac1f6b6435d0</strong>'s label is Nucleoplasm and Nuclear membrane, but previous train set is only Nuclear membrane.</li>\n<li>ID: <strong>cdcfdff6-bb9a-11e8-b2b9-ac1f6b6435d0</strong>'s label is Negative, but previous train set is Nucleoplasm.</li>\n</ol>\n<p>Which is correct?</p>\n<p>Please take a look into this.<br>\n<a href=\"https://www.kaggle.com/onodera/compare-with-previous-hpa-competition\" target=\"_blank\">https://www.kaggle.com/onodera/compare-with-previous-hpa-competition</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1249250,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-03-23T07:48:34.327000",
          "content": "<p>Hi! The current labels are correct. These are updated through different HPA releases. (1) <code>Nucleoplasm</code> is added and (2) is found to be unspecific staining, which belongs to <code>Negative</code> class. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1249255,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2021-03-23T07:54:48.457000",
          "content": "<p>Thank you for quick reply! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1201117,
      "author_name": "satvik patel",
      "author_url": "",
      "post_date": "2021-02-15T07:23:57.597000",
      "content": "<p>Hello M'am,<br>\n Can you describe what is the predictionstring : \"0 1 eNoLCAgIMAEABJkBdQ==\" include with example?<br>\nIs this string encripted?<br>\nWhat is the \"0 1\" , \"eNoLCAgIMAEABJkBdQ==\"?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1201511,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-15T12:57:11.283000",
          "content": "<p>Hi there. Please see <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\"><strong>this page</strong></a> for more information on the submission format details.</p>\n<hr>\n<p><strong><code>0</code></strong> is the predicted class-label corresponding to <strong><code>Nucleoplasm</code></strong><br><br>\n<strong><code>1</code></strong> is the <strong>confidence</strong> of the predicted class label <strong><code>Nucleoplasm</code></strong><br><br>\n<strong><code>eNoLCAgIMA...EABJkBdQ==</code></strong> is the RLE, ZLIB compressed and Base64 encoded segmentation mask for the cell we are predicting on (that we are saying has the <strong><code>Nucleoplasm</code></strong> class).</p>\n<blockquote>\n  <p><strong><em>FROM THE LINK ABOVE</em></strong></p>\n  <p><em>For each image in the test set, you must predict a list of instance segmentation masks and their associated detection score (Confidence)</em></p>\n  <p><em>The binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format as EncodedMask. Specifically, we use the Coco masks RLE encoding/decoding (see the encode method of COCO’s mask API), the zlib compression/decompression (RFC1950), and vanilla base64 encoding.</em></p>\n</blockquote>\n<hr>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1205650,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-16T21:01:46.213000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1174970,
      "author_name": "Konstantin Lopukhin",
      "author_url": "",
      "post_date": "2021-01-28T20:09:17.467000",
      "content": "<blockquote>\n  <p>There are many options of pretrained models that perform quite well on our images (eg. Cellpose or HPACellSegmentation). Please check out this notebook to see how you can make use of public HPA data and HPACellSegmentation model.</p>\n</blockquote>\n<p>That's really helpful, thanks. Was HPACellSegmentation model used to create the ground truth cell segmentation for this competition, with classes assigned by annotators?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1174981,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-28T20:17:22.603000",
          "content": "<p>HPACellSegmentation was used as a basis followed my manual correction of the masks for the ground truth. Classes per single cell were subsequently assigned by annotators.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1177078,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2021-01-30T03:24:01.597000",
          "content": "<p>I also created this notebook that allows to use the HPACellSegmentation with the constraints of the competition with no internet - <a href=\"https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet\" target=\"_blank\">https://www.kaggle.com/rdizzl3/hpa-segmentation-masks-no-internet</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1173418,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-01-27T21:28:11.227000",
      "content": "<p>Thank you for hosting this competition. It is fascinating and I’m excited to learn!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1174983,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-28T20:17:56.233000",
          "content": "<p>Thank you! And best of luck solving the problem.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1173396,
      "author_name": "Nigel A. R. Henry",
      "author_url": "",
      "post_date": "2021-01-27T21:03:07.487000",
      "content": "<p>It is strange to me that:</p>\n<blockquote>\n  <p>External data, freely &amp; publicly available, is allowed. This includes pre-trained models.<br>\n  <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements</a></p>\n</blockquote>\n<p>Since a \"pre-trained model\" could embed learning from non-freely and non-publicly available data sets… how does that work?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1174106,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-01-28T09:14:53.780000",
          "content": "<p>The rule is, as far as I know, standard for these kinds of competitions to make sure there is equal access to data-sets and pre-trained models.</p>\n<p>As long as the pre-trained models are freely available (and their licenses allow for usage in open source contexts) it should be fine to use regardless of them having been embedded information from non-freely available datasets.</p>\n<p>However, I will ask members of Kaggle staff to clarify the idea of the rule.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1176707,
          "author_name": "Maggie",
          "author_url": "",
          "post_date": "2021-01-29T18:32:28.500000",
          "content": "<p>We agree with Casper's response, pre-trained models are allowed but the expectation is that the pre-trained models are freely available to all participants. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1245260,
      "author_name": "Yuriy",
      "author_url": "",
      "post_date": "2021-03-19T16:14:52.860000",
      "content": "<p>Hello everyone, </p>\n<p>I'm a novice on Kaggle, so sorry if that's an obvious question, but :)</p>\n<p>Kaggle coding interface allows uploading files. It calls it 'Create a New Dataset', but I see that I can upload any file including weights of a model that I have pre-trained in different environment. In that environment I did not have any limitations like 30 hours of GPU per week or 20 GB of HDD and so on. </p>\n<p>So the question is - am I really allowed to do this move? Is it fair? Or all the training has to be done in Kaggle notebook strictly and I will be penalized for such a submission?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1249693,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-03-23T13:07:32.190000",
          "content": "<p>Yes, you can certainly train a neural net in a different environment and upload the weights. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1250018,
          "author_name": "Yuriy",
          "author_url": "",
          "post_date": "2021-03-23T17:56:23.490000",
          "content": "<p>Thanks for the info!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1196044,
      "author_name": "Roman Glushko 🦁",
      "author_url": "",
      "post_date": "2021-02-11T08:11:07.537000",
      "content": "<p>Got a question about desired cell segmentation. Here is the case I have noticed (sample <code>0a83291a-bbc5-11e8-b2bc-ac1f6b6435d0</code>):<br>\n<img src=\"https://i.ibb.co/X8vyfKH/Screenshot-2021-02-11-at-09-56-03.png\" alt=\"adjacent cells\"></p>\n<p>On the picture, there are two nucleis (hence, two cells as I get) are tightly adjacent. What would be a segmentation mask for these cells? And how would that be presented in the submission file? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1198647,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-02-13T07:22:00.013000",
          "content": "<p>Hi! in creating the ground-truth for test set, our annotators used the baseline prediction from <a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a> and then manually adjusted every images. So in this case, if the model returns a mask with 2 nuclei that can clearly be separated into 2 cells, an annotator would manually correct the segmentation masks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1199526,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-02-13T21:34:29.970000",
          "content": "<p>I am very much enjoying this competition and learning a lot.</p>\n<p>BUT</p>\n<p>I very confused on the metric and how this example would be scored.  Since we don't get to manually correct the segmentation masks using baseline HPACellSegmentation my submission would show this as a single cell.  Just using my limited eyeball it appears that both cells would report containing the same features if segmented into two cells.</p>\n<p>If my submission is a single cell with the correct features will it hurt my score?</p>\n<p>I ask because it almost appears that in addition to creating a model that gets the feature predictions correct I also need to revise the segmentation so that it matches manually annotated cell output from HPACellSegmentation.  </p>\n<p>So I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1201097,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-02-15T06:53:31.693000",
          "content": "<p>The main purpose of the segmentation task is to identify the cell that you are predicting. And the threshold for matching is only 0.6. By using the HPACellSegmentation, you would already be matching ~90% of the cells in test set. You can choose to spend time improving the rest 10% of segmentation (which could lead to a better generalized model), or spend time developing an efficient model for single cell classification.   </p>\n<p><code>So I am at a loss to decide if a small improvement in my LB score is due to better model or due to better matching of segmentation to the manually annotated test set.</code></p>\n<p>If you didn't change the submitted masks, then your LB improvement will be based on your improved model.</p>\n<p>Hope this helps!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1263843,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2021-04-05T17:55:47.737000",
      "content": "<p>Interesting and important initiative! KTH, sounds like an initiative wholly or partly initiated from Sweden? Extra fun if that’s the case, and first time I come across a Swedish host at Kaggle, on the other hand, I've only been at Kaggle roughly a year. I will do my best to put the extra pressure aside as a swede 😉</p>\n<p>Vi får se om det blir en intern tävling om vem blir den bästa svensken/svenska teamet, finns så många tävlingsmänniskor därute som ska tävla om allt 😉 Samarbete inom Kaggle communityn är såklart också viktigt, är så vi blir bättre tillsammans och finner den bästa lösningen. Trevlig fortsättning på påsken! 😊</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1263876,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-04-05T18:31:45.900000",
          "content": "<p>Yes this is a Kaggle challenge hosted from Sweden! Great to see some local participation :-) <br>\nAlso, this is not the first one hosted from Sweden. We've hosted one before and also the PANDA challenge was in part hosted by researchers at the Karolinska Institute.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1263947,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2021-04-05T19:19:28.587000",
          "content": "<p>So it was a swe. hosted comp., that’s great, glad to hear! Alright, PANDA also, that competition I remember well, learned a lot from it, the other HPA was before I joined Kaggle 😊 Good choice to hosting at Kaggle, many brilliant minds here and SOTA AI developments 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1263365,
      "author_name": "MDL",
      "author_url": "",
      "post_date": "2021-04-05T10:56:42.447000",
      "content": "<p>Dear Organization Team.<br>\nPlease direct us how to link submission notebook to my submission</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1264849,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-06T13:16:56.027000",
          "content": "<p>Hi! I am not sure what you mean.<br>\nThrough my understanding, you can submit solutions with <code>Submit Predictions</code> button, which triggers a pop-up where you can choose which notebook, version and file that you want to submit. This means your submission is already link to the notebook.<br>\nTo make sure, you can first test by copying and submitting a public notebook that works, such as <a href=\"https://www.kaggle.com/samusram/hpa-rgb-model-rgby-cell-level-classification\" target=\"_blank\">this</a><br>\nHope it helps!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1205651,
      "author_name": "Markus Wagner",
      "author_url": "",
      "post_date": "2021-02-16T21:02:32.363000",
      "content": "<p>Hi Emma,</p>\n<p>Thank you for posting this beautiful competition.</p>\n<p>I have a question regarding the \"Mitotic Spindles\" (=Number 11) - if I look in the \"train data\" (i.e. train.csv) I can find 78 IDs which images contain \"Mitotic Spindles\". But if I take a closer look in your database in order to understand how \"Mitotic Spindles\" look like I can on the 78 IDs only detect 14 Spindles. Is there a possibility to get more images which contain \"Mitotic Spindles\"?  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1213853,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-02-22T11:33:25.260000",
          "content": "<p>Hi Markus,</p>\n<p>Yes you can find more images with mitotic spindles in the <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg?sort=recent-comments\" target=\"_blank\">HPA public image dataset</a> (external data).</p>\n<p>I would also like to explain why you only see few mitotic spindle patterns in the images with such labels. The image-level labels are what we refer to as weak or noisy. During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label.</p>\n<p>The test set consists of images where each single cell has been annotated independently. Hence the accuracy of these labels is much better, and will be correct for each cell in every image. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1214410,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-22T20:32:19.927000",
          "content": "<p>Hi Emma,<br>\nThanks, that is very helpful info. A related question: do you use the same hand-labeling accuracy/method  for both the public (31%) and private test sets (69%) ? Or just for the private test set?<br>\nZoltan</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1214418,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-02-22T20:38:10.690000",
          "content": "<p>The entire test set was hand-labeled the same way, and later split up into the public and test set. So the labeling accuracy should be similar in the private and public test set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1198906,
      "author_name": "RSASHWIN",
      "author_url": "",
      "post_date": "2021-02-13T11:40:19.650000",
      "content": "<p>Hi,  </p>\n<p>I am new to the challenge and a beginner in Kaggle. I think a challenge similar to this was hosted 2 years back, What are the points of differences between the two competitions?</p>\n<p>My understanding is that the challenge which was hosted 2 years back had a lot of labels and this challenge may not have that much labels as it is stated weakly labelled. The previous challenge had around 28 labels and in this challenge I can see around 18 labels.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1200159,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-02-14T13:13:32.530000",
          "content": "<p>Hi,</p>\n<p>Welcome to the challenge! You are correct in that we hosted a similar competition two years ago. The similarity is that the competitions include multi-label classification in similar microscope images. The main difference is that you need to classify the label(s) of each cell in every image with only weak image-level training data in this competition, compared to the classification of image-level labels before. You are also correct in that fewer labels are included in this competition. You can learn more about the single cell classifications we are asking for <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">here</a>. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1196809,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2021-02-11T17:09:47.420000",
      "content": "<p>How many test images?</p>\n<p>The public test has 2236/4 images.  All my submissions so far have timed out - so I need to speed up my code.  But I cannot find any reference to the expected number of images in the private test.</p>\n<p>Can you provide a number or a ballpark for the number of images that I will need to process?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1196963,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-11T19:46:05.907000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a>, I had the same problem previously. I ran multiple experiments and it appears that the private test dataset takes 3-4 times longer to process than the public test-dataset.</p>\n<ul>\n<li><em>This is just a theory, but the public leaderboard is calculated on 31% of the test data… if we assume this is the test data that is provided to us, we can calculate that scoring all of the test data will take</em> <strong><em>3.33 times longer</em></strong> <em>than scoring the public test data.</em></li>\n</ul>\n<p>I recently refactored my code to be much faster to combat this challenge (9-hour submission down to 6-hour). So, please don't hesitate to reach out or <a href=\"https://www.kaggle.com/dschettler8845/hpa-cellwise-classification-inference?scriptVersionId=54064021\" target=\"_blank\"><strong>check my submission kernel</strong></a> for more information.</p>\n<p>I hope this helps!</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1197034,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-02-11T21:37:16.553000",
          "content": "<p>Thanks - I am likely close to sneaking under the wire based on your 3.33 - was worried I needed to trim lots of hours.  </p>\n<p>Will grab your code - already using your data sets :)    Pretty sure without your shared stuff I would not be still active in this competition.  </p>\n<p>Thanks again</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1197059,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-02-11T22:39:24.540000",
          "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">Darien</a></p>\n<p>Successfully forked your submission kernel to one of my local machines.  Running my latest model now and it looks like about 1 hr 40 minutes is projected run time for test while it was 2 Hrs 37 minutes on this machine for my previous code.</p>\n<p>In addition to faster yours is so much sweeter - I love the demo images.  Thanks again for the share.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1197062,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-11T22:41:23.060000",
          "content": "<p>You're very welcome! Glad it helped :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1186403,
      "author_name": "KevinLee",
      "author_url": "",
      "post_date": "2021-02-04T19:57:13.737000",
      "content": "<p>Is there any way to get the data available with a wget command? I'd like to run analysis on a server that I don't have root access to install the kaggle api.  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1187033,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2021-02-05T07:15:53.607000",
          "content": "<p>You don't need root access to install the kaggle api, but yes you can use wget - you start download in the browser and then copy the download URL.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1178250,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-01-30T18:54:22.830000",
      "content": "<p>I asked this as a separate discussion topic so feel free to reply there as well.</p>\n<p><b></b></p>\n<blockquote>\n  <p>As a point of curiosity, can we find out (or do we know) what the various \"proteins of interest\" are that we are identifying?</p>\n  <p>For instance, I would assume the proteins responsible for Apoptosis might be of interest.</p>\n</blockquote>\n<p></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1178335,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-30T19:49:44.003000",
          "content": "<p>Good question! We cannot tell you exactly what the proteins are now. All we can say is that in this competition a large part of all human proteins are included in either the train or test set. These proteins are involved in a variety of biological processes.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1177418,
      "author_name": "tito",
      "author_url": "",
      "post_date": "2021-01-30T09:12:49.043000",
      "content": "<p>Thank you for hosting this competition.<br>\nI found a tiny bug.<br>\nYou can find the following information on the <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">evaluation page</a>.</p>\n<pre>ImageID,ImageWidth,ImageHeight,PredictionString\n</pre>\n<p>But It will cause <strong>Submission Scoring Error</strong>.<br>\nIt seems that header should be:</p>\n<pre>ID,ImageWidth,ImageHeight,PredictionString\n</pre>",
      "votes": 0,
      "replies": [
        {
          "id": 1177672,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-30T12:57:03.993000",
          "content": "<p>Thanks for noticing that. The format in <code>sample_submission.csv</code> is correct. I have updated Evaluation page.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1174398,
      "author_name": "rilesdg3",
      "author_url": "",
      "post_date": "2021-01-28T12:50:47.260000",
      "content": "<p>The link from the email sent yesterday 2021-01-27 to join this competition takes you to the wrong page(<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification\" target=\"_blank\">https://www.kaggle.com/c/human-protein-atlas-image-classification</a>)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1174427,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-28T13:21:46.047000",
          "content": "<p>We had not noticed this unfortunate mistake. Thank you for letting us know.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1173711,
      "author_name": "BharatSS",
      "author_url": "",
      "post_date": "2021-01-28T04:35:07.830000",
      "content": "<p>Hi Emma, can a participant use Apache Spark in one's Kaggle Notebook for this competition ? Let me know. <br>\nThank you,<br>\nBharat</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1174493,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-01-28T14:11:40.157000",
          "content": "<p>I see no reason to disallow Spark in the competition. I have not used it in the context of Kaggle notebooks, but assuming the Spark cluster manager can be set up to run for a code competition submission it should be fine.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1193657,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-09T19:12:24.593000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1197538,
      "author_name": "Asuragan",
      "author_url": "",
      "post_date": "2021-02-12T08:23:04.970000",
      "content": "<p>Thank you for hosting this competition</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1187246,
      "author_name": "Maria Dyakova",
      "author_url": "",
      "post_date": "2021-02-05T10:09:22.987000",
      "content": "<p>Thank you very much</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1171981": "Welcome to the Human Protein Atlas single cell classification challenge!\n \nIn this competition, we ask you to help us classify subcellular protein localization patterns of single cells in microscope images. Solving the single-cell image classification challenge will help us determine precise locations for all human proteins in each individual cell in our large collection of open access images. This will contribute to a better understanding of functional differences between otherwise seemingly identical cells.\n\nYou’re given a set of training images, where each image contains multiple cells that may vary in their protein localization patterns. Importantly, we provide rough labels for the protein pattern or patterns for each image, which may not be correct for every cell in the image. Your model will need to figure out how to train with our weakly labeled data and predict precise labels for each cell in the test image. In addition to these data, you’re given access to the publicly available images within the Human Protein Atlas, should you want more training data. In order to perform well in this challenge, you may need a good segmentation model. There are many options of pretrained models that perform quite well on our images (eg. [Cellpose](https://www.cellpose.org/) or [HPACellSegmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)). Please check out [this notebook](https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg) to see how you can make use of public HPA data and HPACellSegmentation model.\n\nPlease also note that this is a code competition where a large part of test images are hidden. To gain a better understanding of the task, check out our [guide explaining the patterns](https://www.kaggle.com/lnhtrang/single-cell-patterns) and cells in our microscope images.\n\nJust like in our [previous competition](https://www.kaggle.com/c/human-protein-atlas-image-classification), we present here a multi-label problem. The main difference is that we are asking you to classify the label(s) of each cell in every image with only weak image-level training data, compared to the classification of image-level labels before. With this unique competition setup, we hope to see novel solutions that are capable of utilizing large amounts of unlabelled or weakly labelled data to generate robust, discriminative features. We envision that these solutions may help us understand protein localization patterns across a multitude of human cells. We have demonstrated the promise of both citizen science and artificial intelligence approaches in describing the location of human proteins in microscope images. However, none map the protein distribution at a single-cell level or match human level performance ([Sullivan et al, 2018;](https://www.nature.com/articles/nbt.4225) [Thul et al, 2017](https://science.sciencemag.org/content/356/6340/eaal3321.full); [Ouyang et al, 2020](https://www.nature.com/articles/s41592-019-0658-6/)), something we believe you can help us with!\n\nWe in the [Human Protein Atlas](https://www.proteinatlas.org/) team are true believers in open science and the power of crowdsourcing. We are excited to see how the Kaggle community will tackle this problem and what new models and innovative solutions will be developed. \n\nAlong with me are the team that helped make this possible. You'll see us around on the Discussion boards to answer any questions you may have. We want to help you to develop the best models possible!\n\nGood luck, and have fun!\n",
    "1248993": "I compared label with previous one, and found weird labels. \ne.g.\n1. ID: **6795f08c-bb99-11e8-b2b9-ac1f6b6435d0**'s label is Nucleoplasm and Nuclear membrane, but previous train set is only Nuclear membrane.\n2. ID: **cdcfdff6-bb9a-11e8-b2b9-ac1f6b6435d0**'s label is Negative, but previous train set is Nucleoplasm.\n\nWhich is correct?\n\nPlease take a look into this.\nhttps://www.kaggle.com/onodera/compare-with-previous-hpa-competition",
    "1201117": "Hello M'am,\n Can you describe what is the predictionstring : \"0 1 eNoLCAgIMAEABJkBdQ==\" include with example?\nIs this string encripted?\nWhat is the \"0 1\" , \"eNoLCAgIMAEABJkBdQ==\"?",
    "1174970": "> There are many options of pretrained models that perform quite well on our images (eg. Cellpose or HPACellSegmentation). Please check out this notebook to see how you can make use of public HPA data and HPACellSegmentation model.\n\nThat's really helpful, thanks. Was HPACellSegmentation model used to create the ground truth cell segmentation for this competition, with classes assigned by annotators?",
    "1173418": "Thank you for hosting this competition. It is fascinating and I’m excited to learn!\n",
    "1173396": "It is strange to me that:\n> External data, freely & publicly available, is allowed. This includes pre-trained models.\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/overview/code-requirements\n\nSince a \"pre-trained model\" could embed learning from non-freely and non-publicly available data sets... how does that work?",
    "1245260": "Hello everyone, \n\nI'm a novice on Kaggle, so sorry if that's an obvious question, but :)\n\nKaggle coding interface allows uploading files. It calls it 'Create a New Dataset', but I see that I can upload any file including weights of a model that I have pre-trained in different environment. In that environment I did not have any limitations like 30 hours of GPU per week or 20 GB of HDD and so on. \n\nSo the question is - am I really allowed to do this move? Is it fair? Or all the training has to be done in Kaggle notebook strictly and I will be penalized for such a submission?",
    "1196044": "Got a question about desired cell segmentation. Here is the case I have noticed (sample `0a83291a-bbc5-11e8-b2bc-ac1f6b6435d0`):\n![adjacent cells](https://i.ibb.co/X8vyfKH/Screenshot-2021-02-11-at-09-56-03.png)\n\nOn the picture, there are two nucleis (hence, two cells as I get) are tightly adjacent. What would be a segmentation mask for these cells? And how would that be presented in the submission file? ",
    "1263843": "Interesting and important initiative! KTH, sounds like an initiative wholly or partly initiated from Sweden? Extra fun if that’s the case, and first time I come across a Swedish host at Kaggle, on the other hand, I've only been at Kaggle roughly a year. I will do my best to put the extra pressure aside as a swede 😉\n\nVi får se om det blir en intern tävling om vem blir den bästa svensken/svenska teamet, finns så många tävlingsmänniskor därute som ska tävla om allt 😉 Samarbete inom Kaggle communityn är såklart också viktigt, är så vi blir bättre tillsammans och finner den bästa lösningen. Trevlig fortsättning på påsken! 😊",
    "1263365": "Dear Organization Team.\nPlease direct us how to link submission notebook to my submission",
    "1205651": "Hi Emma,\n\nThank you for posting this beautiful competition.\n\nI have a question regarding the \"Mitotic Spindles\" (=Number 11) - if I look in the \"train data\" (i.e. train.csv) I can find 78 IDs which images contain \"Mitotic Spindles\". But if I take a closer look in your database in order to understand how \"Mitotic Spindles\" look like I can on the 78 IDs only detect 14 Spindles. Is there a possibility to get more images which contain \"Mitotic Spindles\"?  \n\n",
    "1198906": "Hi,  \n\nI am new to the challenge and a beginner in Kaggle. I think a challenge similar to this was hosted 2 years back, What are the points of differences between the two competitions?\n\nMy understanding is that the challenge which was hosted 2 years back had a lot of labels and this challenge may not have that much labels as it is stated weakly labelled. The previous challenge had around 28 labels and in this challenge I can see around 18 labels.  ",
    "1196809": "How many test images?\n\nThe public test has 2236/4 images.  All my submissions so far have timed out - so I need to speed up my code.  But I cannot find any reference to the expected number of images in the private test.\n\nCan you provide a number or a ballpark for the number of images that I will need to process?",
    "1186403": "Is there any way to get the data available with a wget command? I'd like to run analysis on a server that I don't have root access to install the kaggle api.  ",
    "1178250": "I asked this as a separate discussion topic so feel free to reply there as well.\n\n<b>\n\n>As a point of curiosity, can we find out (or do we know) what the various \"proteins of interest\" are that we are identifying?\n>\n> For instance, I would assume the proteins responsible for Apoptosis might be of interest.\n\n</b>",
    "1177418": "Thank you for hosting this competition.\nI found a tiny bug.\nYou can find the following information on the [evaluation page](https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation).\n<pre>\nImageID,ImageWidth,ImageHeight,PredictionString\n</pre>\n\nBut It will cause **Submission Scoring Error**.\nIt seems that header should be:\n<pre>\nID,ImageWidth,ImageHeight,PredictionString\n</pre>",
    "1174398": "The link from the email sent yesterday 2021-01-27 to join this competition takes you to the wrong page(https://www.kaggle.com/c/human-protein-atlas-image-classification)",
    "1173711": "Hi Emma, can a participant use Apache Spark in one's Kaggle Notebook for this competition ? Let me know. \nThank you,\nBharat",
    "1193657": "",
    "1197538": "Thank you for hosting this competition",
    "1187246": "Thank you very much"
  }
}