{
  "id": 215141,
  "title": "What does the ground truth look like?",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/215141",
  "author_name": "Peiyuan Liao",
  "post_date": "2021-01-28T19:26:20.507000",
  "votes": 26,
  "comment_count": 25,
  "views": 0,
  "content": "<p>e.g. if we want to detect nuclear speckles, do we segment the entire cell and do the labeling, or do we just segment the small dots?</p>\n<p>In addition, should we predict individual masks for individual cells, or a single mask for the entire class?</p>",
  "messages": [
    {
      "id": 1174914,
      "postDate": "2021-01-28T19:26:20.507Z",
      "content": "<p>e.g. if we want to detect nuclear speckles, do we segment the entire cell and do the labeling, or do we just segment the small dots?</p>\n<p>In addition, should we predict individual masks for individual cells, or a single mask for the entire class?</p>",
      "rawMarkdown": "e.g. if we want to detect nuclear speckles, do we segment the entire cell and do the labeling, or do we just segment the small dots?\n\nIn addition, should we predict individual masks for individual cells, or a single mask for the entire class?",
      "votes": 25
    },
    {
      "id": 1174999,
      "postDate": "2021-01-28T20:30:13.753Z",
      "content": "<p>I believe we segment the entire cell as it says in the evaluation that a single instance of segmentation can have multiple labels.</p>\n<p>As it is instance segmentation you will need individual masks for individual cells.</p>",
      "rawMarkdown": "I believe we segment the entire cell as it says in the evaluation that a single instance of segmentation can have multiple labels.\n\nAs it is instance segmentation you will need individual masks for individual cells.",
      "votes": 1,
      "replies": [
        {
          "id": 1207396,
          "postDate": "2021-02-17T20:26:46.437Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, what do you mean by individual masks for individual cells? </p>\n<p>Say in the given image there are 10 cells and we get a segmentation mask for each of them. We will have to denote mask for say cell one as 0 and cell two a 1 and so on in a preferred direction? </p>\n<p>Can we put it like this - do semantic segmentation of the image(binary - background and cell); for each cell predict the label? </p>",
          "rawMarkdown": "Hey @dschettler8845, what do you mean by individual masks for individual cells? \n\nSay in the given image there are 10 cells and we get a segmentation mask for each of them. We will have to denote mask for say cell one as 0 and cell two a 1 and so on in a preferred direction? \n\nCan we put it like this - do semantic segmentation of the image(binary - background and cell); for each cell predict the label? ",
          "votes": 1
        },
        {
          "id": 1207547,
          "postDate": "2021-02-17T23:08:07.910Z",
          "content": "<p>Mostly correct.</p>\n<p>You will segment the image and get 10 masks (let's say each cell has a different value like you indicated… background is 0, cell 1 has a value of 1, cell 2 has a value of 2, etc.).</p>\n<p>Then you will predict, for a given cell mask, what are the classes present in that cell. By individual mask, I am describing the single value of the mask corresponding to that cell.</p>\n<p>Let's see the example through. For our ten cells, we have 10 masks.</p>\n<p>We can get each mask by itself by performing something like this:</p>\n<hr>\n<p><b></b></p>\n<pre><code># all_cell_masks is the numpy array containing all segmented cells\n# where each cell is indicated by a different numeric value\ncell_1 = np.where(all_cell_masks==1, 1, 0)\ncell_2 = np.where(all_cell_masks==2, 1, 0)\n...\ncell_10 = np.where(all_cell_masks==10, 1, 0)\n</code></pre>\n<p></p>\n<hr>\n<p>Now we have ten binary mask arrays, one for each cell. Now we need to make a prediction (possibly multiple predictions), with a confidence score, on each mask. Now the binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format in the prediction string.</p>\n<hr>\n<p><b></b></p>\n<table>\n<thead>\n<tr>\n<th>ImageID</th>\n<th>ImageWidth</th>\n<th>ImageHeight</th>\n<th>PredictionString</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>IMG_ID</code></td>\n<td><code>WIDTH</code></td>\n<td><code>HEIGHT</code></td>\n<td>\" <code>LABEL1</code> <code>CONFIDENCE1</code> <code>MASK1</code> <code>LABEL2</code> <code>CONFIDENCE2</code> <code>MASK2</code> <code>...</code> <code>LABEL10</code> <code>CONFIDENCE10</code> <code>MASK10</code> \"</td>\n</tr>\n</tbody>\n</table>\n<p></p>\n<hr>\n<p>And to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.</p>\n<p>Hopefully, this clears things up. Please let me know if you have any other questions or anything is unclear.</p>",
          "rawMarkdown": "Mostly correct.\n\nYou will segment the image and get 10 masks (let's say each cell has a different value like you indicated... background is 0, cell 1 has a value of 1, cell 2 has a value of 2, etc.).\n\nThen you will predict, for a given cell mask, what are the classes present in that cell. By individual mask, I am describing the single value of the mask corresponding to that cell.\n\nLet's see the example through. For our ten cells, we have 10 masks.\n\nWe can get each mask by itself by performing something like this:\n\n---\n\n<b>\n\n```python\n# all_cell_masks is the numpy array containing all segmented cells\n# where each cell is indicated by a different numeric value\ncell_1 = np.where(all_cell_masks==1, 1, 0)\ncell_2 = np.where(all_cell_masks==2, 1, 0)\n...\ncell_10 = np.where(all_cell_masks==10, 1, 0)\n```\n\n</b>\n\n---\n\nNow we have ten binary mask arrays, one for each cell. Now we need to make a prediction (possibly multiple predictions), with a confidence score, on each mask. Now the binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format in the prediction string.\n\n---\n\n<b>\n\n| ImageID | ImageWidth | ImageHeight | PredictionString |\n| --- | --- | --- | --- |\n| `IMG_ID` | `WIDTH` | `HEIGHT` | \" `LABEL1` `CONFIDENCE1` `MASK1` `LABEL2` `CONFIDENCE2` `MASK2` `...` `LABEL10` `CONFIDENCE10` `MASK10` \" |\n\n</b>\n\n---\n\nAnd to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.\n\nHopefully, this clears things up. Please let me know if you have any other questions or anything is unclear.",
          "votes": 7
        },
        {
          "id": 1208245,
          "postDate": "2021-02-18T08:12:20.157Z",
          "content": "<p>Thank you for the clarification. This clears up my doubts. </p>\n<blockquote>\n  <p>And to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.</p>\n</blockquote>\n<p>This line hits be hard. Indeed instance segmentation is about getting individual masks even for the same object. </p>\n<p>Thanks again. </p>",
          "rawMarkdown": "Thank you for the clarification. This clears up my doubts. \n\n> And to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.\n\nThis line hits be hard. Indeed instance segmentation is about getting individual masks even for the same object. \n\nThanks again. "
        }
      ]
    },
    {
      "id": 1175468,
      "postDate": "2021-01-29T07:31:15.540Z",
      "content": "<p>Could competition host provide several demo result of cell segmentation, or give us some hint how segmentation done, by existing tool or human labeling? Thanks.</p>",
      "rawMarkdown": "Could competition host provide several demo result of cell segmentation, or give us some hint how segmentation done, by existing tool or human labeling? Thanks.",
      "votes": 2,
      "replies": [
        {
          "id": 1175536,
          "postDate": "2021-01-29T08:11:02.770Z",
          "content": "<p>Hi! The segmentation example can be found here (the 2nd half of the notebook) <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a></p>\n<p>When creating the ground truth, we used this model as baseline, and then our annotators go through all masks and adjust manually. I can tell you that just by using the baseline, you will match ~90% of the cells in test set.  </p>",
          "rawMarkdown": "Hi! The segmentation example can be found here (the 2nd half of the notebook) https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\n\nWhen creating the ground truth, we used this model as baseline, and then our annotators go through all masks and adjust manually. I can tell you that just by using the baseline, you will match ~90% of the cells in test set.  ",
          "votes": 6
        },
        {
          "id": 1175554,
          "postDate": "2021-01-29T08:20:17.803Z",
          "content": "<p>The grount truth segmentation was done as outlined in the example above (<a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a>), and subsequently manually corrected by human labeling.</p>",
          "rawMarkdown": "The grount truth segmentation was done as outlined in the example above ([HPACellSegmentation](https://github.com/CellProfiling/HPA-Cell-Segmentation)), and subsequently manually corrected by human labeling.",
          "votes": 5
        },
        {
          "id": 1175624,
          "postDate": "2021-01-29T08:58:32.710Z",
          "content": "<p>Thanks for your reply, now I get better understanding of this competition.</p>",
          "rawMarkdown": "Thanks for your reply, now I get better understanding of this competition."
        },
        {
          "id": 1175627,
          "postDate": "2021-01-29T09:00:45.310Z",
          "content": "<p>A really beginner question for someone not familiar with biology at all: do we need to segment nuclei or full cell?</p>",
          "rawMarkdown": "A really beginner question for someone not familiar with biology at all: do we need to segment nuclei or full cell?",
          "votes": 1
        },
        {
          "id": 1175633,
          "postDate": "2021-01-29T09:05:21.767Z",
          "content": "<p>Full cell! <br>\nUsually in biological images, cells are crowded and cell borders are hard to tell (especially in tissue), so a common protocol is to segment the nuclei first, and use them as seeds to segment the cells.</p>",
          "rawMarkdown": "Full cell! \nUsually in biological images, cells are crowded and cell borders are hard to tell (especially in tissue), so a common protocol is to segment the nuclei first, and use them as seeds to segment the cells.",
          "votes": 6
        },
        {
          "id": 1175634,
          "postDate": "2021-01-29T09:05:24.437Z",
          "content": "<p>You should segment the full cell.</p>",
          "rawMarkdown": "You should segment the full cell.",
          "votes": 5
        },
        {
          "id": 1175665,
          "postDate": "2021-01-29T09:22:01.213Z",
          "content": "<p>I find that the train set of 2021 and 2019 competition have large overlap. Are the private test newly produced  and labelled or derived from exist data? Will there be any potential risk of leakage or private test set exposed to public like hubmap competition?</p>",
          "rawMarkdown": "I find that the train set of 2021 and 2019 competition have large overlap. Are the private test newly produced  and labelled or derived from exist data? Will there be any potential risk of leakage or private test set exposed to public like hubmap competition?",
          "votes": 1
        },
        {
          "id": 1175746,
          "postDate": "2021-01-29T10:12:04.333Z",
          "content": "<p>You are correct, the train set 2021 is a part of train set 2019. The train set of this competition was filtered for 17 cell lines from the last competition. <br>\nThe single cell labels (test set ground truth) are newly produced for this competition. As stated <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">here</a> :</p>\n<p><code>For this competition, we specifically annotated every single cell in a subset of images (mostly private/never published, and some public images) with higher SCV compared to public HPA images, to act as test set</code></p>\n<p>You can find the image-level label for part of the test set that is public. But the risk for leakage of single cell label is very close to 0 (We will guard this :) ). </p>",
          "rawMarkdown": "You are correct, the train set 2021 is a part of train set 2019. The train set of this competition was filtered for 17 cell lines from the last competition. \nThe single cell labels (test set ground truth) are newly produced for this competition. As stated [here](https://www.kaggle.com/lnhtrang/single-cell-patterns) :\n\n` For this competition, we specifically annotated every single cell in a subset of images (mostly private/never published, and some public images) with higher SCV compared to public HPA images, to act as test set`\n\nYou can find the image-level label for part of the test set that is public. But the risk for leakage of single cell label is very close to 0 (We will guard this :) ). ",
          "votes": 6
        },
        {
          "id": 1176148,
          "postDate": "2021-01-29T14:22:14.217Z",
          "content": "<p>Just to be very clear. No images in the private test set have been made public in any way before, neither in the 2019 competition, or in the HPA database.</p>",
          "rawMarkdown": "Just to be very clear. No images in the private test set have been made public in any way before, neither in the 2019 competition, or in the HPA database.",
          "votes": 5
        },
        {
          "id": 1176591,
          "postDate": "2021-01-29T17:31:34.317Z",
          "content": "<p>And what's the difference between the 2021 public data and the one provided in the Data page?</p>",
          "rawMarkdown": "And what's the difference between the 2021 public data and the one provided in the Data page?"
        },
        {
          "id": 1176627,
          "postDate": "2021-01-29T17:45:47.257Z",
          "content": "<p>I'm not sure I fully understand the question. Are you wondering about the difference between the train data provided on the data page for this competition, and the external public HPA data? If so, these are different datasets (a few images may be overlapping, but the majority are different).</p>",
          "rawMarkdown": "I'm not sure I fully understand the question. Are you wondering about the difference between the train data provided on the data page for this competition, and the external public HPA data? If so, these are different datasets (a few images may be overlapping, but the majority are different).",
          "votes": 4
        },
        {
          "id": 1176641,
          "postDate": "2021-01-29T17:51:06.470Z",
          "content": "<p>That answered my question, and just to confirm, the external public HPA data is referring to the one uploaded by <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> , right? Thank you very much!</p>",
          "rawMarkdown": "That answered my question, and just to confirm, the external public HPA data is referring to the one uploaded by @lnhtrang , right? Thank you very much!"
        },
        {
          "id": 1176644,
          "postDate": "2021-01-29T17:52:10.030Z",
          "content": "<p>A difference between the train dataset in this competition and the external public HPA data is (as <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> mentioned above) that there are more cell lines and more labels in the public HPA data than what is included in this challenge.</p>",
          "rawMarkdown": "A difference between the train dataset in this competition and the external public HPA data is (as @lnhtrang mentioned above) that there are more cell lines and more labels in the public HPA data than what is included in this challenge.",
          "votes": 2
        },
        {
          "id": 1176657,
          "postDate": "2021-01-29T17:58:27.447Z",
          "content": "<p>yes it's referring to that one</p>",
          "rawMarkdown": "yes it's referring to that one",
          "votes": 1
        }
      ]
    },
    {
      "id": 1175039,
      "postDate": "2021-01-28T21:23:49.913Z",
      "content": "<p>Hi, thanks for the question. Yes, <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> is right. </p>\n<p>You only need to segment the cell and the cell mask will be shared by all the labels applied to the same cell. </p>\n<p>When annotating the dataset (i.e. the test set), we first did cell segmentation to get the cell masks, then for each cell in the image, our experts at HAP assigned one or multiple labels to the cell.</p>",
      "rawMarkdown": "Hi, thanks for the question. Yes, @dschettler8845 is right. \n\nYou only need to segment the cell and the cell mask will be shared by all the labels applied to the same cell. \n\nWhen annotating the dataset (i.e. the test set), we first did cell segmentation to get the cell masks, then for each cell in the image, our experts at HAP assigned one or multiple labels to the cell.",
      "votes": 2,
      "replies": [
        {
          "id": 1175056,
          "postDate": "2021-01-28T21:56:42.713Z",
          "content": "<p>Hey great <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a> that you for that explanation. So can I verify one more item about the metric that each cell that we segment will have its own (label, confidence, rle_encoded_mask)? I want to use a more complicated example to see if I generally understand.</p>\n<p>Let's say there are two cells in an image, cell 1 has labels 0, 14 and cell 2 has label 0 would a possible submission string be:</p>\n<p><code>0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask</code>?</p>\n<p>and also for our own validation purposes we will actually be creating the ground truth cell masks using some method / model / etc.</p>",
          "rawMarkdown": "Hey great @weiouyang that you for that explanation. So can I verify one more item about the metric that each cell that we segment will have its own (label, confidence, rle_encoded_mask)? I want to use a more complicated example to see if I generally understand.\n\nLet's say there are two cells in an image, cell 1 has labels 0, 14 and cell 2 has label 0 would a possible submission string be:\n\n`0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask`?\n\nand also for our own validation purposes we will actually be creating the ground truth cell masks using some method / model / etc.",
          "votes": 3
        },
        {
          "id": 1175553,
          "postDate": "2021-01-29T08:19:39.620Z",
          "content": "<p>You are correct, but remember the whole line to be compatible with the metrics:</p>\n<blockquote>\n  <p>ImageAID,ImageAWidth,ImageAHeight,0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask</p>\n</blockquote>\n<p> Now since the validation metrics is mAP, confidence score affects scoring.<br>\nYou can read more here <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation</a></p>",
          "rawMarkdown": "You are correct, but remember the whole line to be compatible with the metrics:\n> ImageAID,ImageAWidth,ImageAHeight,0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask\n\n~~Confidence is just ranking of scoring (which mask/label got scored first), so you can set them all to 1 like this and it won't affect your score.~~ Now since the validation metrics is mAP, confidence score affects scoring.\nYou can read more here https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation",
          "votes": 8
        },
        {
          "id": 1176689,
          "postDate": "2021-01-29T18:16:08.293Z",
          "content": "<p>Thanks for that clarification <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>! This will be helpful moving forward</p>",
          "rawMarkdown": "Thanks for that clarification @lnhtrang! This will be helpful moving forward"
        },
        {
          "id": 1198195,
          "postDate": "2021-02-12T19:36:58.027Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1193080,
      "postDate": "2021-02-09T12:56:25.540Z",
      "content": "<p>Thanks for the post.<br>\nBased on what I understand, if we use the model which segments the whole part of every cell for prediction, there's gonna be many FP when true label locates only inside the nucleus (I mean, IoU &lt; 0.6), for example.<br>\nIs this a big problem??</p>\n<p>EDIT: asked here<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309</a></p>",
      "rawMarkdown": "Thanks for the post.\nBased on what I understand, if we use the model which segments the whole part of every cell for prediction, there's gonna be many FP when true label locates only inside the nucleus (I mean, IoU < 0.6), for example.\nIs this a big problem??\n\nEDIT: asked here\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309"
    }
  ],
  "comments": [
    {
      "id": 1174999,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2021-01-28T20:30:13.753000",
      "content": "<p>I believe we segment the entire cell as it says in the evaluation that a single instance of segmentation can have multiple labels.</p>\n<p>As it is instance segmentation you will need individual masks for individual cells.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1207396,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-17T20:26:46.437000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>, what do you mean by individual masks for individual cells? </p>\n<p>Say in the given image there are 10 cells and we get a segmentation mask for each of them. We will have to denote mask for say cell one as 0 and cell two a 1 and so on in a preferred direction? </p>\n<p>Can we put it like this - do semantic segmentation of the image(binary - background and cell); for each cell predict the label? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1207547,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-02-17T23:08:07.910000",
          "content": "<p>Mostly correct.</p>\n<p>You will segment the image and get 10 masks (let's say each cell has a different value like you indicated… background is 0, cell 1 has a value of 1, cell 2 has a value of 2, etc.).</p>\n<p>Then you will predict, for a given cell mask, what are the classes present in that cell. By individual mask, I am describing the single value of the mask corresponding to that cell.</p>\n<p>Let's see the example through. For our ten cells, we have 10 masks.</p>\n<p>We can get each mask by itself by performing something like this:</p>\n<hr>\n<p><b></b></p>\n<pre><code># all_cell_masks is the numpy array containing all segmented cells\n# where each cell is indicated by a different numeric value\ncell_1 = np.where(all_cell_masks==1, 1, 0)\ncell_2 = np.where(all_cell_masks==2, 1, 0)\n...\ncell_10 = np.where(all_cell_masks==10, 1, 0)\n</code></pre>\n<p></p>\n<hr>\n<p>Now we have ten binary mask arrays, one for each cell. Now we need to make a prediction (possibly multiple predictions), with a confidence score, on each mask. Now the binary segmentation masks are run-length encoded (RLE), zlib compressed, and base64 encoded to be used in text format in the prediction string.</p>\n<hr>\n<p><b></b></p>\n<table>\n<thead>\n<tr>\n<th>ImageID</th>\n<th>ImageWidth</th>\n<th>ImageHeight</th>\n<th>PredictionString</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><code>IMG_ID</code></td>\n<td><code>WIDTH</code></td>\n<td><code>HEIGHT</code></td>\n<td>\" <code>LABEL1</code> <code>CONFIDENCE1</code> <code>MASK1</code> <code>LABEL2</code> <code>CONFIDENCE2</code> <code>MASK2</code> <code>...</code> <code>LABEL10</code> <code>CONFIDENCE10</code> <code>MASK10</code> \"</td>\n</tr>\n</tbody>\n</table>\n<p></p>\n<hr>\n<p>And to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.</p>\n<p>Hopefully, this clears things up. Please let me know if you have any other questions or anything is unclear.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1208245,
          "author_name": "Ayush Thakur",
          "author_url": "",
          "post_date": "2021-02-18T08:12:20.157000",
          "content": "<p>Thank you for the clarification. This clears up my doubts. </p>\n<blockquote>\n  <p>And to your final point, doing binary semantic segmentation on each cell in exclusion (and track which is which) is pretty much the definition of performing instance segmentation.</p>\n</blockquote>\n<p>This line hits be hard. Indeed instance segmentation is about getting individual masks even for the same object. </p>\n<p>Thanks again. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1175468,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2021-01-29T07:31:15.540000",
      "content": "<p>Could competition host provide several demo result of cell segmentation, or give us some hint how segmentation done, by existing tool or human labeling? Thanks.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1175536,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-29T08:11:02.770000",
          "content": "<p>Hi! The segmentation example can be found here (the 2nd half of the notebook) <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a></p>\n<p>When creating the ground truth, we used this model as baseline, and then our annotators go through all masks and adjust manually. I can tell you that just by using the baseline, you will match ~90% of the cells in test set.  </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1175554,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T08:20:17.803000",
          "content": "<p>The grount truth segmentation was done as outlined in the example above (<a href=\"https://github.com/CellProfiling/HPA-Cell-Segmentation\" target=\"_blank\">HPACellSegmentation</a>), and subsequently manually corrected by human labeling.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1175624,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-01-29T08:58:32.710000",
          "content": "<p>Thanks for your reply, now I get better understanding of this competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1175627,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2021-01-29T09:00:45.310000",
          "content": "<p>A really beginner question for someone not familiar with biology at all: do we need to segment nuclei or full cell?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1175633,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-29T09:05:21.767000",
          "content": "<p>Full cell! <br>\nUsually in biological images, cells are crowded and cell borders are hard to tell (especially in tissue), so a common protocol is to segment the nuclei first, and use them as seeds to segment the cells.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1175634,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T09:05:24.437000",
          "content": "<p>You should segment the full cell.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1175665,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-01-29T09:22:01.213000",
          "content": "<p>I find that the train set of 2021 and 2019 competition have large overlap. Are the private test newly produced  and labelled or derived from exist data? Will there be any potential risk of leakage or private test set exposed to public like hubmap competition?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1175746,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-29T10:12:04.333000",
          "content": "<p>You are correct, the train set 2021 is a part of train set 2019. The train set of this competition was filtered for 17 cell lines from the last competition. <br>\nThe single cell labels (test set ground truth) are newly produced for this competition. As stated <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">here</a> :</p>\n<p><code>For this competition, we specifically annotated every single cell in a subset of images (mostly private/never published, and some public images) with higher SCV compared to public HPA images, to act as test set</code></p>\n<p>You can find the image-level label for part of the test set that is public. But the risk for leakage of single cell label is very close to 0 (We will guard this :) ). </p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1176148,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T14:22:14.217000",
          "content": "<p>Just to be very clear. No images in the private test set have been made public in any way before, neither in the 2019 competition, or in the HPA database.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1176591,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2021-01-29T17:31:34.317000",
          "content": "<p>And what's the difference between the 2021 public data and the one provided in the Data page?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1176627,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T17:45:47.257000",
          "content": "<p>I'm not sure I fully understand the question. Are you wondering about the difference between the train data provided on the data page for this competition, and the external public HPA data? If so, these are different datasets (a few images may be overlapping, but the majority are different).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1176641,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2021-01-29T17:51:06.470000",
          "content": "<p>That answered my question, and just to confirm, the external public HPA data is referring to the one uploaded by <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> , right? Thank you very much!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1176644,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T17:52:10.030000",
          "content": "<p>A difference between the train dataset in this competition and the external public HPA data is (as <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> mentioned above) that there are more cell lines and more labels in the public HPA data than what is included in this challenge.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1176657,
          "author_name": "Emma Lundberg",
          "author_url": "",
          "post_date": "2021-01-29T17:58:27.447000",
          "content": "<p>yes it's referring to that one</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1175039,
      "author_name": "Wei Ouyang",
      "author_url": "",
      "post_date": "2021-01-28T21:23:49.913000",
      "content": "<p>Hi, thanks for the question. Yes, <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> is right. </p>\n<p>You only need to segment the cell and the cell mask will be shared by all the labels applied to the same cell. </p>\n<p>When annotating the dataset (i.e. the test set), we first did cell segmentation to get the cell masks, then for each cell in the image, our experts at HAP assigned one or multiple labels to the cell.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1175056,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2021-01-28T21:56:42.713000",
          "content": "<p>Hey great <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a> that you for that explanation. So can I verify one more item about the metric that each cell that we segment will have its own (label, confidence, rle_encoded_mask)? I want to use a more complicated example to see if I generally understand.</p>\n<p>Let's say there are two cells in an image, cell 1 has labels 0, 14 and cell 2 has label 0 would a possible submission string be:</p>\n<p><code>0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask</code>?</p>\n<p>and also for our own validation purposes we will actually be creating the ground truth cell masks using some method / model / etc.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1175553,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-01-29T08:19:39.620000",
          "content": "<p>You are correct, but remember the whole line to be compatible with the metrics:</p>\n<blockquote>\n  <p>ImageAID,ImageAWidth,ImageAHeight,0 1 rle_encoded_cell_1_mask 14 1 rle_encoded_cell_1_mask 0 1 rle encoded_cell_2_mask</p>\n</blockquote>\n<p> Now since the validation metrics is mAP, confidence score affects scoring.<br>\nYou can read more here <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/overview/evaluation</a></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1176689,
          "author_name": "RDizzl3",
          "author_url": "",
          "post_date": "2021-01-29T18:16:08.293000",
          "content": "<p>Thanks for that clarification <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>! This will be helpful moving forward</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1198195,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-12T19:36:58.027000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1193080,
      "author_name": "cool_rabbit",
      "author_url": "",
      "post_date": "2021-02-09T12:56:25.540000",
      "content": "<p>Thanks for the post.<br>\nBased on what I understand, if we use the model which segments the whole part of every cell for prediction, there's gonna be many FP when true label locates only inside the nucleus (I mean, IoU &lt; 0.6), for example.<br>\nIs this a big problem??</p>\n<p>EDIT: asked here<br>\n<a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1174914": "e.g. if we want to detect nuclear speckles, do we segment the entire cell and do the labeling, or do we just segment the small dots?\n\nIn addition, should we predict individual masks for individual cells, or a single mask for the entire class?",
    "1174999": "I believe we segment the entire cell as it says in the evaluation that a single instance of segmentation can have multiple labels.\n\nAs it is instance segmentation you will need individual masks for individual cells.",
    "1175468": "Could competition host provide several demo result of cell segmentation, or give us some hint how segmentation done, by existing tool or human labeling? Thanks.",
    "1175039": "Hi, thanks for the question. Yes, @dschettler8845 is right. \n\nYou only need to segment the cell and the cell mask will be shared by all the labels applied to the same cell. \n\nWhen annotating the dataset (i.e. the test set), we first did cell segmentation to get the cell masks, then for each cell in the image, our experts at HAP assigned one or multiple labels to the cell.",
    "1193080": "Thanks for the post.\nBased on what I understand, if we use the model which segments the whole part of every cell for prediction, there's gonna be many FP when true label locates only inside the nucleus (I mean, IoU < 0.6), for example.\nIs this a big problem??\n\nEDIT: asked here\nhttps://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/218309"
  }
}