{
  "id": 223645,
  "title": "AMA 4-5pm CET on March 8th [still monitoring]",
  "url": "/competitions/hpa-single-cell-image-classification/discussion/223645",
  "author_name": "Trang Le",
  "post_date": "2021-03-04T20:55:35.433000",
  "votes": 29,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Update: Although the AMA has passed, we will check the thread occasionally so please ask if you have any Qs!</p>\n<p>Hi All!</p>\n<p>We have received many great questions in this competition. We want to help you in understanding the data and the set-up, which ultimately leads to a better solution. However, It is a bit hard to keep track of all discussions, and there are many questions that have been asked multiple times across different threads. So we would like to host an AMA (Ask Me/Us Anything) session at <strong>4-5pm CET on Monday March 8th</strong>. At this time is when many people in our team will be online (especially our expert annotators who created the groundtruth) and we will try to answer as many questions as possible.</p>\n<p>You can also post questions here beforehand, and we will answer them in order of upvotes (we will of course try to answer them all). </p>\n<p>Before you ask questions, please check out these discussions first to see if your questions have already been answered here:</p>\n<p>How individual single cell pattern looks like, what is SCV <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns</a> <br>\nExternal HPA data, overlap between trainset and external HPA data, convert external HPA classes to 19 classes in this competition <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a> <br>\nCellSegmentator <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a><br>\nHow does groundtruth look like <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141</a> <br>\nAbout Negative: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748</a> <br>\nAbout confidence: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885</a><br>\nWhy there are duplicate labels for the same image, merged labels: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562</a> </p>\n<p>All the bests,<br>\n<a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a>, <a href=\"https://www.kaggle.com/uaxelsson\" target=\"_blank\">@uaxelsson</a>, <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> and <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a></p>",
  "messages": [
    {
      "id": 1226778,
      "postDate": "2021-03-04T20:55:35.433Z",
      "content": "<p>Update: Although the AMA has passed, we will check the thread occasionally so please ask if you have any Qs!</p>\n<p>Hi All!</p>\n<p>We have received many great questions in this competition. We want to help you in understanding the data and the set-up, which ultimately leads to a better solution. However, It is a bit hard to keep track of all discussions, and there are many questions that have been asked multiple times across different threads. So we would like to host an AMA (Ask Me/Us Anything) session at <strong>4-5pm CET on Monday March 8th</strong>. At this time is when many people in our team will be online (especially our expert annotators who created the groundtruth) and we will try to answer as many questions as possible.</p>\n<p>You can also post questions here beforehand, and we will answer them in order of upvotes (we will of course try to answer them all). </p>\n<p>Before you ask questions, please check out these discussions first to see if your questions have already been answered here:</p>\n<p>How individual single cell pattern looks like, what is SCV <a href=\"https://www.kaggle.com/lnhtrang/single-cell-patterns\" target=\"_blank\">https://www.kaggle.com/lnhtrang/single-cell-patterns</a> <br>\nExternal HPA data, overlap between trainset and external HPA data, convert external HPA classes to 19 classes in this competition <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a> <br>\nCellSegmentator <a href=\"https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\" target=\"_blank\">https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg</a><br>\nHow does groundtruth look like <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141</a> <br>\nAbout Negative: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748</a> <br>\nAbout confidence: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885</a><br>\nWhy there are duplicate labels for the same image, merged labels: <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562</a> </p>\n<p>All the bests,<br>\n<a href=\"https://www.kaggle.com/emmalumpan\" target=\"_blank\">@emmalumpan</a>, <a href=\"https://www.kaggle.com/weiouyang\" target=\"_blank\">@weiouyang</a>, <a href=\"https://www.kaggle.com/uaxelsson\" target=\"_blank\">@uaxelsson</a>, <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> and <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a></p>",
      "rawMarkdown": "Update: Although the AMA has passed, we will check the thread occasionally so please ask if you have any Qs!\n\nHi All!\n\nWe have received many great questions in this competition. We want to help you in understanding the data and the set-up, which ultimately leads to a better solution. However, It is a bit hard to keep track of all discussions, and there are many questions that have been asked multiple times across different threads. So we would like to host an AMA (Ask Me/Us Anything) session at **4-5pm CET on Monday March 8th**. At this time is when many people in our team will be online (especially our expert annotators who created the groundtruth) and we will try to answer as many questions as possible.\n\nYou can also post questions here beforehand, and we will answer them in order of upvotes (we will of course try to answer them all). \n\nBefore you ask questions, please check out these discussions first to see if your questions have already been answered here:\n\nHow individual single cell pattern looks like, what is SCV https://www.kaggle.com/lnhtrang/single-cell-patterns \nExternal HPA data, overlap between trainset and external HPA data, convert external HPA classes to 19 classes in this competition https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg \nCellSegmentator https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\nHow does groundtruth look like https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141 \nAbout Negative: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748 \nAbout confidence: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885\nWhy there are duplicate labels for the same image, merged labels: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562 \n\nAll the bests,\n@emmalumpan, @weiouyang, @uaxelsson, @cwinsnes and @lnhtrang",
      "votes": 29
    },
    {
      "id": 1296060,
      "postDate": "2021-05-06T23:49:20.847Z",
      "content": "<p>In the sample_submission.csv … can we expect the width and height to be correct, so we don't need to check it at all? Seems to be correct on the public test set? Thanks :)</p>",
      "rawMarkdown": "In the sample_submission.csv ... can we expect the width and height to be correct, so we don't need to check it at all? Seems to be correct on the public test set? Thanks :)",
      "votes": 1,
      "replies": [
        {
          "id": 1297288,
          "postDate": "2021-05-07T22:00:13.167Z",
          "content": "<p>It would be a very nasty surprise if that were not the case… hope the hosts can confirm :) </p>",
          "rawMarkdown": "It would be a very nasty surprise if that were not the case... hope the hosts can confirm :) ",
          "votes": 1
        },
        {
          "id": 1297600,
          "postDate": "2021-05-08T06:59:55.713Z",
          "content": "<p>I agree :D ha ha</p>",
          "rawMarkdown": "I agree :D ha ha"
        },
        {
          "id": 1297987,
          "postDate": "2021-05-08T13:28:38.740Z",
          "content": "<p>The sample_submission.csv only contains 559 images from the public test set right, the width and height are correct in that file for public test set.</p>",
          "rawMarkdown": "The sample_submission.csv only contains 559 images from the public test set right, ~~so you still want to get the right image size from unseen images from private test set. That being said, ~~the width and height are correct in that file for public test set."
        },
        {
          "id": 1298004,
          "postDate": "2021-05-08T13:40:41.547Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> the question is whether the width and the height are correct in that file for private test set? We assume it is, so we use that data to accelerate segmentation. If it is incorrect, then that would impact the score in negative way. </p>",
          "rawMarkdown": "Thanks @lnhtrang the question is whether the width and the height are correct in that file for private test set? We assume it is, so we use that data to accelerate segmentation. If it is incorrect, then that would impact the score in negative way. ",
          "votes": 1
        },
        {
          "id": 1298008,
          "postDate": "2021-05-08T13:44:00.550Z",
          "content": "<p>dumb question but where can I see that file? I only see the sample_submission.csv for public test set :) </p>",
          "rawMarkdown": "dumb question but where can I see that file? I only see the sample_submission.csv for public test set :) "
        },
        {
          "id": 1298010,
          "postDate": "2021-05-08T13:47:26.100Z",
          "content": "<p>There are no record of such issue in perious competition I joined. But, doesn't the public/priave split provided by host? (you may ignore my question if anything sensitive)</p>",
          "rawMarkdown": "There are no record of such issue in perious competition I joined. But, doesn't the public/priave split provided by host? (you may ignore my question if anything sensitive)",
          "votes": 1
        },
        {
          "id": 1298014,
          "postDate": "2021-05-08T13:50:28.533Z",
          "content": "<p>We can't see that file as participants - our models are evaluated on it but it is not available. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> might be able to confirm? </p>",
          "rawMarkdown": "We can't see that file as participants - our models are evaluated on it but it is not available. @philculliton might be able to confirm? ",
          "votes": 1
        },
        {
          "id": 1298018,
          "postDate": "2021-05-08T13:57:13.840Z",
          "content": "<p>probably it's me being unfamiliar with kaggle platform. I just checked the <code>sample_submission.csv</code> file in the <code>Data</code> tab, which only contains 559 images, and their width and height are all correct. If it's a publicly available file, then it's not competition-sensitive. This file was produced by <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> so maybe he is a better person to confirm. I assume it's produced by the same code, so if the public test set is correct, the private test set should be as well. </p>",
          "rawMarkdown": "probably it's me being unfamiliar with kaggle platform. I just checked the `sample_submission.csv` file in the `Data` tab, which only contains 559 images, and their width and height are all correct. If it's a publicly available file, then it's not competition-sensitive. This file was produced by @philculliton so maybe he is a better person to confirm. I assume it's produced by the same code, so if the public test set is correct, the private test set should be as well. "
        },
        {
          "id": 1298051,
          "postDate": "2021-05-08T14:21:20.933Z",
          "content": "<p>Does the sample_submission.csv when evaluted as a full submission contain all test images (public+private)? And is the widht and height information correct for both private and public test images?</p>\n<p>I assume sample_submission.csv file is replaced when the full submission is made in the same way as it is with the test/ folder (test images)?</p>",
          "rawMarkdown": "Does the sample_submission.csv when evaluted as a full submission contain all test images (public+private)? And is the widht and height information correct for both private and public test images?\n\nI assume sample_submission.csv file is replaced when the full submission is made in the same way as it is with the test/ folder (test images)?"
        },
        {
          "id": 1298059,
          "postDate": "2021-05-08T14:26:03.197Z",
          "content": "<p>I found another <code>sample_submission.csv</code> file as host, in which all width and height are correct for all test set. Hope that <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can see this and have a final confirmation, but I think your assumption is correct <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> so please proceed as normal.</p>",
          "rawMarkdown": "I found another `sample_submission.csv` file as host, in which all width and height are correct for all test set. Hope that @philculliton can see this and have a final confirmation, but I think your assumption is correct @thedrcat @crodoc so please proceed as normal.",
          "votes": 2
        },
        {
          "id": 1298075,
          "postDate": "2021-05-08T14:34:35.807Z",
          "content": "<p>This sounds good! Thanks <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>.<br>\nI guess a lot of us are just loading the <code>sample_submission.csv</code> file and just replacing the <code>PredictionString</code> part. If the file was not updated at full submission time, I think a lot of teams would score 0 on the private LB.</p>\n<p>I guess you are doing the same thing <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> :D</p>\n<p>Would be great if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> also confirms, otherwise a lot of us need to do some fast updates to our notebooks.</p>",
          "rawMarkdown": "This sounds good! Thanks @lnhtrang.\nI guess a lot of us are just loading the `sample_submission.csv` file and just replacing the `PredictionString` part. If the file was not updated at full submission time, I think a lot of teams would score 0 on the private LB.\n\nI guess you are doing the same thing @thedrcat :D\n\nWould be great if @philculliton also confirms, otherwise a lot of us need to do some fast updates to our notebooks."
        },
        {
          "id": 1298851,
          "postDate": "2021-05-09T09:16:29.943Z",
          "content": "<p>Hey, I just tested it by reading the files from the test directory into a dataframe and comparing this dataframe to the one provided as <code>sample_submission.csv</code>. It's not giving any error…</p>\n<pre><code>import glob\nimport re\n\ndata_df_sample_submission = pd.read_csv('../input/hpa-single-cell-image-classification/sample_submission.csv')\n\ntest_files = os.listdir(\"../input/hpa-single-cell-image-classification/test\")\ncolor_list = [\"_red.png\", \"_green.png\", \"_yellow.png\", \"_blue.png\"]\ntest_files_names =  [re.sub(r'|'.join(map(re.escape, color_list)), '', elem) for elem in test_files]\ntest_files_names = list(set(test_files_names))\n\nd = []\nfor i, file in enumerate(test_files_names):\n    img = cv2.imread(f\"../input/hpa-single-cell-image-classification/test/{file}_red.png\")\n    height, width, channels = img.shape\n    d.append({\n        \"ID\" : file,\n        \"ImageWidth\" : width,\n        \"ImageHeight\": height,\n        \"PredictionString\" : \"0 1 eNoLCAgIMAEABJkBdQ==\"\n    })\n    if i%50 == 0:\n        print(i)\n\ndf_from_files = pd.DataFrame(d).sort_values(\"ID\").reset_index(drop=True)\n</code></pre>\n<p>and last line!:</p>\n<pre><code>if all(df_from_files == data_df_sample_submission):\n    sub.to_csv(\"submission.csv\",index=False)\n</code></pre>",
          "rawMarkdown": "Hey, I just tested it by reading the files from the test directory into a dataframe and comparing this dataframe to the one provided as `sample_submission.csv`. It's not giving any error...\n\n```\nimport glob\nimport re\n\ndata_df_sample_submission = pd.read_csv('../input/hpa-single-cell-image-classification/sample_submission.csv')\n\ntest_files = os.listdir(\"../input/hpa-single-cell-image-classification/test\")\ncolor_list = [\"_red.png\", \"_green.png\", \"_yellow.png\", \"_blue.png\"]\ntest_files_names =  [re.sub(r'|'.join(map(re.escape, color_list)), '', elem) for elem in test_files]\ntest_files_names = list(set(test_files_names))\n\nd = []\nfor i, file in enumerate(test_files_names):\n    img = cv2.imread(f\"../input/hpa-single-cell-image-classification/test/{file}_red.png\")\n    height, width, channels = img.shape\n    d.append({\n        \"ID\" : file,\n        \"ImageWidth\" : width,\n        \"ImageHeight\": height,\n        \"PredictionString\" : \"0 1 eNoLCAgIMAEABJkBdQ==\"\n    })\n    if i%50 == 0:\n        print(i)\n\ndf_from_files = pd.DataFrame(d).sort_values(\"ID\").reset_index(drop=True)\n```\n\nand last line!:\n```\nif all(df_from_files == data_df_sample_submission):\n    sub.to_csv(\"submission.csv\",index=False)\n```",
          "votes": 4
        },
        {
          "id": 1298989,
          "postDate": "2021-05-09T11:47:04.727Z",
          "content": "<p>Great job!!!!</p>",
          "rawMarkdown": "Great job!!!!",
          "votes": 1
        },
        {
          "id": 1298997,
          "postDate": "2021-05-09T11:51:40.747Z",
          "content": "<p>Still I would be happy to have this confirmed by a host 😄</p>",
          "rawMarkdown": "Still I would be happy to have this confirmed by a host 😄"
        },
        {
          "id": 1302256,
          "postDate": "2021-05-11T13:06:18.960Z",
          "content": "<p>I did a search on the kaggle site … seems the <code>sample_submission.csv</code> file will be replaced on full submission … yaaay :)</p>\n<p>Check out a few links I picked from the search:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222037#1217676</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202223#1107145</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177356#985410</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/77241#453866</a></p>",
          "rawMarkdown": "I did a search on the kaggle site ... seems the `sample_submission.csv` file will be replaced on full submission ... yaaay :)\n\nCheck out a few links I picked from the search:\n[https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222037#1217676](url)\n[https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202223#1107145](url)\n[https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177356#985410](url)\n[https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/77241#453866](url)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1231013,
      "postDate": "2021-03-08T16:00:35.787Z",
      "content": "<p>Some of the cells appear to be undergoing (or recently completed) cytokinesis. At what point in this process would human annotators consider them to be two distinct cells for the purposes of the competition? I.e. if two nuclei with complete nuclear membranes are physically in contact and are identified as a single cell by the HPA cell segmenter would a human annotator typically split this into two distinct cell masks?</p>",
      "rawMarkdown": "Some of the cells appear to be undergoing (or recently completed) cytokinesis. At what point in this process would human annotators consider them to be two distinct cells for the purposes of the competition? I.e. if two nuclei with complete nuclear membranes are physically in contact and are identified as a single cell by the HPA cell segmenter would a human annotator typically split this into two distinct cell masks?",
      "votes": 1,
      "replies": [
        {
          "id": 1231104,
          "postDate": "2021-03-08T17:08:11.083Z",
          "content": "<p>If 2 nuclei and membrane can be distinguished, HPACellSegmentator will post process it as 2 separate cells. If the Segmentator fails to do so, human annotator will do it for the test set.</p>",
          "rawMarkdown": "If 2 nuclei and membrane can be distinguished, HPACellSegmentator will post process it as 2 separate cells. If the Segmentator fails to do so, human annotator will do it for the test set."
        }
      ]
    },
    {
      "id": 1268967,
      "postDate": "2021-04-10T02:21:26.410Z",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a></p>\n<p>A quick question about the proportion of negative cells in slides that come with unique label ( non-negative label ).</p>\n<p>Based on the domain knowledge ( of how these slides are produced and how different cell types cluster in the slides )</p>\n<ol>\n<li>Should the average negative proportion be small  below 10%) or large (above 80 percent) or somewhere in between ? </li>\n<li>Should I expect the variance of negative proportion to be large ?  </li>\n</ol>",
      "rawMarkdown": "@lnhtrang\n\nA quick question about the proportion of negative cells in slides that come with unique label ( non-negative label ).\n\nBased on the domain knowledge ( of how these slides are produced and how different cell types cluster in the slides )\n\n1.  Should the average negative proportion be small  below 10%) or large (above 80 percent) or somewhere in between ? \n2. Should I expect the variance of negative proportion to be large ?  \n\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1271074,
          "postDate": "2021-04-12T09:20:57.333Z",
          "content": "<p>Hi! The distribution of <code>Negative</code> labels does vary a lot, based on the tagged protein, so it can be anywhere from &lt;10 to 90%. A clue is that high <code>Negative</code> % for proteins related to mitosis. I don't exactly know the proportion on average for train set.</p>",
          "rawMarkdown": "Hi! The distribution of `Negative` labels does vary a lot, based on the tagged protein, so it can be anywhere from <10 to 90%. A clue is that high `Negative` % for proteins related to mitosis. I don't exactly know the proportion on average for train set.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1231169,
      "postDate": "2021-03-08T18:13:40.993Z",
      "content": "<p>In one post in this forum, someone said that's it's not possible to have, for one cell, class 18 (negative) mixed with any other classes for the same cell. Does it matter for the scoring if we submit such combination?</p>",
      "rawMarkdown": "In one post in this forum, someone said that's it's not possible to have, for one cell, class 18 (negative) mixed with any other classes for the same cell. Does it matter for the scoring if we submit such combination?",
      "votes": 2,
      "replies": [
        {
          "id": 1231441,
          "postDate": "2021-03-09T01:48:26.557Z",
          "content": "<p>I have the same question!!<br>\nTheoretically impossible because of Negative, but waiting for the official answer.</p>",
          "rawMarkdown": "I have the same question!!\nTheoretically impossible because of Negative, but waiting for the official answer.",
          "votes": 1
        },
        {
          "id": 1232166,
          "postDate": "2021-03-09T14:30:41.893Z",
          "content": "<p>The <code>Negative</code> class is treated as any other class in terms of the scoring, so if you combine it with other classes for the same cell it will be scored in the same manner as any other combination of classes.</p>\n<p>However, as stated before: there will be no instances of a cell labeled <code>Negative</code> that is also labeled as something else.</p>",
          "rawMarkdown": "The `Negative` class is treated as any other class in terms of the scoring, so if you combine it with other classes for the same cell it will be scored in the same manner as any other combination of classes.\n\nHowever, as stated before: there will be no instances of a cell labeled `Negative` that is also labeled as something else.",
          "votes": 4
        }
      ]
    },
    {
      "id": 1230548,
      "postDate": "2021-03-08T08:21:39.017Z",
      "content": "<p>Thanks a lot for this offer! I have a question related to this comment:</p>\n<blockquote>\n  <p>During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label. </p>\n</blockquote>\n<p>Does this mean that we can have train/public set images with mitotic spindle label but actually no mitotic spindle represented, because it only appeared in other images from the same sample? </p>",
      "rawMarkdown": "Thanks a lot for this offer! I have a question related to this comment:\n> During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label. \n\nDoes this mean that we can have train/public set images with mitotic spindle label but actually no mitotic spindle represented, because it only appeared in other images from the same sample? \n",
      "votes": 2,
      "replies": [
        {
          "id": 1230967,
          "postDate": "2021-03-08T15:19:42.053Z",
          "content": "<p>You are correct. There could be images tagged with <code>Mitotic spindle</code> but actually no mitotic spindle present. But these should be rare cases, because the production team took up to 6 images and display the best 2 on the atlas (So the public HPA are the better quality ones, and will almost always have the pattern). This problem is most pronounced for mitotic structures (that appears in about 5% of all cells at any given timepoint). Other organelles don't suffer from this problem.</p>",
          "rawMarkdown": "You are correct. There could be images tagged with `Mitotic spindle` but actually no mitotic spindle present. But these should be rare cases, because the production team took up to 6 images and display the best 2 on the atlas (So the public HPA are the better quality ones, and will almost always have the pattern). This problem is most pronounced for mitotic structures (that appears in about 5% of all cells at any given timepoint). Other organelles don't suffer from this problem.",
          "votes": 2
        },
        {
          "id": 1286177,
          "postDate": "2021-04-27T16:35:06.087Z",
          "content": "<p>hi . i have a small doubt . what do you mean by mitotic structures. what are all the classes that have relation to mitotic structures. i am trying to understand how to handle these classes.</p>",
          "rawMarkdown": "hi . i have a small doubt . what do you mean by mitotic structures. what are all the classes that have relation to mitotic structures. i am trying to understand how to handle these classes."
        },
        {
          "id": 1287551,
          "postDate": "2021-04-29T07:11:01.173Z",
          "content": "<p>Hi! for this competition, the only class is <code>Mitotic spindle</code>.</p>",
          "rawMarkdown": "Hi! for this competition, the only class is `Mitotic spindle`."
        }
      ]
    },
    {
      "id": 1299909,
      "postDate": "2021-05-10T06:25:56.037Z",
      "content": "<p>Hi, can a host please take a look at <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745</a></p>",
      "rawMarkdown": "Hi, can a host please take a look at https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745"
    },
    {
      "id": 1289326,
      "postDate": "2021-04-30T21:09:02.103Z",
      "content": "<p>Dear organizers,</p>\n<p>when you have a moment could you please clarify the mitotic spindle labeling of the following cells:</p>\n<ol>\n<li><p>This cell comes from an image labeled as <em>Mitotic spindle</em>, the cell seems to be the only one in the mitotic phase in the image.<br>\n<img src=\"https://i.ibb.co/HYxjdrX/positive-mitotic.png\" alt=\"\"></p></li>\n<li><p>The following cells are from images <strong>without</strong> the <em>Mitotic spindle</em> label.<br>\n<img src=\"https://i.ibb.co/7QJph0y/negative-mitotic.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/Pt8xNsM/q2.png\" alt=\"\"></p></li>\n</ol>\n<p><strong>What bothers me:</strong><br>\nDoes the cell from 1. have an apparent mitotic spindle pattern? Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?</p>\n<p>Thank you in advance!</p>",
      "rawMarkdown": "Dear organizers,\n\nwhen you have a moment could you please clarify the mitotic spindle labeling of the following cells:\n\n1. This cell comes from an image labeled as *Mitotic spindle*, the cell seems to be the only one in the mitotic phase in the image.\n![](https://i.ibb.co/HYxjdrX/positive-mitotic.png)\n\n2. The following cells are from images **without** the *Mitotic spindle* label.\n![](https://i.ibb.co/7QJph0y/negative-mitotic.png)\n![](https://i.ibb.co/Pt8xNsM/q2.png)\n\n**What bothers me:**\nDoes the cell from 1. have an apparent mitotic spindle pattern? Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?\n\nThank you in advance!",
      "replies": [
        {
          "id": 1297942,
          "postDate": "2021-05-08T12:49:36.570Z",
          "content": "<p>Hi! Sorry for the late reply. I actually don't see the images you attached, even after switching browsers.</p>\n<blockquote>\n  <p>Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?</p>\n</blockquote>\n<p>Well mitotic spindle appears in cells that are going through mitosis, so they do have different shapes and organization of nucleus/microtubules. Neural networks are effective at picking these signals up, so they may do well with just green signal, but I suspect red and blue channels would help alot.</p>",
          "rawMarkdown": "Hi! Sorry for the late reply. I actually don't see the images you attached, even after switching browsers.\n> Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?\n\nWell mitotic spindle appears in cells that are going through mitosis, so they do have different shapes and organization of nucleus/microtubules. Neural networks are effective at picking these signals up, so they may do well with just green signal, but I suspect red and blue channels would help alot.",
          "votes": 1
        },
        {
          "id": 1298137,
          "postDate": "2021-05-08T15:20:57.513Z",
          "content": "<p>Hi! Thank you for the reply!</p>",
          "rawMarkdown": "Hi! Thank you for the reply!"
        }
      ]
    },
    {
      "id": 1263442,
      "postDate": "2021-04-05T12:26:56.040Z",
      "content": "<p>Hi, is the label in trainset, image level, labeled by experts by eye and image or based on antibody and experiment setup?</p>",
      "rawMarkdown": "Hi, is the label in trainset, image level, labeled by experts by eye and image or based on antibody and experiment setup?",
      "replies": [
        {
          "id": 1264799,
          "postDate": "2021-04-06T12:39:41.867Z",
          "content": "<p>You can read a bit on our annotation process <a href=\"https://www.proteinatlas.org/about/assays+annotation#if_annotation\" target=\"_blank\">here</a>. In short: the annotation is primarily done by eye per image, but we use external information to make sure that the information in the images make sense and is accurate.</p>",
          "rawMarkdown": "You can read a bit on our annotation process [here](https://www.proteinatlas.org/about/assays+annotation#if_annotation). In short: the annotation is primarily done by eye per image, but we use external information to make sure that the information in the images make sense and is accurate.",
          "votes": 1
        },
        {
          "id": 1275529,
          "postDate": "2021-04-16T12:01:52.027Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> , thanks for the info.</p>\n<p>Is the annotation process the same for the trainset and the (external) publichpa?</p>\n<ol>\n<li>Can we expect the same quality of annotations in both datasets?</li>\n<li>Is the data from EVE online used in any way?</li>\n<li>Who picked the trainset and how? (is it a random sample from the bigger dataset or was it selected based on some criteria?)</li>\n</ol>\n<p>Thanks again!</p>",
          "rawMarkdown": "Hey @cwinsnes , thanks for the info.\n\nIs the annotation process the same for the trainset and the (external) publichpa?\n2. Can we expect the same quality of annotations in both datasets?\n3. Is the data from EVE online used in any way?\n4. Who picked the trainset and how? (is it a random sample from the bigger dataset or was it selected based on some criteria?)\n\nThanks again!\n"
        },
        {
          "id": 1277755,
          "postDate": "2021-04-19T07:03:59.817Z",
          "content": "<ol>\n<li><p>The quality of the annotations should be similar in both datasets as the process is the same for them in terms of label annotation. The steps that differ are that images that are not on the proteinatlas yet may not have gone through the validation steps which prove the biological correctness, but that should not affect this challenge as you don't work with the antibody information.</p></li>\n<li><p>The EVE online players helped us out tremendously and many annotations in the publichpa set was discovered thanks to them. We don't have any new annotations from them, however, as Project Discovery has moved on to other things. <a href=\"https://www.eveonline.com/discovery\" target=\"_blank\">They are currently helping with COVID19 research. Check it out here</a>! For the specifics of what the Eve players helped us with at the HPA, <a href=\"https://www.nature.com/articles/nbt.4225.epdf\" target=\"_blank\">you can read the article that we wrote here</a>.</p></li>\n<li><p>Not sure I can provide all the details but there is random sampling from a big dataset of high quality images, weighted to make sure all classes are present.</p></li>\n</ol>",
          "rawMarkdown": "1. The quality of the annotations should be similar in both datasets as the process is the same for them in terms of label annotation. The steps that differ are that images that are not on the proteinatlas yet may not have gone through the validation steps which prove the biological correctness, but that should not affect this challenge as you don't work with the antibody information.\n\n2. The EVE online players helped us out tremendously and many annotations in the publichpa set was discovered thanks to them. We don't have any new annotations from them, however, as Project Discovery has moved on to other things. [They are currently helping with COVID19 research. Check it out here](https://www.eveonline.com/discovery)! For the specifics of what the Eve players helped us with at the HPA, [you can read the article that we wrote here](https://www.nature.com/articles/nbt.4225.epdf).\n\n3. Not sure I can provide all the details but there is random sampling from a big dataset of high quality images, weighted to make sure all classes are present.",
          "votes": 1
        },
        {
          "id": 1277864,
          "postDate": "2021-04-19T10:06:27.680Z",
          "content": "<p>This helps a lot! Thanks!</p>",
          "rawMarkdown": "This helps a lot! Thanks!"
        }
      ]
    },
    {
      "id": 1230977,
      "postDate": "2021-03-08T15:32:15.670Z",
      "content": "<p>There are a number of instances where the HPA cell segmenter identifies cells with minimal staining in any channel. E.g. all four channels are nearly black when viewed as images--difficult to see if there's actually a cell present or not. Would these instances typically be deleted by human annotators or treated as \"Negative\" class examples.</p>",
      "rawMarkdown": "There are a number of instances where the HPA cell segmenter identifies cells with minimal staining in any channel. E.g. all four channels are nearly black when viewed as images--difficult to see if there's actually a cell present or not. Would these instances typically be deleted by human annotators or treated as \"Negative\" class examples.",
      "replies": [
        {
          "id": 1230984,
          "postDate": "2021-03-08T15:36:57.723Z",
          "content": "<p>I actually haven't seen this. Can you please show an example?</p>",
          "rawMarkdown": "I actually haven't seen this. Can you please show an example?"
        },
        {
          "id": 1230992,
          "postDate": "2021-03-08T15:41:04.040Z",
          "content": "<p>e559ff24-bbba-11e8-b2ba-ac1f6b6435d0 is a representative example. The staining is quite weak overall, and some of the masks found are visually difficult to distinguish from the background.</p>",
          "rawMarkdown": "e559ff24-bbba-11e8-b2ba-ac1f6b6435d0 is a representative example. The staining is quite weak overall, and some of the masks found are visually difficult to distinguish from the background."
        },
        {
          "id": 1231011,
          "postDate": "2021-03-08T15:58:59.007Z",
          "content": "<p>Hi! we have checked this sample. It is true that the other channels are weak. The nuclei can still be distinguished, hence the model can still make prediction. In this case, human annotator will remove them (delete). </p>",
          "rawMarkdown": "Hi! we have checked this sample. It is true that the other channels are weak. The nuclei can still be distinguished, hence the model can still make prediction. In this case, human annotator will remove them (delete). "
        },
        {
          "id": 1252205,
          "postDate": "2021-03-25T13:52:19.083Z",
          "content": "<p>Sorry for the late reply to this <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> - I was confused by your response.</p>\n<p>Should the indicated slide be labelled as Negative? Or is it simply that many of the cells within the image will be labelled as Negative?</p>",
          "rawMarkdown": "Sorry for the late reply to this @lnhtrang - I was confused by your response.\n\nShould the indicated slide be labelled as Negative? Or is it simply that many of the cells within the image will be labelled as Negative?"
        },
        {
          "id": 1265862,
          "postDate": "2021-04-07T09:06:30.810Z",
          "content": "<p>Hi! In this case, the annotators will just remove the image, so there won't be such example in the test set.</p>",
          "rawMarkdown": "Hi! In this case, the annotators will just remove the image, so there won't be such example in the test set."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1296060,
      "author_name": "CroDoc",
      "author_url": "",
      "post_date": "2021-05-06T23:49:20.847000",
      "content": "<p>In the sample_submission.csv … can we expect the width and height to be correct, so we don't need to check it at all? Seems to be correct on the public test set? Thanks :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1297288,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-07T22:00:13.167000",
          "content": "<p>It would be a very nasty surprise if that were not the case… hope the hosts can confirm :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1297600,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-08T06:59:55.713000",
          "content": "<p>I agree :D ha ha</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1297987,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-05-08T13:28:38.740000",
          "content": "<p>The sample_submission.csv only contains 559 images from the public test set right, the width and height are correct in that file for public test set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298004,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-08T13:40:41.547000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> the question is whether the width and the height are correct in that file for private test set? We assume it is, so we use that data to accelerate segmentation. If it is incorrect, then that would impact the score in negative way. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298008,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-05-08T13:44:00.550000",
          "content": "<p>dumb question but where can I see that file? I only see the sample_submission.csv for public test set :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298010,
          "author_name": "sheep",
          "author_url": "",
          "post_date": "2021-05-08T13:47:26.100000",
          "content": "<p>There are no record of such issue in perious competition I joined. But, doesn't the public/priave split provided by host? (you may ignore my question if anything sensitive)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298014,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2021-05-08T13:50:28.533000",
          "content": "<p>We can't see that file as participants - our models are evaluated on it but it is not available. <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> might be able to confirm? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298018,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-05-08T13:57:13.840000",
          "content": "<p>probably it's me being unfamiliar with kaggle platform. I just checked the <code>sample_submission.csv</code> file in the <code>Data</code> tab, which only contains 559 images, and their width and height are all correct. If it's a publicly available file, then it's not competition-sensitive. This file was produced by <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> so maybe he is a better person to confirm. I assume it's produced by the same code, so if the public test set is correct, the private test set should be as well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298051,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-08T14:21:20.933000",
          "content": "<p>Does the sample_submission.csv when evaluted as a full submission contain all test images (public+private)? And is the widht and height information correct for both private and public test images?</p>\n<p>I assume sample_submission.csv file is replaced when the full submission is made in the same way as it is with the test/ folder (test images)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298059,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-05-08T14:26:03.197000",
          "content": "<p>I found another <code>sample_submission.csv</code> file as host, in which all width and height are correct for all test set. Hope that <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> can see this and have a final confirmation, but I think your assumption is correct <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> <a href=\"https://www.kaggle.com/crodoc\" target=\"_blank\">@crodoc</a> so please proceed as normal.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1298075,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-08T14:34:35.807000",
          "content": "<p>This sounds good! Thanks <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a>.<br>\nI guess a lot of us are just loading the <code>sample_submission.csv</code> file and just replacing the <code>PredictionString</code> part. If the file was not updated at full submission time, I think a lot of teams would score 0 on the private LB.</p>\n<p>I guess you are doing the same thing <a href=\"https://www.kaggle.com/thedrcat\" target=\"_blank\">@thedrcat</a> :D</p>\n<p>Would be great if <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> also confirms, otherwise a lot of us need to do some fast updates to our notebooks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1298851,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-09T09:16:29.943000",
          "content": "<p>Hey, I just tested it by reading the files from the test directory into a dataframe and comparing this dataframe to the one provided as <code>sample_submission.csv</code>. It's not giving any error…</p>\n<pre><code>import glob\nimport re\n\ndata_df_sample_submission = pd.read_csv('../input/hpa-single-cell-image-classification/sample_submission.csv')\n\ntest_files = os.listdir(\"../input/hpa-single-cell-image-classification/test\")\ncolor_list = [\"_red.png\", \"_green.png\", \"_yellow.png\", \"_blue.png\"]\ntest_files_names =  [re.sub(r'|'.join(map(re.escape, color_list)), '', elem) for elem in test_files]\ntest_files_names = list(set(test_files_names))\n\nd = []\nfor i, file in enumerate(test_files_names):\n    img = cv2.imread(f\"../input/hpa-single-cell-image-classification/test/{file}_red.png\")\n    height, width, channels = img.shape\n    d.append({\n        \"ID\" : file,\n        \"ImageWidth\" : width,\n        \"ImageHeight\": height,\n        \"PredictionString\" : \"0 1 eNoLCAgIMAEABJkBdQ==\"\n    })\n    if i%50 == 0:\n        print(i)\n\ndf_from_files = pd.DataFrame(d).sort_values(\"ID\").reset_index(drop=True)\n</code></pre>\n<p>and last line!:</p>\n<pre><code>if all(df_from_files == data_df_sample_submission):\n    sub.to_csv(\"submission.csv\",index=False)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1298989,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-09T11:47:04.727000",
          "content": "<p>Great job!!!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298997,
          "author_name": "Alexander Riedel",
          "author_url": "",
          "post_date": "2021-05-09T11:51:40.747000",
          "content": "<p>Still I would be happy to have this confirmed by a host 😄</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1302256,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-05-11T13:06:18.960000",
          "content": "<p>I did a search on the kaggle site … seems the <code>sample_submission.csv</code> file will be replaced on full submission … yaaay :)</p>\n<p>Check out a few links I picked from the search:<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222037#1217676</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/202223#1107145</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/177356#985410</a><br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/petfinder-adoption-prediction/discussion/77241#453866</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1231013,
      "author_name": "Andrew Tratz",
      "author_url": "",
      "post_date": "2021-03-08T16:00:35.787000",
      "content": "<p>Some of the cells appear to be undergoing (or recently completed) cytokinesis. At what point in this process would human annotators consider them to be two distinct cells for the purposes of the competition? I.e. if two nuclei with complete nuclear membranes are physically in contact and are identified as a single cell by the HPA cell segmenter would a human annotator typically split this into two distinct cell masks?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1231104,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-03-08T17:08:11.083000",
          "content": "<p>If 2 nuclei and membrane can be distinguished, HPACellSegmentator will post process it as 2 separate cells. If the Segmentator fails to do so, human annotator will do it for the test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1268967,
      "author_name": "NakedKoala",
      "author_url": "",
      "post_date": "2021-04-10T02:21:26.410000",
      "content": "<p><a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a></p>\n<p>A quick question about the proportion of negative cells in slides that come with unique label ( non-negative label ).</p>\n<p>Based on the domain knowledge ( of how these slides are produced and how different cell types cluster in the slides )</p>\n<ol>\n<li>Should the average negative proportion be small  below 10%) or large (above 80 percent) or somewhere in between ? </li>\n<li>Should I expect the variance of negative proportion to be large ?  </li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 1271074,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-12T09:20:57.333000",
          "content": "<p>Hi! The distribution of <code>Negative</code> labels does vary a lot, based on the tagged protein, so it can be anywhere from &lt;10 to 90%. A clue is that high <code>Negative</code> % for proteins related to mitosis. I don't exactly know the proportion on average for train set.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1231169,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2021-03-08T18:13:40.993000",
      "content": "<p>In one post in this forum, someone said that's it's not possible to have, for one cell, class 18 (negative) mixed with any other classes for the same cell. Does it matter for the scoring if we submit such combination?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1231441,
          "author_name": "cool_rabbit",
          "author_url": "",
          "post_date": "2021-03-09T01:48:26.557000",
          "content": "<p>I have the same question!!<br>\nTheoretically impossible because of Negative, but waiting for the official answer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1232166,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-03-09T14:30:41.893000",
          "content": "<p>The <code>Negative</code> class is treated as any other class in terms of the scoring, so if you combine it with other classes for the same cell it will be scored in the same manner as any other combination of classes.</p>\n<p>However, as stated before: there will be no instances of a cell labeled <code>Negative</code> that is also labeled as something else.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1230548,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2021-03-08T08:21:39.017000",
      "content": "<p>Thanks a lot for this offer! I have a question related to this comment:</p>\n<blockquote>\n  <p>During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label. </p>\n</blockquote>\n<p>Does this mean that we can have train/public set images with mitotic spindle label but actually no mitotic spindle represented, because it only appeared in other images from the same sample? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1230967,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-03-08T15:19:42.053000",
          "content": "<p>You are correct. There could be images tagged with <code>Mitotic spindle</code> but actually no mitotic spindle present. But these should be rare cases, because the production team took up to 6 images and display the best 2 on the atlas (So the public HPA are the better quality ones, and will almost always have the pattern). This problem is most pronounced for mitotic structures (that appears in about 5% of all cells at any given timepoint). Other organelles don't suffer from this problem.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1286177,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2021-04-27T16:35:06.087000",
          "content": "<p>hi . i have a small doubt . what do you mean by mitotic structures. what are all the classes that have relation to mitotic structures. i am trying to understand how to handle these classes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1287551,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-29T07:11:01.173000",
          "content": "<p>Hi! for this competition, the only class is <code>Mitotic spindle</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1299909,
      "author_name": "novice03",
      "author_url": "",
      "post_date": "2021-05-10T06:25:56.037000",
      "content": "<p>Hi, can a host please take a look at <a href=\"https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745\" target=\"_blank\">https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1289326,
      "author_name": "Raman",
      "author_url": "",
      "post_date": "2021-04-30T21:09:02.103000",
      "content": "<p>Dear organizers,</p>\n<p>when you have a moment could you please clarify the mitotic spindle labeling of the following cells:</p>\n<ol>\n<li><p>This cell comes from an image labeled as <em>Mitotic spindle</em>, the cell seems to be the only one in the mitotic phase in the image.<br>\n<img src=\"https://i.ibb.co/HYxjdrX/positive-mitotic.png\" alt=\"\"></p></li>\n<li><p>The following cells are from images <strong>without</strong> the <em>Mitotic spindle</em> label.<br>\n<img src=\"https://i.ibb.co/7QJph0y/negative-mitotic.png\" alt=\"\"><br>\n<img src=\"https://i.ibb.co/Pt8xNsM/q2.png\" alt=\"\"></p></li>\n</ol>\n<p><strong>What bothers me:</strong><br>\nDoes the cell from 1. have an apparent mitotic spindle pattern? Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?</p>\n<p>Thank you in advance!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1297942,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-05-08T12:49:36.570000",
          "content": "<p>Hi! Sorry for the late reply. I actually don't see the images you attached, even after switching browsers.</p>\n<blockquote>\n  <p>Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?</p>\n</blockquote>\n<p>Well mitotic spindle appears in cells that are going through mitosis, so they do have different shapes and organization of nucleus/microtubules. Neural networks are effective at picking these signals up, so they may do well with just green signal, but I suspect red and blue channels would help alot.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1298137,
          "author_name": "Raman",
          "author_url": "",
          "post_date": "2021-05-08T15:20:57.513000",
          "content": "<p>Hi! Thank you for the reply!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1263442,
      "author_name": "sheep",
      "author_url": "",
      "post_date": "2021-04-05T12:26:56.040000",
      "content": "<p>Hi, is the label in trainset, image level, labeled by experts by eye and image or based on antibody and experiment setup?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1264799,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-04-06T12:39:41.867000",
          "content": "<p>You can read a bit on our annotation process <a href=\"https://www.proteinatlas.org/about/assays+annotation#if_annotation\" target=\"_blank\">here</a>. In short: the annotation is primarily done by eye per image, but we use external information to make sure that the information in the images make sense and is accurate.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1275529,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-04-16T12:01:52.027000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/cwinsnes\" target=\"_blank\">@cwinsnes</a> , thanks for the info.</p>\n<p>Is the annotation process the same for the trainset and the (external) publichpa?</p>\n<ol>\n<li>Can we expect the same quality of annotations in both datasets?</li>\n<li>Is the data from EVE online used in any way?</li>\n<li>Who picked the trainset and how? (is it a random sample from the bigger dataset or was it selected based on some criteria?)</li>\n</ol>\n<p>Thanks again!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1277755,
          "author_name": "Casper Winsnes",
          "author_url": "",
          "post_date": "2021-04-19T07:03:59.817000",
          "content": "<ol>\n<li><p>The quality of the annotations should be similar in both datasets as the process is the same for them in terms of label annotation. The steps that differ are that images that are not on the proteinatlas yet may not have gone through the validation steps which prove the biological correctness, but that should not affect this challenge as you don't work with the antibody information.</p></li>\n<li><p>The EVE online players helped us out tremendously and many annotations in the publichpa set was discovered thanks to them. We don't have any new annotations from them, however, as Project Discovery has moved on to other things. <a href=\"https://www.eveonline.com/discovery\" target=\"_blank\">They are currently helping with COVID19 research. Check it out here</a>! For the specifics of what the Eve players helped us with at the HPA, <a href=\"https://www.nature.com/articles/nbt.4225.epdf\" target=\"_blank\">you can read the article that we wrote here</a>.</p></li>\n<li><p>Not sure I can provide all the details but there is random sampling from a big dataset of high quality images, weighted to make sure all classes are present.</p></li>\n</ol>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1277864,
          "author_name": "CroDoc",
          "author_url": "",
          "post_date": "2021-04-19T10:06:27.680000",
          "content": "<p>This helps a lot! Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1230977,
      "author_name": "Andrew Tratz",
      "author_url": "",
      "post_date": "2021-03-08T15:32:15.670000",
      "content": "<p>There are a number of instances where the HPA cell segmenter identifies cells with minimal staining in any channel. E.g. all four channels are nearly black when viewed as images--difficult to see if there's actually a cell present or not. Would these instances typically be deleted by human annotators or treated as \"Negative\" class examples.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1230984,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-03-08T15:36:57.723000",
          "content": "<p>I actually haven't seen this. Can you please show an example?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1230992,
          "author_name": "Andrew Tratz",
          "author_url": "",
          "post_date": "2021-03-08T15:41:04.040000",
          "content": "<p>e559ff24-bbba-11e8-b2ba-ac1f6b6435d0 is a representative example. The staining is quite weak overall, and some of the masks found are visually difficult to distinguish from the background.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1231011,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-03-08T15:58:59.007000",
          "content": "<p>Hi! we have checked this sample. It is true that the other channels are weak. The nuclei can still be distinguished, hence the model can still make prediction. In this case, human annotator will remove them (delete). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1252205,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2021-03-25T13:52:19.083000",
          "content": "<p>Sorry for the late reply to this <a href=\"https://www.kaggle.com/lnhtrang\" target=\"_blank\">@lnhtrang</a> - I was confused by your response.</p>\n<p>Should the indicated slide be labelled as Negative? Or is it simply that many of the cells within the image will be labelled as Negative?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1265862,
          "author_name": "Trang Le",
          "author_url": "",
          "post_date": "2021-04-07T09:06:30.810000",
          "content": "<p>Hi! In this case, the annotators will just remove the image, so there won't be such example in the test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1226778": "Update: Although the AMA has passed, we will check the thread occasionally so please ask if you have any Qs!\n\nHi All!\n\nWe have received many great questions in this competition. We want to help you in understanding the data and the set-up, which ultimately leads to a better solution. However, It is a bit hard to keep track of all discussions, and there are many questions that have been asked multiple times across different threads. So we would like to host an AMA (Ask Me/Us Anything) session at **4-5pm CET on Monday March 8th**. At this time is when many people in our team will be online (especially our expert annotators who created the groundtruth) and we will try to answer as many questions as possible.\n\nYou can also post questions here beforehand, and we will answer them in order of upvotes (we will of course try to answer them all). \n\nBefore you ask questions, please check out these discussions first to see if your questions have already been answered here:\n\nHow individual single cell pattern looks like, what is SCV https://www.kaggle.com/lnhtrang/single-cell-patterns \nExternal HPA data, overlap between trainset and external HPA data, convert external HPA classes to 19 classes in this competition https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg \nCellSegmentator https://www.kaggle.com/lnhtrang/hpa-public-data-download-and-hpacellseg\nHow does groundtruth look like https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/215141 \nAbout Negative: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/220748 \nAbout confidence: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/219885\nWhy there are duplicate labels for the same image, merged labels: https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/217806#1193562 \n\nAll the bests,\n@emmalumpan, @weiouyang, @uaxelsson, @cwinsnes and @lnhtrang",
    "1296060": "In the sample_submission.csv ... can we expect the width and height to be correct, so we don't need to check it at all? Seems to be correct on the public test set? Thanks :)",
    "1231013": "Some of the cells appear to be undergoing (or recently completed) cytokinesis. At what point in this process would human annotators consider them to be two distinct cells for the purposes of the competition? I.e. if two nuclei with complete nuclear membranes are physically in contact and are identified as a single cell by the HPA cell segmenter would a human annotator typically split this into two distinct cell masks?",
    "1268967": "@lnhtrang\n\nA quick question about the proportion of negative cells in slides that come with unique label ( non-negative label ).\n\nBased on the domain knowledge ( of how these slides are produced and how different cell types cluster in the slides )\n\n1.  Should the average negative proportion be small  below 10%) or large (above 80 percent) or somewhere in between ? \n2. Should I expect the variance of negative proportion to be large ?  \n\n\n\n",
    "1231169": "In one post in this forum, someone said that's it's not possible to have, for one cell, class 18 (negative) mixed with any other classes for the same cell. Does it matter for the scoring if we submit such combination?",
    "1230548": "Thanks a lot for this offer! I have a question related to this comment:\n> During annotation, the image-level labels are set per sample (i.e per a group of up to 6 images from the same sample). This means that common labels present will be annotated. Mitotic spindle is a cellular structure that appears when cells divide, which happens approximately once per 24 hours. This means that only 1 in 20-50 cells in a population will at any given time point show mitotic spindles. Hence, it is not uncommon that we see a spindle in 2 images from a sample but not in 4, yet all images would get the mitotic spindle label. \n\nDoes this mean that we can have train/public set images with mitotic spindle label but actually no mitotic spindle represented, because it only appeared in other images from the same sample? \n",
    "1299909": "Hi, can a host please take a look at https://www.kaggle.com/c/hpa-single-cell-image-classification/discussion/237745",
    "1289326": "Dear organizers,\n\nwhen you have a moment could you please clarify the mitotic spindle labeling of the following cells:\n\n1. This cell comes from an image labeled as *Mitotic spindle*, the cell seems to be the only one in the mitotic phase in the image.\n![](https://i.ibb.co/HYxjdrX/positive-mitotic.png)\n\n2. The following cells are from images **without** the *Mitotic spindle* label.\n![](https://i.ibb.co/7QJph0y/negative-mitotic.png)\n![](https://i.ibb.co/Pt8xNsM/q2.png)\n\n**What bothers me:**\nDoes the cell from 1. have an apparent mitotic spindle pattern? Is the green channel enough to proclaim mitotic spindle label or do we also need obvious microtubules in the red channel?\n\nThank you in advance!",
    "1263442": "Hi, is the label in trainset, image level, labeled by experts by eye and image or based on antibody and experiment setup?",
    "1230977": "There are a number of instances where the HPA cell segmenter identifies cells with minimal staining in any channel. E.g. all four channels are nearly black when viewed as images--difficult to see if there's actually a cell present or not. Would these instances typically be deleted by human annotators or treated as \"Negative\" class examples."
  }
}