{
  "id": 227616,
  "title": "Are we allowed to manually draw pseudo label for training?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/227616",
  "author_name": "",
  "post_date": "2021-03-21T12:21:36.489314200Z",
  "votes": 47,
  "comment_count": 48,
  "views": 0,
  "content": "<p>I found that my model often miss dark glomeruli in public test data. And stronger data augmentation didn't work. So I manually draw pseudo label of some of these dark glomeruli and use them in training.</p>\n<p>The model trained with pseudo label gave me a great boost (LB from 0.914 to 0.932)</p>\n<p>But I'm not sure whether it is allowed in this competition.</p>\n<p>ps. I don't mean to break the rules😭</p>\n<p><strong>UPDATE:</strong><br>\nMy problem and my idea are shown below:</p>\n<p>I found my score of d488c759a single image was much lower than other 4 images. I checked my prediction result and found that there are some darker, smaller things which are not similar with normal glomeruli. They distributed more densely at the boundary of the slice and there seems to be less cell nucleus in these structures.<br>\n<img src=\"https://i.postimg.cc/pLbNLQF9/tmp.png\" alt=\"\"></p>\n<p>I checked all images from training data, public test data and external data from HUBMAP, and found that these things only exist in d488c759a image. So I assume these things cause my low score of d488c759a.</p>\n<p>I tried some data augmentation tricks but didn't help. Finally, I hand-labelled some of them (not whole image, not whole dataset) as pseudo-label and trained a model with the pseudo-label. </p>\n<p>I am really sorry breaking rules unintentionally. I hope someone could solve this problem with normal deep learning tricks!!</p>\n<p><strong>UPDATE2</strong></p>\n<p>My hand-labelled (for dark glomeruli) + pseudo-labelled (for normal glomeruli) mask of d488c759a </p>\n<p><a href=\"https://www.kaggle.com/carnozhao/d48-hand-labelled\" target=\"_blank\">https://www.kaggle.com/carnozhao/d48-hand-labelled</a></p>",
  "messages": [
    {
      "id": "1247107",
      "postDate": "03/21/2021 12:21:36",
      "content": "<p>I found that my model often miss dark glomeruli in public test data. And stronger data augmentation didn't work. So I manually draw pseudo label of some of these dark glomeruli and use them in training.</p>\n<p>The model trained with pseudo label gave me a great boost (LB from 0.914 to 0.932)</p>\n<p>But I'm not sure whether it is allowed in this competition.</p>\n<p>ps. I don't mean to break the rules😭</p>\n<p><strong>UPDATE:</strong><br>\nMy problem and my idea are shown below:</p>\n<p>I found my score of d488c759a single image was much lower than other 4 images. I checked my prediction result and found that there are some darker, smaller things which are not similar with normal glomeruli. They distributed more densely at the boundary of the slice and there seems to be less cell nucleus in these structures.<br>\n<img src=\"https://i.postimg.cc/pLbNLQF9/tmp.png\" alt=\"\"></p>\n<p>I checked all images from training data, public test data and external data from HUBMAP, and found that these things only exist in d488c759a image. So I assume these things cause my low score of d488c759a.</p>\n<p>I tried some data augmentation tricks but didn't help. Finally, I hand-labelled some of them (not whole image, not whole dataset) as pseudo-label and trained a model with the pseudo-label. </p>\n<p>I am really sorry breaking rules unintentionally. I hope someone could solve this problem with normal deep learning tricks!!</p>\n<p><strong>UPDATE2</strong></p>\n<p>My hand-labelled (for dark glomeruli) + pseudo-labelled (for normal glomeruli) mask of d488c759a </p>\n<p><a href=\"https://www.kaggle.com/carnozhao/d48-hand-labelled\" target=\"_blank\">https://www.kaggle.com/carnozhao/d48-hand-labelled</a></p>",
      "rawMarkdown": "I found that my model often miss dark glomeruli in public test data. And stronger data augmentation didn't work. So I manually draw pseudo label of some of these dark glomeruli and use them in training.\n\nThe model trained with pseudo label gave me a great boost (LB from 0.914 to 0.932)\n\nBut I'm not sure whether it is allowed in this competition.\n\nps. I don't mean to break the rules😭\n\n**UPDATE:**\nMy problem and my idea are shown below:\n\nI found my score of d488c759a single image was much lower than other 4 images. I checked my prediction result and found that there are some darker, smaller things which are not similar with normal glomeruli. They distributed more densely at the boundary of the slice and there seems to be less cell nucleus in these structures.\n![](https://i.postimg.cc/pLbNLQF9/tmp.png)\n\nI checked all images from training data, public test data and external data from HUBMAP, and found that these things only exist in d488c759a image. So I assume these things cause my low score of d488c759a.\n\nI tried some data augmentation tricks but didn't help. Finally, I hand-labelled some of them (not whole image, not whole dataset) as pseudo-label and trained a model with the pseudo-label. \n\nI am really sorry breaking rules unintentionally. I hope someone could solve this problem with normal deep learning tricks!!\n\n**UPDATE2**\n\nMy hand-labelled (for dark glomeruli) + pseudo-labelled (for normal glomeruli) mask of d488c759a \n\n[https://www.kaggle.com/carnozhao/d48-hand-labelled](https://www.kaggle.com/carnozhao/d48-hand-labelled)",
      "votes": null
    },
    {
      "id": "1247112",
      "postDate": "03/21/2021 12:32:01",
      "content": "<p>Good question.</p>\n<p>From <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884</a> : </p>\n<blockquote>\n  <p>Given that the new private data is not available via the HuBMAP portal, hand-labeling of the data is now allowed.</p>\n</blockquote>\n<p>From <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/rules\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/rules</a> : </p>\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>The tricky part comes from that you're hand labelling the public test set. I cannot say for sure but my guess is that it's not allowed. You can however hand label data that is not in the test data.</p>",
      "rawMarkdown": "Good question.\n\nFrom https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884 : \n> Given that the new private data is not available via the HuBMAP portal, hand-labeling of the data is now allowed.\n\nFrom https://www.kaggle.com/c/hubmap-kidney-segmentation/rules : \n> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nThe tricky part comes from that you're hand labelling the public test set. I cannot say for sure but my guess is that it's not allowed. You can however hand label data that is not in the test data.",
      "votes": null
    },
    {
      "id": "1247114",
      "postDate": "03/21/2021 12:40:12",
      "content": "<p>Thanks! BTW, how can I withdraw my submission</p>",
      "rawMarkdown": "Thanks! BTW, how can I withdraw my submission",
      "votes": null
    },
    {
      "id": "1247123",
      "postDate": "03/21/2021 13:01:05",
      "content": "<p>I don't think it's possible ^^<br>\nThere's still a lot of time so people should catch-up to your score, don't worry about it.</p>",
      "rawMarkdown": "I don't think it's possible ^^\nThere's still a lot of time so people should catch-up to your score, don't worry about it.",
      "votes": null
    },
    {
      "id": "1247253",
      "postDate": "03/21/2021 15:14:40",
      "content": "<p>would you mind showing your method for handing draw. :) I am curious about it.</p>",
      "rawMarkdown": "would you mind showing your method for handing draw. :) I am curious about it.",
      "votes": null
    },
    {
      "id": "1247294",
      "postDate": "03/21/2021 15:46:15",
      "content": "<p>I draw yellow mask over the image in windows paint. After loading the image, mask is where the pixel equals to (255, 255, 0)</p>\n<p>Of course you can try other color.</p>",
      "rawMarkdown": "I draw yellow mask over the image in windows paint. After loading the image, mask is where the pixel equals to (255, 255, 0)\n\nOf course you can try other color.",
      "votes": null
    },
    {
      "id": "1247360",
      "postDate": "03/21/2021 16:52:01",
      "content": "<p>Got it! windows paint. nice method hhh.</p>",
      "rawMarkdown": "Got it! windows paint. nice method hhh.",
      "votes": null
    },
    {
      "id": "1247703",
      "postDate": "03/22/2021 01:31:53",
      "content": "<p>Nice score, but honestly I did not really figure it out. Did you mean you pseudo labled some of the public test dataset and use those to train your model, and finally used this model to predict the public test dataset itself? Won't it be data leakage?</p>",
      "rawMarkdown": "Nice score, but honestly I did not really figure it out. Did you mean you pseudo labled some of the public test dataset and use those to train your model, and finally used this model to predict the public test dataset itself? Won't it be data leakage?",
      "votes": null
    },
    {
      "id": "1247711",
      "postDate": "03/22/2021 01:45:20",
      "content": "<p>Most of my pseudo labels are generated by model prediction. Besides that, I only draw pseudo labels on part of one image, about 10~20 glomeruli. I think there is data leakage but not too much. </p>",
      "rawMarkdown": "Most of my pseudo labels are generated by model prediction. Besides that, I only draw pseudo labels on part of one image, about 10~20 glomeruli. I think there is data leakage but not too much.",
      "votes": null
    },
    {
      "id": "1247972",
      "postDate": "03/22/2021 08:27:51",
      "content": "<p>So in summary hand labeling is allowed as long as it isn't the test dataset (i.e. the public part). <br>\nShould hand labeling external datasets be declared or are we good even if not acknowledging this?</p>",
      "rawMarkdown": "So in summary hand labeling is allowed as long as it isn't the test dataset (i.e. the public part). \nShould hand labeling external datasets be declared or are we good even if not acknowledging this?",
      "votes": null
    },
    {
      "id": "1249808",
      "postDate": "03/23/2021 14:33:12",
      "content": "<p>Hi all - I'll confirm Yassine's comment here - Hand labeling is allowed per the host's ruling, as long as it isn't the test dataset.</p>\n<p>Carno - please contact <a href=\"https://www.kaggle.com/compliance\" target=\"_blank\">compliance</a> and we'll see if we can work with you to invalidate the submissions.</p>",
      "rawMarkdown": "Hi all - I'll confirm Yassine's comment here - Hand labeling is allowed per the host's ruling, as long as it isn't the test dataset.\n\nCarno - please contact [compliance](https://www.kaggle.com/compliance) and we'll see if we can work with you to invalidate the submissions.",
      "votes": null
    },
    {
      "id": "1250108",
      "postDate": "03/23/2021 19:46:20",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> thanks for confirming! </p>",
      "rawMarkdown": "addisonhoward thanks for confirming!",
      "votes": null
    },
    {
      "id": "1250357",
      "postDate": "03/24/2021 01:34:58",
      "content": "<p>thanks, I have contacted the compliance</p>",
      "rawMarkdown": "thanks, I have contacted the compliance",
      "votes": null
    },
    {
      "id": "1250442",
      "postDate": "03/24/2021 03:47:32",
      "content": "<p>Hi all,</p>\n<p>We've discussed this internally and will note that both hand-labeling and pseudo-labeling of the data is allowed per the competition host's comments. As this is a code competition, there is no ability to label the unseen, private test set, we do not expect it to gain any advantage on the models that determine the final winners. While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.</p>\n<p>We don't invalidate submissions by request (my apologies for stating such above). In any future competition where you are concerned about a submission you made that may violate certain rules, your best course of action is to <em>not select those submissions for final scoring.</em> This does not apply to private code sharing or use of multiple accounts, in which case it does not matter which submissions you select, it is still considered a violation.</p>\n<p>I've updated the rules accordingly for avoidance of doubt regarding such.</p>\n<p>Thanks,</p>\n<p>Addison</p>\n<p>cc/ <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>",
      "rawMarkdown": "Hi all,\n\nWe've discussed this internally and will note that both hand-labeling and pseudo-labeling of the data is allowed per the competition host's comments. As this is a code competition, there is no ability to label the unseen, private test set, we do not expect it to gain any advantage on the models that determine the final winners. While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.\n\nWe don't invalidate submissions by request (my apologies for stating such above). In any future competition where you are concerned about a submission you made that may violate certain rules, your best course of action is to *not select those submissions for final scoring.* This does not apply to private code sharing or use of multiple accounts, in which case it does not matter which submissions you select, it is still considered a violation.\n\nI've updated the rules accordingly for avoidance of doubt regarding such.\n\nThanks,\n\nAddison\n\ncc/ @carnozhao",
      "votes": null
    },
    {
      "id": "1250582",
      "postDate": "03/24/2021 06:12:48",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <br>\nIs there any reason why host published the public test data?<br>\nOther code competitions did't publish public test data. </p>\n<p>Currently, I feel that the host is asking the participants to hand labeling.</p>\n<blockquote>\n  <p>While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.</p>\n</blockquote>\n<p>If the public test data doesn't affect the unseen test set, shouldn't the label of the public test data be public?</p>",
      "rawMarkdown": "addisonhoward \nIs there any reason why host published the public test data?\nOther code competitions did't publish public test data. \n\nCurrently, I feel that the host is asking the participants to hand labeling.\n\n> While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.\n\nIf the public test data doesn't affect the unseen test set, shouldn't the label of the public test data be public?",
      "votes": null
    },
    {
      "id": "1250593",
      "postDate": "03/24/2021 06:27:21",
      "content": "<p>Should we compete for hand-labeling skills with more and more train data…?😨</p>",
      "rawMarkdown": "Should we compete for hand-labeling skills with more and more train data...?😨",
      "votes": null
    },
    {
      "id": "1250598",
      "postDate": "03/24/2021 06:32:54",
      "content": "<p>Welcome to Kaggle, where you can attend the best hand-labeling competitions in the world!</p>",
      "rawMarkdown": "Welcome to Kaggle, where you can attend the best hand-labeling competitions in the world!",
      "votes": null
    },
    {
      "id": "1250603",
      "postDate": "03/24/2021 06:36:12",
      "content": "<p>Any pathologist? lol</p>",
      "rawMarkdown": "Any pathologist? lol",
      "votes": null
    },
    {
      "id": "1251153",
      "postDate": "03/24/2021 14:25:56",
      "content": "<p>For research organizations with public funding, often times there are requirements about how and when the data is released. For a project this large, it can be difficult to amass a truly hidden test set for scoring. </p>\n<p>We do not release the public test set labels, as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score. </p>",
      "rawMarkdown": "For research organizations with public funding, often times there are requirements about how and when the data is released. For a project this large, it can be difficult to amass a truly hidden test set for scoring. \n\nWe do not release the public test set labels, as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.",
      "votes": null
    },
    {
      "id": "1251865",
      "postDate": "03/25/2021 08:04:28",
      "content": "<blockquote>\n  <p>as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.</p>\n</blockquote>\n<p>We should validate model with train data, not leaderboard. We can use leaderboard for comparing other kaggler's model and competing kagglers. I think that is the main purpose of participating in the competition.<br>\nHand labeling just break public leaderboard, and also,  it pour cold water on the enthusiasm of the participants, especially the beginners.</p>",
      "rawMarkdown": "> as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.\n\nWe should validate model with train data, not leaderboard. We can use leaderboard for comparing other kaggler's model and competing kagglers. I think that is the main purpose of participating in the competition.\nHand labeling just break public leaderboard, and also,  it pour cold water on the enthusiasm of the participants, especially the beginners.",
      "votes": null
    },
    {
      "id": "1252081",
      "postDate": "03/25/2021 12:13:49",
      "content": "<p>What? hand labeling of public test data is allowed?</p>",
      "rawMarkdown": "What? hand labeling of public test data is allowed?",
      "votes": null
    },
    {
      "id": "1252159",
      "postDate": "03/25/2021 13:15:39",
      "content": "<p>The updated rule means we should compete for hand labeling skills on Public LB? That is nonsense…</p>",
      "rawMarkdown": "The updated rule means we should compete for hand labeling skills on Public LB? That is nonsense...",
      "votes": null
    },
    {
      "id": "1252389",
      "postDate": "03/25/2021 16:40:23",
      "content": "<p>Did that mean we can use hand-labeled samples to strengthen our model? If so, the significance will go bad. I suggest making the public test samples into training dataset. using a part of the other (&lt;67%) to be the real public test set therefore this competition could have a chance to be a typical machine learning competition rather than a painter game.</p>",
      "rawMarkdown": "Did that mean we can use hand-labeled samples to strengthen our model? If so, the significance will go bad. I suggest making the public test samples into training dataset. using a part of the other (<67%) to be the real public test set therefore this competition could have a chance to be a typical machine learning competition rather than a painter game.",
      "votes": null
    },
    {
      "id": "1252393",
      "postDate": "03/25/2021 16:44:09",
      "content": "<p>Agreed. <a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> </p>",
      "rawMarkdown": "Agreed. @southsakura",
      "votes": null
    },
    {
      "id": "1252751",
      "postDate": "03/26/2021 02:16:34",
      "content": "<p>Agreed with <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>. Don't know why the organizer of this challenge does not look at other challenges that have been hosted in Kaggle. No one set a rule like this to hand-labelling the test data in the leaderboard. I think this competition should focus more on quality of the methods and tackling real-world problem with limited data rather than reaching a Dice Score of 99.0%. </p>\n<p>On the other hand, the Kagglers rather focusing on innovative approaches, would focus more on labelling the data which reduce the impact of this challenge. </p>",
      "rawMarkdown": "Agreed with @phalanx. Don't know why the organizer of this challenge does not look at other challenges that have been hosted in Kaggle. No one set a rule like this to hand-labelling the test data in the leaderboard. I think this competition should focus more on quality of the methods and tackling real-world problem with limited data rather than reaching a Dice Score of 99.0%. \n\nOn the other hand, the Kagglers rather focusing on innovative approaches, would focus more on labelling the data which reduce the impact of this challenge.",
      "votes": null
    },
    {
      "id": "1252775",
      "postDate": "03/26/2021 03:24:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>,</p>\n<p>You manualy painted the image with a pain program and SAVE it ?, and use it for training? I don't understand,and where is the mask? did you used it? are you using the lafoss title notebook to break the image in tiles?</p>\n<p>Can you explain a little bit more about how and what you used to do this.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @carnozhao,\n\nYou manualy painted the image with a pain program and SAVE it ?, and use it for training? I don't understand,and where is the mask? did you used it? are you using the lafoss title notebook to break the image in tiles?\n\nCan you explain a little bit more about how and what you used to do this.\n\nThanks!",
      "votes": null
    },
    {
      "id": "1253856",
      "postDate": "03/27/2021 05:23:27",
      "content": "<p>Would like to add this from the RANZCR competitions experience and believe competition rules here should be amended in a similar way but clarifying the public test data set since there are now old test that could be considered external data perhaps -</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644</a>  <br>\n\"There have been some questions from competitors regarding hand-labelled annotations. To clarify the rules as per section A2:</p>\n<pre><code>Publicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"\n</code></pre>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> this is with respect to this public test annotation but could apply to other external data.  There is no way to know if the private test data also has \"dark glomeruli\" as <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> discovered in the public test data.  If so this confers an advantage to those that have these annotations. </p>",
      "rawMarkdown": "Would like to add this from the RANZCR competitions experience and believe competition rules here should be amended in a similar way but clarifying the public test data set since there are now old test that could be considered external data perhaps -\n \nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644  \n\"There have been some questions from competitors regarding hand-labelled annotations. To clarify the rules as per section A2:\n\n    Publicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"\n\n@addisonhoward this is with respect to this public test annotation but could apply to other external data.  There is no way to know if the private test data also has \"dark glomeruli\" as @carnozhao discovered in the public test data.  If so this confers an advantage to those that have these annotations.",
      "votes": null
    },
    {
      "id": "1254597",
      "postDate": "03/27/2021 21:11:59",
      "content": "<p>Not sure, but shouldn't a hand-labeled dataset be regarded as a dataset different from the competitions dataset and, thus, it must be classified as external.  In that case, it should be publicly available to all.</p>\n<p>Of course, it would not be fair to those who had done the hard part of hand-labeling, so I am not advocating this. On the contrary, the disclosure better be voluntary.</p>\n<p>Another point of consideration: if this type of glomeruli (is it 'sclerotic' glomeruli?) exists only in one out of 20 images in train and public test set, what are the chances that there are any in the private test set.</p>",
      "rawMarkdown": "Not sure, but shouldn't a hand-labeled dataset be regarded as a dataset different from the competitions dataset and, thus, it must be classified as external.  In that case, it should be publicly available to all.\n\nOf course, it would not be fair to those who had done the hard part of hand-labeling, so I am not advocating this. On the contrary, the disclosure better be voluntary.\n\nAnother point of consideration: if this type of glomeruli (is it 'sclerotic' glomeruli?) exists only in one out of 20 images in train and public test set, what are the chances that there are any in the private test set.",
      "votes": null
    },
    {
      "id": "1254632",
      "postDate": "03/27/2021 22:19:10",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, but the ruling says that the datasets that were used for hand labeling has to be shared, why you are not sharing the DS?</p>",
      "rawMarkdown": "carnozhao, but the ruling says that the datasets that were used for hand labeling has to be shared, why you are not sharing the DS?",
      "votes": null
    },
    {
      "id": "1254685",
      "postDate": "03/28/2021 01:18:21",
      "content": "<p>I updated the dataset link.</p>",
      "rawMarkdown": "I updated the dataset link.",
      "votes": null
    },
    {
      "id": "1254686",
      "postDate": "03/28/2021 01:18:27",
      "content": "<p>I updated the dataset link.</p>",
      "rawMarkdown": "I updated the dataset link.",
      "votes": null
    },
    {
      "id": "1254768",
      "postDate": "03/28/2021 05:33:07",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> - thanks for your update but it is unusual for hand labels of test data to be allowed.</p>\n<p>wrt this d488c759a image, think it relates to this post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993\" target=\"_blank\">both non-sclerotic and sclerotic glomeruli on this competion?</a>  as well as <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116#1250326\" target=\"_blank\">Annotations Updated</a> <br>\nThese question train images  aaa6a05cc and b9a3865fc missing labels, and those missing being sclerotic glomeruli.</p>\n<p>Just a suggestion, you may want to see if your model predicts those train images sclerotic glomeruli, so areas missing labels currently.  You could also add labels for them if you are retraining and then see how your model predicts for  d488c759a.  </p>\n<p>Getting an answer on sclerotic glomeruli  in this competition would be helpful.  It may resolve this issue of hand labels. </p>",
      "rawMarkdown": "carnozhao - thanks for your update but it is unusual for hand labels of test data to be allowed.\n\nwrt this d488c759a image, think it relates to this post [both non-sclerotic and sclerotic glomeruli on this competion?](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993)  as well as [Annotations Updated](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116#1250326) \nThese question train images  aaa6a05cc and b9a3865fc missing labels, and those missing being sclerotic glomeruli.\n\nJust a suggestion, you may want to see if your model predicts those train images sclerotic glomeruli, so areas missing labels currently.  You could also add labels for them if you are retraining and then see how your model predicts for  d488c759a.  \n\nGetting an answer on sclerotic glomeruli  in this competition would be helpful.  It may resolve this issue of hand labels.",
      "votes": null
    },
    {
      "id": "1254955",
      "postDate": "03/28/2021 09:26:14",
      "content": "<p>Good point, <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> could you confirm that any hand labels (both train set and public set) fall in the same category as <strong>external data</strong> and must be made available for all around one week before the end of competition (entry dead line). Thanks.</p>\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>",
      "rawMarkdown": "Good point, @addisonhoward could you confirm that any hand labels (both train set and public set) fall in the same category as **external data** and must be made available for all around one week before the end of competition (entry dead line). Thanks.\n\n> C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
      "votes": null
    },
    {
      "id": "1255153",
      "postDate": "03/28/2021 13:57:02",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, but it is the dataset, which it definition is a collection of imges, in computer vision, not a csv file with the results of a bunch of tiles …</p>",
      "rawMarkdown": "Thanks @carnozhao, but it is the dataset, which it definition is a collection of imges, in computer vision, not a csv file with the results of a bunch of tiles ...",
      "votes": null
    },
    {
      "id": "1255237",
      "postDate": "03/28/2021 15:49:08",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> Thanks for sharing your masks. Your CSV is the resulting masks of your re-trained model predictions after hand labelling or initial model prediction + the hand labels you've done. I don't think you need to share all your pseudo labels but only what you've added/updated/removed (i.e. only a few shapes).</p>\n<p><img src=\"https://nsa40.casimages.com/img/2021/03/28/21032805511139924.png\" alt=\"\"></p>",
      "rawMarkdown": "carnozhao Thanks for sharing your masks. Your CSV is the resulting masks of your re-trained model predictions after hand labelling or initial model prediction + the hand labels you've done. I don't think you need to share all your pseudo labels but only what you've added/updated/removed (i.e. only a few shapes).\n\n![](https://nsa40.casimages.com/img/2021/03/28/21032805511139924.png)",
      "votes": null
    },
    {
      "id": "1255310",
      "postDate": "03/28/2021 17:09:24",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> the the csv file are the positions of the glomeruli? and he is using a third party software to label the images and the produce the .jason file ?</p>\n<p>Thanks</p>",
      "rawMarkdown": "mpware the the csv file are the positions of the glomeruli? and he is using a third party software to label the images and the produce the .jason file ?\n\nThanks",
      "votes": null
    },
    {
      "id": "1255318",
      "postDate": "03/28/2021 17:14:13",
      "content": "<p>CSV includes the RLE encoded masks. That's the format of the submission file in this competition. He said he used paint to hand label (one single specified color) and then I guess he just used CV2 to open labelled image and extract area with the specified color.</p>",
      "rawMarkdown": "CSV includes the RLE encoded masks. That's the format of the submission file in this competition. He said he used paint to hand label (one single specified color) and then I guess he just used CV2 to open labelled image and extract area with the specified color.",
      "votes": null
    },
    {
      "id": "1255402",
      "postDate": "03/28/2021 19:00:14",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , maybe it is a language barrier, but the way I understand is that he needs to share the Dataset not a csv, right?</p>",
      "rawMarkdown": "Thanks @mpware , maybe it is a language barrier, but the way I understand is that he needs to share the Dataset not a csv, right?",
      "votes": null
    },
    {
      "id": "1255403",
      "postDate": "03/28/2021 19:01:58",
      "content": "<p><a href=\"https://www.kaggle.com/oscarrangel\" target=\"_blank\">@oscarrangel</a> CSV is the dataset…</p>",
      "rawMarkdown": "oscarrangel CSV is the dataset...",
      "votes": null
    },
    {
      "id": "1255409",
      "postDate": "03/28/2021 19:11:05",
      "content": "<p>It's the same, with the CSV you can generate the image with masks:</p>\n<pre><code>def rle2mask(mask_rle, shape=(29020, 46660)):\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n</code></pre>",
      "rawMarkdown": "It's the same, with the CSV you can generate the image with masks:\n\n```\ndef rle2mask(mask_rle, shape=(29020, 46660)):\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n```",
      "votes": null
    },
    {
      "id": "1255526",
      "postDate": "03/28/2021 21:51:57",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>",
      "rawMarkdown": "thanks @mpware",
      "votes": null
    },
    {
      "id": "1255603",
      "postDate": "03/29/2021 00:49:39",
      "content": "<p>The mask is my initial model prediction and my hand-labelled part, not the re-trained model. I have mixed them up so it's too complex to split them😂</p>",
      "rawMarkdown": "The mask is my initial model prediction and my hand-labelled part, not the re-trained model. I have mixed them up so it's too complex to split them😂",
      "votes": null
    },
    {
      "id": "1255605",
      "postDate": "03/29/2021 00:54:17",
      "content": "<p>Thanks for your information! </p>",
      "rawMarkdown": "Thanks for your information!",
      "votes": null
    },
    {
      "id": "1255687",
      "postDate": "03/29/2021 04:15:28",
      "content": "<p>No worries <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> and did not mean to come across too hard on you and your work.  Thought maybe if there are ways around hand labels for test even if it is public it might be better in the long run.</p>",
      "rawMarkdown": "No worries @carnozhao and did not mean to come across too hard on you and your work.  Thought maybe if there are ways around hand labels for test even if it is public it might be better in the long run.",
      "votes": null
    },
    {
      "id": "1256154",
      "postDate": "03/29/2021 15:32:54",
      "content": "<p>Hey there - the data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.</p>",
      "rawMarkdown": "Hey there - the data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.",
      "votes": null
    },
    {
      "id": "1256214",
      "postDate": "03/29/2021 16:52:04",
      "content": "<p>Thanks for the clarification, so all hand labelling can remain private.</p>",
      "rawMarkdown": "Thanks for the clarification, so all hand labelling can remain private.",
      "votes": null
    },
    {
      "id": "1256484",
      "postDate": "03/29/2021 23:27:06",
      "content": "<p>Kaggle, where you can attend the best hand-labeling competitions in the world! forget about ComputerVision AI…. when you can just hand label your data and win………</p>",
      "rawMarkdown": "Kaggle, where you can attend the best hand-labeling competitions in the world! forget about ComputerVision AI.... when you can just hand label your data and win.........",
      "votes": null
    },
    {
      "id": "1262127",
      "postDate": "04/03/2021 20:26:01",
      "content": "<p>It's very interesting. I analysed this hand-labelled in detail and majority of hand labelled cells are sclerosed glomeruli type. This might just imply that there are some sclerosed glomeruli included in ground truth, otherwise it shouldn't boost your score. </p>",
      "rawMarkdown": "It's very interesting. I analysed this hand-labelled in detail and majority of hand labelled cells are sclerosed glomeruli type. This might just imply that there are some sclerosed glomeruli included in ground truth, otherwise it shouldn't boost your score.",
      "votes": null
    },
    {
      "id": "1262152",
      "postDate": "04/03/2021 21:26:53",
      "content": "<p>As the competition hosts have said, sclerosed glomeruli are not included in the annotation. Those are all normal glomeruli. The difference in their appearance is caused by the imaging process. Some of the samples were obtained from fresh frozen tissue and others are Formalin Fixed Paraffin Embedded (FFPE). The following shows how the two preparation methodologies effect the same type of tissue:<br>\n<img src=\"https://onlinelibrary.wiley.com/cms/asset/3374bc9f-b17e-45ad-8c69-ec2670e640db/prca2006-fig-0002-m.jpg\" alt=\"\"><br>\nI'll be ignoring the leaderboard for now and focus on making sure my model can generalize beyond the training data. If he's including some of the public test data with his training data, then obviously his score will improve because the model memorizes the answer for those specific observations. However, that indicates that the model doesn't generalize and most likely won't perform as well on the private test set. For comparison, below is the prediction of my latest submission for the same patch as above. I trained my model only on the training data and it scored 0.877. As you can see it correctly recognizes the normal glomeruli.<br>\n<img src=\"https://i.postimg.cc/dDJPfS78/Screen-Shot-2021-04-03-at-5-06-09-PM.png\" alt=\"\"></p>",
      "rawMarkdown": "As the competition hosts have said, sclerosed glomeruli are not included in the annotation. Those are all normal glomeruli. The difference in their appearance is caused by the imaging process. Some of the samples were obtained from fresh frozen tissue and others are Formalin Fixed Paraffin Embedded (FFPE). The following shows how the two preparation methodologies effect the same type of tissue:\n![](https://onlinelibrary.wiley.com/cms/asset/3374bc9f-b17e-45ad-8c69-ec2670e640db/prca2006-fig-0002-m.jpg)\nI'll be ignoring the leaderboard for now and focus on making sure my model can generalize beyond the training data. If he's including some of the public test data with his training data, then obviously his score will improve because the model memorizes the answer for those specific observations. However, that indicates that the model doesn't generalize and most likely won't perform as well on the private test set. For comparison, below is the prediction of my latest submission for the same patch as above. I trained my model only on the training data and it scored 0.877. As you can see it correctly recognizes the normal glomeruli.\n![](https://i.postimg.cc/dDJPfS78/Screen-Shot-2021-04-03-at-5-06-09-PM.png)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1247112,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "03/21/2021 12:32:01",
      "content": "<p>Good question.</p>\n<p>From <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884</a> : </p>\n<blockquote>\n  <p>Given that the new private data is not available via the HuBMAP portal, hand-labeling of the data is now allowed.</p>\n</blockquote>\n<p>From <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/rules\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/rules</a> : </p>\n<blockquote>\n  <p>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n</blockquote>\n<p>The tricky part comes from that you're hand labelling the public test set. I cannot say for sure but my guess is that it's not allowed. You can however hand label data that is not in the test data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1247114,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/21/2021 12:40:12",
          "content": "<p>Thanks! BTW, how can I withdraw my submission</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1247123,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "03/21/2021 13:01:05",
          "content": "<p>I don't think it's possible ^^<br>\nThere's still a lot of time so people should catch-up to your score, don't worry about it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1247972,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/22/2021 08:27:51",
          "content": "<p>So in summary hand labeling is allowed as long as it isn't the test dataset (i.e. the public part). <br>\nShould hand labeling external datasets be declared or are we good even if not acknowledging this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1249808,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "03/23/2021 14:33:12",
          "content": "<p>Hi all - I'll confirm Yassine's comment here - Hand labeling is allowed per the host's ruling, as long as it isn't the test dataset.</p>\n<p>Carno - please contact <a href=\"https://www.kaggle.com/compliance\" target=\"_blank\">compliance</a> and we'll see if we can work with you to invalidate the submissions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250108,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/23/2021 19:46:20",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> thanks for confirming! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250357,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/24/2021 01:34:58",
          "content": "<p>thanks, I have contacted the compliance</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250442,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "03/24/2021 03:47:32",
          "content": "<p>Hi all,</p>\n<p>We've discussed this internally and will note that both hand-labeling and pseudo-labeling of the data is allowed per the competition host's comments. As this is a code competition, there is no ability to label the unseen, private test set, we do not expect it to gain any advantage on the models that determine the final winners. While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.</p>\n<p>We don't invalidate submissions by request (my apologies for stating such above). In any future competition where you are concerned about a submission you made that may violate certain rules, your best course of action is to <em>not select those submissions for final scoring.</em> This does not apply to private code sharing or use of multiple accounts, in which case it does not matter which submissions you select, it is still considered a violation.</p>\n<p>I've updated the rules accordingly for avoidance of doubt regarding such.</p>\n<p>Thanks,</p>\n<p>Addison</p>\n<p>cc/ <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250582,
          "author_name": "yukkyo",
          "author_url": "",
          "post_date": "03/24/2021 06:12:48",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <br>\nIs there any reason why host published the public test data?<br>\nOther code competitions did't publish public test data. </p>\n<p>Currently, I feel that the host is asking the participants to hand labeling.</p>\n<blockquote>\n  <p>While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.</p>\n</blockquote>\n<p>If the public test data doesn't affect the unseen test set, shouldn't the label of the public test data be public?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251153,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "03/24/2021 14:25:56",
          "content": "<p>For research organizations with public funding, often times there are requirements about how and when the data is released. For a project this large, it can be difficult to amass a truly hidden test set for scoring. </p>\n<p>We do not release the public test set labels, as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1251865,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "03/25/2021 08:04:28",
          "content": "<blockquote>\n  <p>as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.</p>\n</blockquote>\n<p>We should validate model with train data, not leaderboard. We can use leaderboard for comparing other kaggler's model and competing kagglers. I think that is the main purpose of participating in the competition.<br>\nHand labeling just break public leaderboard, and also,  it pour cold water on the enthusiasm of the participants, especially the beginners.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252081,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "03/25/2021 12:13:49",
          "content": "<p>What? hand labeling of public test data is allowed?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252159,
          "author_name": "ttahara",
          "author_url": "",
          "post_date": "03/25/2021 13:15:39",
          "content": "<p>The updated rule means we should compete for hand labeling skills on Public LB? That is nonsense…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252389,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "03/25/2021 16:40:23",
          "content": "<p>Did that mean we can use hand-labeled samples to strengthen our model? If so, the significance will go bad. I suggest making the public test samples into training dataset. using a part of the other (&lt;67%) to be the real public test set therefore this competition could have a chance to be a typical machine learning competition rather than a painter game.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252393,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "03/25/2021 16:44:09",
          "content": "<p>Agreed. <a href=\"https://www.kaggle.com/southsakura\" target=\"_blank\">@southsakura</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1252751,
          "author_name": "aramisvesal",
          "author_url": "",
          "post_date": "03/26/2021 02:16:34",
          "content": "<p>Agreed with <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a>. Don't know why the organizer of this challenge does not look at other challenges that have been hosted in Kaggle. No one set a rule like this to hand-labelling the test data in the leaderboard. I think this competition should focus more on quality of the methods and tackling real-world problem with limited data rather than reaching a Dice Score of 99.0%. </p>\n<p>On the other hand, the Kagglers rather focusing on innovative approaches, would focus more on labelling the data which reduce the impact of this challenge. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1247253,
      "author_name": "southsakura",
      "author_url": "",
      "post_date": "03/21/2021 15:14:40",
      "content": "<p>would you mind showing your method for handing draw. :) I am curious about it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1247294,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/21/2021 15:46:15",
          "content": "<p>I draw yellow mask over the image in windows paint. After loading the image, mask is where the pixel equals to (255, 255, 0)</p>\n<p>Of course you can try other color.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1247360,
          "author_name": "southsakura",
          "author_url": "",
          "post_date": "03/21/2021 16:52:01",
          "content": "<p>Got it! windows paint. nice method hhh.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1247703,
      "author_name": "louieshao",
      "author_url": "",
      "post_date": "03/22/2021 01:31:53",
      "content": "<p>Nice score, but honestly I did not really figure it out. Did you mean you pseudo labled some of the public test dataset and use those to train your model, and finally used this model to predict the public test dataset itself? Won't it be data leakage?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1247711,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/22/2021 01:45:20",
          "content": "<p>Most of my pseudo labels are generated by model prediction. Besides that, I only draw pseudo labels on part of one image, about 10~20 glomeruli. I think there is data leakage but not too much. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1250593,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "03/24/2021 06:27:21",
      "content": "<p>Should we compete for hand-labeling skills with more and more train data…?😨</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250598,
          "author_name": "louieshao",
          "author_url": "",
          "post_date": "03/24/2021 06:32:54",
          "content": "<p>Welcome to Kaggle, where you can attend the best hand-labeling competitions in the world!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250603,
          "author_name": "drtausamaru",
          "author_url": "",
          "post_date": "03/24/2021 06:36:12",
          "content": "<p>Any pathologist? lol</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1252775,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "03/26/2021 03:24:33",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>,</p>\n<p>You manualy painted the image with a pain program and SAVE it ?, and use it for training? I don't understand,and where is the mask? did you used it? are you using the lafoss title notebook to break the image in tiles?</p>\n<p>Can you explain a little bit more about how and what you used to do this.</p>\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1253856,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/27/2021 05:23:27",
      "content": "<p>Would like to add this from the RANZCR competitions experience and believe competition rules here should be amended in a similar way but clarifying the public test data set since there are now old test that could be considered external data perhaps -</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644</a>  <br>\n\"There have been some questions from competitors regarding hand-labelled annotations. To clarify the rules as per section A2:</p>\n<pre><code>Publicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"\n</code></pre>\n<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> this is with respect to this public test annotation but could apply to other external data.  There is no way to know if the private test data also has \"dark glomeruli\" as <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> discovered in the public test data.  If so this confers an advantage to those that have these annotations. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1254597,
      "author_name": "isakev",
      "author_url": "",
      "post_date": "03/27/2021 21:11:59",
      "content": "<p>Not sure, but shouldn't a hand-labeled dataset be regarded as a dataset different from the competitions dataset and, thus, it must be classified as external.  In that case, it should be publicly available to all.</p>\n<p>Of course, it would not be fair to those who had done the hard part of hand-labeling, so I am not advocating this. On the contrary, the disclosure better be voluntary.</p>\n<p>Another point of consideration: if this type of glomeruli (is it 'sclerotic' glomeruli?) exists only in one out of 20 images in train and public test set, what are the chances that there are any in the private test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1254686,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/28/2021 01:18:27",
          "content": "<p>I updated the dataset link.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1254632,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "03/27/2021 22:19:10",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, but the ruling says that the datasets that were used for hand labeling has to be shared, why you are not sharing the DS?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1254685,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/28/2021 01:18:21",
          "content": "<p>I updated the dataset link.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1254955,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/28/2021 09:26:14",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> could you confirm that any hand labels (both train set and public set) fall in the same category as <strong>external data</strong> and must be made available for all around one week before the end of competition (entry dead line). Thanks.</p>\n<blockquote>\n  <p>C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255153,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "03/28/2021 13:57:02",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a>, but it is the dataset, which it definition is a collection of imges, in computer vision, not a csv file with the results of a bunch of tiles …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255237,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/28/2021 15:49:08",
          "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> Thanks for sharing your masks. Your CSV is the resulting masks of your re-trained model predictions after hand labelling or initial model prediction + the hand labels you've done. I don't think you need to share all your pseudo labels but only what you've added/updated/removed (i.e. only a few shapes).</p>\n<p><img src=\"https://nsa40.casimages.com/img/2021/03/28/21032805511139924.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255310,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "03/28/2021 17:09:24",
          "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> the the csv file are the positions of the glomeruli? and he is using a third party software to label the images and the produce the .jason file ?</p>\n<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255318,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/28/2021 17:14:13",
          "content": "<p>CSV includes the RLE encoded masks. That's the format of the submission file in this competition. He said he used paint to hand label (one single specified color) and then I guess he just used CV2 to open labelled image and extract area with the specified color.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255402,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "03/28/2021 19:00:14",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> , maybe it is a language barrier, but the way I understand is that he needs to share the Dataset not a csv, right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255403,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "03/28/2021 19:01:58",
          "content": "<p><a href=\"https://www.kaggle.com/oscarrangel\" target=\"_blank\">@oscarrangel</a> CSV is the dataset…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255409,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/28/2021 19:11:05",
          "content": "<p>It's the same, with the CSV you can generate the image with masks:</p>\n<pre><code>def rle2mask(mask_rle, shape=(29020, 46660)):\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255526,
          "author_name": "oscarrangel",
          "author_url": "",
          "post_date": "03/28/2021 21:51:57",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255603,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/29/2021 00:49:39",
          "content": "<p>The mask is my initial model prediction and my hand-labelled part, not the re-trained model. I have mixed them up so it's too complex to split them😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1256154,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "03/29/2021 15:32:54",
          "content": "<p>Hey there - the data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1256214,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "03/29/2021 16:52:04",
          "content": "<p>Thanks for the clarification, so all hand labelling can remain private.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1254768,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "03/28/2021 05:33:07",
      "content": "<p><a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> - thanks for your update but it is unusual for hand labels of test data to be allowed.</p>\n<p>wrt this d488c759a image, think it relates to this post <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993\" target=\"_blank\">both non-sclerotic and sclerotic glomeruli on this competion?</a>  as well as <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116#1250326\" target=\"_blank\">Annotations Updated</a> <br>\nThese question train images  aaa6a05cc and b9a3865fc missing labels, and those missing being sclerotic glomeruli.</p>\n<p>Just a suggestion, you may want to see if your model predicts those train images sclerotic glomeruli, so areas missing labels currently.  You could also add labels for them if you are retraining and then see how your model predicts for  d488c759a.  </p>\n<p>Getting an answer on sclerotic glomeruli  in this competition would be helpful.  It may resolve this issue of hand labels. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1255605,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "03/29/2021 00:54:17",
          "content": "<p>Thanks for your information! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1255687,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "03/29/2021 04:15:28",
          "content": "<p>No worries <a href=\"https://www.kaggle.com/carnozhao\" target=\"_blank\">@carnozhao</a> and did not mean to come across too hard on you and your work.  Thought maybe if there are ways around hand labels for test even if it is public it might be better in the long run.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1256484,
      "author_name": "oscarrangel",
      "author_url": "",
      "post_date": "03/29/2021 23:27:06",
      "content": "<p>Kaggle, where you can attend the best hand-labeling competitions in the world! forget about ComputerVision AI…. when you can just hand label your data and win………</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1262127,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "04/03/2021 20:26:01",
      "content": "<p>It's very interesting. I analysed this hand-labelled in detail and majority of hand labelled cells are sclerosed glomeruli type. This might just imply that there are some sclerosed glomeruli included in ground truth, otherwise it shouldn't boost your score. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1262152,
          "author_name": "erikdali",
          "author_url": "",
          "post_date": "04/03/2021 21:26:53",
          "content": "<p>As the competition hosts have said, sclerosed glomeruli are not included in the annotation. Those are all normal glomeruli. The difference in their appearance is caused by the imaging process. Some of the samples were obtained from fresh frozen tissue and others are Formalin Fixed Paraffin Embedded (FFPE). The following shows how the two preparation methodologies effect the same type of tissue:<br>\n<img src=\"https://onlinelibrary.wiley.com/cms/asset/3374bc9f-b17e-45ad-8c69-ec2670e640db/prca2006-fig-0002-m.jpg\" alt=\"\"><br>\nI'll be ignoring the leaderboard for now and focus on making sure my model can generalize beyond the training data. If he's including some of the public test data with his training data, then obviously his score will improve because the model memorizes the answer for those specific observations. However, that indicates that the model doesn't generalize and most likely won't perform as well on the private test set. For comparison, below is the prediction of my latest submission for the same patch as above. I trained my model only on the training data and it scored 0.877. As you can see it correctly recognizes the normal glomeruli.<br>\n<img src=\"https://i.postimg.cc/dDJPfS78/Screen-Shot-2021-04-03-at-5-06-09-PM.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1247107": "I found that my model often miss dark glomeruli in public test data. And stronger data augmentation didn't work. So I manually draw pseudo label of some of these dark glomeruli and use them in training.\n\nThe model trained with pseudo label gave me a great boost (LB from 0.914 to 0.932)\n\nBut I'm not sure whether it is allowed in this competition.\n\nps. I don't mean to break the rules😭\n\n**UPDATE:**\nMy problem and my idea are shown below:\n\nI found my score of d488c759a single image was much lower than other 4 images. I checked my prediction result and found that there are some darker, smaller things which are not similar with normal glomeruli. They distributed more densely at the boundary of the slice and there seems to be less cell nucleus in these structures.\n![](https://i.postimg.cc/pLbNLQF9/tmp.png)\n\nI checked all images from training data, public test data and external data from HUBMAP, and found that these things only exist in d488c759a image. So I assume these things cause my low score of d488c759a.\n\nI tried some data augmentation tricks but didn't help. Finally, I hand-labelled some of them (not whole image, not whole dataset) as pseudo-label and trained a model with the pseudo-label. \n\nI am really sorry breaking rules unintentionally. I hope someone could solve this problem with normal deep learning tricks!!\n\n**UPDATE2**\n\nMy hand-labelled (for dark glomeruli) + pseudo-labelled (for normal glomeruli) mask of d488c759a \n\n[https://www.kaggle.com/carnozhao/d48-hand-labelled](https://www.kaggle.com/carnozhao/d48-hand-labelled)",
    "1247112": "Good question.\n\nFrom https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/207884 : \n> Given that the new private data is not available via the HuBMAP portal, hand-labeling of the data is now allowed.\n\nFrom https://www.kaggle.com/c/hubmap-kidney-segmentation/rules : \n> Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nThe tricky part comes from that you're hand labelling the public test set. I cannot say for sure but my guess is that it's not allowed. You can however hand label data that is not in the test data.",
    "1247114": "Thanks! BTW, how can I withdraw my submission",
    "1247123": "I don't think it's possible ^^\nThere's still a lot of time so people should catch-up to your score, don't worry about it.",
    "1247253": "would you mind showing your method for handing draw. :) I am curious about it.",
    "1247294": "I draw yellow mask over the image in windows paint. After loading the image, mask is where the pixel equals to (255, 255, 0)\n\nOf course you can try other color.",
    "1247360": "Got it! windows paint. nice method hhh.",
    "1247703": "Nice score, but honestly I did not really figure it out. Did you mean you pseudo labled some of the public test dataset and use those to train your model, and finally used this model to predict the public test dataset itself? Won't it be data leakage?",
    "1247711": "Most of my pseudo labels are generated by model prediction. Besides that, I only draw pseudo labels on part of one image, about 10~20 glomeruli. I think there is data leakage but not too much.",
    "1247972": "So in summary hand labeling is allowed as long as it isn't the test dataset (i.e. the public part). \nShould hand labeling external datasets be declared or are we good even if not acknowledging this?",
    "1249808": "Hi all - I'll confirm Yassine's comment here - Hand labeling is allowed per the host's ruling, as long as it isn't the test dataset.\n\nCarno - please contact [compliance](https://www.kaggle.com/compliance) and we'll see if we can work with you to invalidate the submissions.",
    "1250108": "addisonhoward thanks for confirming!",
    "1250357": "thanks, I have contacted the compliance",
    "1250442": "Hi all,\n\nWe've discussed this internally and will note that both hand-labeling and pseudo-labeling of the data is allowed per the competition host's comments. As this is a code competition, there is no ability to label the unseen, private test set, we do not expect it to gain any advantage on the models that determine the final winners. While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.\n\nWe don't invalidate submissions by request (my apologies for stating such above). In any future competition where you are concerned about a submission you made that may violate certain rules, your best course of action is to *not select those submissions for final scoring.* This does not apply to private code sharing or use of multiple accounts, in which case it does not matter which submissions you select, it is still considered a violation.\n\nI've updated the rules accordingly for avoidance of doubt regarding such.\n\nThanks,\n\nAddison\n\ncc/ @carnozhao",
    "1250582": "addisonhoward \nIs there any reason why host published the public test data?\nOther code competitions did't publish public test data. \n\nCurrently, I feel that the host is asking the participants to hand labeling.\n\n> While this may impact the public leaderboard during the competition and may violate the spirit of the rules, it will not determine the final winners, which are based solely on an unseen test set.\n\nIf the public test data doesn't affect the unseen test set, shouldn't the label of the public test data be public?",
    "1250593": "Should we compete for hand-labeling skills with more and more train data...?😨",
    "1250598": "Welcome to Kaggle, where you can attend the best hand-labeling competitions in the world!",
    "1250603": "Any pathologist? lol",
    "1251153": "For research organizations with public funding, often times there are requirements about how and when the data is released. For a project this large, it can be difficult to amass a truly hidden test set for scoring. \n\nWe do not release the public test set labels, as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.",
    "1251865": "> as one of the great benefits of the leaderboard is to use it as a vehicle for validation and checking your score.\n\nWe should validate model with train data, not leaderboard. We can use leaderboard for comparing other kaggler's model and competing kagglers. I think that is the main purpose of participating in the competition.\nHand labeling just break public leaderboard, and also,  it pour cold water on the enthusiasm of the participants, especially the beginners.",
    "1252081": "What? hand labeling of public test data is allowed?",
    "1252159": "The updated rule means we should compete for hand labeling skills on Public LB? That is nonsense...",
    "1252389": "Did that mean we can use hand-labeled samples to strengthen our model? If so, the significance will go bad. I suggest making the public test samples into training dataset. using a part of the other (<67%) to be the real public test set therefore this competition could have a chance to be a typical machine learning competition rather than a painter game.",
    "1252393": "Agreed. @southsakura",
    "1252751": "Agreed with @phalanx. Don't know why the organizer of this challenge does not look at other challenges that have been hosted in Kaggle. No one set a rule like this to hand-labelling the test data in the leaderboard. I think this competition should focus more on quality of the methods and tackling real-world problem with limited data rather than reaching a Dice Score of 99.0%. \n\nOn the other hand, the Kagglers rather focusing on innovative approaches, would focus more on labelling the data which reduce the impact of this challenge.",
    "1252775": "Hi @carnozhao,\n\nYou manualy painted the image with a pain program and SAVE it ?, and use it for training? I don't understand,and where is the mask? did you used it? are you using the lafoss title notebook to break the image in tiles?\n\nCan you explain a little bit more about how and what you used to do this.\n\nThanks!",
    "1253856": "Would like to add this from the RANZCR competitions experience and believe competition rules here should be amended in a similar way but clarifying the public test data set since there are now old test that could be considered external data perhaps -\n \nhttps://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/222644  \n\"There have been some questions from competitors regarding hand-labelled annotations. To clarify the rules as per section A2:\n\n    Publicly, freely available external data is permitted. Entrants may re-annotate images in the training set, however Entrants will (i) ensure the re-annotated data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the re-annotated data for the participants to the official competition forum prior to the Entry Deadline. Entrants may not hand-label predictions in the test data set, including having human observers rate and evaluate the test data set.\"\n\n@addisonhoward this is with respect to this public test annotation but could apply to other external data.  There is no way to know if the private test data also has \"dark glomeruli\" as @carnozhao discovered in the public test data.  If so this confers an advantage to those that have these annotations.",
    "1254597": "Not sure, but shouldn't a hand-labeled dataset be regarded as a dataset different from the competitions dataset and, thus, it must be classified as external.  In that case, it should be publicly available to all.\n\nOf course, it would not be fair to those who had done the hard part of hand-labeling, so I am not advocating this. On the contrary, the disclosure better be voluntary.\n\nAnother point of consideration: if this type of glomeruli (is it 'sclerotic' glomeruli?) exists only in one out of 20 images in train and public test set, what are the chances that there are any in the private test set.",
    "1254632": "carnozhao, but the ruling says that the datasets that were used for hand labeling has to be shared, why you are not sharing the DS?",
    "1254685": "I updated the dataset link.",
    "1254686": "I updated the dataset link.",
    "1254768": "carnozhao - thanks for your update but it is unusual for hand labels of test data to be allowed.\n\nwrt this d488c759a image, think it relates to this post [both non-sclerotic and sclerotic glomeruli on this competion?](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/228993)  as well as [Annotations Updated](https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/198116#1250326) \nThese question train images  aaa6a05cc and b9a3865fc missing labels, and those missing being sclerotic glomeruli.\n\nJust a suggestion, you may want to see if your model predicts those train images sclerotic glomeruli, so areas missing labels currently.  You could also add labels for them if you are retraining and then see how your model predicts for  d488c759a.  \n\nGetting an answer on sclerotic glomeruli  in this competition would be helpful.  It may resolve this issue of hand labels.",
    "1254955": "Good point, @addisonhoward could you confirm that any hand labels (both train set and public set) fall in the same category as **external data** and must be made available for all around one week before the end of competition (entry dead line). Thanks.\n\n> C. External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants and (ii) post such access to the External Data for the participants to the official competition forum prior to the Entry Deadline.",
    "1255153": "Thanks @carnozhao, but it is the dataset, which it definition is a collection of imges, in computer vision, not a csv file with the results of a bunch of tiles ...",
    "1255237": "carnozhao Thanks for sharing your masks. Your CSV is the resulting masks of your re-trained model predictions after hand labelling or initial model prediction + the hand labels you've done. I don't think you need to share all your pseudo labels but only what you've added/updated/removed (i.e. only a few shapes).\n\n![](https://nsa40.casimages.com/img/2021/03/28/21032805511139924.png)",
    "1255310": "mpware the the csv file are the positions of the glomeruli? and he is using a third party software to label the images and the produce the .jason file ?\n\nThanks",
    "1255318": "CSV includes the RLE encoded masks. That's the format of the submission file in this competition. He said he used paint to hand label (one single specified color) and then I guess he just used CV2 to open labelled image and extract area with the specified color.",
    "1255402": "Thanks @mpware , maybe it is a language barrier, but the way I understand is that he needs to share the Dataset not a csv, right?",
    "1255403": "oscarrangel CSV is the dataset...",
    "1255409": "It's the same, with the CSV you can generate the image with masks:\n\n```\ndef rle2mask(mask_rle, shape=(29020, 46660)):\n    s = mask_rle.split()\n    starts, lengths = [np.asarray(x, dtype=int) for x in (s[0:][::2], s[1:][::2])]\n    starts -= 1\n    ends = starts + lengths\n    img = np.zeros(shape[0]*shape[1], dtype=np.uint8)\n    for lo, hi in zip(starts, ends):\n        img[lo:hi] = 1\n    return img.reshape(shape).T\n```",
    "1255526": "thanks @mpware",
    "1255603": "The mask is my initial model prediction and my hand-labelled part, not the re-trained model. I have mixed them up so it's too complex to split them😂",
    "1255605": "Thanks for your information!",
    "1255687": "No worries @carnozhao and did not mean to come across too hard on you and your work.  Thought maybe if there are ways around hand labels for test even if it is public it might be better in the long run.",
    "1256154": "Hey there - the data itself and notifying of its existence is considered publicly available external data. The hand-labeling performed by any individual is not required to be made public.",
    "1256214": "Thanks for the clarification, so all hand labelling can remain private.",
    "1256484": "Kaggle, where you can attend the best hand-labeling competitions in the world! forget about ComputerVision AI.... when you can just hand label your data and win.........",
    "1262127": "It's very interesting. I analysed this hand-labelled in detail and majority of hand labelled cells are sclerosed glomeruli type. This might just imply that there are some sclerosed glomeruli included in ground truth, otherwise it shouldn't boost your score.",
    "1262152": "As the competition hosts have said, sclerosed glomeruli are not included in the annotation. Those are all normal glomeruli. The difference in their appearance is caused by the imaging process. Some of the samples were obtained from fresh frozen tissue and others are Formalin Fixed Paraffin Embedded (FFPE). The following shows how the two preparation methodologies effect the same type of tissue:\n![](https://onlinelibrary.wiley.com/cms/asset/3374bc9f-b17e-45ad-8c69-ec2670e640db/prca2006-fig-0002-m.jpg)\nI'll be ignoring the leaderboard for now and focus on making sure my model can generalize beyond the training data. If he's including some of the public test data with his training data, then obviously his score will improve because the model memorizes the answer for those specific observations. However, that indicates that the model doesn't generalize and most likely won't perform as well on the private test set. For comparison, below is the prediction of my latest submission for the same patch as above. I trained my model only on the training data and it scored 0.877. As you can see it correctly recognizes the normal glomeruli.\n![](https://i.postimg.cc/dDJPfS78/Screen-Shot-2021-04-03-at-5-06-09-PM.png)"
  },
  "source": "meta"
}