{
  "id": 207884,
  "title": "HuBMAP Hacking the Kidney Competition: Revised Glom Masks, New Private Dataset, and Timeline Update",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/207884",
  "author_name": "",
  "post_date": "2020-12-31T19:24:30.018688600Z",
  "votes": 114,
  "comment_count": 57,
  "views": 0,
  "content": "<p>Dear Teams,<br>\nWe are excited and impressed with all the engagement and participation for this hackathon. There is significant potential for the development of annotation tools that could greatly accelerate impact on human health, and we appreciate the substantial time investment many participants have made. The scientific value of the contest is of the utmost importance.</p>\n<p>Unfortunately, it has come to our attention that unlabeled images in the private test data have been discovered on the NIH HuBMAP data portal, compromising the ability of tools developed to e.g., generalize to new datasets. </p>\n<p>Please accept our sincere apologies. </p>\n<p>We, the hackathon organizers, are taking the necessary step of completely updating the private test data. We have also added new quality control measures to the glomeruli annotation masks. </p>\n<p>Specifically, we will do the following:</p>\n<ul>\n<li>All 20 of the currently used datasets will be made public and the 20 glomeruli annotation mask JSON files will be updated using the new quality controls. </li>\n<li>A new, unpublished dataset of 10 images (5 fresh frozen, 5 FFPE) and their annotation masks (that have undergone the very same new quality controls) will replace the current private test dataset. <br>\nThese data will remain unseen and unpublished until after the competition has concluded. We have added two-factor authentication and restricted access to ensure security of these data.</li>\n<li>All new data will become available in mid-January 2021 and the final submission deadline will be extended to mid to late March 2021. All other deadlines will be extended accordingly. We have temporarily set a new deadline date as a placeholder while the new dataset is under review and preparation, and will update the timeline with finalized dates once the new dataset is released.</li>\n</ul>\n<p>Given that the new private data is <strong>not</strong> available via the HuBMAP portal, hand-labeling of the data <strong>is now allowed</strong>. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.</p>\n<p>We are working closely with the NIH HuBMAP Program Office, the competition sponsors, and the competition judges while implementing these changes and would like to thank all involved for their extensive expert support.</p>\n<p>Last but not least, we wish you a marvelous New Year 2021!</p>\n<p>Sincerely,<br>\nKaty Borner </p>",
  "messages": [
    {
      "id": "1134044",
      "postDate": "12/31/2020 19:24:30",
      "content": "<p>Dear Teams,<br>\nWe are excited and impressed with all the engagement and participation for this hackathon. There is significant potential for the development of annotation tools that could greatly accelerate impact on human health, and we appreciate the substantial time investment many participants have made. The scientific value of the contest is of the utmost importance.</p>\n<p>Unfortunately, it has come to our attention that unlabeled images in the private test data have been discovered on the NIH HuBMAP data portal, compromising the ability of tools developed to e.g., generalize to new datasets. </p>\n<p>Please accept our sincere apologies. </p>\n<p>We, the hackathon organizers, are taking the necessary step of completely updating the private test data. We have also added new quality control measures to the glomeruli annotation masks. </p>\n<p>Specifically, we will do the following:</p>\n<ul>\n<li>All 20 of the currently used datasets will be made public and the 20 glomeruli annotation mask JSON files will be updated using the new quality controls. </li>\n<li>A new, unpublished dataset of 10 images (5 fresh frozen, 5 FFPE) and their annotation masks (that have undergone the very same new quality controls) will replace the current private test dataset. <br>\nThese data will remain unseen and unpublished until after the competition has concluded. We have added two-factor authentication and restricted access to ensure security of these data.</li>\n<li>All new data will become available in mid-January 2021 and the final submission deadline will be extended to mid to late March 2021. All other deadlines will be extended accordingly. We have temporarily set a new deadline date as a placeholder while the new dataset is under review and preparation, and will update the timeline with finalized dates once the new dataset is released.</li>\n</ul>\n<p>Given that the new private data is <strong>not</strong> available via the HuBMAP portal, hand-labeling of the data <strong>is now allowed</strong>. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.</p>\n<p>We are working closely with the NIH HuBMAP Program Office, the competition sponsors, and the competition judges while implementing these changes and would like to thank all involved for their extensive expert support.</p>\n<p>Last but not least, we wish you a marvelous New Year 2021!</p>\n<p>Sincerely,<br>\nKaty Borner </p>",
      "rawMarkdown": "Dear Teams,\nWe are excited and impressed with all the engagement and participation for this hackathon. There is significant potential for the development of annotation tools that could greatly accelerate impact on human health, and we appreciate the substantial time investment many participants have made. The scientific value of the contest is of the utmost importance.\n\nUnfortunately, it has come to our attention that unlabeled images in the private test data have been discovered on the NIH HuBMAP data portal, compromising the ability of tools developed to e.g., generalize to new datasets. \n\nPlease accept our sincere apologies. \n\nWe, the hackathon organizers, are taking the necessary step of completely updating the private test data. We have also added new quality control measures to the glomeruli annotation masks. \n\nSpecifically, we will do the following:\n- All 20 of the currently used datasets will be made public and the 20 glomeruli annotation mask JSON files will be updated using the new quality controls. \n- A new, unpublished dataset of 10 images (5 fresh frozen, 5 FFPE) and their annotation masks (that have undergone the very same new quality controls) will replace the current private test dataset. \nThese data will remain unseen and unpublished until after the competition has concluded. We have added two-factor authentication and restricted access to ensure security of these data.\n- All new data will become available in mid-January 2021 and the final submission deadline will be extended to mid to late March 2021. All other deadlines will be extended accordingly. We have temporarily set a new deadline date as a placeholder while the new dataset is under review and preparation, and will update the timeline with finalized dates once the new dataset is released.\n\nGiven that the new private data is **not** available via the HuBMAP portal, hand-labeling of the data **is now allowed**. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.\n\nWe are working closely with the NIH HuBMAP Program Office, the competition sponsors, and the competition judges while implementing these changes and would like to thank all involved for their extensive expert support.\n\nLast but not least, we wish you a marvelous New Year 2021!\n\nSincerely,\nKaty Borner",
      "votes": null
    },
    {
      "id": "1134083",
      "postDate": "12/31/2020 20:16:25",
      "content": "<p>Thanks for the update.  If the 20 current images (8 train + 5 public test + 7 private test) will be made public with annotations released and then the 10 new images become the new private test, then what will be the new public test set?</p>",
      "rawMarkdown": "Thanks for the update.  If the 20 current images (8 train + 5 public test + 7 private test) will be made public with annotations released and then the 10 new images become the new private test, then what will be the new public test set?",
      "votes": null
    },
    {
      "id": "1134242",
      "postDate": "01/01/2021 03:54:55",
      "content": "<p>Would the ground truth shift in afa5e8098 be corrected in the revised dataset?</p>",
      "rawMarkdown": "Would the ground truth shift in afa5e8098 be corrected in the revised dataset?",
      "votes": null
    },
    {
      "id": "1134356",
      "postDate": "01/01/2021 07:38:41",
      "content": "<p>Great question - we'll announce the split with the upcoming data release in a few weeks. The split itself is still TBD on our end.</p>",
      "rawMarkdown": "Great question - we'll announce the split with the upcoming data release in a few weeks. The split itself is still TBD on our end.",
      "votes": null
    },
    {
      "id": "1134954",
      "postDate": "01/01/2021 18:19:28",
      "content": "<p>Wow those are great news.<br>\nThat is a lot of well appreciated work you are putting into this to make the competition as fair as possible and the results as robust as possible (usable for further research and production).</p>",
      "rawMarkdown": "Wow those are great news.\nThat is a lot of well appreciated work you are putting into this to make the competition as fair as possible and the results as robust as possible (usable for further research and production).",
      "votes": null
    },
    {
      "id": "1134985",
      "postDate": "01/01/2021 19:12:20",
      "content": "<p>My concern would be that if only 10 new images are going to be annotated, it may be tough to make a public/private split out of those 10 that leaves a meaningful LB.  I realize 10 isn't a lot less than the 12 we have now but it places more emphasis on fewer images.<br>\n Just an idea: you could hold back the annotations on a few of the current test images and still use that for the new public test set.  This would also provide some continuity of the current public LB rather than have it completely reset.  There's lots of ways to slice things up and I'm sure Kaggle and the sponsors will figure it out and make it fair and meaningful.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2Fad7aa5606f79dcbfc8137a728bcfcc5a%2Fhubmap_data.png?generation=1609528269340385&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "My concern would be that if only 10 new images are going to be annotated, it may be tough to make a public/private split out of those 10 that leaves a meaningful LB.  I realize 10 isn't a lot less than the 12 we have now but it places more emphasis on fewer images.\n Just an idea: you could hold back the annotations on a few of the current test images and still use that for the new public test set.  This would also provide some continuity of the current public LB rather than have it completely reset.  There's lots of ways to slice things up and I'm sure Kaggle and the sponsors will figure it out and make it fair and meaningful.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2Fad7aa5606f79dcbfc8137a728bcfcc5a%2Fhubmap_data.png?generation=1609528269340385&alt=media)",
      "votes": null
    },
    {
      "id": "1135319",
      "postDate": "01/02/2021 06:04:03",
      "content": "<p>Dear organizers and kaggle team,<br>\nThank you for taking actions to resolve issues with publicly available private test set and the leak of the private score that potentially could allow some participants to discover the average shift in the private data that maximizes the score. So now the competition might work.<br>\nMy only concern is careful checking for the GT shift in the new test set before updating it (to avoid something we saw for the previous test set).</p>",
      "rawMarkdown": "Dear organizers and kaggle team,\nThank you for taking actions to resolve issues with publicly available private test set and the leak of the private score that potentially could allow some participants to discover the average shift in the private data that maximizes the score. So now the competition might work.\nMy only concern is careful checking for the GT shift in the new test set before updating it (to avoid something we saw for the previous test set).",
      "votes": null
    },
    {
      "id": "1136280",
      "postDate": "01/02/2021 23:12:28",
      "content": "<p>I wanted to confirm if the current format of the anatomical_structures_segmention_file<br>\nglomerulus_segmentation_file jsons and RLE encoding column of thetrain.csv would be the same for the new dataset?<br>\nThank you for the information.</p>",
      "rawMarkdown": "I wanted to confirm if the current format of the anatomical_structures_segmention_file\nglomerulus_segmentation_file jsons and RLE encoding column of thetrain.csv would be the same for the new dataset?\nThank you for the information.",
      "votes": null
    },
    {
      "id": "1136956",
      "postDate": "01/03/2021 15:08:10",
      "content": "<p>That's great! A very responsible approach and well-timed decision. Thank you very much!</p>",
      "rawMarkdown": "That's great! A very responsible approach and well-timed decision. Thank you very much!",
      "votes": null
    },
    {
      "id": "1138312",
      "postDate": "01/04/2021 15:09:16",
      "content": "<p>Yes, it would.</p>",
      "rawMarkdown": "Yes, it would.",
      "votes": null
    },
    {
      "id": "1138313",
      "postDate": "01/04/2021 15:09:44",
      "content": "<p>Yes, the format will remain the same.</p>",
      "rawMarkdown": "Yes, the format will remain the same.",
      "votes": null
    },
    {
      "id": "1138468",
      "postDate": "01/04/2021 17:49:29",
      "content": "<p>this also means we need to re-run the submissions to be used for the final score. Are we going to get extra GPU/TPU time, or will you re-run those? Thanks!!</p>",
      "rawMarkdown": "this also means we need to re-run the submissions to be used for the final score. Are we going to get extra GPU/TPU time, or will you re-run those? Thanks!!",
      "votes": null
    },
    {
      "id": "1138756",
      "postDate": "01/05/2021 00:08:17",
      "content": "<p>Thanks for bring that up <a href=\"https://www.kaggle.com/pratikkumar\" target=\"_blank\">@pratikkumar</a>.</p>",
      "rawMarkdown": "Thanks for bring that up @pratikkumar.",
      "votes": null
    },
    {
      "id": "1139256",
      "postDate": "01/05/2021 09:22:01",
      "content": "<p>thank you very much!<br>\nthanks for the effort to create new test dataset in such a short time.</p>",
      "rawMarkdown": "thank you very much!\nthanks for the effort to create new test dataset in such a short time.",
      "votes": null
    },
    {
      "id": "1140730",
      "postDate": "01/06/2021 08:18:51",
      "content": "<p>Thank you for your honesty and quick response!</p>\n<p>By the way, do you plan to create an <code>external data thread</code>?<br>\nI'm confused by the fact that various threads contain information about various datasets.</p>\n<p>If possible, I'd like host to create an external data thread, where people can post if they want to use datasets other than this competition.</p>",
      "rawMarkdown": "Thank you for your honesty and quick response!\n\nBy the way, do you plan to create an `external data thread`?\nI'm confused by the fact that various threads contain information about various datasets.\n\nIf possible, I'd like host to create an external data thread, where people can post if they want to use datasets other than this competition.",
      "votes": null
    },
    {
      "id": "1140821",
      "postDate": "01/06/2021 09:46:24",
      "content": "<p>About availability of public dataset,<br>\nnow in this competition we're allowed to label data manually, then I guess some teams are already preparing for hand labeling,<br>\nif new public tiff images are available like before, some teams will have additional advantage by hand-labeling public images, and I guess those situation is not what hosts want (I guess diversity of modeling is what they want).<br>\nThen I recommend not to share public images, make it only available for kernel submission inference, how do you think?</p>",
      "rawMarkdown": "About availability of public dataset,\nnow in this competition we're allowed to label data manually, then I guess some teams are already preparing for hand labeling,\nif new public tiff images are available like before, some teams will have additional advantage by hand-labeling public images, and I guess those situation is not what hosts want (I guess diversity of modeling is what they want).\nThen I recommend not to share public images, make it only available for kernel submission inference, how do you think?",
      "votes": null
    },
    {
      "id": "1149519",
      "postDate": "01/11/2021 22:49:31",
      "content": "<p>Just to confirm, has the dataset been updated?</p>",
      "rawMarkdown": "Just to confirm, has the dataset been updated?",
      "votes": null
    },
    {
      "id": "1149554",
      "postDate": "01/12/2021 00:04:37",
      "content": "<p>same question</p>",
      "rawMarkdown": "same question",
      "votes": null
    },
    {
      "id": "1150440",
      "postDate": "01/12/2021 15:33:42",
      "content": "<p>Not yet, but we will let everyone know as soon as it is!</p>",
      "rawMarkdown": "Not yet, but we will let everyone know as soon as it is!",
      "votes": null
    },
    {
      "id": "1150441",
      "postDate": "01/12/2021 15:33:42",
      "content": "<p>Not yet, but we will let everyone know as soon as it is!</p>",
      "rawMarkdown": "Not yet, but we will let everyone know as soon as it is!",
      "votes": null
    },
    {
      "id": "1150544",
      "postDate": "01/12/2021 16:51:50",
      "content": "<p>The amount of false positives ( non glomeruli) is very high in this version.</p>",
      "rawMarkdown": "The amount of false positives ( non glomeruli) is very high in this version.",
      "votes": null
    },
    {
      "id": "1150694",
      "postDate": "01/12/2021 19:17:27",
      "content": "<p>Awesome, thanks!</p>",
      "rawMarkdown": "Awesome, thanks!",
      "votes": null
    },
    {
      "id": "1157676",
      "postDate": "01/18/2021 04:58:53",
      "content": "<p><a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> Thank you for the actions taken.Its been long. When the dataset will have an update? Its already end of January.</p>",
      "rawMarkdown": "katyborner Thank you for the actions taken.Its been long. When the dataset will have an update? Its already end of January.",
      "votes": null
    },
    {
      "id": "1159633",
      "postDate": "01/19/2021 11:29:19",
      "content": "<p>I am having the same issue, especially with tiff 095bf7a1f, 1e2425f28, e79de561c which have the worst recall/ highest FN for me. most FNs don`t like gloms to my untrained eye. <a href=\"https://www.kaggle.com/rosuluc/error-analysis#Plot-worst-predictions\" target=\"_blank\">error analysis notebook here</a></p>",
      "rawMarkdown": "I am having the same issue, especially with tiff 095bf7a1f, 1e2425f28, e79de561c which have the worst recall/ highest FN for me. most FNs don`t like gloms to my untrained eye. [error analysis notebook here](https://www.kaggle.com/rosuluc/error-analysis#Plot-worst-predictions)",
      "votes": null
    },
    {
      "id": "1163863",
      "postDate": "01/22/2021 02:04:11",
      "content": "<p>Just to confirm, has the dataset been updated and when to update timeline？^_^</p>",
      "rawMarkdown": "Just to confirm, has the dataset been updated and when to update timeline？^_^",
      "votes": null
    },
    {
      "id": "1167780",
      "postDate": "01/24/2021 13:30:22",
      "content": "<p>Same question here, mid Jan is passed ;)</p>",
      "rawMarkdown": "Same question here, mid Jan is passed ;)",
      "votes": null
    },
    {
      "id": "1167831",
      "postDate": "01/24/2021 14:02:36",
      "content": "<p>It is indeed and still no update yet. :D<br>\nI have started looking at another competition in the meantime. ;)</p>",
      "rawMarkdown": "It is indeed and still no update yet. :D\nI have started looking at another competition in the meantime. ;)",
      "votes": null
    },
    {
      "id": "1168902",
      "postDate": "01/25/2021 08:35:12",
      "content": "<p>I have the same question: when will the new dataset be released? (:D)</p>",
      "rawMarkdown": "I have the same question: when will the new dataset be released? (:D)",
      "votes": null
    },
    {
      "id": "1169808",
      "postDate": "01/25/2021 18:40:29",
      "content": "<p>Check this answer here: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568</a></p>\n<p>In short: still no date yet and the deadline will be extended by 2 months once the new dataset is released.</p>",
      "rawMarkdown": "Check this answer here: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568\n\nIn short: still no date yet and the deadline will be extended by 2 months once the new dataset is released.",
      "votes": null
    },
    {
      "id": "1171839",
      "postDate": "01/27/2021 06:04:38",
      "content": "<p>When will the dataset be updated?</p>",
      "rawMarkdown": "When will the dataset be updated?",
      "votes": null
    },
    {
      "id": "1179903",
      "postDate": "02/01/2021 00:40:39",
      "content": "<p>Thanks so much for the updating efforts, but…as some people posted below or elsewhere, when will the new dataset be available?<br>\nWe’ve been waiting for 1 month…it’s early-February and no statement about the delay 😢</p>",
      "rawMarkdown": "Thanks so much for the updating efforts, but...as some people posted below or elsewhere, when will the new dataset be available?\nWe’ve been waiting for 1 month...it’s early-February and no statement about the delay 😢",
      "votes": null
    },
    {
      "id": "1180009",
      "postDate": "02/01/2021 03:38:21",
      "content": "<p>Thanks for the efforts of organizers from HuBMAP and Kaggle. I understand that a lot of works should be done to update the dataset, but is there any expected time for us to get the new data?</p>",
      "rawMarkdown": "Thanks for the efforts of organizers from HuBMAP and Kaggle. I understand that a lot of works should be done to update the dataset, but is there any expected time for us to get the new data?",
      "votes": null
    },
    {
      "id": "1180090",
      "postDate": "02/01/2021 04:50:20",
      "content": "<p>Hi all,</p>\n<p>Thanks for your patience! We're still preparing the new dataset. Standby!</p>",
      "rawMarkdown": "Hi all,\n\nThanks for your patience! We're still preparing the new dataset. Standby!",
      "votes": null
    },
    {
      "id": "1180120",
      "postDate": "02/01/2021 05:26:11",
      "content": "<p>Thank you for the announcement.</p>\n<p>However, our concern is the schedule, not the host’s status.<br>\nIs it difficult to announce the schedule?</p>",
      "rawMarkdown": "Thank you for the announcement.\n\nHowever, our concern is the schedule, not the host’s status.\nIs it difficult to announce the schedule?",
      "votes": null
    },
    {
      "id": "1180137",
      "postDate": "02/01/2021 05:49:07",
      "content": "<p>It's a bit difficult because the schedule is relative to the data release date. Once the data is released (date still TBD), we'll extend the competition two months from that time period to provide ample time to continue to work on the problem. </p>",
      "rawMarkdown": "It's a bit difficult because the schedule is relative to the data release date. Once the data is released (date still TBD), we'll extend the competition two months from that time period to provide ample time to continue to work on the problem.",
      "votes": null
    },
    {
      "id": "1180153",
      "postDate": "02/01/2021 06:03:02",
      "content": "<p>Thank you for your answer.<br>\nI'm looking forward to it :)</p>",
      "rawMarkdown": "Thank you for your answer.\nI'm looking forward to it :)",
      "votes": null
    },
    {
      "id": "1190823",
      "postDate": "02/08/2021 03:38:16",
      "content": "<p>Hi team</p>\n<p>May we please update <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/overview/timeline\" target=\"_blank\">the timeline section of the competition</a> to reflect the updated timeline -26th March. </p>\n<p>Thanks!!</p>",
      "rawMarkdown": "Hi team\n\nMay we please update [the timeline section of the competition](https://www.kaggle.com/c/hubmap-kidney-segmentation/overview/timeline) to reflect the updated timeline -26th March. \n\nThanks!!",
      "votes": null
    },
    {
      "id": "1206156",
      "postDate": "02/17/2021 07:40:57",
      "content": "<p>Does anyone know when the new train/public test datasets will be released?</p>",
      "rawMarkdown": "Does anyone know when the new train/public test datasets will be released?",
      "votes": null
    },
    {
      "id": "1206336",
      "postDate": "02/17/2021 08:25:52",
      "content": "<p>Pretty soon I guess, <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> posted an update in another discussion. </p>",
      "rawMarkdown": "Pretty soon I guess, @addisonhoward posted an update in another discussion.",
      "votes": null
    },
    {
      "id": "1207212",
      "postDate": "02/17/2021 18:45:07",
      "content": "<p>I'm not as optimistic.  There's still no committed date for the data which to me means we're still in TBD mode.  Medical segmentation data is often difficult and expensive to put together so the delay is understandable.</p>",
      "rawMarkdown": "I'm not as optimistic.  There's still no committed date for the data which to me means we're still in TBD mode.  Medical segmentation data is often difficult and expensive to put together so the delay is understandable.",
      "votes": null
    },
    {
      "id": "1212036",
      "postDate": "02/20/2021 19:55:11",
      "content": "<p>I'm concerned that the training set that I used to train my models is not of the same population as the test set that my models are applied to.   So, could someone please specify what are the up-to-date (as of 20 Jan 2021) Kaggle DataSets containing the Train images and the public Test images?   URLs preferred since names are pretty close to each other.</p>",
      "rawMarkdown": "I'm concerned that the training set that I used to train my models is not of the same population as the test set that my models are applied to.   So, could someone please specify what are the up-to-date (as of 20 Jan 2021) Kaggle DataSets containing the Train images and the public Test images?   URLs preferred since names are pretty close to each other.",
      "votes": null
    },
    {
      "id": "1214010",
      "postDate": "02/22/2021 14:30:02",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "1214521",
      "postDate": "02/23/2021 00:21:29",
      "content": "<p>Hi team,<br>\nDoes the data update, or when will the new data release? <br>\nThanks!!</p>",
      "rawMarkdown": "Hi team,\nDoes the data update, or when will the new data release? \nThanks!!",
      "votes": null
    },
    {
      "id": "1217974",
      "postDate": "02/25/2021 13:08:17",
      "content": "<p>I also wonder if we can also expect the structure to be more or less reliable in the new version of the public/private test dataset? For example, the 2f6ecfcdf-anatomical-structure.json looks to be messed-up for the test set, however, for training data, it looks much more consistent, with just a few glomeruli appear in the medulla part (by mistake?)</p>\n<p><img src=\"https://i.ibb.co/P6qhxd2/image.png\" alt=\"\"></p>",
      "rawMarkdown": "I also wonder if we can also expect the structure to be more or less reliable in the new version of the public/private test dataset? For example, the 2f6ecfcdf-anatomical-structure.json looks to be messed-up for the test set, however, for training data, it looks much more consistent, with just a few glomeruli appear in the medulla part (by mistake?)\n\n![](https://i.ibb.co/P6qhxd2/image.png)",
      "votes": null
    },
    {
      "id": "1218150",
      "postDate": "02/25/2021 15:41:06",
      "content": "<p>This may be a manual labeling problem, similar to noise? I guess this problem should still exist on the new data.</p>",
      "rawMarkdown": "This may be a manual labeling problem, similar to noise? I guess this problem should still exist on the new data.",
      "votes": null
    },
    {
      "id": "1218247",
      "postDate": "02/25/2021 16:58:13",
      "content": "<p>I don't remember where I read it but the organizers seem to be aware of the poor quality of the annotation and are working on improving this as well. Will see I guess.</p>",
      "rawMarkdown": "I don't remember where I read it but the organizers seem to be aware of the poor quality of the annotation and are working on improving this as well. Will see I guess.",
      "votes": null
    },
    {
      "id": "1218367",
      "postDate": "02/25/2021 19:02:07",
      "content": "<p>We're aware of this! Hence some of the delay in getting the latest version out - we want to make sure we perform extra checks on the data labeling.</p>",
      "rawMarkdown": "We're aware of this! Hence some of the delay in getting the latest version out - we want to make sure we perform extra checks on the data labeling.",
      "votes": null
    },
    {
      "id": "1218393",
      "postDate": "02/25/2021 19:55:03",
      "content": "<p>Thank you for the update, Addison! Looking forward to the new dataset.</p>\n<p>Hopefully, this kind of glomerulus-level labels will be fixed as well to reduce the chance of a lottery 🙂: image 0486052bb, index 2</p>\n<p><img src=\"https://i.ibb.co/JmdGzNk/image.png\" alt=\"\"></p>",
      "rawMarkdown": "Thank you for the update, Addison! Looking forward to the new dataset.\n\nHopefully, this kind of glomerulus-level labels will be fixed as well to reduce the chance of a lottery 🙂: image 0486052bb, index 2\n\n![](https://i.ibb.co/JmdGzNk/image.png)",
      "votes": null
    },
    {
      "id": "1221552",
      "postDate": "03/01/2021 04:57:52",
      "content": "<p>new data published ?</p>",
      "rawMarkdown": "new data published ?",
      "votes": null
    },
    {
      "id": "1221615",
      "postDate": "03/01/2021 06:01:30",
      "content": "<p>Looks like a D letter. 😁 </p>",
      "rawMarkdown": "Looks like a D letter. 😁",
      "votes": null
    },
    {
      "id": "1222137",
      "postDate": "03/01/2021 15:24:47",
      "content": "<p>Any news on the status of this competition?</p>",
      "rawMarkdown": "Any news on the status of this competition?",
      "votes": null
    },
    {
      "id": "1222358",
      "postDate": "03/01/2021 18:09:27",
      "content": "<p>Maybe a dying glomerulus 😏</p>",
      "rawMarkdown": "Maybe a dying glomerulus 😏",
      "votes": null
    },
    {
      "id": "1226941",
      "postDate": "03/05/2021 03:36:03",
      "content": "<p>Does anyone know when the new train/public test datasets will be released?</p>",
      "rawMarkdown": "Does anyone know when the new train/public test datasets will be released?",
      "votes": null
    },
    {
      "id": "1227460",
      "postDate": "03/05/2021 14:59:19",
      "content": "<p>I have been saying soon for a long time. So now I prefer to say \"wait and see\" :D</p>",
      "rawMarkdown": "I have been saying soon for a long time. So now I prefer to say \"wait and see\" :D",
      "votes": null
    },
    {
      "id": "1227945",
      "postDate": "03/06/2021 00:41:25",
      "content": "<p><a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> Any update on the new dataset? we know that the annotation takes more time but please just give us some updates. </p>",
      "rawMarkdown": "katyborner Any update on the new dataset? we know that the annotation takes more time but please just give us some updates.",
      "votes": null
    },
    {
      "id": "1230663",
      "postDate": "03/08/2021 10:42:55",
      "content": "<p>coule any one can tell me how to get the private test set? thanks a lot!</p>",
      "rawMarkdown": "coule any one can tell me how to get the private test set? thanks a lot!",
      "votes": null
    },
    {
      "id": "1253248",
      "postDate": "03/26/2021 14:17:08",
      "content": "<p>Dear organizers, thank you for releasing the new and improved data set. We've noticed that globally sclerotic glomeruli are not included in the ground truth anymore. Do we understand correctly that the networks should only learn to segment non-sclerotic glomeruli?</p>",
      "rawMarkdown": "Dear organizers, thank you for releasing the new and improved data set. We've noticed that globally sclerotic glomeruli are not included in the ground truth anymore. Do we understand correctly that the networks should only learn to segment non-sclerotic glomeruli?",
      "votes": null
    },
    {
      "id": "1294792",
      "postDate": "05/05/2021 23:37:54",
      "content": "<p>Given that the new private data is <strong>not</strong> available via the HuBMAP portal, hand-labeling of the data <strong>is now allowed</strong>. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.</p>\n<p>Just a quick question. Although the private test set is not seen, it can be seen by the model during notebook submission and used for pseudo-labeling and retraining. And this new model can be used to make the prediction. Is this allowed or forbidden?  <a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> </p>\n<p>I couldn't find anything on this in the discussion forum.</p>\n<p>thanks</p>",
      "rawMarkdown": "Given that the new private data is **not** available via the HuBMAP portal, hand-labeling of the data **is now allowed**. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.\n\n\n\n\n\n\nJust a quick question. Although the private test set is not seen, it can be seen by the model during notebook submission and used for pseudo-labeling and retraining. And this new model can be used to make the prediction. Is this allowed or forbidden?  @katyborner \n\nI couldn't find anything on this in the discussion forum.\n\nthanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1134083,
      "author_name": "tivfrvqhs5",
      "author_url": "",
      "post_date": "12/31/2020 20:16:25",
      "content": "<p>Thanks for the update.  If the 20 current images (8 train + 5 public test + 7 private test) will be made public with annotations released and then the 10 new images become the new private test, then what will be the new public test set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1134356,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "01/01/2021 07:38:41",
          "content": "<p>Great question - we'll announce the split with the upcoming data release in a few weeks. The split itself is still TBD on our end.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1134985,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "01/01/2021 19:12:20",
          "content": "<p>My concern would be that if only 10 new images are going to be annotated, it may be tough to make a public/private split out of those 10 that leaves a meaningful LB.  I realize 10 isn't a lot less than the 12 we have now but it places more emphasis on fewer images.<br>\n Just an idea: you could hold back the annotations on a few of the current test images and still use that for the new public test set.  This would also provide some continuity of the current public LB rather than have it completely reset.  There's lots of ways to slice things up and I'm sure Kaggle and the sponsors will figure it out and make it fair and meaningful.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2Fad7aa5606f79dcbfc8137a728bcfcc5a%2Fhubmap_data.png?generation=1609528269340385&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1134242,
      "author_name": "pratikkumar",
      "author_url": "",
      "post_date": "01/01/2021 03:54:55",
      "content": "<p>Would the ground truth shift in afa5e8098 be corrected in the revised dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1138312,
          "author_name": "leahscherschel",
          "author_url": "",
          "post_date": "01/04/2021 15:09:16",
          "content": "<p>Yes, it would.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1138468,
          "author_name": "andrasferenczi",
          "author_url": "",
          "post_date": "01/04/2021 17:49:29",
          "content": "<p>this also means we need to re-run the submissions to be used for the final score. Are we going to get extra GPU/TPU time, or will you re-run those? Thanks!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1138756,
          "author_name": "jamzing",
          "author_url": "",
          "post_date": "01/05/2021 00:08:17",
          "content": "<p>Thanks for bring that up <a href=\"https://www.kaggle.com/pratikkumar\" target=\"_blank\">@pratikkumar</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1134954,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "01/01/2021 18:19:28",
      "content": "<p>Wow those are great news.<br>\nThat is a lot of well appreciated work you are putting into this to make the competition as fair as possible and the results as robust as possible (usable for further research and production).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1135319,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "01/02/2021 06:04:03",
      "content": "<p>Dear organizers and kaggle team,<br>\nThank you for taking actions to resolve issues with publicly available private test set and the leak of the private score that potentially could allow some participants to discover the average shift in the private data that maximizes the score. So now the competition might work.<br>\nMy only concern is careful checking for the GT shift in the new test set before updating it (to avoid something we saw for the previous test set).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1136280,
      "author_name": "kulsooma",
      "author_url": "",
      "post_date": "01/02/2021 23:12:28",
      "content": "<p>I wanted to confirm if the current format of the anatomical_structures_segmention_file<br>\nglomerulus_segmentation_file jsons and RLE encoding column of thetrain.csv would be the same for the new dataset?<br>\nThank you for the information.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1138313,
          "author_name": "leahscherschel",
          "author_url": "",
          "post_date": "01/04/2021 15:09:44",
          "content": "<p>Yes, the format will remain the same.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1136956,
      "author_name": "purplejester",
      "author_url": "",
      "post_date": "01/03/2021 15:08:10",
      "content": "<p>That's great! A very responsible approach and well-timed decision. Thank you very much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1139256,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/05/2021 09:22:01",
      "content": "<p>thank you very much!<br>\nthanks for the effort to create new test dataset in such a short time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1140730,
      "author_name": "yukkyo",
      "author_url": "",
      "post_date": "01/06/2021 08:18:51",
      "content": "<p>Thank you for your honesty and quick response!</p>\n<p>By the way, do you plan to create an <code>external data thread</code>?<br>\nI'm confused by the fact that various threads contain information about various datasets.</p>\n<p>If possible, I'd like host to create an external data thread, where people can post if they want to use datasets other than this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1140821,
      "author_name": "ryunosukeishizaki",
      "author_url": "",
      "post_date": "01/06/2021 09:46:24",
      "content": "<p>About availability of public dataset,<br>\nnow in this competition we're allowed to label data manually, then I guess some teams are already preparing for hand labeling,<br>\nif new public tiff images are available like before, some teams will have additional advantage by hand-labeling public images, and I guess those situation is not what hosts want (I guess diversity of modeling is what they want).<br>\nThen I recommend not to share public images, make it only available for kernel submission inference, how do you think?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1149519,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "01/11/2021 22:49:31",
      "content": "<p>Just to confirm, has the dataset been updated?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1149554,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "01/12/2021 00:04:37",
          "content": "<p>same question</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150440,
          "author_name": "leahscherschel",
          "author_url": "",
          "post_date": "01/12/2021 15:33:42",
          "content": "<p>Not yet, but we will let everyone know as soon as it is!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150441,
          "author_name": "leahscherschel",
          "author_url": "",
          "post_date": "01/12/2021 15:33:42",
          "content": "<p>Not yet, but we will let everyone know as soon as it is!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1150694,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/12/2021 19:17:27",
          "content": "<p>Awesome, thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1150544,
      "author_name": "rpsantosakaggle",
      "author_url": "",
      "post_date": "01/12/2021 16:51:50",
      "content": "<p>The amount of false positives ( non glomeruli) is very high in this version.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1159633,
          "author_name": "rosuluc",
          "author_url": "",
          "post_date": "01/19/2021 11:29:19",
          "content": "<p>I am having the same issue, especially with tiff 095bf7a1f, 1e2425f28, e79de561c which have the worst recall/ highest FN for me. most FNs don`t like gloms to my untrained eye. <a href=\"https://www.kaggle.com/rosuluc/error-analysis#Plot-worst-predictions\" target=\"_blank\">error analysis notebook here</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1157676,
      "author_name": "arunmohan003",
      "author_url": "",
      "post_date": "01/18/2021 04:58:53",
      "content": "<p><a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> Thank you for the actions taken.Its been long. When the dataset will have an update? Its already end of January.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1163863,
      "author_name": "er9213",
      "author_url": "",
      "post_date": "01/22/2021 02:04:11",
      "content": "<p>Just to confirm, has the dataset been updated and when to update timeline？^_^</p>",
      "votes": null,
      "replies": [
        {
          "id": 1167780,
          "author_name": "siavashh",
          "author_url": "",
          "post_date": "01/24/2021 13:30:22",
          "content": "<p>Same question here, mid Jan is passed ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1167831,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/24/2021 14:02:36",
          "content": "<p>It is indeed and still no update yet. :D<br>\nI have started looking at another competition in the meantime. ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1168902,
      "author_name": "whitegg",
      "author_url": "",
      "post_date": "01/25/2021 08:35:12",
      "content": "<p>I have the same question: when will the new dataset be released? (:D)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1169808,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "01/25/2021 18:40:29",
          "content": "<p>Check this answer here: <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568\" target=\"_blank\">https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568</a></p>\n<p>In short: still no date yet and the deadline will be extended by 2 months once the new dataset is released.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1171839,
      "author_name": "sakuraikazumi",
      "author_url": "",
      "post_date": "01/27/2021 06:04:38",
      "content": "<p>When will the dataset be updated?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1179903,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "02/01/2021 00:40:39",
      "content": "<p>Thanks so much for the updating efforts, but…as some people posted below or elsewhere, when will the new dataset be available?<br>\nWe’ve been waiting for 1 month…it’s early-February and no statement about the delay 😢</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180009,
      "author_name": "lijiaqi96",
      "author_url": "",
      "post_date": "02/01/2021 03:38:21",
      "content": "<p>Thanks for the efforts of organizers from HuBMAP and Kaggle. I understand that a lot of works should be done to update the dataset, but is there any expected time for us to get the new data?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1180090,
      "author_name": "addisonhoward",
      "author_url": "",
      "post_date": "02/01/2021 04:50:20",
      "content": "<p>Hi all,</p>\n<p>Thanks for your patience! We're still preparing the new dataset. Standby!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1180120,
          "author_name": "yukkyo",
          "author_url": "",
          "post_date": "02/01/2021 05:26:11",
          "content": "<p>Thank you for the announcement.</p>\n<p>However, our concern is the schedule, not the host’s status.<br>\nIs it difficult to announce the schedule?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1180137,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "02/01/2021 05:49:07",
          "content": "<p>It's a bit difficult because the schedule is relative to the data release date. Once the data is released (date still TBD), we'll extend the competition two months from that time period to provide ample time to continue to work on the problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1180153,
          "author_name": "yukkyo",
          "author_url": "",
          "post_date": "02/01/2021 06:03:02",
          "content": "<p>Thank you for your answer.<br>\nI'm looking forward to it :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1190823,
      "author_name": "kmldas",
      "author_url": "",
      "post_date": "02/08/2021 03:38:16",
      "content": "<p>Hi team</p>\n<p>May we please update <a href=\"https://www.kaggle.com/c/hubmap-kidney-segmentation/overview/timeline\" target=\"_blank\">the timeline section of the competition</a> to reflect the updated timeline -26th March. </p>\n<p>Thanks!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1206156,
      "author_name": "fattane",
      "author_url": "",
      "post_date": "02/17/2021 07:40:57",
      "content": "<p>Does anyone know when the new train/public test datasets will be released?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1206336,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/17/2021 08:25:52",
          "content": "<p>Pretty soon I guess, <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> posted an update in another discussion. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207212,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "02/17/2021 18:45:07",
          "content": "<p>I'm not as optimistic.  There's still no committed date for the data which to me means we're still in TBD mode.  Medical segmentation data is often difficult and expensive to put together so the delay is understandable.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1212036,
      "author_name": "markalavin",
      "author_url": "",
      "post_date": "02/20/2021 19:55:11",
      "content": "<p>I'm concerned that the training set that I used to train my models is not of the same population as the test set that my models are applied to.   So, could someone please specify what are the up-to-date (as of 20 Jan 2021) Kaggle DataSets containing the Train images and the public Test images?   URLs preferred since names are pretty close to each other.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214010,
      "author_name": "shilei2403",
      "author_url": "",
      "post_date": "02/22/2021 14:30:02",
      "content": "<p>thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1214521,
      "author_name": "aikeyz",
      "author_url": "",
      "post_date": "02/23/2021 00:21:29",
      "content": "<p>Hi team,<br>\nDoes the data update, or when will the new data release? <br>\nThanks!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1217974,
      "author_name": "gdonchyts",
      "author_url": "",
      "post_date": "02/25/2021 13:08:17",
      "content": "<p>I also wonder if we can also expect the structure to be more or less reliable in the new version of the public/private test dataset? For example, the 2f6ecfcdf-anatomical-structure.json looks to be messed-up for the test set, however, for training data, it looks much more consistent, with just a few glomeruli appear in the medulla part (by mistake?)</p>\n<p><img src=\"https://i.ibb.co/P6qhxd2/image.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1218150,
          "author_name": "whitegg",
          "author_url": "",
          "post_date": "02/25/2021 15:41:06",
          "content": "<p>This may be a manual labeling problem, similar to noise? I guess this problem should still exist on the new data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218247,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/25/2021 16:58:13",
          "content": "<p>I don't remember where I read it but the organizers seem to be aware of the poor quality of the annotation and are working on improving this as well. Will see I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218367,
          "author_name": "addisonhoward",
          "author_url": "",
          "post_date": "02/25/2021 19:02:07",
          "content": "<p>We're aware of this! Hence some of the delay in getting the latest version out - we want to make sure we perform extra checks on the data labeling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1218393,
          "author_name": "gdonchyts",
          "author_url": "",
          "post_date": "02/25/2021 19:55:03",
          "content": "<p>Thank you for the update, Addison! Looking forward to the new dataset.</p>\n<p>Hopefully, this kind of glomerulus-level labels will be fixed as well to reduce the chance of a lottery 🙂: image 0486052bb, index 2</p>\n<p><img src=\"https://i.ibb.co/JmdGzNk/image.png\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221615,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/01/2021 06:01:30",
          "content": "<p>Looks like a D letter. 😁 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222358,
          "author_name": "gdonchyts",
          "author_url": "",
          "post_date": "03/01/2021 18:09:27",
          "content": "<p>Maybe a dying glomerulus 😏</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1221552,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "03/01/2021 04:57:52",
      "content": "<p>new data published ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222137,
      "author_name": "antoni4040",
      "author_url": "",
      "post_date": "03/01/2021 15:24:47",
      "content": "<p>Any news on the status of this competition?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1226941,
      "author_name": "er9213",
      "author_url": "",
      "post_date": "03/05/2021 03:36:03",
      "content": "<p>Does anyone know when the new train/public test datasets will be released?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1227460,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/05/2021 14:59:19",
          "content": "<p>I have been saying soon for a long time. So now I prefer to say \"wait and see\" :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1227945,
      "author_name": "aramisvesal",
      "author_url": "",
      "post_date": "03/06/2021 00:41:25",
      "content": "<p><a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> Any update on the new dataset? we know that the annotation takes more time but please just give us some updates. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1230663,
          "author_name": "ambitionkingo",
          "author_url": "",
          "post_date": "03/08/2021 10:42:55",
          "content": "<p>coule any one can tell me how to get the private test set? thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1253248,
      "author_name": "mhermsen",
      "author_url": "",
      "post_date": "03/26/2021 14:17:08",
      "content": "<p>Dear organizers, thank you for releasing the new and improved data set. We've noticed that globally sclerotic glomeruli are not included in the ground truth anymore. Do we understand correctly that the networks should only learn to segment non-sclerotic glomeruli?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1294792,
      "author_name": "mustafaa",
      "author_url": "",
      "post_date": "05/05/2021 23:37:54",
      "content": "<p>Given that the new private data is <strong>not</strong> available via the HuBMAP portal, hand-labeling of the data <strong>is now allowed</strong>. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.</p>\n<p>Just a quick question. Although the private test set is not seen, it can be seen by the model during notebook submission and used for pseudo-labeling and retraining. And this new model can be used to make the prediction. Is this allowed or forbidden?  <a href=\"https://www.kaggle.com/katyborner\" target=\"_blank\">@katyborner</a> </p>\n<p>I couldn't find anything on this in the discussion forum.</p>\n<p>thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1134044": "Dear Teams,\nWe are excited and impressed with all the engagement and participation for this hackathon. There is significant potential for the development of annotation tools that could greatly accelerate impact on human health, and we appreciate the substantial time investment many participants have made. The scientific value of the contest is of the utmost importance.\n\nUnfortunately, it has come to our attention that unlabeled images in the private test data have been discovered on the NIH HuBMAP data portal, compromising the ability of tools developed to e.g., generalize to new datasets. \n\nPlease accept our sincere apologies. \n\nWe, the hackathon organizers, are taking the necessary step of completely updating the private test data. We have also added new quality control measures to the glomeruli annotation masks. \n\nSpecifically, we will do the following:\n- All 20 of the currently used datasets will be made public and the 20 glomeruli annotation mask JSON files will be updated using the new quality controls. \n- A new, unpublished dataset of 10 images (5 fresh frozen, 5 FFPE) and their annotation masks (that have undergone the very same new quality controls) will replace the current private test dataset. \nThese data will remain unseen and unpublished until after the competition has concluded. We have added two-factor authentication and restricted access to ensure security of these data.\n- All new data will become available in mid-January 2021 and the final submission deadline will be extended to mid to late March 2021. All other deadlines will be extended accordingly. We have temporarily set a new deadline date as a placeholder while the new dataset is under review and preparation, and will update the timeline with finalized dates once the new dataset is released.\n\nGiven that the new private data is **not** available via the HuBMAP portal, hand-labeling of the data **is now allowed**. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.\n\nWe are working closely with the NIH HuBMAP Program Office, the competition sponsors, and the competition judges while implementing these changes and would like to thank all involved for their extensive expert support.\n\nLast but not least, we wish you a marvelous New Year 2021!\n\nSincerely,\nKaty Borner",
    "1134083": "Thanks for the update.  If the 20 current images (8 train + 5 public test + 7 private test) will be made public with annotations released and then the 10 new images become the new private test, then what will be the new public test set?",
    "1134242": "Would the ground truth shift in afa5e8098 be corrected in the revised dataset?",
    "1134356": "Great question - we'll announce the split with the upcoming data release in a few weeks. The split itself is still TBD on our end.",
    "1134954": "Wow those are great news.\nThat is a lot of well appreciated work you are putting into this to make the competition as fair as possible and the results as robust as possible (usable for further research and production).",
    "1134985": "My concern would be that if only 10 new images are going to be annotated, it may be tough to make a public/private split out of those 10 that leaves a meaningful LB.  I realize 10 isn't a lot less than the 12 we have now but it places more emphasis on fewer images.\n Just an idea: you could hold back the annotations on a few of the current test images and still use that for the new public test set.  This would also provide some continuity of the current public LB rather than have it completely reset.  There's lots of ways to slice things up and I'm sure Kaggle and the sponsors will figure it out and make it fair and meaningful.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F888191%2Fad7aa5606f79dcbfc8137a728bcfcc5a%2Fhubmap_data.png?generation=1609528269340385&alt=media)",
    "1135319": "Dear organizers and kaggle team,\nThank you for taking actions to resolve issues with publicly available private test set and the leak of the private score that potentially could allow some participants to discover the average shift in the private data that maximizes the score. So now the competition might work.\nMy only concern is careful checking for the GT shift in the new test set before updating it (to avoid something we saw for the previous test set).",
    "1136280": "I wanted to confirm if the current format of the anatomical_structures_segmention_file\nglomerulus_segmentation_file jsons and RLE encoding column of thetrain.csv would be the same for the new dataset?\nThank you for the information.",
    "1136956": "That's great! A very responsible approach and well-timed decision. Thank you very much!",
    "1138312": "Yes, it would.",
    "1138313": "Yes, the format will remain the same.",
    "1138468": "this also means we need to re-run the submissions to be used for the final score. Are we going to get extra GPU/TPU time, or will you re-run those? Thanks!!",
    "1138756": "Thanks for bring that up @pratikkumar.",
    "1139256": "thank you very much!\nthanks for the effort to create new test dataset in such a short time.",
    "1140730": "Thank you for your honesty and quick response!\n\nBy the way, do you plan to create an `external data thread`?\nI'm confused by the fact that various threads contain information about various datasets.\n\nIf possible, I'd like host to create an external data thread, where people can post if they want to use datasets other than this competition.",
    "1140821": "About availability of public dataset,\nnow in this competition we're allowed to label data manually, then I guess some teams are already preparing for hand labeling,\nif new public tiff images are available like before, some teams will have additional advantage by hand-labeling public images, and I guess those situation is not what hosts want (I guess diversity of modeling is what they want).\nThen I recommend not to share public images, make it only available for kernel submission inference, how do you think?",
    "1149519": "Just to confirm, has the dataset been updated?",
    "1149554": "same question",
    "1150440": "Not yet, but we will let everyone know as soon as it is!",
    "1150441": "Not yet, but we will let everyone know as soon as it is!",
    "1150544": "The amount of false positives ( non glomeruli) is very high in this version.",
    "1150694": "Awesome, thanks!",
    "1157676": "katyborner Thank you for the actions taken.Its been long. When the dataset will have an update? Its already end of January.",
    "1159633": "I am having the same issue, especially with tiff 095bf7a1f, 1e2425f28, e79de561c which have the worst recall/ highest FN for me. most FNs don`t like gloms to my untrained eye. [error analysis notebook here](https://www.kaggle.com/rosuluc/error-analysis#Plot-worst-predictions)",
    "1163863": "Just to confirm, has the dataset been updated and when to update timeline？^_^",
    "1167780": "Same question here, mid Jan is passed ;)",
    "1167831": "It is indeed and still no update yet. :D\nI have started looking at another competition in the meantime. ;)",
    "1168902": "I have the same question: when will the new dataset be released? (:D)",
    "1169808": "Check this answer here: https://www.kaggle.com/c/hubmap-kidney-segmentation/discussion/211446#1169568\n\nIn short: still no date yet and the deadline will be extended by 2 months once the new dataset is released.",
    "1171839": "When will the dataset be updated?",
    "1179903": "Thanks so much for the updating efforts, but...as some people posted below or elsewhere, when will the new dataset be available?\nWe’ve been waiting for 1 month...it’s early-February and no statement about the delay 😢",
    "1180009": "Thanks for the efforts of organizers from HuBMAP and Kaggle. I understand that a lot of works should be done to update the dataset, but is there any expected time for us to get the new data?",
    "1180090": "Hi all,\n\nThanks for your patience! We're still preparing the new dataset. Standby!",
    "1180120": "Thank you for the announcement.\n\nHowever, our concern is the schedule, not the host’s status.\nIs it difficult to announce the schedule?",
    "1180137": "It's a bit difficult because the schedule is relative to the data release date. Once the data is released (date still TBD), we'll extend the competition two months from that time period to provide ample time to continue to work on the problem.",
    "1180153": "Thank you for your answer.\nI'm looking forward to it :)",
    "1190823": "Hi team\n\nMay we please update [the timeline section of the competition](https://www.kaggle.com/c/hubmap-kidney-segmentation/overview/timeline) to reflect the updated timeline -26th March. \n\nThanks!!",
    "1206156": "Does anyone know when the new train/public test datasets will be released?",
    "1206336": "Pretty soon I guess, @addisonhoward posted an update in another discussion.",
    "1207212": "I'm not as optimistic.  There's still no committed date for the data which to me means we're still in TBD mode.  Medical segmentation data is often difficult and expensive to put together so the delay is understandable.",
    "1212036": "I'm concerned that the training set that I used to train my models is not of the same population as the test set that my models are applied to.   So, could someone please specify what are the up-to-date (as of 20 Jan 2021) Kaggle DataSets containing the Train images and the public Test images?   URLs preferred since names are pretty close to each other.",
    "1214010": "thank you!",
    "1214521": "Hi team,\nDoes the data update, or when will the new data release? \nThanks!!",
    "1217974": "I also wonder if we can also expect the structure to be more or less reliable in the new version of the public/private test dataset? For example, the 2f6ecfcdf-anatomical-structure.json looks to be messed-up for the test set, however, for training data, it looks much more consistent, with just a few glomeruli appear in the medulla part (by mistake?)\n\n![](https://i.ibb.co/P6qhxd2/image.png)",
    "1218150": "This may be a manual labeling problem, similar to noise? I guess this problem should still exist on the new data.",
    "1218247": "I don't remember where I read it but the organizers seem to be aware of the poor quality of the annotation and are working on improving this as well. Will see I guess.",
    "1218367": "We're aware of this! Hence some of the delay in getting the latest version out - we want to make sure we perform extra checks on the data labeling.",
    "1218393": "Thank you for the update, Addison! Looking forward to the new dataset.\n\nHopefully, this kind of glomerulus-level labels will be fixed as well to reduce the chance of a lottery 🙂: image 0486052bb, index 2\n\n![](https://i.ibb.co/JmdGzNk/image.png)",
    "1221552": "new data published ?",
    "1221615": "Looks like a D letter. 😁",
    "1222137": "Any news on the status of this competition?",
    "1222358": "Maybe a dying glomerulus 😏",
    "1226941": "Does anyone know when the new train/public test datasets will be released?",
    "1227460": "I have been saying soon for a long time. So now I prefer to say \"wait and see\" :D",
    "1227945": "katyborner Any update on the new dataset? we know that the annotation takes more time but please just give us some updates.",
    "1230663": "coule any one can tell me how to get the private test set? thanks a lot!",
    "1253248": "Dear organizers, thank you for releasing the new and improved data set. We've noticed that globally sclerotic glomeruli are not included in the ground truth anymore. Do we understand correctly that the networks should only learn to segment non-sclerotic glomeruli?",
    "1294792": "Given that the new private data is **not** available via the HuBMAP portal, hand-labeling of the data **is now allowed**. Pseudo-labeling (e.g., using a model to label additional data beyond the provided dataset) is also allowed.\n\n\n\n\n\n\nJust a quick question. Although the private test set is not seen, it can be seen by the model during notebook submission and used for pseudo-labeling and retraining. And this new model can be used to make the prediction. Is this allowed or forbidden?  @katyborner \n\nI couldn't find anything on this in the discussion forum.\n\nthanks"
  },
  "source": "meta"
}