{
  "id": 664184,
  "title": "Dataset update with improved label quality",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/664184",
  "author_name": "Sohier Dane",
  "post_date": "2025-12-22T23:37:41.766000",
  "votes": 33,
  "comment_count": 51,
  "views": 0,
  "content": "<p>We're in the process of posting an updated copy of the dataset with improved label quality. The update will cover:</p>\n<ul>\n<li>Incorrectly merged components were split. In a few cases this was impractical and the images have been moved to a new <code>deprecated</code> folder. We do not recommend continuing to use those files but wanted to ensure equal access in case someone does find a use for them.</li>\n<li>Very small connected components which were incorrectly labeled as foreground have been removed.</li>\n<li>Holes and voids have been manually fixed where required, programmatically addressed with binary hole filling or other morphological operations. A limited number of samples were also removed for these defects in cases where fixing would take an extended period of time.</li>\n<li>Label thickness should now have a more uniform 3vx thick labels. The dataset had a 3vx padding in all dimensions that was supposed to be an ignore label (value = 2), but it was erroneously labeled as background (value = 0). This padding was not reported in the original dataset description. We are ensuring now that it is correctly labeled as 2.</li>\n</ul>\n<p>I will be rescoring all submissions shortly as a few test set images are now ignored for scoring purposes. Note that this may take a day or so to complete due to the long metric runtime.</p>\n<p>Please note that we may need to make additional updates after the holidays, but this update is expected to cover most of the fixes.</p>\n<p>Edit: The rescore is complete. </p>\n<p>Edit 2: The initial rescore <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/664184#3383353\" target=\"_blank\">did not properly reflect the update</a>. I have resolved the issue and started a fresh rescore. Apologies for the confusion.</p>\n<p>Edit 3: The second rescore is now complete.</p>",
  "messages": [
    {
      "id": 3380737,
      "postDate": "2025-12-22T23:37:41.767Z",
      "content": "<p>We're in the process of posting an updated copy of the dataset with improved label quality. The update will cover:</p>\n<ul>\n<li>Incorrectly merged components were split. In a few cases this was impractical and the images have been moved to a new <code>deprecated</code> folder. We do not recommend continuing to use those files but wanted to ensure equal access in case someone does find a use for them.</li>\n<li>Very small connected components which were incorrectly labeled as foreground have been removed.</li>\n<li>Holes and voids have been manually fixed where required, programmatically addressed with binary hole filling or other morphological operations. A limited number of samples were also removed for these defects in cases where fixing would take an extended period of time.</li>\n<li>Label thickness should now have a more uniform 3vx thick labels. The dataset had a 3vx padding in all dimensions that was supposed to be an ignore label (value = 2), but it was erroneously labeled as background (value = 0). This padding was not reported in the original dataset description. We are ensuring now that it is correctly labeled as 2.</li>\n</ul>\n<p>I will be rescoring all submissions shortly as a few test set images are now ignored for scoring purposes. Note that this may take a day or so to complete due to the long metric runtime.</p>\n<p>Please note that we may need to make additional updates after the holidays, but this update is expected to cover most of the fixes.</p>\n<p>Edit: The rescore is complete. </p>\n<p>Edit 2: The initial rescore <a href=\"https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/664184#3383353\" target=\"_blank\">did not properly reflect the update</a>. I have resolved the issue and started a fresh rescore. Apologies for the confusion.</p>\n<p>Edit 3: The second rescore is now complete.</p>",
      "rawMarkdown": "We're in the process of posting an updated copy of the dataset with improved label quality. The update will cover:\n\n- Incorrectly merged components were split. In a few cases this was impractical and the images have been moved to a new `deprecated` folder. We do not recommend continuing to use those files but wanted to ensure equal access in case someone does find a use for them.\n- Very small connected components which were incorrectly labeled as foreground have been removed.\n- Holes and voids have been manually fixed where required, programmatically addressed with binary hole filling or other morphological operations. A limited number of samples were also removed for these defects in cases where fixing would take an extended period of time.\n- Label thickness should now have a more uniform 3vx thick labels. The dataset had a 3vx padding in all dimensions that was supposed to be an ignore label (value = 2), but it was erroneously labeled as background (value = 0). This padding was not reported in the original dataset description. We are ensuring now that it is correctly labeled as 2.\n\nI will be rescoring all submissions shortly as a few test set images are now ignored for scoring purposes. Note that this may take a day or so to complete due to the long metric runtime.\n\nPlease note that we may need to make additional updates after the holidays, but this update is expected to cover most of the fixes.\n\nEdit: The rescore is complete. \n\nEdit 2: The initial rescore [did not properly reflect the update](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/664184#3383353). I have resolved the issue and started a fresh rescore. Apologies for the confusion.\n\nEdit 3: The second rescore is now complete.",
      "votes": 33
    },
    {
      "id": 3381743,
      "postDate": "2025-12-25T12:36:02.407Z",
      "content": "<p>We’re investigating reports that the recent dataset update may not be fully reflected in the evaluation (e.g., unchanged scores after rescoring and lower public LB scores for models trained on the updated data). We’re auditing the update process and coordinating with Kaggle support as needed.\nGiven the holiday period, responses may be a bit slower than usual — thanks for your patience. In the meantime, please keep iterating using your own validation/CV on the updated training data; that’s the most reliable signal while we verify the leaderboard side. We’ll post an update as soon as we know more, and if a fix is needed we’ll ensure scoring is handled fairly for everyone (including re-scoring if required).</p>",
      "rawMarkdown": "We’re investigating reports that the recent dataset update may not be fully reflected in the evaluation (e.g., unchanged scores after rescoring and lower public LB scores for models trained on the updated data). We’re auditing the update process and coordinating with Kaggle support as needed.\nGiven the holiday period, responses may be a bit slower than usual — thanks for your patience. In the meantime, please keep iterating using your own validation/CV on the updated training data; that’s the most reliable signal while we verify the leaderboard side. We’ll post an update as soon as we know more, and if a fix is needed we’ll ensure scoring is handled fairly for everyone (including re-scoring if required).",
      "votes": 14
    },
    {
      "id": 3380779,
      "postDate": "2025-12-23T03:13:32.913Z",
      "content": "<p>Before:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F8195ae7f8885e188ec57628c4bb26674%2F2025-12-23%2011.54.55.png?generation=1766458516230057&amp;alt=media\" alt=\"\"> \nAfter:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fb2116191a724afb70a1721ab5311632e%2F2025-12-23%2011.55.04.png?generation=1766458533613557&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Before:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F8195ae7f8885e188ec57628c4bb26674%2F2025-12-23%2011.54.55.png?generation=1766458516230057&alt=media) \nAfter:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fb2116191a724afb70a1721ab5311632e%2F2025-12-23%2011.55.04.png?generation=1766458533613557&alt=media)",
      "votes": 10,
      "replies": [
        {
          "id": 3380824,
          "postDate": "2025-12-23T05:58:09.120Z",
          "content": "<p>Hmmm, will probably have to retrain a lot if all the labels have been reduced to thin sheets like this.</p>",
          "rawMarkdown": "Hmmm, will probably have to retrain a lot if all the labels have been reduced to thin sheets like this.",
          "replies": [
            {
              "id": 3380839,
              "postDate": "2025-12-23T07:04:48.637Z",
              "content": "<p>Fact is that the previous labels were already thin sheets like these in multiple places. This round we tried to keep the thickness consistent across the dataset.</p>",
              "rawMarkdown": "Fact is that the previous labels were already thin sheets like these in multiple places. This round we tried to keep the thickness consistent across the dataset."
            }
          ]
        }
      ]
    },
    {
      "id": 3386879,
      "postDate": "2026-01-06T02:16:25.923Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F0b7abe81673d70d4d80557c7a420bf75%2FScreenshot%202024-05-05%20222202.png?generation=1767665765991708&amp;alt=media\" alt=\"\"></p>\n<p>id = 1006462223</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F0b7abe81673d70d4d80557c7a420bf75%2FScreenshot%202024-05-05%20222202.png?generation=1767665765991708&alt=media)\n\nid = 1006462223",
      "votes": 3
    },
    {
      "id": 3381497,
      "postDate": "2025-12-24T18:34:37.020Z",
      "content": "<p>Thank you very much for updating the dataset.</p>\n<p>I still feel like a Psyduck and want some help here.</p>\n<table>\n  <tbody><tr>\n    <td>\n      <p><b>Old Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Fc104695a25692638dbf9efdcb8d65ea1%2Fvesuvius_404970490_old.png?generation=1766600021359526&amp;alt=media\">\n    </td>\n    <td>\n      <p><b>New Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Ff1ac3c464ac074ff54ff5870d0668b80%2Fvesuvius_404970490_new.png?generation=1766600046174188&amp;alt=media\">\n    </td>\n  </tr>\n</tbody></table>\n<p>For each layer, suppose it has two surfaces.</p>\n<p>It seems only one surface is preferred to be the label.</p>\n<p>I wonder how is the surface selected during data annotation?</p>\n<p>I also wonder why not choose the \"center-face\", similar to the concept of centerline?</p>\n<p>Now with the surface thinner, being 3 voxels, if the surface is randomly chosen, it is more unpredictable compared to the previous thicker version of surface.</p>\n<p>For the same prediction by a previous model, my local validation weighted topo-3d score drops from 0.58 to 0.49, for old and new labels respectively.</p>\n<p>BTW, my model is based on <a href=\"https://github.com/MIC-DKFZ/nnUNet\" target=\"_blank\">https://github.com/MIC-DKFZ/nnUNet</a></p>\n<p>And my submission notebook is based on <a href=\"https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46\" target=\"_blank\">https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46</a></p>\n<p>Many thanks to the original authors for their kind sharing.</p>",
      "rawMarkdown": "Thank you very much for updating the dataset.\n\nI still feel like a Psyduck and want some help here.\n\n<table>\n  <tr>\n    <td style=\"text-align: center;\">\n      <p><b>Old Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Fc104695a25692638dbf9efdcb8d65ea1%2Fvesuvius_404970490_old.png?generation=1766600021359526&alt=media\" width=\"500\">\n    </td>\n    <td style=\"text-align: center;\">\n      <p><b>New Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Ff1ac3c464ac074ff54ff5870d0668b80%2Fvesuvius_404970490_new.png?generation=1766600046174188&alt=media\" width=\"500\">\n    </td>\n  </tr>\n</table>\n\nFor each layer, suppose it has two surfaces.\n\nIt seems only one surface is preferred to be the label.\n\nI wonder how is the surface selected during data annotation?\n\nI also wonder why not choose the \"center-face\", similar to the concept of centerline?\n\nNow with the surface thinner, being 3 voxels, if the surface is randomly chosen, it is more unpredictable compared to the previous thicker version of surface.\n\nFor the same prediction by a previous model, my local validation weighted topo-3d score drops from 0.58 to 0.49, for old and new labels respectively.\n\nBTW, my model is based on https://github.com/MIC-DKFZ/nnUNet\n\nAnd my submission notebook is based on https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46\n\nMany thanks to the original authors for their kind sharing.",
      "votes": 3,
      "replies": [
        {
          "id": 3381571,
          "postDate": "2025-12-24T23:05:29.763Z",
          "content": "<p>The surface chosen is the “recto” or front surface , as this is the surface which contains the text (in most instances, there are some scrolls written on both sides but this is much less common) </p>\n<p>To the idea of using a centerline — the reason this is not chosen is it is not infrequent for the recto and verso to be split , sometimes for a rather long period of time and at reasonably large distances, and in these cases the concept of a centerline is somewhat ambiguous. </p>\n<p>So the choice is more of a pragmatic one: first , because that’s where the text we seek lies, and second because it’s just more practical. There is no real “inner” part of a papyrus sheet ; its horizontal strips on the front and vertical ones on the back, glued together. </p>\n<p>The labels should all lie on this recto surface. </p>",
          "rawMarkdown": "The surface chosen is the “recto” or front surface , as this is the surface which contains the text (in most instances, there are some scrolls written on both sides but this is much less common) \n\nTo the idea of using a centerline — the reason this is not chosen is it is not infrequent for the recto and verso to be split , sometimes for a rather long period of time and at reasonably large distances, and in these cases the concept of a centerline is somewhat ambiguous. \n\n\nSo the choice is more of a pragmatic one: first , because that’s where the text we seek lies, and second because it’s just more practical. There is no real “inner” part of a papyrus sheet ; its horizontal strips on the front and vertical ones on the back, glued together. \n\nThe labels should all lie on this recto surface. ",
          "votes": 3,
          "replies": [
            {
              "id": 3381574,
              "postDate": "2025-12-24T23:22:32.963Z",
              "content": "<p>Thank you Sean.\nThis is very informative.\nI hope the model can capture the difference between the recto and verso.</p>",
              "rawMarkdown": "Thank you Sean.\nThis is very informative.\nI hope the model can capture the difference between the recto and verso.",
              "votes": 2
            },
            {
              "id": 3381575,
              "postDate": "2025-12-24T23:41:28.717Z",
              "content": "<p>In most cases I've found it capable. the model does not really struggle to detect the recto side. the primary struggle it faces in when a sheet is frayed or torn, and when the sheets are very tightly packed. </p>\n<p>here is an example of a model i have trained with inference on the entirety of Scroll 1 (PHerc Paris 4) :</p>\n<ul>\n<li><a href=\"https://dl.ash2txt.org/community-uploads/bruniss/scrolls/s1/surfaces/s1_059_ome.zarr/\" target=\"_blank\">https://dl.ash2txt.org/community-uploads/bruniss/scrolls/s1/surfaces/s1_059_ome.zarr/</a> </li>\n</ul>\n<p>the primary way to help the model get better at recto vs verso is heavy spatial augmentations (flips, rotations, mirrors, etc) </p>",
              "rawMarkdown": "In most cases I've found it capable. the model does not really struggle to detect the recto side. the primary struggle it faces in when a sheet is frayed or torn, and when the sheets are very tightly packed. \n\n here is an example of a model i have trained with inference on the entirety of Scroll 1 (PHerc Paris 4) :\n- https://dl.ash2txt.org/community-uploads/bruniss/scrolls/s1/surfaces/s1_059_ome.zarr/ \n\nthe primary way to help the model get better at recto vs verso is heavy spatial augmentations (flips, rotations, mirrors, etc) ",
              "votes": 2
            }
          ]
        },
        {
          "id": 3381577,
          "postDate": "2025-12-24T23:56:26.800Z",
          "content": "<p>Also if you use that notebook I would highly recommend against using the morphological closing that is used by it , it absolutely will cause components to merge. Even the open is somewhat risky as any 1vx thick line will be removed and any 2-3vx line is at risk of becoming multiple components </p>",
          "rawMarkdown": "Also if you use that notebook I would highly recommend against using the morphological closing that is used by it , it absolutely will cause components to merge. Even the open is somewhat risky as any 1vx thick line will be removed and any 2-3vx line is at risk of becoming multiple components ",
          "votes": 2,
          "replies": [
            {
              "id": 3381585,
              "postDate": "2025-12-25T01:08:28.027Z",
              "content": "<p>I agree with you on morphological closing, and I think opening is more preferable.</p>\n<p>I just looked at my trained results from the new labels, which is now much better and requires less post-processing efforts, since the predicted sheets are already quite far away from each other.</p>",
              "rawMarkdown": "I agree with you on morphological closing, and I think opening is more preferable.\n\nI just looked at my trained results from the new labels, which is now much better and requires less post-processing efforts, since the predicted sheets are already quite far away from each other."
            },
            {
              "id": 3381586,
              "postDate": "2025-12-25T01:16:50.317Z",
              "content": "<p>The nnUNet baseline has quite good performance out of the box.</p>\n<p>For the submission notebook, I mostly borrowed the logic about test.csv, visualization before submission, and the zip file format required. This part I tried to find an official one, yet none exists….</p>\n<p>For post-processing, the suitable operators are also heavily dependent on the first stage predictions, and users may need some trial-and-error tweaks for their cases.</p>",
              "rawMarkdown": "The nnUNet baseline has quite good performance out of the box.\n\nFor the submission notebook, I mostly borrowed the logic about test.csv, visualization before submission, and the zip file format required. This part I tried to find an official one, yet none exists....\n\nFor post-processing, the suitable operators are also heavily dependent on the first stage predictions, and users may need some trial-and-error tweaks for their cases."
            },
            {
              "id": 3381934,
              "postDate": "2025-12-26T01:34:28.180Z",
              "content": "<p>Yes nnUNetv2 baseline nnUNetTrainer is an excellent starting point. I would not be surprised if DKFZ single submission is just the regular ResEncM preset with the DA5 trainer and no other modifications other than to max epochs. I'd expect the baseline nnUNet to come in around 53-55 lb score on prev data, unsure of what it would be since the update as i have not yet trained a model with it. </p>",
              "rawMarkdown": "Yes nnUNetv2 baseline nnUNetTrainer is an excellent starting point. I would not be surprised if DKFZ single submission is just the regular ResEncM preset with the DA5 trainer and no other modifications other than to max epochs. I'd expect the baseline nnUNet to come in around 53-55 lb score on prev data, unsure of what it would be since the update as i have not yet trained a model with it. \n\n",
              "votes": 1
            },
            {
              "id": 3382190,
              "postDate": "2025-12-26T19:48:56.583Z",
              "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> \nThe nnUNet is sure a strong codebase and it's highly likely that the final top solutions would be using this. I planned to add this <strong>no-new-Unet</strong> to <a href=\"https://github.com/innat/medic-ai/issues/30\" target=\"_blank\">medicai</a>. Though I still need to evaluate the complexities and challenges of it. However, I like to ask</p>\n<ul>\n<li>I learned that in official vesuvius website, you guys were using nnUNet. May I ask the specific strength and challenges you guys faced while working with it for your dataset?</li>\n<li>In MONAI, there is <code>full_res</code> version of this model, called <code>DynUNet</code>, any thoughts on this?</li>\n</ul>",
              "rawMarkdown": "@seanjohnsonsp \nThe nnUNet is sure a strong codebase and it's highly likely that the final top solutions would be using this. I planned to add this **no-new-Unet** to [medicai](https://github.com/innat/medic-ai/issues/30). Though I still need to evaluate the complexities and challenges of it. However, I like to ask\n\n- I learned that in official vesuvius website, you guys were using nnUNet. May I ask the specific strength and challenges you guys faced while working with it for your dataset?\n- In MONAI, there is `full_res` version of this model, called `DynUNet`, any thoughts on this?",
              "votes": 2
            },
            {
              "id": 3382790,
              "postDate": "2025-12-28T15:11:08.123Z",
              "content": "<p>The only issue we've faced with nnUNet with our data is that our \"canonical\" storage method for large scroll volumes is zarr, which is unsupported in nnUNet. this can be solved by chunking the data out for training, but chunking the data out for inference in this way means you've got to get creative at inference time for blending. </p>\n<p>It's also rather opinionated on formats/has lots of abstraction, but these are somewhat necessary evils if you want your framework to work well consistently, especially for users who just want to put data in and get predictions out, and dont know (or want to know) the ins and outs. </p>\n<p>the monai version is fine, but i think nnUNet without the rest of the framework is not really a great comparison. if you run some training runs in your own pipeline but instantiate the \"nnunet\" model,  and then compare something else to it and say \"x is better than nnUNet\", you're not making a fair comparison, in my opinion. </p>\n<p>The strength of nnUNet comes from the <em>entire framework</em> , the model architecture itself is not exactly groundbreaking (the resunet is mostly basic block D from the resnet bag of tricks paper from 7 years ago).  the training framework around this architecture is what makes nnUNet such a consistently strong performer. i dont mean this as a hit on the arch choice either, even very recent and complex transformer based models trained in exactly the same framework struggle to beat a basic resencunet in lots of tasks.</p>",
              "rawMarkdown": "The only issue we've faced with nnUNet with our data is that our \"canonical\" storage method for large scroll volumes is zarr, which is unsupported in nnUNet. this can be solved by chunking the data out for training, but chunking the data out for inference in this way means you've got to get creative at inference time for blending. \n\nIt's also rather opinionated on formats/has lots of abstraction, but these are somewhat necessary evils if you want your framework to work well consistently, especially for users who just want to put data in and get predictions out, and dont know (or want to know) the ins and outs. \n\nthe monai version is fine, but i think nnUNet without the rest of the framework is not really a great comparison. if you run some training runs in your own pipeline but instantiate the \"nnunet\" model,  and then compare something else to it and say \"x is better than nnUNet\", you're not making a fair comparison, in my opinion. \n\nThe strength of nnUNet comes from the _entire framework_ , the model architecture itself is not exactly groundbreaking (the resunet is mostly basic block D from the resnet bag of tricks paper from 7 years ago).  the training framework around this architecture is what makes nnUNet such a consistently strong performer. i dont mean this as a hit on the arch choice either, even very recent and complex transformer based models trained in exactly the same framework struggle to beat a basic resencunet in lots of tasks.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3381615,
      "postDate": "2025-12-25T03:28:48.593Z",
      "content": "<p>I would like to ask why the scores on the public leaderboard have not changed? After the dataset was updated, my local test scores changed drastically, but the leaderboard remains unchanged. This is very strange</p>",
      "rawMarkdown": "I would like to ask why the scores on the public leaderboard have not changed? After the dataset was updated, my local test scores changed drastically, but the leaderboard remains unchanged. This is very strange",
      "votes": 4
    },
    {
      "id": 3382487,
      "postDate": "2025-12-27T16:52:57.167Z",
      "content": "<p>Has anyone managed to improve lb score with the new labels? Really hoping we get more clarity on this soon as I’m not sure how it is possible that we get improved labels but all cv and lb goes down. It would only make sense to me if lb was still not right but I can’t imagine top teams are chasing what would likely become outdated lb?</p>",
      "rawMarkdown": "Has anyone managed to improve lb score with the new labels? Really hoping we get more clarity on this soon as I’m not sure how it is possible that we get improved labels but all cv and lb goes down. It would only make sense to me if lb was still not right but I can’t imagine top teams are chasing what would likely become outdated lb?",
      "votes": 1,
      "replies": [
        {
          "id": 3382493,
          "postDate": "2025-12-27T17:02:21.173Z",
          "content": "<p>I saw a few people get higher on leaderboard in past day or two, don't know if they submitted according to the old dataset or new. Like <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, he is always ranked higher than the last time you saw the lb.</p>",
          "rawMarkdown": "I saw a few people get higher on leaderboard in past day or two, don't know if they submitted according to the old dataset or new. Like @christofhenkel, he is always ranked higher than the last time you saw the lb.",
          "replies": [
            {
              "id": 3382495,
              "postDate": "2025-12-27T17:04:57.450Z",
              "content": "<p>I’m thinking it has to be old but I am not sure I would understand the point. Trying to make sure this shift hasn’t put us behind</p>",
              "rawMarkdown": "I’m thinking it has to be old but I am not sure I would understand the point. Trying to make sure this shift hasn’t put us behind"
            }
          ]
        },
        {
          "id": 3382846,
          "postDate": "2025-12-28T17:28:13.487Z",
          "content": "<p>I'm in the same boat - currently training on the new dataset and seeing similar\n  patterns. My preliminary results (mid-training) aren't trending up either.</p>\n<p>One thing I've noticed is what feels like drift from the actual ground truth -\n  almost like we're learning the annotator's interpretation rather than true\n  surface detection. The old dataset had thicker, more variable annotations,\n  while the new one is uniform 3vx thickness.</p>\n<p>It makes me wonder if the public LB is still based on the old annotation style,\n  which would explain why models trained on the new cleaner labels aren't scoring\n  as well. We might be optimizing for a different \"truth\" than what the LB is\n  measuring.</p>\n<p>Really hope the organizers can clarify whether the LB evaluation uses the old\n  or new ground truth. Otherwise we're kind of flying blind on which direction\n  to optimize for.</p>",
          "rawMarkdown": "I'm in the same boat - currently training on the new dataset and seeing similar\n  patterns. My preliminary results (mid-training) aren't trending up either.\n\n  One thing I've noticed is what feels like drift from the actual ground truth -\n  almost like we're learning the annotator's interpretation rather than true\n  surface detection. The old dataset had thicker, more variable annotations,\n  while the new one is uniform 3vx thickness.\n\n  It makes me wonder if the public LB is still based on the old annotation style,\n  which would explain why models trained on the new cleaner labels aren't scoring\n  as well. We might be optimizing for a different \"truth\" than what the LB is\n  measuring.\n\n  Really hope the organizers can clarify whether the LB evaluation uses the old\n  or new ground truth. Otherwise we're kind of flying blind on which direction\n  to optimize for.",
          "votes": 1,
          "replies": [
            {
              "id": 3382909,
              "postDate": "2025-12-28T19:40:03.413Z",
              "content": "<p>Yep on the same line of thinking. Hope it gets fixed or clarified soon!</p>",
              "rawMarkdown": "Yep on the same line of thinking. Hope it gets fixed or clarified soon!"
            }
          ]
        }
      ]
    },
    {
      "id": 3381900,
      "postDate": "2025-12-25T21:07:51.943Z",
      "content": "<p>Even though the dice score has reduced significantly, topology has improved a lot, the lb metric I calculated locally have improved despite fall in simple dice.</p>",
      "rawMarkdown": "Even though the dice score has reduced significantly, topology has improved a lot, the lb metric I calculated locally have improved despite fall in simple dice.",
      "votes": 1
    },
    {
      "id": 3381307,
      "postDate": "2025-12-24T08:13:18.847Z",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThank you for sharing. Did you already update from test dataset to new dataset?</p>",
      "rawMarkdown": "@sohier @giorgioangelotti \nThank you for sharing. Did you already update from test dataset to new dataset?",
      "votes": 1
    },
    {
      "id": 3381252,
      "postDate": "2025-12-24T05:18:15.813Z",
      "content": "<p>Has there been any update to the test set labels?</p>",
      "rawMarkdown": "Has there been any update to the test set labels?",
      "votes": 1,
      "replies": [
        {
          "id": 3381294,
          "postDate": "2025-12-24T07:12:18.290Z",
          "content": "<p>Yes, to some of them. But they were already cleaner than the train set labels in general.</p>",
          "rawMarkdown": "Yes, to some of them. But they were already cleaner than the train set labels in general.",
          "votes": 2,
          "replies": [
            {
              "id": 3381296,
              "postDate": "2025-12-24T07:25:42.383Z",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> Sorry, we mean labels in the hidden test set for evaluating LB. All of our score still the same as previous result.</p>",
              "rawMarkdown": "@giorgioangelotti Sorry, we mean labels in the hidden test set for evaluating LB. All of our score still the same as previous result.",
              "votes": 1
            },
            {
              "id": 3381297,
              "postDate": "2025-12-24T07:28:15.137Z",
              "content": "<p>same here, I also didn't see change in any of the scores</p>",
              "rawMarkdown": "same here, I also didn't see change in any of the scores",
              "votes": 1
            },
            {
              "id": 3381298,
              "postDate": "2025-12-24T07:29:06.147Z",
              "content": "<p>Also with regards to the thickness in the test set used to evaluate LB scores, is it similar to what we have in the updated training data? Simply put, the exact model that scored ~0.54 LB retrained on new data scores only ~0.45 now. Is it because test set has different thickness?</p>",
              "rawMarkdown": "Also with regards to the thickness in the test set used to evaluate LB scores, is it similar to what we have in the updated training data? Simply put, the exact model that scored ~0.54 LB retrained on new data scores only ~0.45 now. Is it because test set has different thickness?",
              "votes": 1
            },
            {
              "id": 3381301,
              "postDate": "2025-12-24T07:43:48.870Z",
              "content": "<p>Yes, the thickness is the same as the new training data. I agree that it is suspicious that scores did not change. Could it be that the labels in the test set were not overwritten during the update? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "rawMarkdown": "Yes, the thickness is the same as the new training data. I agree that it is suspicious that scores did not change. Could it be that the labels in the test set were not overwritten during the update? @sohier ",
              "votes": 1
            },
            {
              "id": 3381333,
              "postDate": "2025-12-24T09:05:19.347Z",
              "content": "<p><a href=\"https://www.kaggle.com/p4rallax\" target=\"_blank\">@p4rallax</a> We have observed a similar situation. Did your local validation score drop as well?\"</p>",
              "rawMarkdown": "@p4rallax We have observed a similar situation. Did your local validation score drop as well?\""
            },
            {
              "id": 3381335,
              "postDate": "2025-12-24T09:08:53.297Z",
              "content": "<p>Yep, ~0.78 val_dice -&gt; 0.55</p>",
              "rawMarkdown": "Yep, ~0.78 val_dice -> 0.55",
              "votes": 1
            },
            {
              "id": 3381338,
              "postDate": "2025-12-24T09:14:23.717Z",
              "content": "<p>This val dice thing happening with me too.</p>",
              "rawMarkdown": "This val dice thing happening with me too.",
              "votes": 1
            },
            {
              "id": 3381339,
              "postDate": "2025-12-24T09:14:24.550Z",
              "content": "<p>We are also👀</p>",
              "rawMarkdown": "We are also👀",
              "votes": 2
            },
            {
              "id": 3381387,
              "postDate": "2025-12-24T12:34:31.587Z",
              "content": "<p>😅me too…</p>",
              "rawMarkdown": "😅me too...",
              "votes": 2
            },
            {
              "id": 3381400,
              "postDate": "2025-12-24T13:25:28.950Z",
              "content": "<p>I also noticed that the test dataset for the public leaderboard hasn't been updated yet.\nIt seems that other people around me are experiencing the same issue, so I hope the host will address this.</p>\n<p>For reference, I have trained and submission on new data, so I will share it here.</p>\n<table>\n<thead>\n<tr>\n<th>Submission TS</th>\n<th>TrainDataset</th>\n<th>Local Validation(DiceScore)</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Sun Dec 21 2025 23:32:44 GMT+0900</td>\n<td><strong>Old Dataset</strong></td>\n<td><strong>0.7679</strong></td>\n<td><strong>0.522</strong></td>\n</tr>\n<tr>\n<td>Wed Dec 24 2025 16:17:51 GMT+0900</td>\n<td><strong>New Dataset</strong></td>\n<td><strong>0.6658</strong></td>\n<td><strong>0.380</strong></td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "I also noticed that the test dataset for the public leaderboard hasn't been updated yet.\nIt seems that other people around me are experiencing the same issue, so I hope the host will address this.\n\nFor reference, I have trained and submission on new data, so I will share it here.\nSubmission TS | TrainDataset | Local Validation(DiceScore) | Public LB |\n| --- | --- | --- | --- |\n|Sun Dec 21 2025 23:32:44 GMT+0900  | **Old Dataset**  | **0.7679**  | **0.522** |\n|Wed Dec 24 2025 16:17:51 GMT+0900   | **New Dataset**  | **0.6658**  | **0.380** |\n",
              "votes": 4
            },
            {
              "id": 3381738,
              "postDate": "2025-12-25T11:57:11.947Z",
              "content": "<p>Same issue here with almost identical training conditions:</p>\n<table>\n<thead>\n<tr>\n<th>train_dataset</th>\n<th>(local) val_dice</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>old</td>\n<td>0.8269</td>\n<td>0.526</td>\n</tr>\n<tr>\n<td>new</td>\n<td>0.7287</td>\n<td>0.394</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "Same issue here with almost identical training conditions:\n\n| train_dataset | (local) val_dice | Public LB\n| --- | --- | --- |\n| old | 0.8269 | 0.526\n| new | 0.7287 | 0.394\n"
            },
            {
              "id": 3383353,
              "postDate": "2025-12-30T00:34:53.580Z",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I identified the issue. I had updated the solution file and metric label dataset, but missed the additional step of manually updating the metric notebook to use the latest metric label dataset. The problem is now fixed and I have kicked off a fresh rescore. It should take roughly another day for the new scores to be fully populated.</p>",
              "rawMarkdown": "@giorgioangelotti I identified the issue. I had updated the solution file and metric label dataset, but missed the additional step of manually updating the metric notebook to use the latest metric label dataset. The problem is now fixed and I have kicked off a fresh rescore. It should take roughly another day for the new scores to be fully populated.",
              "votes": 6
            }
          ]
        }
      ]
    },
    {
      "id": 3380800,
      "postDate": "2025-12-23T04:53:20.483Z",
      "content": "<p>Thank you! And thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for finding and visualizing many of these issues</p>",
      "rawMarkdown": "Thank you! And thank you @hengck23 for finding and visualizing many of these issues",
      "votes": 1
    },
    {
      "id": 3380832,
      "postDate": "2025-12-23T06:39:50.907Z",
      "content": "<p>Really thanks host effort <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>! Found that model learns a better topology from the new mask. Beyond the competition itself, I’m really eager to see the entire approach work end-to-end: from a carbonized scroll to actual text, with its meaning successfully extracted.</p>",
      "rawMarkdown": "Really thanks host effort @giorgioangelotti! Found that model learns a better topology from the new mask. Beyond the competition itself, I’m really eager to see the entire approach work end-to-end: from a carbonized scroll to actual text, with its meaning successfully extracted.",
      "votes": 2
    },
    {
      "id": 3386679,
      "postDate": "2026-01-05T17:05:19.693Z",
      "content": "<p>Dear all,\nI see that the volumes now have the label 2 in the border. Is there a reason for it?\nIt seems that the volumes were padded with the value 0, could you share the code use for it? \nI'm assuming that volumes were padded with a fixed number of voxels in each edge. Perhaps we could simply crop back the borders, and train the models on those new cropped volumes.\nAlso, as far as I understand, we don't need to predict the unlabelled regions, therefore, they are not really important, although I believe the output label must have the same shape as the input volume.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Dear all,\nI see that the volumes now have the label 2 in the border. Is there a reason for it?\nIt seems that the volumes were padded with the value 0, could you share the code use for it? \nI'm assuming that volumes were padded with a fixed number of voxels in each edge. Perhaps we could simply crop back the borders, and train the models on those new cropped volumes.\nAlso, as far as I understand, we don't need to predict the unlabelled regions, therefore, they are not really important, although I believe the output label must have the same shape as the input volume.\n\nThanks!"
    },
    {
      "id": 3385107,
      "postDate": "2026-01-02T14:53:59.623Z",
      "content": "<p>This is, likely, a question to <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> - could you please share a link to full zarr volumes with updated labels (if available)?</p>",
      "rawMarkdown": "This is, likely, a question to @giorgioangelotti - could you please share a link to full zarr volumes with updated labels (if available)?",
      "replies": [
        {
          "id": 3386125,
          "postDate": "2026-01-04T16:48:25.203Z",
          "content": "<p>Hi,\nThey are not available!</p>",
          "rawMarkdown": "Hi,\nThey are not available!",
          "votes": 1
        }
      ]
    },
    {
      "id": 3381983,
      "postDate": "2025-12-26T06:02:22.270Z",
      "content": "<p>Will the commit records prior to the test data update be revoked?</p>",
      "rawMarkdown": "Will the commit records prior to the test data update be revoked?"
    },
    {
      "id": 3381181,
      "postDate": "2025-12-24T01:09:09.767Z",
      "content": "<p>Hello!  Just getting started.  It seems that for the training labels, all pixels are \"2\" for unlabeled, and the interesting labels - images that have 0, 1, and 2 values - are all under deprecated labels.  Demonstrative notebook is <a href=\"https://www.kaggle.com/code/iyersk/interesting-labels-deprecated\" target=\"_blank\">here</a>.  Am I missing something?  If this is true, then we have to use the deprecated labels if our models are to learn anything.</p>",
      "rawMarkdown": "Hello!  Just getting started.  It seems that for the training labels, all pixels are \"2\" for unlabeled, and the interesting labels - images that have 0, 1, and 2 values - are all under deprecated labels.  Demonstrative notebook is [here](https://www.kaggle.com/code/iyersk/interesting-labels-deprecated).  Am I missing something?  If this is true, then we have to use the deprecated labels if our models are to learn anything.",
      "replies": [
        {
          "id": 3381216,
          "postDate": "2025-12-24T03:33:44.940Z",
          "content": "<p>I too just looked at it, and maybe you're correct or we're too dumb.</p>",
          "rawMarkdown": "I too just looked at it, and maybe you're correct or we're too dumb.",
          "replies": [
            {
              "id": 3381232,
              "postDate": "2025-12-24T04:24:57.907Z",
              "content": "<p>Image.open(path) treats this 3D data as a 2D image.  Since it defaults to loading only the 0th layer (Z=0), np. unique reports that only the value 2 is present.</p>\n<p><code>def get_image_array(path):\n    img_array = tifffile.imread(path)\n    return img_array</code></p>",
              "rawMarkdown": "Image.open(path) treats this 3D data as a 2D image.  Since it defaults to loading only the 0th layer (Z=0), np. unique reports that only the value 2 is present.\n\n`def get_image_array(path):\n    img_array = tifffile.imread(path)\n    return img_array`"
            },
            {
              "id": 3381236,
              "postDate": "2025-12-24T04:31:22.190Z",
              "content": "<blockquote>\n  <p>we're too dumb.</p>\n</blockquote>\n<p>spot on, lol </p>",
              "rawMarkdown": ">  we're too dumb.\n\nspot on, lol \n\n"
            },
            {
              "id": 3381405,
              "postDate": "2025-12-24T13:41:28.570Z",
              "content": "<p>Got it, thank you!</p>",
              "rawMarkdown": "Got it, thank you!"
            },
            {
              "id": 3381407,
              "postDate": "2025-12-24T13:42:16.350Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3381408,
              "postDate": "2025-12-24T13:42:38.957Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3380834,
      "postDate": "2025-12-23T06:52:11.023Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3381743,
      "author_name": "Giorgio Angelotti",
      "author_url": "",
      "post_date": "2025-12-25T12:36:02.407000",
      "content": "<p>We’re investigating reports that the recent dataset update may not be fully reflected in the evaluation (e.g., unchanged scores after rescoring and lower public LB scores for models trained on the updated data). We’re auditing the update process and coordinating with Kaggle support as needed.\nGiven the holiday period, responses may be a bit slower than usual — thanks for your patience. In the meantime, please keep iterating using your own validation/CV on the updated training data; that’s the most reliable signal while we verify the leaderboard side. We’ll post an update as soon as we know more, and if a fix is needed we’ll ensure scoring is handled fairly for everyone (including re-scoring if required).</p>",
      "votes": 14,
      "replies": []
    },
    {
      "id": 3380779,
      "author_name": "tattaka",
      "author_url": "",
      "post_date": "2025-12-23T03:13:32.913000",
      "content": "<p>Before:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F8195ae7f8885e188ec57628c4bb26674%2F2025-12-23%2011.54.55.png?generation=1766458516230057&amp;alt=media\" alt=\"\"> \nAfter:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fb2116191a724afb70a1721ab5311632e%2F2025-12-23%2011.55.04.png?generation=1766458533613557&amp;alt=media\" alt=\"\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 3380824,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2025-12-23T05:58:09.120000",
          "content": "<p>Hmmm, will probably have to retrain a lot if all the labels have been reduced to thin sheets like this.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3380839,
              "author_name": "Giorgio Angelotti",
              "author_url": "",
              "post_date": "2025-12-23T07:04:48.637000",
              "content": "<p>Fact is that the previous labels were already thin sheets like these in multiple places. This round we tried to keep the thickness consistent across the dataset.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3386879,
      "author_name": "Manas Choudhary",
      "author_url": "",
      "post_date": "2026-01-06T02:16:25.923000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F0b7abe81673d70d4d80557c7a420bf75%2FScreenshot%202024-05-05%20222202.png?generation=1767665765991708&amp;alt=media\" alt=\"\"></p>\n<p>id = 1006462223</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3381497,
      "author_name": "HOUJING HUANG",
      "author_url": "",
      "post_date": "2025-12-24T18:34:37.020000",
      "content": "<p>Thank you very much for updating the dataset.</p>\n<p>I still feel like a Psyduck and want some help here.</p>\n<table>\n  <tbody><tr>\n    <td>\n      <p><b>Old Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Fc104695a25692638dbf9efdcb8d65ea1%2Fvesuvius_404970490_old.png?generation=1766600021359526&amp;alt=media\">\n    </td>\n    <td>\n      <p><b>New Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Ff1ac3c464ac074ff54ff5870d0668b80%2Fvesuvius_404970490_new.png?generation=1766600046174188&amp;alt=media\">\n    </td>\n  </tr>\n</tbody></table>\n<p>For each layer, suppose it has two surfaces.</p>\n<p>It seems only one surface is preferred to be the label.</p>\n<p>I wonder how is the surface selected during data annotation?</p>\n<p>I also wonder why not choose the \"center-face\", similar to the concept of centerline?</p>\n<p>Now with the surface thinner, being 3 voxels, if the surface is randomly chosen, it is more unpredictable compared to the previous thicker version of surface.</p>\n<p>For the same prediction by a previous model, my local validation weighted topo-3d score drops from 0.58 to 0.49, for old and new labels respectively.</p>\n<p>BTW, my model is based on <a href=\"https://github.com/MIC-DKFZ/nnUNet\" target=\"_blank\">https://github.com/MIC-DKFZ/nnUNet</a></p>\n<p>And my submission notebook is based on <a href=\"https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46\" target=\"_blank\">https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46</a></p>\n<p>Many thanks to the original authors for their kind sharing.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3381571,
          "author_name": "Sean Johnson_SP",
          "author_url": "",
          "post_date": "2025-12-24T23:05:29.763000",
          "content": "<p>The surface chosen is the “recto” or front surface , as this is the surface which contains the text (in most instances, there are some scrolls written on both sides but this is much less common) </p>\n<p>To the idea of using a centerline — the reason this is not chosen is it is not infrequent for the recto and verso to be split , sometimes for a rather long period of time and at reasonably large distances, and in these cases the concept of a centerline is somewhat ambiguous. </p>\n<p>So the choice is more of a pragmatic one: first , because that’s where the text we seek lies, and second because it’s just more practical. There is no real “inner” part of a papyrus sheet ; its horizontal strips on the front and vertical ones on the back, glued together. </p>\n<p>The labels should all lie on this recto surface. </p>",
          "votes": 3,
          "replies": [
            {
              "id": 3381574,
              "author_name": "HOUJING HUANG",
              "author_url": "",
              "post_date": "2025-12-24T23:22:32.963000",
              "content": "<p>Thank you Sean.\nThis is very informative.\nI hope the model can capture the difference between the recto and verso.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3381575,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2025-12-24T23:41:28.717000",
              "content": "<p>In most cases I've found it capable. the model does not really struggle to detect the recto side. the primary struggle it faces in when a sheet is frayed or torn, and when the sheets are very tightly packed. </p>\n<p>here is an example of a model i have trained with inference on the entirety of Scroll 1 (PHerc Paris 4) :</p>\n<ul>\n<li><a href=\"https://dl.ash2txt.org/community-uploads/bruniss/scrolls/s1/surfaces/s1_059_ome.zarr/\" target=\"_blank\">https://dl.ash2txt.org/community-uploads/bruniss/scrolls/s1/surfaces/s1_059_ome.zarr/</a> </li>\n</ul>\n<p>the primary way to help the model get better at recto vs verso is heavy spatial augmentations (flips, rotations, mirrors, etc) </p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3381577,
          "author_name": "Sean Johnson_SP",
          "author_url": "",
          "post_date": "2025-12-24T23:56:26.800000",
          "content": "<p>Also if you use that notebook I would highly recommend against using the morphological closing that is used by it , it absolutely will cause components to merge. Even the open is somewhat risky as any 1vx thick line will be removed and any 2-3vx line is at risk of becoming multiple components </p>",
          "votes": 2,
          "replies": [
            {
              "id": 3381585,
              "author_name": "HOUJING HUANG",
              "author_url": "",
              "post_date": "2025-12-25T01:08:28.027000",
              "content": "<p>I agree with you on morphological closing, and I think opening is more preferable.</p>\n<p>I just looked at my trained results from the new labels, which is now much better and requires less post-processing efforts, since the predicted sheets are already quite far away from each other.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381586,
              "author_name": "HOUJING HUANG",
              "author_url": "",
              "post_date": "2025-12-25T01:16:50.317000",
              "content": "<p>The nnUNet baseline has quite good performance out of the box.</p>\n<p>For the submission notebook, I mostly borrowed the logic about test.csv, visualization before submission, and the zip file format required. This part I tried to find an official one, yet none exists….</p>\n<p>For post-processing, the suitable operators are also heavily dependent on the first stage predictions, and users may need some trial-and-error tweaks for their cases.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381934,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2025-12-26T01:34:28.180000",
              "content": "<p>Yes nnUNetv2 baseline nnUNetTrainer is an excellent starting point. I would not be surprised if DKFZ single submission is just the regular ResEncM preset with the DA5 trainer and no other modifications other than to max epochs. I'd expect the baseline nnUNet to come in around 53-55 lb score on prev data, unsure of what it would be since the update as i have not yet trained a model with it. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3382190,
              "author_name": "Innat",
              "author_url": "",
              "post_date": "2025-12-26T19:48:56.583000",
              "content": "<p><a href=\"https://www.kaggle.com/seanjohnsonsp\" target=\"_blank\">@seanjohnsonsp</a> \nThe nnUNet is sure a strong codebase and it's highly likely that the final top solutions would be using this. I planned to add this <strong>no-new-Unet</strong> to <a href=\"https://github.com/innat/medic-ai/issues/30\" target=\"_blank\">medicai</a>. Though I still need to evaluate the complexities and challenges of it. However, I like to ask</p>\n<ul>\n<li>I learned that in official vesuvius website, you guys were using nnUNet. May I ask the specific strength and challenges you guys faced while working with it for your dataset?</li>\n<li>In MONAI, there is <code>full_res</code> version of this model, called <code>DynUNet</code>, any thoughts on this?</li>\n</ul>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3382790,
              "author_name": "Sean Johnson_SP",
              "author_url": "",
              "post_date": "2025-12-28T15:11:08.123000",
              "content": "<p>The only issue we've faced with nnUNet with our data is that our \"canonical\" storage method for large scroll volumes is zarr, which is unsupported in nnUNet. this can be solved by chunking the data out for training, but chunking the data out for inference in this way means you've got to get creative at inference time for blending. </p>\n<p>It's also rather opinionated on formats/has lots of abstraction, but these are somewhat necessary evils if you want your framework to work well consistently, especially for users who just want to put data in and get predictions out, and dont know (or want to know) the ins and outs. </p>\n<p>the monai version is fine, but i think nnUNet without the rest of the framework is not really a great comparison. if you run some training runs in your own pipeline but instantiate the \"nnunet\" model,  and then compare something else to it and say \"x is better than nnUNet\", you're not making a fair comparison, in my opinion. </p>\n<p>The strength of nnUNet comes from the <em>entire framework</em> , the model architecture itself is not exactly groundbreaking (the resunet is mostly basic block D from the resnet bag of tricks paper from 7 years ago).  the training framework around this architecture is what makes nnUNet such a consistently strong performer. i dont mean this as a hit on the arch choice either, even very recent and complex transformer based models trained in exactly the same framework struggle to beat a basic resencunet in lots of tasks.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3381615,
      "author_name": "Wang Zhiyao (王致尧)",
      "author_url": "",
      "post_date": "2025-12-25T03:28:48.593000",
      "content": "<p>I would like to ask why the scores on the public leaderboard have not changed? After the dataset was updated, my local test scores changed drastically, but the leaderboard remains unchanged. This is very strange</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3382487,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2025-12-27T16:52:57.167000",
      "content": "<p>Has anyone managed to improve lb score with the new labels? Really hoping we get more clarity on this soon as I’m not sure how it is possible that we get improved labels but all cv and lb goes down. It would only make sense to me if lb was still not right but I can’t imagine top teams are chasing what would likely become outdated lb?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3382493,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2025-12-27T17:02:21.173000",
          "content": "<p>I saw a few people get higher on leaderboard in past day or two, don't know if they submitted according to the old dataset or new. Like <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>, he is always ranked higher than the last time you saw the lb.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3382495,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2025-12-27T17:04:57.450000",
              "content": "<p>I’m thinking it has to be old but I am not sure I would understand the point. Trying to make sure this shift hasn’t put us behind</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3382846,
          "author_name": "TWEAK",
          "author_url": "",
          "post_date": "2025-12-28T17:28:13.487000",
          "content": "<p>I'm in the same boat - currently training on the new dataset and seeing similar\n  patterns. My preliminary results (mid-training) aren't trending up either.</p>\n<p>One thing I've noticed is what feels like drift from the actual ground truth -\n  almost like we're learning the annotator's interpretation rather than true\n  surface detection. The old dataset had thicker, more variable annotations,\n  while the new one is uniform 3vx thickness.</p>\n<p>It makes me wonder if the public LB is still based on the old annotation style,\n  which would explain why models trained on the new cleaner labels aren't scoring\n  as well. We might be optimizing for a different \"truth\" than what the LB is\n  measuring.</p>\n<p>Really hope the organizers can clarify whether the LB evaluation uses the old\n  or new ground truth. Otherwise we're kind of flying blind on which direction\n  to optimize for.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3382909,
              "author_name": "Cody_Null",
              "author_url": "",
              "post_date": "2025-12-28T19:40:03.413000",
              "content": "<p>Yep on the same line of thinking. Hope it gets fixed or clarified soon!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3381900,
      "author_name": "Manas Choudhary",
      "author_url": "",
      "post_date": "2025-12-25T21:07:51.943000",
      "content": "<p>Even though the dice score has reduced significantly, topology has improved a lot, the lb metric I calculated locally have improved despite fall in simple dice.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3381307,
      "author_name": "tereka",
      "author_url": "",
      "post_date": "2025-12-24T08:13:18.847000",
      "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> \nThank you for sharing. Did you already update from test dataset to new dataset?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3381252,
      "author_name": "Satwik",
      "author_url": "",
      "post_date": "2025-12-24T05:18:15.813000",
      "content": "<p>Has there been any update to the test set labels?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3381294,
          "author_name": "Giorgio Angelotti",
          "author_url": "",
          "post_date": "2025-12-24T07:12:18.290000",
          "content": "<p>Yes, to some of them. But they were already cleaner than the train set labels in general.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3381296,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-12-24T07:25:42.383000",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> Sorry, we mean labels in the hidden test set for evaluating LB. All of our score still the same as previous result.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381297,
              "author_name": "Manas Choudhary",
              "author_url": "",
              "post_date": "2025-12-24T07:28:15.137000",
              "content": "<p>same here, I also didn't see change in any of the scores</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381298,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-12-24T07:29:06.147000",
              "content": "<p>Also with regards to the thickness in the test set used to evaluate LB scores, is it similar to what we have in the updated training data? Simply put, the exact model that scored ~0.54 LB retrained on new data scores only ~0.45 now. Is it because test set has different thickness?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381301,
              "author_name": "Giorgio Angelotti",
              "author_url": "",
              "post_date": "2025-12-24T07:43:48.870000",
              "content": "<p>Yes, the thickness is the same as the new training data. I agree that it is suspicious that scores did not change. Could it be that the labels in the test set were not overwritten during the update? <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381333,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2025-12-24T09:05:19.347000",
              "content": "<p><a href=\"https://www.kaggle.com/p4rallax\" target=\"_blank\">@p4rallax</a> We have observed a similar situation. Did your local validation score drop as well?\"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381335,
              "author_name": "Satwik",
              "author_url": "",
              "post_date": "2025-12-24T09:08:53.297000",
              "content": "<p>Yep, ~0.78 val_dice -&gt; 0.55</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381338,
              "author_name": "Manas Choudhary",
              "author_url": "",
              "post_date": "2025-12-24T09:14:23.717000",
              "content": "<p>This val dice thing happening with me too.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3381339,
              "author_name": "tingyi",
              "author_url": "",
              "post_date": "2025-12-24T09:14:24.550000",
              "content": "<p>We are also👀</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3381387,
              "author_name": "Boredom",
              "author_url": "",
              "post_date": "2025-12-24T12:34:31.587000",
              "content": "<p>😅me too…</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3381400,
              "author_name": "yukiZ",
              "author_url": "",
              "post_date": "2025-12-24T13:25:28.950000",
              "content": "<p>I also noticed that the test dataset for the public leaderboard hasn't been updated yet.\nIt seems that other people around me are experiencing the same issue, so I hope the host will address this.</p>\n<p>For reference, I have trained and submission on new data, so I will share it here.</p>\n<table>\n<thead>\n<tr>\n<th>Submission TS</th>\n<th>TrainDataset</th>\n<th>Local Validation(DiceScore)</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Sun Dec 21 2025 23:32:44 GMT+0900</td>\n<td><strong>Old Dataset</strong></td>\n<td><strong>0.7679</strong></td>\n<td><strong>0.522</strong></td>\n</tr>\n<tr>\n<td>Wed Dec 24 2025 16:17:51 GMT+0900</td>\n<td><strong>New Dataset</strong></td>\n<td><strong>0.6658</strong></td>\n<td><strong>0.380</strong></td>\n</tr>\n</tbody>\n</table>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3381738,
              "author_name": "rootenter",
              "author_url": "",
              "post_date": "2025-12-25T11:57:11.947000",
              "content": "<p>Same issue here with almost identical training conditions:</p>\n<table>\n<thead>\n<tr>\n<th>train_dataset</th>\n<th>(local) val_dice</th>\n<th>Public LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>old</td>\n<td>0.8269</td>\n<td>0.526</td>\n</tr>\n<tr>\n<td>new</td>\n<td>0.7287</td>\n<td>0.394</td>\n</tr>\n</tbody>\n</table>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3383353,
              "author_name": "Sohier Dane",
              "author_url": "",
              "post_date": "2025-12-30T00:34:53.580000",
              "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> I identified the issue. I had updated the solution file and metric label dataset, but missed the additional step of manually updating the metric notebook to use the latest metric label dataset. The problem is now fixed and I have kicked off a fresh rescore. It should take roughly another day for the new scores to be fully populated.</p>",
              "votes": 6,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3380800,
      "author_name": "Cody_Null",
      "author_url": "",
      "post_date": "2025-12-23T04:53:20.483000",
      "content": "<p>Thank you! And thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for finding and visualizing many of these issues</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3380832,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-12-23T06:39:50.907000",
      "content": "<p>Really thanks host effort <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>! Found that model learns a better topology from the new mask. Beyond the competition itself, I’m really eager to see the entire approach work end-to-end: from a carbonized scroll to actual text, with its meaning successfully extracted.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3386679,
      "author_name": "Andre Filipe Ferreira",
      "author_url": "",
      "post_date": "2026-01-05T17:05:19.693000",
      "content": "<p>Dear all,\nI see that the volumes now have the label 2 in the border. Is there a reason for it?\nIt seems that the volumes were padded with the value 0, could you share the code use for it? \nI'm assuming that volumes were padded with a fixed number of voxels in each edge. Perhaps we could simply crop back the borders, and train the models on those new cropped volumes.\nAlso, as far as I understand, we don't need to predict the unlabelled regions, therefore, they are not really important, although I believe the output label must have the same shape as the input volume.</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3385107,
      "author_name": "Victor Shlepov",
      "author_url": "",
      "post_date": "2026-01-02T14:53:59.623000",
      "content": "<p>This is, likely, a question to <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> - could you please share a link to full zarr volumes with updated labels (if available)?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3386125,
          "author_name": "Giorgio Angelotti",
          "author_url": "",
          "post_date": "2026-01-04T16:48:25.203000",
          "content": "<p>Hi,\nThey are not available!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3381983,
      "author_name": "GG Ayo (AyoGG)",
      "author_url": "",
      "post_date": "2025-12-26T06:02:22.270000",
      "content": "<p>Will the commit records prior to the test data update be revoked?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3381181,
      "author_name": "RishabhIyer",
      "author_url": "",
      "post_date": "2025-12-24T01:09:09.767000",
      "content": "<p>Hello!  Just getting started.  It seems that for the training labels, all pixels are \"2\" for unlabeled, and the interesting labels - images that have 0, 1, and 2 values - are all under deprecated labels.  Demonstrative notebook is <a href=\"https://www.kaggle.com/code/iyersk/interesting-labels-deprecated\" target=\"_blank\">here</a>.  Am I missing something?  If this is true, then we have to use the deprecated labels if our models are to learn anything.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3381216,
          "author_name": "Manas Choudhary",
          "author_url": "",
          "post_date": "2025-12-24T03:33:44.940000",
          "content": "<p>I too just looked at it, and maybe you're correct or we're too dumb.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3381232,
              "author_name": "Boredom",
              "author_url": "",
              "post_date": "2025-12-24T04:24:57.907000",
              "content": "<p>Image.open(path) treats this 3D data as a 2D image.  Since it defaults to loading only the 0th layer (Z=0), np. unique reports that only the value 2 is present.</p>\n<p><code>def get_image_array(path):\n    img_array = tifffile.imread(path)\n    return img_array</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381236,
              "author_name": "Manas Choudhary",
              "author_url": "",
              "post_date": "2025-12-24T04:31:22.190000",
              "content": "<blockquote>\n  <p>we're too dumb.</p>\n</blockquote>\n<p>spot on, lol </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381405,
              "author_name": "RishabhIyer",
              "author_url": "",
              "post_date": "2025-12-24T13:41:28.570000",
              "content": "<p>Got it, thank you!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381407,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-12-24T13:42:16.350000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3381408,
              "author_name": "",
              "author_url": "",
              "post_date": "2025-12-24T13:42:38.957000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3380834,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-12-23T06:52:11.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3380737": "We're in the process of posting an updated copy of the dataset with improved label quality. The update will cover:\n\n- Incorrectly merged components were split. In a few cases this was impractical and the images have been moved to a new `deprecated` folder. We do not recommend continuing to use those files but wanted to ensure equal access in case someone does find a use for them.\n- Very small connected components which were incorrectly labeled as foreground have been removed.\n- Holes and voids have been manually fixed where required, programmatically addressed with binary hole filling or other morphological operations. A limited number of samples were also removed for these defects in cases where fixing would take an extended period of time.\n- Label thickness should now have a more uniform 3vx thick labels. The dataset had a 3vx padding in all dimensions that was supposed to be an ignore label (value = 2), but it was erroneously labeled as background (value = 0). This padding was not reported in the original dataset description. We are ensuring now that it is correctly labeled as 2.\n\nI will be rescoring all submissions shortly as a few test set images are now ignored for scoring purposes. Note that this may take a day or so to complete due to the long metric runtime.\n\nPlease note that we may need to make additional updates after the holidays, but this update is expected to cover most of the fixes.\n\nEdit: The rescore is complete. \n\nEdit 2: The initial rescore [did not properly reflect the update](https://www.kaggle.com/competitions/vesuvius-challenge-surface-detection/discussion/664184#3383353). I have resolved the issue and started a fresh rescore. Apologies for the confusion.\n\nEdit 3: The second rescore is now complete.",
    "3381743": "We’re investigating reports that the recent dataset update may not be fully reflected in the evaluation (e.g., unchanged scores after rescoring and lower public LB scores for models trained on the updated data). We’re auditing the update process and coordinating with Kaggle support as needed.\nGiven the holiday period, responses may be a bit slower than usual — thanks for your patience. In the meantime, please keep iterating using your own validation/CV on the updated training data; that’s the most reliable signal while we verify the leaderboard side. We’ll post an update as soon as we know more, and if a fix is needed we’ll ensure scoring is handled fairly for everyone (including re-scoring if required).",
    "3380779": "Before:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2F8195ae7f8885e188ec57628c4bb26674%2F2025-12-23%2011.54.55.png?generation=1766458516230057&alt=media) \nAfter:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1745801%2Fb2116191a724afb70a1721ab5311632e%2F2025-12-23%2011.55.04.png?generation=1766458533613557&alt=media)",
    "3386879": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F26230365%2F0b7abe81673d70d4d80557c7a420bf75%2FScreenshot%202024-05-05%20222202.png?generation=1767665765991708&alt=media)\n\nid = 1006462223",
    "3381497": "Thank you very much for updating the dataset.\n\nI still feel like a Psyduck and want some help here.\n\n<table>\n  <tr>\n    <td style=\"text-align: center;\">\n      <p><b>Old Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Fc104695a25692638dbf9efdcb8d65ea1%2Fvesuvius_404970490_old.png?generation=1766600021359526&alt=media\" width=\"500\">\n    </td>\n    <td style=\"text-align: center;\">\n      <p><b>New Label</b></p>\n      <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F17269532%2Ff1ac3c464ac074ff54ff5870d0668b80%2Fvesuvius_404970490_new.png?generation=1766600046174188&alt=media\" width=\"500\">\n    </td>\n  </tr>\n</table>\n\nFor each layer, suppose it has two surfaces.\n\nIt seems only one surface is preferred to be the label.\n\nI wonder how is the surface selected during data annotation?\n\nI also wonder why not choose the \"center-face\", similar to the concept of centerline?\n\nNow with the surface thinner, being 3 voxels, if the surface is randomly chosen, it is more unpredictable compared to the previous thicker version of surface.\n\nFor the same prediction by a previous model, my local validation weighted topo-3d score drops from 0.58 to 0.49, for old and new labels respectively.\n\nBTW, my model is based on https://github.com/MIC-DKFZ/nnUNet\n\nAnd my submission notebook is based on https://www.kaggle.com/code/manish756/unet2-5segmentation-model-0-46\n\nMany thanks to the original authors for their kind sharing.",
    "3381615": "I would like to ask why the scores on the public leaderboard have not changed? After the dataset was updated, my local test scores changed drastically, but the leaderboard remains unchanged. This is very strange",
    "3382487": "Has anyone managed to improve lb score with the new labels? Really hoping we get more clarity on this soon as I’m not sure how it is possible that we get improved labels but all cv and lb goes down. It would only make sense to me if lb was still not right but I can’t imagine top teams are chasing what would likely become outdated lb?",
    "3381900": "Even though the dice score has reduced significantly, topology has improved a lot, the lb metric I calculated locally have improved despite fall in simple dice.",
    "3381307": "@sohier @giorgioangelotti \nThank you for sharing. Did you already update from test dataset to new dataset?",
    "3381252": "Has there been any update to the test set labels?",
    "3380800": "Thank you! And thank you @hengck23 for finding and visualizing many of these issues",
    "3380832": "Really thanks host effort @giorgioangelotti! Found that model learns a better topology from the new mask. Beyond the competition itself, I’m really eager to see the entire approach work end-to-end: from a carbonized scroll to actual text, with its meaning successfully extracted.",
    "3386679": "Dear all,\nI see that the volumes now have the label 2 in the border. Is there a reason for it?\nIt seems that the volumes were padded with the value 0, could you share the code use for it? \nI'm assuming that volumes were padded with a fixed number of voxels in each edge. Perhaps we could simply crop back the borders, and train the models on those new cropped volumes.\nAlso, as far as I understand, we don't need to predict the unlabelled regions, therefore, they are not really important, although I believe the output label must have the same shape as the input volume.\n\nThanks!",
    "3385107": "This is, likely, a question to @giorgioangelotti - could you please share a link to full zarr volumes with updated labels (if available)?",
    "3381983": "Will the commit records prior to the test data update be revoked?",
    "3381181": "Hello!  Just getting started.  It seems that for the training labels, all pixels are \"2\" for unlabeled, and the interesting labels - images that have 0, 1, and 2 values - are all under deprecated labels.  Demonstrative notebook is [here](https://www.kaggle.com/code/iyersk/interesting-labels-deprecated).  Am I missing something?  If this is true, then we have to use the deprecated labels if our models are to learn anything.",
    "3380834": ""
  }
}