{
  "id": 651700,
  "title": "Few questions to host",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/651700",
  "author_name": "",
  "post_date": "2025-12-04T18:45:38.765500900Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, nice rumble we’ve got going here! I was checking the dataset and have a couple of questions, would appreciate your comments:</p>\n<ol>\n<li><p>All training images seem to have 3-pixel margins (zero-valued pixels) in the xy-plane. Could you confirm if this is also the case for the test set?</p></li>\n<li><p>I’ve seen the discussion about 3D patch sizes, but I haven’t seen a definitive answer yet - are all test volumes expected to be 256, 320, or 384?</p></li>\n<li><p>As a side note, if I understand correctly, the training volumes are sub-volumes of larger CT scans. If so, I would recommend providing the full (or masked) array next time - for example as a Zarr file or another suitable format. Otherwise, border-alignment issues will remain.</p></li>\n<li><p>CT volumes from scanners or PACS are almost always stored as int16 or uint16, with meaningful Hounsfield units (HU). Is there a particular reason they were cast to uint8 here?</p></li>\n</ol>",
  "messages": [
    {
      "id": "3362485",
      "postDate": "12/04/2025 18:45:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a>, nice rumble we’ve got going here! I was checking the dataset and have a couple of questions, would appreciate your comments:</p>\n<ol>\n<li><p>All training images seem to have 3-pixel margins (zero-valued pixels) in the xy-plane. Could you confirm if this is also the case for the test set?</p></li>\n<li><p>I’ve seen the discussion about 3D patch sizes, but I haven’t seen a definitive answer yet - are all test volumes expected to be 256, 320, or 384?</p></li>\n<li><p>As a side note, if I understand correctly, the training volumes are sub-volumes of larger CT scans. If so, I would recommend providing the full (or masked) array next time - for example as a Zarr file or another suitable format. Otherwise, border-alignment issues will remain.</p></li>\n<li><p>CT volumes from scanners or PACS are almost always stored as int16 or uint16, with meaningful Hounsfield units (HU). Is there a particular reason they were cast to uint8 here?</p></li>\n</ol>",
      "rawMarkdown": "Hi @giorgioangelotti, nice rumble we’ve got going here! I was checking the dataset and have a couple of questions, would appreciate your comments:\n\n1. All training images seem to have 3-pixel margins (zero-valued pixels) in the xy-plane. Could you confirm if this is also the case for the test set?\n\n2. I’ve seen the discussion about 3D patch sizes, but I haven’t seen a definitive answer yet - are all test volumes expected to be 256, 320, or 384?\n\n3. As a side note, if I understand correctly, the training volumes are sub-volumes of larger CT scans. If so, I would recommend providing the full (or masked) array next time - for example as a Zarr file or another suitable format. Otherwise, border-alignment issues will remain.\n\n4. CT volumes from scanners or PACS are almost always stored as int16 or uint16, with meaningful Hounsfield units (HU). Is there a particular reason they were cast to uint8 here?",
      "votes": null
    },
    {
      "id": "3362970",
      "postDate": "12/05/2025 17:18:48",
      "content": "<p>Thank you for the questions and welcome to the competition!</p>\n<ol>\n<li>Thanks for noticing this. Indeed, that padding was supposed to be an ignore-label. We will investigate with our team of annotators whether this happens uniformly both in the training and the test set, and consider whether a data update should be necessary.</li>\n<li>Yes.</li>\n<li>Full CT scans are available as unlabeled data as OME-Zarr on the repositories of the Vesuvius Challenge. The repositories don't include all the volumes used to create the training and the test set. I would recommend starting from these <a href=\"https://data.aws.ash2txt.org/samples\" target=\"_blank\">https://data.aws.ash2txt.org/samples</a> . But it could be worthy also having a look at those <a href=\"https://dl.ash2txt.org/full-scrolls/\" target=\"_blank\">https://dl.ash2txt.org/full-scrolls/</a></li>\n<li>The Hounsfield scale doesn't necessarily make sense here because everything is carbon based and live approximately in the same range of values. We are talking about carbonized scrolls of papyrus written with carbonized Roman ink (extremely thin, and almost entirely carbon based). In our new scans (the one performed at the ESRF synchrotron) we are picking the windowing in order to maximize what seems to be the range of useful available information. In the old scans (acquired at Diamond Light Source) the choice of windowing has been less accurate because at the time it was not well known whether useful information to detect the ink could be hidden in the \"noise\" or very small values. We are using uint8 instead of uint16 because this, at least so far, has not been proved to be detrimental for the ink detection task, which requires to detect thin features practically invisible to the naked eye and should be thus be more sensible to a change in precision than the task tackled in this active competition!</li>\n</ol>\n<p>Hope this helps! </p>",
      "rawMarkdown": "Thank you for the questions and welcome to the competition!\n\n1. Thanks for noticing this. Indeed, that padding was supposed to be an ignore-label. We will investigate with our team of annotators whether this happens uniformly both in the training and the test set, and consider whether a data update should be necessary.\n2. Yes.\n3. Full CT scans are available as unlabeled data as OME-Zarr on the repositories of the Vesuvius Challenge. The repositories don't include all the volumes used to create the training and the test set. I would recommend starting from these https://data.aws.ash2txt.org/samples . But it could be worthy also having a look at those https://dl.ash2txt.org/full-scrolls/\n4. The Hounsfield scale doesn't necessarily make sense here because everything is carbon based and live approximately in the same range of values. We are talking about carbonized scrolls of papyrus written with carbonized Roman ink (extremely thin, and almost entirely carbon based). In our new scans (the one performed at the ESRF synchrotron) we are picking the windowing in order to maximize what seems to be the range of useful available information. In the old scans (acquired at Diamond Light Source) the choice of windowing has been less accurate because at the time it was not well known whether useful information to detect the ink could be hidden in the \"noise\" or very small values. We are using uint8 instead of uint16 because this, at least so far, has not been proved to be detrimental for the ink detection task, which requires to detect thin features practically invisible to the naked eye and should be thus be more sensible to a change in precision than the task tackled in this active competition!\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "3364400",
      "postDate": "12/06/2025 13:14:21",
      "content": "<p>Thank you for the detailed answer. Jus a couple of follow-up questions:</p>\n<p>Say, I use this one for the pre-training: <a href=\"https://dl.ash2txt.org/full-scrolls/\" target=\"_blank\">https://dl.ash2txt.org/full-scrolls/</a></p>\n<p>What is a better start: volume_grids,  volumes, volumes_masked or volumes_zarr_standardized?</p>\n<p>Volumes are uint16 - maybe you could share a pipeline how the competition data (images) was processed from one of those sources so we could replicate it?</p>",
      "rawMarkdown": "Thank you for the detailed answer. Jus a couple of follow-up questions:\n\nSay, I use this one for the pre-training: https://dl.ash2txt.org/full-scrolls/\n\nWhat is a better start: volume_grids,  volumes, volumes_masked or volumes_zarr_standardized?\n\nVolumes are uint16 - maybe you could share a pipeline how the competition data (images) was processed from one of those sources so we could replicate it?",
      "votes": null
    },
    {
      "id": "3365718",
      "postDate": "12/07/2025 10:17:53",
      "content": "<p>I prefer using the OME-zarrs because they are easier to access. However, the \"standardization\" procedure used to create those volumes is not \"orthodox\" (it's overly compressing the information imho). This is why I would recommend using these volumes instead: <a href=\"https://data.aws.ash2txt.org/samples/\" target=\"_blank\">https://data.aws.ash2txt.org/samples/</a></p>",
      "rawMarkdown": "I prefer using the OME-zarrs because they are easier to access. However, the \"standardization\" procedure used to create those volumes is not \"orthodox\" (it's overly compressing the information imho). This is why I would recommend using these volumes instead: https://data.aws.ash2txt.org/samples/",
      "votes": null
    },
    {
      "id": "3392907",
      "postDate": "01/17/2026 18:46:34",
      "content": "<p>What was the result of the investigation in 1.? Is this border region uniform in the test region? Is it ignored for scoring?</p>",
      "rawMarkdown": "What was the result of the investigation in 1.? Is this border region uniform in the test region? Is it ignored for scoring?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3362970,
      "author_name": "giorgioangelotti",
      "author_url": "",
      "post_date": "12/05/2025 17:18:48",
      "content": "<p>Thank you for the questions and welcome to the competition!</p>\n<ol>\n<li>Thanks for noticing this. Indeed, that padding was supposed to be an ignore-label. We will investigate with our team of annotators whether this happens uniformly both in the training and the test set, and consider whether a data update should be necessary.</li>\n<li>Yes.</li>\n<li>Full CT scans are available as unlabeled data as OME-Zarr on the repositories of the Vesuvius Challenge. The repositories don't include all the volumes used to create the training and the test set. I would recommend starting from these <a href=\"https://data.aws.ash2txt.org/samples\" target=\"_blank\">https://data.aws.ash2txt.org/samples</a> . But it could be worthy also having a look at those <a href=\"https://dl.ash2txt.org/full-scrolls/\" target=\"_blank\">https://dl.ash2txt.org/full-scrolls/</a></li>\n<li>The Hounsfield scale doesn't necessarily make sense here because everything is carbon based and live approximately in the same range of values. We are talking about carbonized scrolls of papyrus written with carbonized Roman ink (extremely thin, and almost entirely carbon based). In our new scans (the one performed at the ESRF synchrotron) we are picking the windowing in order to maximize what seems to be the range of useful available information. In the old scans (acquired at Diamond Light Source) the choice of windowing has been less accurate because at the time it was not well known whether useful information to detect the ink could be hidden in the \"noise\" or very small values. We are using uint8 instead of uint16 because this, at least so far, has not been proved to be detrimental for the ink detection task, which requires to detect thin features practically invisible to the naked eye and should be thus be more sensible to a change in precision than the task tackled in this active competition!</li>\n</ol>\n<p>Hope this helps! </p>",
      "votes": null,
      "replies": [
        {
          "id": 3364400,
          "author_name": "victorshlepov",
          "author_url": "",
          "post_date": "12/06/2025 13:14:21",
          "content": "<p>Thank you for the detailed answer. Jus a couple of follow-up questions:</p>\n<p>Say, I use this one for the pre-training: <a href=\"https://dl.ash2txt.org/full-scrolls/\" target=\"_blank\">https://dl.ash2txt.org/full-scrolls/</a></p>\n<p>What is a better start: volume_grids,  volumes, volumes_masked or volumes_zarr_standardized?</p>\n<p>Volumes are uint16 - maybe you could share a pipeline how the competition data (images) was processed from one of those sources so we could replicate it?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3365718,
              "author_name": "giorgioangelotti",
              "author_url": "",
              "post_date": "12/07/2025 10:17:53",
              "content": "<p>I prefer using the OME-zarrs because they are easier to access. However, the \"standardization\" procedure used to create those volumes is not \"orthodox\" (it's overly compressing the information imho). This is why I would recommend using these volumes instead: <a href=\"https://data.aws.ash2txt.org/samples/\" target=\"_blank\">https://data.aws.ash2txt.org/samples/</a></p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 3392907,
          "author_name": "tims457",
          "author_url": "",
          "post_date": "01/17/2026 18:46:34",
          "content": "<p>What was the result of the investigation in 1.? Is this border region uniform in the test region? Is it ignored for scoring?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3362485": "Hi @giorgioangelotti, nice rumble we’ve got going here! I was checking the dataset and have a couple of questions, would appreciate your comments:\n\n1. All training images seem to have 3-pixel margins (zero-valued pixels) in the xy-plane. Could you confirm if this is also the case for the test set?\n\n2. I’ve seen the discussion about 3D patch sizes, but I haven’t seen a definitive answer yet - are all test volumes expected to be 256, 320, or 384?\n\n3. As a side note, if I understand correctly, the training volumes are sub-volumes of larger CT scans. If so, I would recommend providing the full (or masked) array next time - for example as a Zarr file or another suitable format. Otherwise, border-alignment issues will remain.\n\n4. CT volumes from scanners or PACS are almost always stored as int16 or uint16, with meaningful Hounsfield units (HU). Is there a particular reason they were cast to uint8 here?",
    "3362970": "Thank you for the questions and welcome to the competition!\n\n1. Thanks for noticing this. Indeed, that padding was supposed to be an ignore-label. We will investigate with our team of annotators whether this happens uniformly both in the training and the test set, and consider whether a data update should be necessary.\n2. Yes.\n3. Full CT scans are available as unlabeled data as OME-Zarr on the repositories of the Vesuvius Challenge. The repositories don't include all the volumes used to create the training and the test set. I would recommend starting from these https://data.aws.ash2txt.org/samples . But it could be worthy also having a look at those https://dl.ash2txt.org/full-scrolls/\n4. The Hounsfield scale doesn't necessarily make sense here because everything is carbon based and live approximately in the same range of values. We are talking about carbonized scrolls of papyrus written with carbonized Roman ink (extremely thin, and almost entirely carbon based). In our new scans (the one performed at the ESRF synchrotron) we are picking the windowing in order to maximize what seems to be the range of useful available information. In the old scans (acquired at Diamond Light Source) the choice of windowing has been less accurate because at the time it was not well known whether useful information to detect the ink could be hidden in the \"noise\" or very small values. We are using uint8 instead of uint16 because this, at least so far, has not been proved to be detrimental for the ink detection task, which requires to detect thin features practically invisible to the naked eye and should be thus be more sensible to a change in precision than the task tackled in this active competition!\n\nHope this helps!",
    "3364400": "Thank you for the detailed answer. Jus a couple of follow-up questions:\n\nSay, I use this one for the pre-training: https://dl.ash2txt.org/full-scrolls/\n\nWhat is a better start: volume_grids,  volumes, volumes_masked or volumes_zarr_standardized?\n\nVolumes are uint16 - maybe you could share a pipeline how the competition data (images) was processed from one of those sources so we could replicate it?",
    "3365718": "I prefer using the OME-zarrs because they are easier to access. However, the \"standardization\" procedure used to create those volumes is not \"orthodox\" (it's overly compressing the information imho). This is why I would recommend using these volumes instead: https://data.aws.ash2txt.org/samples/",
    "3392907": "What was the result of the investigation in 1.? Is this border region uniform in the test region? Is it ignored for scoring?"
  },
  "source": "meta"
}