{
  "id": 619251,
  "title": "Welcome to Vesuvius Challenge - Surface Detection competition",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/619251",
  "author_name": "Giorgio Angelotti",
  "post_date": "2025-11-13T19:13:33.291000",
  "votes": 24,
  "comment_count": 5,
  "views": 0,
  "content": "<p><strong>AVE!</strong></p>\n<p>The Vesuvius Challenge Team is excited to present this challenge to you!</p>\n<p>We are a team of researchers who are trying to achieve the unthinkable: we want to resurrect an ancient library of Roman scrolls that were carbonized during a volcanic eruption about 2,000 years ago.</p>\n<p>We have already shown that these scrolls are readable… but we need to unwrap them first.\nThat’s where you come in: we are building a semi-automatic pipeline to virtually unwrap these sealed carbonized scrolls from their X-ray CT scans. The scrolls are extremely fragile and were scanned at a synchrotron (a kind of particle accelerator!).</p>\n<p>The first step of our pipeline involves <strong>machine learning detection of the rolled papyrus surface</strong>.\nWe need model predictions that closely follow the papyrus and are free of topological mistakes (annoying mergers and splits).</p>\n<p>Will you help us, and rewrite history forever?</p>\n<p>Some useful links:</p>\n<ol>\n<li><a href=\"https://scrollprize.org/\" target=\"_blank\">Our official website</a></li>\n<li><a href=\"https://discord.gg/V4fJhvtaQn\" target=\"_blank\">Our official Discord channel</a> (this is not the official Kaggle competition channel, but you can come here to ask questions about the broader project)</li>\n<li><a href=\"https://scrollprize.org/unwrapping\" target=\"_blank\">Information about Virtual Unwrapping</a></li>\n</ol>\n<p>Feel free to also ask questions in the competition Discussion forum. We also strongly encourage you to share your code and ideas!\nWe’re happy to answer and brainstorm!</p>",
  "messages": [
    {
      "id": 3322859,
      "postDate": "2025-11-13T19:13:33.290Z",
      "content": "<p><strong>AVE!</strong></p>\n<p>The Vesuvius Challenge Team is excited to present this challenge to you!</p>\n<p>We are a team of researchers who are trying to achieve the unthinkable: we want to resurrect an ancient library of Roman scrolls that were carbonized during a volcanic eruption about 2,000 years ago.</p>\n<p>We have already shown that these scrolls are readable… but we need to unwrap them first.\nThat’s where you come in: we are building a semi-automatic pipeline to virtually unwrap these sealed carbonized scrolls from their X-ray CT scans. The scrolls are extremely fragile and were scanned at a synchrotron (a kind of particle accelerator!).</p>\n<p>The first step of our pipeline involves <strong>machine learning detection of the rolled papyrus surface</strong>.\nWe need model predictions that closely follow the papyrus and are free of topological mistakes (annoying mergers and splits).</p>\n<p>Will you help us, and rewrite history forever?</p>\n<p>Some useful links:</p>\n<ol>\n<li><a href=\"https://scrollprize.org/\" target=\"_blank\">Our official website</a></li>\n<li><a href=\"https://discord.gg/V4fJhvtaQn\" target=\"_blank\">Our official Discord channel</a> (this is not the official Kaggle competition channel, but you can come here to ask questions about the broader project)</li>\n<li><a href=\"https://scrollprize.org/unwrapping\" target=\"_blank\">Information about Virtual Unwrapping</a></li>\n</ol>\n<p>Feel free to also ask questions in the competition Discussion forum. We also strongly encourage you to share your code and ideas!\nWe’re happy to answer and brainstorm!</p>",
      "rawMarkdown": "**AVE!**\n\nThe Vesuvius Challenge Team is excited to present this challenge to you!\n\nWe are a team of researchers who are trying to achieve the unthinkable: we want to resurrect an ancient library of Roman scrolls that were carbonized during a volcanic eruption about 2,000 years ago.\n\nWe have already shown that these scrolls are readable… but we need to unwrap them first.\nThat’s where you come in: we are building a semi-automatic pipeline to virtually unwrap these sealed carbonized scrolls from their X-ray CT scans. The scrolls are extremely fragile and were scanned at a synchrotron (a kind of particle accelerator!).\n\nThe first step of our pipeline involves **machine learning detection of the rolled papyrus surface**.\nWe need model predictions that closely follow the papyrus and are free of topological mistakes (annoying mergers and splits).\n\nWill you help us, and rewrite history forever?\n\nSome useful links:\n\n1. [Our official website](https://scrollprize.org/)\n2. [Our official Discord channel](https://discord.gg/V4fJhvtaQn) (this is not the official Kaggle competition channel, but you can come here to ask questions about the broader project)\n3. [Information about Virtual Unwrapping](https://scrollprize.org/unwrapping)\n\nFeel free to also ask questions in the competition Discussion forum. We also strongly encourage you to share your code and ideas!\nWe’re happy to answer and brainstorm!\n",
      "votes": 24
    },
    {
      "id": 3346499,
      "postDate": "2025-11-24T12:21:03.873Z",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>\n<p>From the official website, I found the recto-surface dataset: <a href=\"https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/\" target=\"_blank\">https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/</a>. According to the documentation, the labels use 0 for background and 1 for surface. I have two questions:</p>\n<ol>\n<li><p>Does the label 1 here correspond to the same foreground annotation we are trying to segment in the competition? In other words, is it essentially the same target class?</p></li>\n<li><p>The README mentions the following:</p></li>\n</ol>\n<blockquote>\n  <p>They were created by first voxelizing merged meshes created by <a href=\"https://www.kaggle.com/richi\" target=\"_blank\">@richi</a> from the scroll 1 gp banner here and the scroll 4 segments here, with some autogenerated segments used as the labels from 5. The script used to perform this voxelization into zarr format is here</p>\n</blockquote>\n<p>From my understanding, the datasets referenced here scrolls 1, 4, and 5 are different from the Kaggle competition dataset. If that’s correct, it means there is no data leakage between this recto-surface dataset and the official Kaggle competition data.</p>\n<p>sample-view\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F2016b0c24ac144fcbef6fbc54c514f26%2FScreenshot%202025-11-24%20183911.png?generation=1763987964007958&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "@giorgioangelotti \n\nFrom the official website, I found the recto-surface dataset: https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/. According to the documentation, the labels use 0 for background and 1 for surface. I have two questions:\n\n1. Does the label 1 here correspond to the same foreground annotation we are trying to segment in the competition? In other words, is it essentially the same target class?\n\n2. The README mentions the following:\n\n> They were created by first voxelizing merged meshes created by @richi from the scroll 1 gp banner here and the scroll 4 segments here, with some autogenerated segments used as the labels from 5. The script used to perform this voxelization into zarr format is here\n\nFrom my understanding, the datasets referenced here scrolls 1, 4, and 5 are different from the Kaggle competition dataset. If that’s correct, it means there is no data leakage between this recto-surface dataset and the official Kaggle competition data.\n\nsample-view\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F2016b0c24ac144fcbef6fbc54c514f26%2FScreenshot%202025-11-24%20183911.png?generation=1763987964007958&alt=media)",
      "votes": 1,
      "replies": [
        {
          "id": 3350975,
          "postDate": "2025-11-28T03:47:12.327Z",
          "content": "<p>I attempted to train on this dataset and verify it with the official dataset, only to find that the metrics were very low. The annotation of this dataset might not be accurate</p>",
          "rawMarkdown": "I attempted to train on this dataset and verify it with the official dataset, only to find that the metrics were very low. The annotation of this dataset might not be accurate",
          "votes": 1
        },
        {
          "id": 3351111,
          "postDate": "2025-11-28T06:49:11.330Z",
          "content": "<p>This dataset was created last year, and we used as labels voxelized meshes produced by our annotation team with the tools that were available at the time, which were not as good as those we have now. Moreover, these samples were created by merging separate \"segmentations\" of meshes in 3D, and next to the boundaries of these meshes the accuracy of the annotations can be lower.</p>\n<p>Nevertheless, these annotations are mostly good but they are not as accurate as the ones we produced for the Kaggle dataset, which underwent a massive effort of tightening to the actual surface and foolproofing.</p>\n<p>The scrolls used in this dataset are PHerc Paris 4 (Scroll 1), PHerc 1667 (Scroll 4) and PHerc 172 (Scroll 5) which were scanned at Diamond Light Source.\nI don't exclude that some of these scrolls aren't used in the training set for Kaggle. In that case, the same region of interest from the Kaggle dataset is the most accurate.</p>\n<p>For sure, training only on this dataset won't lead to good metrics (as <a href=\"https://www.kaggle.com/cudacoding\" target=\"_blank\">@cudacoding</a> remarked) because this dataset lacks not only a decent amount of samples in region of interests with sheets very packed together, which are the most affected by possible mergers, but also lacks samples from scrolls scanned at the ESRF synchrotron, which have intensity values more similar to the ones shared in this notebook <a href=\"https://www.kaggle.com/code/giorgioangelotti/access-to-additional-unlabeled-data\" target=\"_blank\">here</a>. We expect that a good model should be able to work on both data like the ones from the scans at Diamond Light Source and ESRF, and indeed our training set for Kaggle is a mixture of both.</p>\n<p>Hope this helps clarifying your doubts! I will add also a note on the dataset page on the website! Thanks for pulling this up!</p>",
          "rawMarkdown": "This dataset was created last year, and we used as labels voxelized meshes produced by our annotation team with the tools that were available at the time, which were not as good as those we have now. Moreover, these samples were created by merging separate \"segmentations\" of meshes in 3D, and next to the boundaries of these meshes the accuracy of the annotations can be lower.\n\nNevertheless, these annotations are mostly good but they are not as accurate as the ones we produced for the Kaggle dataset, which underwent a massive effort of tightening to the actual surface and foolproofing.\n\nThe scrolls used in this dataset are PHerc Paris 4 (Scroll 1), PHerc 1667 (Scroll 4) and PHerc 172 (Scroll 5) which were scanned at Diamond Light Source.\nI don't exclude that some of these scrolls aren't used in the training set for Kaggle. In that case, the same region of interest from the Kaggle dataset is the most accurate.\n\nFor sure, training only on this dataset won't lead to good metrics (as @cudacoding remarked) because this dataset lacks not only a decent amount of samples in region of interests with sheets very packed together, which are the most affected by possible mergers, but also lacks samples from scrolls scanned at the ESRF synchrotron, which have intensity values more similar to the ones shared in this notebook [here](https://www.kaggle.com/code/giorgioangelotti/access-to-additional-unlabeled-data). We expect that a good model should be able to work on both data like the ones from the scans at Diamond Light Source and ESRF, and indeed our training set for Kaggle is a mixture of both.\n\nHope this helps clarifying your doubts! I will add also a note on the dataset page on the website! Thanks for pulling this up!"
        }
      ]
    },
    {
      "id": 3324020,
      "postDate": "2025-11-14T16:18:02.080Z",
      "content": "<p>kaggle competitions download -c vesuvius-challenge-surface-detection</p>",
      "rawMarkdown": "kaggle competitions download -c vesuvius-challenge-surface-detection"
    },
    {
      "id": 3405252,
      "postDate": "2026-02-12T14:19:11.567Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3346499,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2025-11-24T12:21:03.873000",
      "content": "<p><a href=\"https://www.kaggle.com/giorgioangelotti\" target=\"_blank\">@giorgioangelotti</a> </p>\n<p>From the official website, I found the recto-surface dataset: <a href=\"https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/\" target=\"_blank\">https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/</a>. According to the documentation, the labels use 0 for background and 1 for surface. I have two questions:</p>\n<ol>\n<li><p>Does the label 1 here correspond to the same foreground annotation we are trying to segment in the competition? In other words, is it essentially the same target class?</p></li>\n<li><p>The README mentions the following:</p></li>\n</ol>\n<blockquote>\n  <p>They were created by first voxelizing merged meshes created by <a href=\"https://www.kaggle.com/richi\" target=\"_blank\">@richi</a> from the scroll 1 gp banner here and the scroll 4 segments here, with some autogenerated segments used as the labels from 5. The script used to perform this voxelization into zarr format is here</p>\n</blockquote>\n<p>From my understanding, the datasets referenced here scrolls 1, 4, and 5 are different from the Kaggle competition dataset. If that’s correct, it means there is no data leakage between this recto-surface dataset and the official Kaggle competition data.</p>\n<p>sample-view\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F2016b0c24ac144fcbef6fbc54c514f26%2FScreenshot%202025-11-24%20183911.png?generation=1763987964007958&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 3350975,
          "author_name": "Boredom",
          "author_url": "",
          "post_date": "2025-11-28T03:47:12.327000",
          "content": "<p>I attempted to train on this dataset and verify it with the official dataset, only to find that the metrics were very low. The annotation of this dataset might not be accurate</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3351111,
          "author_name": "Giorgio Angelotti",
          "author_url": "",
          "post_date": "2025-11-28T06:49:11.330000",
          "content": "<p>This dataset was created last year, and we used as labels voxelized meshes produced by our annotation team with the tools that were available at the time, which were not as good as those we have now. Moreover, these samples were created by merging separate \"segmentations\" of meshes in 3D, and next to the boundaries of these meshes the accuracy of the annotations can be lower.</p>\n<p>Nevertheless, these annotations are mostly good but they are not as accurate as the ones we produced for the Kaggle dataset, which underwent a massive effort of tightening to the actual surface and foolproofing.</p>\n<p>The scrolls used in this dataset are PHerc Paris 4 (Scroll 1), PHerc 1667 (Scroll 4) and PHerc 172 (Scroll 5) which were scanned at Diamond Light Source.\nI don't exclude that some of these scrolls aren't used in the training set for Kaggle. In that case, the same region of interest from the Kaggle dataset is the most accurate.</p>\n<p>For sure, training only on this dataset won't lead to good metrics (as <a href=\"https://www.kaggle.com/cudacoding\" target=\"_blank\">@cudacoding</a> remarked) because this dataset lacks not only a decent amount of samples in region of interests with sheets very packed together, which are the most affected by possible mergers, but also lacks samples from scrolls scanned at the ESRF synchrotron, which have intensity values more similar to the ones shared in this notebook <a href=\"https://www.kaggle.com/code/giorgioangelotti/access-to-additional-unlabeled-data\" target=\"_blank\">here</a>. We expect that a good model should be able to work on both data like the ones from the scans at Diamond Light Source and ESRF, and indeed our training set for Kaggle is a mixture of both.</p>\n<p>Hope this helps clarifying your doubts! I will add also a note on the dataset page on the website! Thanks for pulling this up!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3324020,
      "author_name": "bharath kumar monditoka",
      "author_url": "",
      "post_date": "2025-11-14T16:18:02.080000",
      "content": "<p>kaggle competitions download -c vesuvius-challenge-surface-detection</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3405252,
      "author_name": "",
      "author_url": "",
      "post_date": "2026-02-12T14:19:11.567000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3322859": "**AVE!**\n\nThe Vesuvius Challenge Team is excited to present this challenge to you!\n\nWe are a team of researchers who are trying to achieve the unthinkable: we want to resurrect an ancient library of Roman scrolls that were carbonized during a volcanic eruption about 2,000 years ago.\n\nWe have already shown that these scrolls are readable… but we need to unwrap them first.\nThat’s where you come in: we are building a semi-automatic pipeline to virtually unwrap these sealed carbonized scrolls from their X-ray CT scans. The scrolls are extremely fragile and were scanned at a synchrotron (a kind of particle accelerator!).\n\nThe first step of our pipeline involves **machine learning detection of the rolled papyrus surface**.\nWe need model predictions that closely follow the papyrus and are free of topological mistakes (annoying mergers and splits).\n\nWill you help us, and rewrite history forever?\n\nSome useful links:\n\n1. [Our official website](https://scrollprize.org/)\n2. [Our official Discord channel](https://discord.gg/V4fJhvtaQn) (this is not the official Kaggle competition channel, but you can come here to ask questions about the broader project)\n3. [Information about Virtual Unwrapping](https://scrollprize.org/unwrapping)\n\nFeel free to also ask questions in the competition Discussion forum. We also strongly encourage you to share your code and ideas!\nWe’re happy to answer and brainstorm!\n",
    "3346499": "@giorgioangelotti \n\nFrom the official website, I found the recto-surface dataset: https://dl.ash2txt.org/datasets/seg-derived-recto-surfaces/. According to the documentation, the labels use 0 for background and 1 for surface. I have two questions:\n\n1. Does the label 1 here correspond to the same foreground annotation we are trying to segment in the competition? In other words, is it essentially the same target class?\n\n2. The README mentions the following:\n\n> They were created by first voxelizing merged meshes created by @richi from the scroll 1 gp banner here and the scroll 4 segments here, with some autogenerated segments used as the labels from 5. The script used to perform this voxelization into zarr format is here\n\nFrom my understanding, the datasets referenced here scrolls 1, 4, and 5 are different from the Kaggle competition dataset. If that’s correct, it means there is no data leakage between this recto-surface dataset and the official Kaggle competition data.\n\nsample-view\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1984321%2F2016b0c24ac144fcbef6fbc54c514f26%2FScreenshot%202025-11-24%20183911.png?generation=1763987964007958&alt=media)",
    "3324020": "kaggle competitions download -c vesuvius-challenge-surface-detection",
    "3405252": ""
  }
}