{
  "id": 641165,
  "title": "Dataset creation",
  "url": "/competitions/vesuvius-challenge-surface-detection/discussion/641165",
  "author_name": "",
  "post_date": "2025-11-26T15:33:44.818460900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello all,\nIs there any information about the creation of the dataset. \nI know they are crops from much larger volumes. I would like to know if we have access to the entire volumes and how was the crop selection done as well as the segmentation process.\nMaybe the data comes from here: <a href=\"https://scrollprize.org/unwrapping\" target=\"_blank\">https://scrollprize.org/unwrapping</a>\nThank you for your time!</p>\n<p>Best regards,\nAndré Ferreira</p>",
  "messages": [
    {
      "id": "3349302",
      "postDate": "11/26/2025 15:33:44",
      "content": "<p>Hello all,\nIs there any information about the creation of the dataset. \nI know they are crops from much larger volumes. I would like to know if we have access to the entire volumes and how was the crop selection done as well as the segmentation process.\nMaybe the data comes from here: <a href=\"https://scrollprize.org/unwrapping\" target=\"_blank\">https://scrollprize.org/unwrapping</a>\nThank you for your time!</p>\n<p>Best regards,\nAndré Ferreira</p>",
      "rawMarkdown": "Hello all,\nIs there any information about the creation of the dataset. \nI know they are crops from much larger volumes. I would like to know if we have access to the entire volumes and how was the crop selection done as well as the segmentation process.\nMaybe the data comes from here: https://scrollprize.org/unwrapping\nThank you for your time!\n\nBest regards,\nAndré Ferreira",
      "votes": null
    },
    {
      "id": "3349594",
      "postDate": "11/26/2025 18:25:04",
      "content": "<p>I found a case where the segmentation seems to be incorrect, although I'm not really sure if the problem was interpolation. File with name 1407735.tif, slice 160.</p>",
      "rawMarkdown": "I found a case where the segmentation seems to be incorrect, although I'm not really sure if the problem was interpolation. File with name 1407735.tif, slice 160.",
      "votes": null
    },
    {
      "id": "3349705",
      "postDate": "11/26/2025 22:50:46",
      "content": "<p>Hello! ,</p>\n<p>The dataset is created by a couple methods, but all are derived from 'segmentations'. I put segmentations in quotes here as we use the term a bit weird. A segmentation in this context is a single mesh from a portion of the scroll which was created either manually (using khartes or legacy volume-cartographer) or semi/fully automatically using vc3d. </p>\n<p>The recto portion of the scroll is traced and meshed and then parameterized into a 2d surface using slim or abf. To generate labels, these meshes are voxelized (by rasterizing the 3d triangles and connecting disconnected pieces with a simple line algorithm). From there, I went through each zarr one chunk of size 320^3 at a time and identified areas which had \"clean\" labels, ideally grouped together. </p>\n<p>The volumes in the training set come from a combination of public and yet to be released data. We currently host data in two places: </p>\n<p><a href=\"https://dl.ash2txt.org/full-scrolls\" target=\"_blank\">ash2txt file server</a>:\nin this directory you will find directories for five scrolls, and in the volumes/ directory within them are either zarrs or single 2d tif slices</p>\n<p><a href=\"https://dl.ash2txt.org/fragments/\" target=\"_blank\">fragments</a>:\nthis directory contains fragments (pieces of unrolling attempts)</p>\n<p>we also have two more volumes <a href=\"https://data.aws.ash2txt.org/samples/\" target=\"_blank\">here</a> , which will be where we our moving our data in the future. </p>\n<p>the labeled data represents and infinitely small fraction of the total volumes</p>\n<p>the segmentation guide on <a href=\"www.scrollprize.org/segmentation\" target=\"_blank\">our website</a> should give you a decent introduction into segmentation with vc3d, at least enough to start producing these. if you wish to get into segmentation and run into any issues don't hesitate to ask</p>",
      "rawMarkdown": "Hello! ,\n\nThe dataset is created by a couple methods, but all are derived from 'segmentations'. I put segmentations in quotes here as we use the term a bit weird. A segmentation in this context is a single mesh from a portion of the scroll which was created either manually (using khartes or legacy volume-cartographer) or semi/fully automatically using vc3d. \n\nThe recto portion of the scroll is traced and meshed and then parameterized into a 2d surface using slim or abf. To generate labels, these meshes are voxelized (by rasterizing the 3d triangles and connecting disconnected pieces with a simple line algorithm). From there, I went through each zarr one chunk of size 320^3 at a time and identified areas which had \"clean\" labels, ideally grouped together. \n\nThe volumes in the training set come from a combination of public and yet to be released data. We currently host data in two places: \n\n[ash2txt file server](https://dl.ash2txt.org/full-scrolls):\nin this directory you will find directories for five scrolls, and in the volumes/ directory within them are either zarrs or single 2d tif slices\n\n[fragments](https://dl.ash2txt.org/fragments/):\nthis directory contains fragments (pieces of unrolling attempts)\n\nwe also have two more volumes [here](https://data.aws.ash2txt.org/samples/) , which will be where we our moving our data in the future. \n\nthe labeled data represents and infinitely small fraction of the total volumes\n\nthe segmentation guide on [our website](www.scrollprize.org/segmentation) should give you a decent introduction into segmentation with vc3d, at least enough to start producing these. if you wish to get into segmentation and run into any issues don't hesitate to ask",
      "votes": null
    },
    {
      "id": "3350366",
      "postDate": "11/27/2025 14:17:08",
      "content": "<p>Thank you for the detailed explanation! It's definitly a lot of work to create the labels, thanks!</p>\n<p>Just one more thing to make sure I have the correct mental image of the process. When we load a standard chunk (e.g., $320^3$ voxels), is this effectively a direct, 1:1 crop from the original high-resolution reconstruction? Meaning: If the native resolution is ~8µm, this chunk represents a tiny physical cube of roughly 2.5mm x 2.5mm x 2.5mm cut from the scroll, without any resizing or downsampling? I.e. basically zooming in and crop? I just want to ensure these aren't downscaled overviews, but rather full-resolution 'micro-crops'.</p>",
      "rawMarkdown": "Thank you for the detailed explanation! It's definitly a lot of work to create the labels, thanks!\n\nJust one more thing to make sure I have the correct mental image of the process. When we load a standard chunk (e.g., $320^3$ voxels), is this effectively a direct, 1:1 crop from the original high-resolution reconstruction? Meaning: If the native resolution is ~8µm, this chunk represents a tiny physical cube of roughly 2.5mm x 2.5mm x 2.5mm cut from the scroll, without any resizing or downsampling? I.e. basically zooming in and crop? I just want to ensure these aren't downscaled overviews, but rather full-resolution 'micro-crops'.",
      "votes": null
    },
    {
      "id": "3350801",
      "postDate": "11/27/2025 20:00:19",
      "content": "<p>Correct, these have not been resampled at all , so the physical dimensions are accurate for the scroll they come from. They are chunks directly from the raw volumes. </p>",
      "rawMarkdown": "Correct, these have not been resampled at all , so the physical dimensions are accurate for the scroll they come from. They are chunks directly from the raw volumes.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3349594,
      "author_name": "andrefilipeferreira",
      "author_url": "",
      "post_date": "11/26/2025 18:25:04",
      "content": "<p>I found a case where the segmentation seems to be incorrect, although I'm not really sure if the problem was interpolation. File with name 1407735.tif, slice 160.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3349705,
      "author_name": "seanjohnsonsp",
      "author_url": "",
      "post_date": "11/26/2025 22:50:46",
      "content": "<p>Hello! ,</p>\n<p>The dataset is created by a couple methods, but all are derived from 'segmentations'. I put segmentations in quotes here as we use the term a bit weird. A segmentation in this context is a single mesh from a portion of the scroll which was created either manually (using khartes or legacy volume-cartographer) or semi/fully automatically using vc3d. </p>\n<p>The recto portion of the scroll is traced and meshed and then parameterized into a 2d surface using slim or abf. To generate labels, these meshes are voxelized (by rasterizing the 3d triangles and connecting disconnected pieces with a simple line algorithm). From there, I went through each zarr one chunk of size 320^3 at a time and identified areas which had \"clean\" labels, ideally grouped together. </p>\n<p>The volumes in the training set come from a combination of public and yet to be released data. We currently host data in two places: </p>\n<p><a href=\"https://dl.ash2txt.org/full-scrolls\" target=\"_blank\">ash2txt file server</a>:\nin this directory you will find directories for five scrolls, and in the volumes/ directory within them are either zarrs or single 2d tif slices</p>\n<p><a href=\"https://dl.ash2txt.org/fragments/\" target=\"_blank\">fragments</a>:\nthis directory contains fragments (pieces of unrolling attempts)</p>\n<p>we also have two more volumes <a href=\"https://data.aws.ash2txt.org/samples/\" target=\"_blank\">here</a> , which will be where we our moving our data in the future. </p>\n<p>the labeled data represents and infinitely small fraction of the total volumes</p>\n<p>the segmentation guide on <a href=\"www.scrollprize.org/segmentation\" target=\"_blank\">our website</a> should give you a decent introduction into segmentation with vc3d, at least enough to start producing these. if you wish to get into segmentation and run into any issues don't hesitate to ask</p>",
      "votes": null,
      "replies": [
        {
          "id": 3350366,
          "author_name": "andrefilipeferreira",
          "author_url": "",
          "post_date": "11/27/2025 14:17:08",
          "content": "<p>Thank you for the detailed explanation! It's definitly a lot of work to create the labels, thanks!</p>\n<p>Just one more thing to make sure I have the correct mental image of the process. When we load a standard chunk (e.g., $320^3$ voxels), is this effectively a direct, 1:1 crop from the original high-resolution reconstruction? Meaning: If the native resolution is ~8µm, this chunk represents a tiny physical cube of roughly 2.5mm x 2.5mm x 2.5mm cut from the scroll, without any resizing or downsampling? I.e. basically zooming in and crop? I just want to ensure these aren't downscaled overviews, but rather full-resolution 'micro-crops'.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3350801,
              "author_name": "seanjohnsonsp",
              "author_url": "",
              "post_date": "11/27/2025 20:00:19",
              "content": "<p>Correct, these have not been resampled at all , so the physical dimensions are accurate for the scroll they come from. They are chunks directly from the raw volumes. </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3349302": "Hello all,\nIs there any information about the creation of the dataset. \nI know they are crops from much larger volumes. I would like to know if we have access to the entire volumes and how was the crop selection done as well as the segmentation process.\nMaybe the data comes from here: https://scrollprize.org/unwrapping\nThank you for your time!\n\nBest regards,\nAndré Ferreira",
    "3349594": "I found a case where the segmentation seems to be incorrect, although I'm not really sure if the problem was interpolation. File with name 1407735.tif, slice 160.",
    "3349705": "Hello! ,\n\nThe dataset is created by a couple methods, but all are derived from 'segmentations'. I put segmentations in quotes here as we use the term a bit weird. A segmentation in this context is a single mesh from a portion of the scroll which was created either manually (using khartes or legacy volume-cartographer) or semi/fully automatically using vc3d. \n\nThe recto portion of the scroll is traced and meshed and then parameterized into a 2d surface using slim or abf. To generate labels, these meshes are voxelized (by rasterizing the 3d triangles and connecting disconnected pieces with a simple line algorithm). From there, I went through each zarr one chunk of size 320^3 at a time and identified areas which had \"clean\" labels, ideally grouped together. \n\nThe volumes in the training set come from a combination of public and yet to be released data. We currently host data in two places: \n\n[ash2txt file server](https://dl.ash2txt.org/full-scrolls):\nin this directory you will find directories for five scrolls, and in the volumes/ directory within them are either zarrs or single 2d tif slices\n\n[fragments](https://dl.ash2txt.org/fragments/):\nthis directory contains fragments (pieces of unrolling attempts)\n\nwe also have two more volumes [here](https://data.aws.ash2txt.org/samples/) , which will be where we our moving our data in the future. \n\nthe labeled data represents and infinitely small fraction of the total volumes\n\nthe segmentation guide on [our website](www.scrollprize.org/segmentation) should give you a decent introduction into segmentation with vc3d, at least enough to start producing these. if you wish to get into segmentation and run into any issues don't hesitate to ask",
    "3350366": "Thank you for the detailed explanation! It's definitly a lot of work to create the labels, thanks!\n\nJust one more thing to make sure I have the correct mental image of the process. When we load a standard chunk (e.g., $320^3$ voxels), is this effectively a direct, 1:1 crop from the original high-resolution reconstruction? Meaning: If the native resolution is ~8µm, this chunk represents a tiny physical cube of roughly 2.5mm x 2.5mm x 2.5mm cut from the scroll, without any resizing or downsampling? I.e. basically zooming in and crop? I just want to ensure these aren't downscaled overviews, but rather full-resolution 'micro-crops'.",
    "3350801": "Correct, these have not been resampled at all , so the physical dimensions are accurate for the scroll they come from. They are chunks directly from the raw volumes."
  },
  "source": "meta"
}