{
  "id": 396764,
  "title": "Curiosity question about the scans fragments / papyrus rolls",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/396764",
  "author_name": "",
  "post_date": "2023-03-22T20:46:19.112501100Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi there!</p>\n<p>I have a question about the scanning process used for the competition. It's more out of curiosity than a question for the competition itself.</p>\n<p>The competition's dataset consists of <strong>fragments</strong> that have a relatively flat shape. These fragments appear to be parts of complete rolls that have been destroyed. However, in the <a href=\"https://youtu.be/PpNq2cFotyY?t=759\" target=\"_blank\">video provided</a>, we see that the <strong>rolls are completely scanned</strong> in their entirety.</p>\n<p>As I understand it, the main challenge is to be able to read the content without damaging the rolls. However, it seems that the paper used for the competition is not in the same condition.</p>\n<p>I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.</p>\n<p>Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?</p>\n<p>Does anyone have an explanation for this? 🤔</p>\n<p>I know these answers are not needed for the competition but it is something I am really interested in!</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "2192690",
      "postDate": "03/22/2023 20:46:19",
      "content": "<p>Hi there!</p>\n<p>I have a question about the scanning process used for the competition. It's more out of curiosity than a question for the competition itself.</p>\n<p>The competition's dataset consists of <strong>fragments</strong> that have a relatively flat shape. These fragments appear to be parts of complete rolls that have been destroyed. However, in the <a href=\"https://youtu.be/PpNq2cFotyY?t=759\" target=\"_blank\">video provided</a>, we see that the <strong>rolls are completely scanned</strong> in their entirety.</p>\n<p>As I understand it, the main challenge is to be able to read the content without damaging the rolls. However, it seems that the paper used for the competition is not in the same condition.</p>\n<p>I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.</p>\n<p>Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?</p>\n<p>Does anyone have an explanation for this? 🤔</p>\n<p>I know these answers are not needed for the competition but it is something I am really interested in!</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Hi there!\n\nI have a question about the scanning process used for the competition. It's more out of curiosity than a question for the competition itself.\n\nThe competition's dataset consists of **fragments** that have a relatively flat shape. These fragments appear to be parts of complete rolls that have been destroyed. However, in the [video provided](https://youtu.be/PpNq2cFotyY?t=759), we see that the **rolls are completely scanned** in their entirety.\n\nAs I understand it, the main challenge is to be able to read the content without damaging the rolls. However, it seems that the paper used for the competition is not in the same condition.\n\nI'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.\n\nAlso, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?\n\nDoes anyone have an explanation for this? 🤔\n\nI know these answers are not needed for the competition but it is something I am really interested in!\n\nThank you!",
      "votes": null
    },
    {
      "id": "2192699",
      "postDate": "03/22/2023 20:56:52",
      "content": "<p>The fragments are from scrolls that have been physically opened in the past to reveal visible text on their surfaces.  They were scanned in full, almost flattened fragmentary form. The datasets available here in the competition have been post-processed using the method outlined <a href=\"https://scrollprize.org/tutorial3\" target=\"_blank\">here</a> to create much smaller volumes that behave more like 2D images of each fragment while still containing the local 3D depth information. They are, in that respect, an optimization for people doing machine learning.</p>\n<p>There are many scrolls that have not been physically unrolled. The goal of the Grand Prize on scrollprize.org is to read text from one of those. The goal of the Kaggle competition is to improve ink detection on the fragments. The idea is that an ink detection method that can correctly identify the location of ink in the CT scans of the fragments, where we know the answer, will help us to identify ink from inside the still rolled scrolls, where we don't know the answer.</p>\n<p>The difference in condition (rolled vs unrolled) will certainly be important for transferring knowledge from the fragments to the scrolls. However, it is worth nothing that the fragments themselves consist of multiple layers of text which are stuck together. This makes them <em>locally</em> very similar to rolled scrolls.</p>",
      "rawMarkdown": "The fragments are from scrolls that have been physically opened in the past to reveal visible text on their surfaces.  They were scanned in full, almost flattened fragmentary form. The datasets available here in the competition have been post-processed using the method outlined [here](https://scrollprize.org/tutorial3) to create much smaller volumes that behave more like 2D images of each fragment while still containing the local 3D depth information. They are, in that respect, an optimization for people doing machine learning.\n\nThere are many scrolls that have not been physically unrolled. The goal of the Grand Prize on scrollprize.org is to read text from one of those. The goal of the Kaggle competition is to improve ink detection on the fragments. The idea is that an ink detection method that can correctly identify the location of ink in the CT scans of the fragments, where we know the answer, will help us to identify ink from inside the still rolled scrolls, where we don't know the answer.\n\nThe difference in condition (rolled vs unrolled) will certainly be important for transferring knowledge from the fragments to the scrolls. However, it is worth nothing that the fragments themselves consist of multiple layers of text which are stuck together. This makes them _locally_ very similar to rolled scrolls.",
      "votes": null
    },
    {
      "id": "2192704",
      "postDate": "03/22/2023 20:59:59",
      "content": "<blockquote>\n  <p>I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.</p>\n</blockquote>\n<p>Yes, we think this will be one of the primary challenges of the Grand Prize. We're running the Ink Detection competition on Kaggle to make the models as good as possible, and hopefully one of the winning models will be able to transfer to the inside of the actual scrolls. But we agree, this will be hard.</p>\n<p>That said, locally the pieces of papyrus in the rolled up scrolls are relatively flat, so we hope that will help when applying the models.</p>\n<blockquote>\n  <p>Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?</p>\n</blockquote>\n<p>The fragments are actually multiple layers stuck together. In <a href=\"https://scrollprize.org/tutorial4\" target=\"_blank\">Tutorial 4</a> we talk about these \"hidden layers\", and even show some characters that we have uncovered in them. The data in the Kaggle competition has been preprocessed to only contain the top layer per fragment (see the other tutorials on scrollprize.org for more information on how we did that), but the original data is available on scrollprize.org/data, so you can try to look at the hidden layers there.</p>",
      "rawMarkdown": "> I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.\n\nYes, we think this will be one of the primary challenges of the Grand Prize. We're running the Ink Detection competition on Kaggle to make the models as good as possible, and hopefully one of the winning models will be able to transfer to the inside of the actual scrolls. But we agree, this will be hard.\n\nThat said, locally the pieces of papyrus in the rolled up scrolls are relatively flat, so we hope that will help when applying the models.\n\n> Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?\n\nThe fragments are actually multiple layers stuck together. In [Tutorial 4](https://scrollprize.org/tutorial4) we talk about these \"hidden layers\", and even show some characters that we have uncovered in them. The data in the Kaggle competition has been preprocessed to only contain the top layer per fragment (see the other tutorials on scrollprize.org for more information on how we did that), but the original data is available on scrollprize.org/data, so you can try to look at the hidden layers there.",
      "votes": null
    },
    {
      "id": "2192746",
      "postDate": "03/22/2023 22:07:36",
      "content": "<p>Thank you very much for your clear answer and for hosting this competition, this makes sense to me now! </p>",
      "rawMarkdown": "Thank you very much for your clear answer and for hosting this competition, this makes sense to me now!",
      "votes": null
    },
    {
      "id": "2192748",
      "postDate": "03/22/2023 22:08:28",
      "content": "<p>Thank you very much! This completes perfectly the other answer.</p>",
      "rawMarkdown": "Thank you very much! This completes perfectly the other answer.",
      "votes": null
    },
    {
      "id": "2199608",
      "postDate": "03/27/2023 21:08:09",
      "content": "<p>But just to confirm, we’re not allowed to use the data at scrollprize.org/data (e.g. data for the hidden layers) in this Kaggle competition?</p>",
      "rawMarkdown": "But just to confirm, we’re not allowed to use the data at scrollprize.org/data (e.g. data for the hidden layers) in this Kaggle competition?",
      "votes": null
    },
    {
      "id": "2199630",
      "postDate": "03/27/2023 21:37:08",
      "content": "<p>You are allowed to use the data at scrollprize.org/data. There is no prohibition on using that data.</p>",
      "rawMarkdown": "You are allowed to use the data at scrollprize.org/data. There is no prohibition on using that data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2192699,
      "author_name": "csparker",
      "author_url": "",
      "post_date": "03/22/2023 20:56:52",
      "content": "<p>The fragments are from scrolls that have been physically opened in the past to reveal visible text on their surfaces.  They were scanned in full, almost flattened fragmentary form. The datasets available here in the competition have been post-processed using the method outlined <a href=\"https://scrollprize.org/tutorial3\" target=\"_blank\">here</a> to create much smaller volumes that behave more like 2D images of each fragment while still containing the local 3D depth information. They are, in that respect, an optimization for people doing machine learning.</p>\n<p>There are many scrolls that have not been physically unrolled. The goal of the Grand Prize on scrollprize.org is to read text from one of those. The goal of the Kaggle competition is to improve ink detection on the fragments. The idea is that an ink detection method that can correctly identify the location of ink in the CT scans of the fragments, where we know the answer, will help us to identify ink from inside the still rolled scrolls, where we don't know the answer.</p>\n<p>The difference in condition (rolled vs unrolled) will certainly be important for transferring knowledge from the fragments to the scrolls. However, it is worth nothing that the fragments themselves consist of multiple layers of text which are stuck together. This makes them <em>locally</em> very similar to rolled scrolls.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2192748,
          "author_name": "paulbacher",
          "author_url": "",
          "post_date": "03/22/2023 22:08:28",
          "content": "<p>Thank you very much! This completes perfectly the other answer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2192704,
      "author_name": "jpposma",
      "author_url": "",
      "post_date": "03/22/2023 20:59:59",
      "content": "<blockquote>\n  <p>I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.</p>\n</blockquote>\n<p>Yes, we think this will be one of the primary challenges of the Grand Prize. We're running the Ink Detection competition on Kaggle to make the models as good as possible, and hopefully one of the winning models will be able to transfer to the inside of the actual scrolls. But we agree, this will be hard.</p>\n<p>That said, locally the pieces of papyrus in the rolled up scrolls are relatively flat, so we hope that will help when applying the models.</p>\n<blockquote>\n  <p>Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?</p>\n</blockquote>\n<p>The fragments are actually multiple layers stuck together. In <a href=\"https://scrollprize.org/tutorial4\" target=\"_blank\">Tutorial 4</a> we talk about these \"hidden layers\", and even show some characters that we have uncovered in them. The data in the Kaggle competition has been preprocessed to only contain the top layer per fragment (see the other tutorials on scrollprize.org for more information on how we did that), but the original data is available on scrollprize.org/data, so you can try to look at the hidden layers there.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2192746,
          "author_name": "paulbacher",
          "author_url": "",
          "post_date": "03/22/2023 22:07:36",
          "content": "<p>Thank you very much for your clear answer and for hosting this competition, this makes sense to me now! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2199608,
          "author_name": "nicholasbailey",
          "author_url": "",
          "post_date": "03/27/2023 21:08:09",
          "content": "<p>But just to confirm, we’re not allowed to use the data at scrollprize.org/data (e.g. data for the hidden layers) in this Kaggle competition?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2199630,
              "author_name": "jpposma",
              "author_url": "",
              "post_date": "03/27/2023 21:37:08",
              "content": "<p>You are allowed to use the data at scrollprize.org/data. There is no prohibition on using that data.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2192690": "Hi there!\n\nI have a question about the scanning process used for the competition. It's more out of curiosity than a question for the competition itself.\n\nThe competition's dataset consists of **fragments** that have a relatively flat shape. These fragments appear to be parts of complete rolls that have been destroyed. However, in the [video provided](https://youtu.be/PpNq2cFotyY?t=759), we see that the **rolls are completely scanned** in their entirety.\n\nAs I understand it, the main challenge is to be able to read the content without damaging the rolls. However, it seems that the paper used for the competition is not in the same condition.\n\nI'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.\n\nAlso, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?\n\nDoes anyone have an explanation for this? 🤔\n\nI know these answers are not needed for the competition but it is something I am really interested in!\n\nThank you!",
    "2192699": "The fragments are from scrolls that have been physically opened in the past to reveal visible text on their surfaces.  They were scanned in full, almost flattened fragmentary form. The datasets available here in the competition have been post-processed using the method outlined [here](https://scrollprize.org/tutorial3) to create much smaller volumes that behave more like 2D images of each fragment while still containing the local 3D depth information. They are, in that respect, an optimization for people doing machine learning.\n\nThere are many scrolls that have not been physically unrolled. The goal of the Grand Prize on scrollprize.org is to read text from one of those. The goal of the Kaggle competition is to improve ink detection on the fragments. The idea is that an ink detection method that can correctly identify the location of ink in the CT scans of the fragments, where we know the answer, will help us to identify ink from inside the still rolled scrolls, where we don't know the answer.\n\nThe difference in condition (rolled vs unrolled) will certainly be important for transferring knowledge from the fragments to the scrolls. However, it is worth nothing that the fragments themselves consist of multiple layers of text which are stuck together. This makes them _locally_ very similar to rolled scrolls.",
    "2192704": "> I'm having difficulty understanding how we could use a model trained on flat scans for a papyrus roll. The text would be oriented in various directions and overlap itself, eventually finishing at the very center of the roll.\n\nYes, we think this will be one of the primary challenges of the Grand Prize. We're running the Ink Detection competition on Kaggle to make the models as good as possible, and hopefully one of the winning models will be able to transfer to the inside of the actual scrolls. But we agree, this will be hard.\n\nThat said, locally the pieces of papyrus in the rolled up scrolls are relatively flat, so we hope that will help when applying the models.\n\n> Also, do you know what the thickness is along the different scan layers? Are they composed of a single page or a stack of pages? I believe they are composed of a single page, so there is no text overlap. How can we be sure that there are no more pages in the fragment?\n\nThe fragments are actually multiple layers stuck together. In [Tutorial 4](https://scrollprize.org/tutorial4) we talk about these \"hidden layers\", and even show some characters that we have uncovered in them. The data in the Kaggle competition has been preprocessed to only contain the top layer per fragment (see the other tutorials on scrollprize.org for more information on how we did that), but the original data is available on scrollprize.org/data, so you can try to look at the hidden layers there.",
    "2192746": "Thank you very much for your clear answer and for hosting this competition, this makes sense to me now!",
    "2192748": "Thank you very much! This completes perfectly the other answer.",
    "2199608": "But just to confirm, we’re not allowed to use the data at scrollprize.org/data (e.g. data for the hidden layers) in this Kaggle competition?",
    "2199630": "You are allowed to use the data at scrollprize.org/data. There is no prohibition on using that data."
  },
  "source": "meta"
}