{
  "id": 414686,
  "title": "Manually refining training data",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/414686",
  "author_name": "",
  "post_date": "2023-06-02T17:42:51.722905500Z",
  "votes": 12,
  "comment_count": 5,
  "views": 0,
  "content": "<p>After a fair bit of poking around in the data, I've noticed that the training data is not as accurate as one might wish.  There was some discussion of this on the discord, but one example is that there is an annotated ink region on fragment 2 that is in fact visible on infrared, but where the surface of the papyrus that would hold that ink goes out of frame.  So there's no possible way to detect ink there:  the papyrus isn't visible on any of the X-ray slices.</p>\n<p>After discovering this, I've started putting some work into manually updating the training data to see if this improves performance.  The first set of data updates is complete and freely available for you to use on Kaggle (or elsewhere): <a href=\"https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks\" target=\"_blank\">https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks</a></p>\n<p>To generate these, I overlaid the infrared image, the fragment mask, and the processed X-ray data for each fragment, then removed any regions in the fragment mask that had no X-ray support at all.  This includes regions around the edges where we see down into the support surface or into fragmentary lower levels of the papyrus, as well as cracks inside the fragments where no data is available on X-ray.  These should be usable as drop-in replacements for the masks provided by the organizers.  I'm hoping that if we avoid training on these bad regions that it can help improve model performance.</p>\n<p>Please try them out for model training and let me know if you see any improvements as a result.</p>",
  "messages": [
    {
      "id": "2285424",
      "postDate": "06/02/2023 17:42:51",
      "content": "<p>After a fair bit of poking around in the data, I've noticed that the training data is not as accurate as one might wish.  There was some discussion of this on the discord, but one example is that there is an annotated ink region on fragment 2 that is in fact visible on infrared, but where the surface of the papyrus that would hold that ink goes out of frame.  So there's no possible way to detect ink there:  the papyrus isn't visible on any of the X-ray slices.</p>\n<p>After discovering this, I've started putting some work into manually updating the training data to see if this improves performance.  The first set of data updates is complete and freely available for you to use on Kaggle (or elsewhere): <a href=\"https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks\" target=\"_blank\">https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks</a></p>\n<p>To generate these, I overlaid the infrared image, the fragment mask, and the processed X-ray data for each fragment, then removed any regions in the fragment mask that had no X-ray support at all.  This includes regions around the edges where we see down into the support surface or into fragmentary lower levels of the papyrus, as well as cracks inside the fragments where no data is available on X-ray.  These should be usable as drop-in replacements for the masks provided by the organizers.  I'm hoping that if we avoid training on these bad regions that it can help improve model performance.</p>\n<p>Please try them out for model training and let me know if you see any improvements as a result.</p>",
      "rawMarkdown": "After a fair bit of poking around in the data, I've noticed that the training data is not as accurate as one might wish.  There was some discussion of this on the discord, but one example is that there is an annotated ink region on fragment 2 that is in fact visible on infrared, but where the surface of the papyrus that would hold that ink goes out of frame.  So there's no possible way to detect ink there:  the papyrus isn't visible on any of the X-ray slices.\n\nAfter discovering this, I've started putting some work into manually updating the training data to see if this improves performance.  The first set of data updates is complete and freely available for you to use on Kaggle (or elsewhere): https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks\n\nTo generate these, I overlaid the infrared image, the fragment mask, and the processed X-ray data for each fragment, then removed any regions in the fragment mask that had no X-ray support at all.  This includes regions around the edges where we see down into the support surface or into fragmentary lower levels of the papyrus, as well as cracks inside the fragments where no data is available on X-ray.  These should be usable as drop-in replacements for the masks provided by the organizers.  I'm hoping that if we avoid training on these bad regions that it can help improve model performance.\n\nPlease try them out for model training and let me know if you see any improvements as a result.",
      "votes": null
    },
    {
      "id": "2285472",
      "postDate": "06/02/2023 18:25:10",
      "content": "<p>mine was very bad( contrast + closing) , thanks!<br>\nmanually you said, wow crazy</p>",
      "rawMarkdown": "mine was very bad( contrast + closing) , thanks!\nmanually you said, wow crazy",
      "votes": null
    },
    {
      "id": "2285479",
      "postDate": "06/02/2023 18:28:57",
      "content": "<p>\"Data refinement with a highly trained generalizable biological neural network\"</p>",
      "rawMarkdown": "\"Data refinement with a highly trained generalizable biological neural network\"",
      "votes": null
    },
    {
      "id": "2288201",
      "postDate": "06/05/2023 08:22:15",
      "content": "<p>How were these images obtained? Can you share some code?</p>",
      "rawMarkdown": "How were these images obtained? Can you share some code?",
      "votes": null
    },
    {
      "id": "2288779",
      "postDate": "06/05/2023 17:18:43",
      "content": "<p>As I said, they were generated by hand in GIMP.  There was some code involved (which I will not be making public right now) in pre-processing the X-ray data to show the papyrus surfaces, but once that was done it's just:</p>\n<ul>\n<li>Load up the infrared image, the fragment mask, and the X-ray surface data in GIMP</li>\n<li>Put the mask on top, drop the opacity</li>\n<li>Swap back and forth between viewing the other two images as underlayers</li>\n<li>Remove any parts of the mask that are not visibly part of the primary papyrus surface with a pencil brush in GIMP</li>\n<li>Save files</li>\n<li>Repeat for each fragment</li>\n</ul>\n<p>I suppose if you tried, you could do some kind of automated approach to modifying the mask, but honestly my eyes are probably going to do better than most automated approaches.  It's easy enough to verify the accuracy and take a look yourself, and we're not getting more training data so there's not a particular problem with taking a few hours to do the work by hand.</p>",
      "rawMarkdown": "As I said, they were generated by hand in GIMP.  There was some code involved (which I will not be making public right now) in pre-processing the X-ray data to show the papyrus surfaces, but once that was done it's just:\n- Load up the infrared image, the fragment mask, and the X-ray surface data in GIMP\n- Put the mask on top, drop the opacity\n- Swap back and forth between viewing the other two images as underlayers\n- Remove any parts of the mask that are not visibly part of the primary papyrus surface with a pencil brush in GIMP\n- Save files\n- Repeat for each fragment\n\nI suppose if you tried, you could do some kind of automated approach to modifying the mask, but honestly my eyes are probably going to do better than most automated approaches.  It's easy enough to verify the accuracy and take a look yourself, and we're not getting more training data so there's not a particular problem with taking a few hours to do the work by hand.",
      "votes": null
    },
    {
      "id": "2289698",
      "postDate": "06/06/2023 10:13:50",
      "content": "<p>OK! Thank you!</p>",
      "rawMarkdown": "OK! Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2285472,
      "author_name": "iraqbot",
      "author_url": "",
      "post_date": "06/02/2023 18:25:10",
      "content": "<p>mine was very bad( contrast + closing) , thanks!<br>\nmanually you said, wow crazy</p>",
      "votes": null,
      "replies": [
        {
          "id": 2285479,
          "author_name": "brettolsen",
          "author_url": "",
          "post_date": "06/02/2023 18:28:57",
          "content": "<p>\"Data refinement with a highly trained generalizable biological neural network\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2288201,
      "author_name": "wangxuc",
      "author_url": "",
      "post_date": "06/05/2023 08:22:15",
      "content": "<p>How were these images obtained? Can you share some code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2288779,
          "author_name": "brettolsen",
          "author_url": "",
          "post_date": "06/05/2023 17:18:43",
          "content": "<p>As I said, they were generated by hand in GIMP.  There was some code involved (which I will not be making public right now) in pre-processing the X-ray data to show the papyrus surfaces, but once that was done it's just:</p>\n<ul>\n<li>Load up the infrared image, the fragment mask, and the X-ray surface data in GIMP</li>\n<li>Put the mask on top, drop the opacity</li>\n<li>Swap back and forth between viewing the other two images as underlayers</li>\n<li>Remove any parts of the mask that are not visibly part of the primary papyrus surface with a pencil brush in GIMP</li>\n<li>Save files</li>\n<li>Repeat for each fragment</li>\n</ul>\n<p>I suppose if you tried, you could do some kind of automated approach to modifying the mask, but honestly my eyes are probably going to do better than most automated approaches.  It's easy enough to verify the accuracy and take a look yourself, and we're not getting more training data so there's not a particular problem with taking a few hours to do the work by hand.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2289698,
              "author_name": "wangxuc",
              "author_url": "",
              "post_date": "06/06/2023 10:13:50",
              "content": "<p>OK! Thank you!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2285424": "After a fair bit of poking around in the data, I've noticed that the training data is not as accurate as one might wish.  There was some discussion of this on the discord, but one example is that there is an annotated ink region on fragment 2 that is in fact visible on infrared, but where the surface of the papyrus that would hold that ink goes out of frame.  So there's no possible way to detect ink there:  the papyrus isn't visible on any of the X-ray slices.\n\nAfter discovering this, I've started putting some work into manually updating the training data to see if this improves performance.  The first set of data updates is complete and freely available for you to use on Kaggle (or elsewhere): https://www.kaggle.com/datasets/brettolsen/updated-fragment-masks\n\nTo generate these, I overlaid the infrared image, the fragment mask, and the processed X-ray data for each fragment, then removed any regions in the fragment mask that had no X-ray support at all.  This includes regions around the edges where we see down into the support surface or into fragmentary lower levels of the papyrus, as well as cracks inside the fragments where no data is available on X-ray.  These should be usable as drop-in replacements for the masks provided by the organizers.  I'm hoping that if we avoid training on these bad regions that it can help improve model performance.\n\nPlease try them out for model training and let me know if you see any improvements as a result.",
    "2285472": "mine was very bad( contrast + closing) , thanks!\nmanually you said, wow crazy",
    "2285479": "\"Data refinement with a highly trained generalizable biological neural network\"",
    "2288201": "How were these images obtained? Can you share some code?",
    "2288779": "As I said, they were generated by hand in GIMP.  There was some code involved (which I will not be making public right now) in pre-processing the X-ray data to show the papyrus surfaces, but once that was done it's just:\n- Load up the infrared image, the fragment mask, and the X-ray surface data in GIMP\n- Put the mask on top, drop the opacity\n- Swap back and forth between viewing the other two images as underlayers\n- Remove any parts of the mask that are not visibly part of the primary papyrus surface with a pencil brush in GIMP\n- Save files\n- Repeat for each fragment\n\nI suppose if you tried, you could do some kind of automated approach to modifying the mask, but honestly my eyes are probably going to do better than most automated approaches.  It's easy enough to verify the accuracy and take a look yourself, and we're not getting more training data so there's not a particular problem with taking a few hours to do the work by hand.",
    "2289698": "OK! Thank you!"
  },
  "source": "meta"
}