{
  "id": 403348,
  "title": "Color analysis: Ink = Ink + Noise. ROI",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/403348",
  "author_name": "",
  "post_date": "2023-04-22T16:01:22.114585700Z",
  "votes": 50,
  "comment_count": 13,
  "views": 0,
  "content": "<p>We noticed that the boundary pixels behave abnormally and do not contribute well to the color analysis. Therefore, we excluded the boundary and instead focused on the inner part of the fragment.</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233792973-1c53dd47-2ee8-4dd4-b460-0248ab7e3fb1.png\" alt=\"Trimming process\"></p>\n<p>After analyzing the ink's color distribution of the first fragment, it is clear that there are two overlapping normal distributions.</p>\n<p>The first Gaussian-like distribution, ranging from 0-130, probably represents the genuine ink distribution, while the second one, ranging from 0-255, seems to account for noise.</p>\n<p>The noise may originate from the annotation process of the letters, where the annotator likely focused on outlining the shape of each letter rather than capturing every individual pixel. Furthermore, ink degradation and sparsity over time could have also contributed to the noise.</p>\n<h2>Fragment 1</h2>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233792717-2e0b1fd0-692d-40c5-91da-c5be40fb61ba.png\" alt=\"Fragment 1\"></p>\n<h2>Fragment 2</h2>\n<p>Fragment 2 is more challenging and its histogram requires a bit of linear transformation (roughly 0.6* x + 42) processing but ultimately conforms to the same pattern.<br>\n<img src=\"https://user-images.githubusercontent.com/16557697/233793079-9b331fc6-c956-4ad6-974e-ee71bc3abf6b.png\" alt=\"Fragment 2\"></p>\n<p>A few insights:</p>\n<p>1) It is reasonable to assume that pixels of <strong>fragment 1</strong> with a color value greater than 130 are unlikely to be ink pixels and can be ignored.</p>\n<p>2) We can generate a refined ground truth (GT) mask for <strong>fragment 1</strong> by eliminating the papyrus pixels (&gt;130).</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233793741-c2c91ff1-27db-4643-96c5-e63ae3231fd7.png\" alt=\"Refined mask\"></p>\n<p>3) A refined mask for <strong>Fragment 2</strong> needs to be generated considering the inverse transformation corresponding to the mentioned histogram transformation.</p>",
  "messages": [
    {
      "id": "2230650",
      "postDate": "04/22/2023 16:01:22",
      "content": "<p>We noticed that the boundary pixels behave abnormally and do not contribute well to the color analysis. Therefore, we excluded the boundary and instead focused on the inner part of the fragment.</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233792973-1c53dd47-2ee8-4dd4-b460-0248ab7e3fb1.png\" alt=\"Trimming process\"></p>\n<p>After analyzing the ink's color distribution of the first fragment, it is clear that there are two overlapping normal distributions.</p>\n<p>The first Gaussian-like distribution, ranging from 0-130, probably represents the genuine ink distribution, while the second one, ranging from 0-255, seems to account for noise.</p>\n<p>The noise may originate from the annotation process of the letters, where the annotator likely focused on outlining the shape of each letter rather than capturing every individual pixel. Furthermore, ink degradation and sparsity over time could have also contributed to the noise.</p>\n<h2>Fragment 1</h2>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233792717-2e0b1fd0-692d-40c5-91da-c5be40fb61ba.png\" alt=\"Fragment 1\"></p>\n<h2>Fragment 2</h2>\n<p>Fragment 2 is more challenging and its histogram requires a bit of linear transformation (roughly 0.6* x + 42) processing but ultimately conforms to the same pattern.<br>\n<img src=\"https://user-images.githubusercontent.com/16557697/233793079-9b331fc6-c956-4ad6-974e-ee71bc3abf6b.png\" alt=\"Fragment 2\"></p>\n<p>A few insights:</p>\n<p>1) It is reasonable to assume that pixels of <strong>fragment 1</strong> with a color value greater than 130 are unlikely to be ink pixels and can be ignored.</p>\n<p>2) We can generate a refined ground truth (GT) mask for <strong>fragment 1</strong> by eliminating the papyrus pixels (&gt;130).</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233793741-c2c91ff1-27db-4643-96c5-e63ae3231fd7.png\" alt=\"Refined mask\"></p>\n<p>3) A refined mask for <strong>Fragment 2</strong> needs to be generated considering the inverse transformation corresponding to the mentioned histogram transformation.</p>",
      "rawMarkdown": "We noticed that the boundary pixels behave abnormally and do not contribute well to the color analysis. Therefore, we excluded the boundary and instead focused on the inner part of the fragment.\n\n![Trimming process](https://user-images.githubusercontent.com/16557697/233792973-1c53dd47-2ee8-4dd4-b460-0248ab7e3fb1.png)\n\nAfter analyzing the ink's color distribution of the first fragment, it is clear that there are two overlapping normal distributions.\n\nThe first Gaussian-like distribution, ranging from 0-130, probably represents the genuine ink distribution, while the second one, ranging from 0-255, seems to account for noise.\n\nThe noise may originate from the annotation process of the letters, where the annotator likely focused on outlining the shape of each letter rather than capturing every individual pixel. Furthermore, ink degradation and sparsity over time could have also contributed to the noise.\n\n\n## Fragment 1\n![Fragment 1](https://user-images.githubusercontent.com/16557697/233792717-2e0b1fd0-692d-40c5-91da-c5be40fb61ba.png)\n\n## Fragment 2\n\nFragment 2 is more challenging and its histogram requires a bit of linear transformation (roughly 0.6* x + 42) processing but ultimately conforms to the same pattern.\n![Fragment 2](https://user-images.githubusercontent.com/16557697/233793079-9b331fc6-c956-4ad6-974e-ee71bc3abf6b.png)\n\n\nA few insights:\n\n1) It is reasonable to assume that pixels of **fragment 1** with a color value greater than 130 are unlikely to be ink pixels and can be ignored.\n\n2) We can generate a refined ground truth (GT) mask for **fragment 1** by eliminating the papyrus pixels (>130).\n\n![Refined mask](https://user-images.githubusercontent.com/16557697/233793741-c2c91ff1-27db-4643-96c5-e63ae3231fd7.png)\n\n3) A refined mask for **Fragment 2** needs to be generated considering the inverse transformation corresponding to the mentioned histogram transformation.",
      "votes": null
    },
    {
      "id": "2230833",
      "postDate": "04/22/2023 19:27:21",
      "content": "<h2>Fragment 3</h2>\n<p>The same pattern applies to Fragment 3, where ink and noise can be distinguished. The spectrum of ink colors spans from 0 to 140.</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233802053-e74b36ac-1497-49dd-803b-74d1762f3a38.png\" alt=\"Fragment 3\"></p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233802776-1aab36bb-f912-448c-94d7-abdb9e28a27d.gif\" alt=\"GIF\"></p>",
      "rawMarkdown": "## Fragment 3\n\nThe same pattern applies to Fragment 3, where ink and noise can be distinguished. The spectrum of ink colors spans from 0 to 140.\n\n![Fragment 3](https://user-images.githubusercontent.com/16557697/233802053-e74b36ac-1497-49dd-803b-74d1762f3a38.png)\n\n\n![GIF](https://user-images.githubusercontent.com/16557697/233802776-1aab36bb-f912-448c-94d7-abdb9e28a27d.gif)",
      "votes": null
    },
    {
      "id": "2231499",
      "postDate": "04/23/2023 11:29:21",
      "content": "<p>This is all very interesting. An obvious but important next step for anyone reading this is to see how this mask refinement affects model performance. This could probably be done with just about any working model, I would imagine.</p>\n<p>A few technical notes/questions:</p>\n<ul>\n<li>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.</li>\n<li>The border shrinking seems fairly extreme. Did you try with thinner borders? I'm curious when the noise floor becomes too high. </li>\n<li>It would also be interesting to see how this total effect is observed across characters. We have noticed that not all characters exhibit the same response in the baseline performance.</li>\n<li>The data is 16bpc, but you're working with a 256 bin histogram. I'm curious if the trend is so cleanly visible at full dynamic range or if it needs the aggregation from being binned down. </li>\n</ul>",
      "rawMarkdown": "This is all very interesting. An obvious but important next step for anyone reading this is to see how this mask refinement affects model performance. This could probably be done with just about any working model, I would imagine.\n\nA few technical notes/questions:\n- What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.\n- The border shrinking seems fairly extreme. Did you try with thinner borders? I'm curious when the noise floor becomes too high. \n- It would also be interesting to see how this total effect is observed across characters. We have noticed that not all characters exhibit the same response in the baseline performance.\n- The data is 16bpc, but you're working with a 256 bin histogram. I'm curious if the trend is so cleanly visible at full dynamic range or if it needs the aggregation from being binned down.",
      "votes": null
    },
    {
      "id": "2235071",
      "postDate": "04/25/2023 17:45:48",
      "content": "<p>We can reconstruct the \"true\" papyrus histogram by moving \"false\" ink pixel frequencies to papyrus pixel frequencies. As a result, the papyrus histogram appears as it should (roughly resembling a normal distribution).</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/234257216-62c25fd2-6218-462d-8308-1cba7f268f63.png\" alt=\"Papyrus Reconstruction\"></p>\n<p>The distance <code>d</code> between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified. A larger distance between these peaks signifies a more significant contrast between the two features.</p>\n<p>When the contrast between the objects is higher, it becomes easier for the human eye and/or neural network to differentiate between them. This is because the color or intensity values of the two objects are more distinct from each other, allowing for better visibility and easier separation.</p>\n<p>A more detailed analysis of all slices leads to an interesting observation. This pattern is not present in every slide. In fact, each fragment has its range of slices where the distance 'd' is sufficient. It is important to note that if both peaks are close, it indicates that the ink is so deeply buried in the noise that it becomes undetectable.</p>\n<p>It seems reasonable that the detectable ink would appear in various locations across different scans. I managed to determine the following <code>good</code> slices. Some are significantly better than others, and I believe that selecting the optimal range for each slice is crucial in this competition.</p>\n<table>\n<thead>\n<tr>\n<th>Fragment</th>\n<th>Slices</th>\n<th>Ink Peak</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fragment 1</td>\n<td>21-34</td>\n<td>65</td>\n</tr>\n<tr>\n<td>Fragment 2</td>\n<td>25-38</td>\n<td>88</td>\n</tr>\n<tr>\n<td>Fragment 3</td>\n<td>20-33</td>\n<td>77</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "We can reconstruct the \"true\" papyrus histogram by moving \"false\" ink pixel frequencies to papyrus pixel frequencies. As a result, the papyrus histogram appears as it should (roughly resembling a normal distribution).\n\n\n![Papyrus Reconstruction](https://user-images.githubusercontent.com/16557697/234257216-62c25fd2-6218-462d-8308-1cba7f268f63.png)\n\n\nThe distance `d` between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified. A larger distance between these peaks signifies a more significant contrast between the two features.\n\nWhen the contrast between the objects is higher, it becomes easier for the human eye and/or neural network to differentiate between them. This is because the color or intensity values of the two objects are more distinct from each other, allowing for better visibility and easier separation.\n\nA more detailed analysis of all slices leads to an interesting observation. This pattern is not present in every slide. In fact, each fragment has its range of slices where the distance 'd' is sufficient. It is important to note that if both peaks are close, it indicates that the ink is so deeply buried in the noise that it becomes undetectable.\n\nIt seems reasonable that the detectable ink would appear in various locations across different scans. I managed to determine the following `good` slices. Some are significantly better than others, and I believe that selecting the optimal range for each slice is crucial in this competition.\n\n| Fragment    | Slices  | Ink Peak      |\n|-------------|---------|-----------|\n| Fragment 1  | 21-34   | 65 |\n| Fragment 2  | 25-38   | 88  |\n| Fragment 3  | 20-33   | 77  |",
      "votes": null
    },
    {
      "id": "2235077",
      "postDate": "04/25/2023 17:52:00",
      "content": "<p>Thank you for your remarks. I truly appreciate them.</p>\n<blockquote>\n  <p>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.</p>\n</blockquote>\n<p>It appears that every fragment has a unique range in which the inks are visible. A thorough analysis of all the slices may help identify the optimal range for each fragment. Kindly refer to my comment below for additional information</p>",
      "rawMarkdown": "Thank you for your remarks. I truly appreciate them.\n\n>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.\n\nIt appears that every fragment has a unique range in which the inks are visible. A thorough analysis of all the slices may help identify the optimal range for each fragment. Kindly refer to my comment below for additional information",
      "votes": null
    },
    {
      "id": "2235081",
      "postDate": "04/25/2023 17:58:49",
      "content": "<p>thanks a lot ! i was trying to find the best slices randomly ! </p>",
      "rawMarkdown": "thanks a lot ! i was trying to find the best slices randomly !",
      "votes": null
    },
    {
      "id": "2237178",
      "postDate": "04/27/2023 12:32:41",
      "content": "<p>You would see a thin edge when viewing a fragment from a lateral perspective. This animation demonstrates how fragment 1 appears from this viewpoint. It also shows the corresponding 1D ink mask. The ROI (Region of Interest) of the fragment (21-34) is highlighted.\"</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/234857771-00283d9a-af24-406d-98e5-f1175dd25a4f.gif\" alt=\"fragment1\"></p>\n<p>It's interesting if the neural net can learn from this 1D input rather than using 2D patches. If so, we could use all 65 channels with coordY convolution to assist the neural network in distinguishing between the channels.</p>",
      "rawMarkdown": "You would see a thin edge when viewing a fragment from a lateral perspective. This animation demonstrates how fragment 1 appears from this viewpoint. It also shows the corresponding 1D ink mask. The ROI (Region of Interest) of the fragment (21-34) is highlighted.\"\n\n![fragment1](https://user-images.githubusercontent.com/16557697/234857771-00283d9a-af24-406d-98e5-f1175dd25a4f.gif)\n\nIt's interesting if the neural net can learn from this 1D input rather than using 2D patches. If so, we could use all 65 channels with coordY convolution to assist the neural network in distinguishing between the channels.",
      "votes": null
    },
    {
      "id": "2240436",
      "postDate": "04/30/2023 14:34:19",
      "content": "<p>It is very interesting. However my analysis of the histogram doesn't show any menaingful difference in distribution. It would be great if you could provide some code and explanation of how did you receive these results.<br>\nThere are also something that I would like you to elaborate:</p>\n<ol>\n<li>Whenever you say <em>channel</em> - do you assume slice number?</li>\n<li>The gray-levels are in range 0..65535 for fragment 1, did you normalize them to 0..255 or you are using some different source of data?<br>\nBelow are two examples of histogram of fragment 1 for slices 26 and 30. Images are cropped to remove boundary pixels.<br>\n<strong>fragment 1, slice 26:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2F8854476a54dbd9445b426dad5ee1079a%2Fhistogram%20f%201%20s%2026.jpg?generation=1682867802742448&amp;alt=media\" alt=\"\"><br>\n<strong>fragment 1, slice 30:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2Fef39e1d7ea39ef3223e609ff9f1cb3bb%2Fhistogram%20f%201%20s%2030.jpg?generation=1682867821586031&amp;alt=media\" alt=\"\"></li>\n</ol>",
      "rawMarkdown": "It is very interesting. However my analysis of the histogram doesn't show any menaingful difference in distribution. It would be great if you could provide some code and explanation of how did you receive these results.\nThere are also something that I would like you to elaborate:\n1. Whenever you say *channel* - do you assume slice number?\n2. The gray-levels are in range 0..65535 for fragment 1, did you normalize them to 0..255 or you are using some different source of data?\nBelow are two examples of histogram of fragment 1 for slices 26 and 30. Images are cropped to remove boundary pixels.\n**fragment 1, slice 26:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2F8854476a54dbd9445b426dad5ee1079a%2Fhistogram%20f%201%20s%2026.jpg?generation=1682867802742448&alt=media)\n**fragment 1, slice 30:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2Fef39e1d7ea39ef3223e609ff9f1cb3bb%2Fhistogram%20f%201%20s%2030.jpg?generation=1682867821586031&alt=media)",
      "votes": null
    },
    {
      "id": "2240522",
      "postDate": "04/30/2023 15:52:14",
      "content": "<p><a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> nice post and interesting insights! 👍 Thanks for sharing!🤝</p>",
      "rawMarkdown": "pavelgonchar nice post and interesting insights! 👍 Thanks for sharing!🤝",
      "votes": null
    },
    {
      "id": "2240864",
      "postDate": "04/30/2023 23:25:34",
      "content": "<p>I believe that 2 components of the signal that we see are scroll and air. There is no distinguishable ink signal.<br>\nAlso there is a high chance that scrolls are scanned from the same direction from which infrared picture is taken, meaning layers 0-10 actually contain most ink information if any present at all.<br>\nI think only features around bottom of the scroll that are useful for letter detection are rare geometrical deformations of scroll surface from excessive pressure applied by tool used to apply ink on it.<br>\nThis deformation is probably what most notebooks focusing on layers 20-45 are learning to recognise.</p>",
      "rawMarkdown": "I believe that 2 components of the signal that we see are scroll and air. There is no distinguishable ink signal.\nAlso there is a high chance that scrolls are scanned from the same direction from which infrared picture is taken, meaning layers 0-10 actually contain most ink information if any present at all.\nI think only features around bottom of the scroll that are useful for letter detection are rare geometrical deformations of scroll surface from excessive pressure applied by tool used to apply ink on it.\nThis deformation is probably what most notebooks focusing on layers 20-45 are learning to recognise.",
      "votes": null
    },
    {
      "id": "2247008",
      "postDate": "05/05/2023 16:02:08",
      "content": "<p>The EDA is really good! But how do we utilise it in training. As you said most ink pixels are beween 0-130 is it any usefull fore some sort of preprocessing? because i tried using only those pixels but didnt work</p>",
      "rawMarkdown": "The EDA is really good! But how do we utilise it in training. As you said most ink pixels are beween 0-130 is it any usefull fore some sort of preprocessing? because i tried using only those pixels but didnt work",
      "votes": null
    },
    {
      "id": "2247810",
      "postDate": "05/06/2023 09:38:21",
      "content": "<p>it is \"obvious\" that the pixel label are not correct.<br>\nbelow shows how the model predicts as training iterations proceed.</p>\n<p>[url=<a href=\"https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url\" target=\"_blank\">https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url</a>]<br>\n<img src=\"https://i.ibb.co/Ms3qv86/Selection-999-1945.png\" alt=\"https://i.ibb.co/Ms3qv86/Selection-999-1945.png\"></p>\n<p>you can see the predictions are flip flopping, i.e. the model \"sometimes learn some background as link\"</p>\n<p>you can relabel using the intensity histogram as discussed or some semi-supervised methods (i.e. selective pesudo label)</p>",
      "rawMarkdown": "it is \"obvious\" that the pixel label are not correct.\nbelow shows how the model predicts as training iterations proceed.\n\n[url=https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url]\n![https://i.ibb.co/Ms3qv86/Selection-999-1945.png](https://i.ibb.co/Ms3qv86/Selection-999-1945.png)\n\nyou can see the predictions are flip flopping, i.e. the model \"sometimes learn some background as link\"\n\nyou can relabel using the intensity histogram as discussed or some semi-supervised methods (i.e. selective pesudo label)",
      "votes": null
    },
    {
      "id": "2269304",
      "postDate": "05/22/2023 11:17:36",
      "content": "<p>what if the private dataset is labelled in the same way (I guess it is) ?<br>\nwouldn't it be better to let the model learn labels 'as is', which in a sense 'fills the letters' in the same way as in the private set?<br>\nwouldn't it be harder to first find the 'corrected ink labels' and then to postprocess to fill the letters, rather then let the model learn all at once? <br>\nI am guessing,</p>",
      "rawMarkdown": "what if the private dataset is labelled in the same way (I guess it is) ?\nwouldn't it be better to let the model learn labels 'as is', which in a sense 'fills the letters' in the same way as in the private set?\nwouldn't it be harder to first find the 'corrected ink labels' and then to postprocess to fill the letters, rather then let the model learn all at once? \nI am guessing,",
      "votes": null
    },
    {
      "id": "2272242",
      "postDate": "05/24/2023 11:11:14",
      "content": "<p>If the model can extract abstract features, does this kind of \"label noise\" successfully compliment by the model's expressiveness? I have a feeling of the difficulty in training with \"noisy\" label is simply coming from the lack of expressiveness of the model.</p>\n<p>In other viewpoint, we can consider this competition's label is \"weak\" label. The provided label only depicts corse segmentation map, instead of precise saliency map.</p>",
      "rawMarkdown": "If the model can extract abstract features, does this kind of \"label noise\" successfully compliment by the model's expressiveness? I have a feeling of the difficulty in training with \"noisy\" label is simply coming from the lack of expressiveness of the model.\n\nIn other viewpoint, we can consider this competition's label is \"weak\" label. The provided label only depicts corse segmentation map, instead of precise saliency map.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2230833,
      "author_name": "pavelgonchar",
      "author_url": "",
      "post_date": "04/22/2023 19:27:21",
      "content": "<h2>Fragment 3</h2>\n<p>The same pattern applies to Fragment 3, where ink and noise can be distinguished. The spectrum of ink colors spans from 0 to 140.</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233802053-e74b36ac-1497-49dd-803b-74d1762f3a38.png\" alt=\"Fragment 3\"></p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/233802776-1aab36bb-f912-448c-94d7-abdb9e28a27d.gif\" alt=\"GIF\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2235071,
          "author_name": "pavelgonchar",
          "author_url": "",
          "post_date": "04/25/2023 17:45:48",
          "content": "<p>We can reconstruct the \"true\" papyrus histogram by moving \"false\" ink pixel frequencies to papyrus pixel frequencies. As a result, the papyrus histogram appears as it should (roughly resembling a normal distribution).</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/234257216-62c25fd2-6218-462d-8308-1cba7f268f63.png\" alt=\"Papyrus Reconstruction\"></p>\n<p>The distance <code>d</code> between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified. A larger distance between these peaks signifies a more significant contrast between the two features.</p>\n<p>When the contrast between the objects is higher, it becomes easier for the human eye and/or neural network to differentiate between them. This is because the color or intensity values of the two objects are more distinct from each other, allowing for better visibility and easier separation.</p>\n<p>A more detailed analysis of all slices leads to an interesting observation. This pattern is not present in every slide. In fact, each fragment has its range of slices where the distance 'd' is sufficient. It is important to note that if both peaks are close, it indicates that the ink is so deeply buried in the noise that it becomes undetectable.</p>\n<p>It seems reasonable that the detectable ink would appear in various locations across different scans. I managed to determine the following <code>good</code> slices. Some are significantly better than others, and I believe that selecting the optimal range for each slice is crucial in this competition.</p>\n<table>\n<thead>\n<tr>\n<th>Fragment</th>\n<th>Slices</th>\n<th>Ink Peak</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Fragment 1</td>\n<td>21-34</td>\n<td>65</td>\n</tr>\n<tr>\n<td>Fragment 2</td>\n<td>25-38</td>\n<td>88</td>\n</tr>\n<tr>\n<td>Fragment 3</td>\n<td>20-33</td>\n<td>77</td>\n</tr>\n</tbody>\n</table>",
          "votes": null,
          "replies": [
            {
              "id": 2235081,
              "author_name": "iraqbot",
              "author_url": "",
              "post_date": "04/25/2023 17:58:49",
              "content": "<p>thanks a lot ! i was trying to find the best slices randomly ! </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2231499,
      "author_name": "csparker",
      "author_url": "",
      "post_date": "04/23/2023 11:29:21",
      "content": "<p>This is all very interesting. An obvious but important next step for anyone reading this is to see how this mask refinement affects model performance. This could probably be done with just about any working model, I would imagine.</p>\n<p>A few technical notes/questions:</p>\n<ul>\n<li>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.</li>\n<li>The border shrinking seems fairly extreme. Did you try with thinner borders? I'm curious when the noise floor becomes too high. </li>\n<li>It would also be interesting to see how this total effect is observed across characters. We have noticed that not all characters exhibit the same response in the baseline performance.</li>\n<li>The data is 16bpc, but you're working with a 256 bin histogram. I'm curious if the trend is so cleanly visible at full dynamic range or if it needs the aggregation from being binned down. </li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2235077,
          "author_name": "pavelgonchar",
          "author_url": "",
          "post_date": "04/25/2023 17:52:00",
          "content": "<p>Thank you for your remarks. I truly appreciate them.</p>\n<blockquote>\n  <p>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.</p>\n</blockquote>\n<p>It appears that every fragment has a unique range in which the inks are visible. A thorough analysis of all the slices may help identify the optimal range for each fragment. Kindly refer to my comment below for additional information</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2237178,
      "author_name": "pavelgonchar",
      "author_url": "",
      "post_date": "04/27/2023 12:32:41",
      "content": "<p>You would see a thin edge when viewing a fragment from a lateral perspective. This animation demonstrates how fragment 1 appears from this viewpoint. It also shows the corresponding 1D ink mask. The ROI (Region of Interest) of the fragment (21-34) is highlighted.\"</p>\n<p><img src=\"https://user-images.githubusercontent.com/16557697/234857771-00283d9a-af24-406d-98e5-f1175dd25a4f.gif\" alt=\"fragment1\"></p>\n<p>It's interesting if the neural net can learn from this 1D input rather than using 2D patches. If so, we could use all 65 channels with coordY convolution to assist the neural network in distinguishing between the channels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2240436,
      "author_name": "yurikreinin",
      "author_url": "",
      "post_date": "04/30/2023 14:34:19",
      "content": "<p>It is very interesting. However my analysis of the histogram doesn't show any menaingful difference in distribution. It would be great if you could provide some code and explanation of how did you receive these results.<br>\nThere are also something that I would like you to elaborate:</p>\n<ol>\n<li>Whenever you say <em>channel</em> - do you assume slice number?</li>\n<li>The gray-levels are in range 0..65535 for fragment 1, did you normalize them to 0..255 or you are using some different source of data?<br>\nBelow are two examples of histogram of fragment 1 for slices 26 and 30. Images are cropped to remove boundary pixels.<br>\n<strong>fragment 1, slice 26:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2F8854476a54dbd9445b426dad5ee1079a%2Fhistogram%20f%201%20s%2026.jpg?generation=1682867802742448&amp;alt=media\" alt=\"\"><br>\n<strong>fragment 1, slice 30:</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2Fef39e1d7ea39ef3223e609ff9f1cb3bb%2Fhistogram%20f%201%20s%2030.jpg?generation=1682867821586031&amp;alt=media\" alt=\"\"></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 2240864,
          "author_name": "elvenmonk",
          "author_url": "",
          "post_date": "04/30/2023 23:25:34",
          "content": "<p>I believe that 2 components of the signal that we see are scroll and air. There is no distinguishable ink signal.<br>\nAlso there is a high chance that scrolls are scanned from the same direction from which infrared picture is taken, meaning layers 0-10 actually contain most ink information if any present at all.<br>\nI think only features around bottom of the scroll that are useful for letter detection are rare geometrical deformations of scroll surface from excessive pressure applied by tool used to apply ink on it.<br>\nThis deformation is probably what most notebooks focusing on layers 20-45 are learning to recognise.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2240522,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "04/30/2023 15:52:14",
      "content": "<p><a href=\"https://www.kaggle.com/pavelgonchar\" target=\"_blank\">@pavelgonchar</a> nice post and interesting insights! 👍 Thanks for sharing!🤝</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247008,
      "author_name": "khailashsanthakumar",
      "author_url": "",
      "post_date": "05/05/2023 16:02:08",
      "content": "<p>The EDA is really good! But how do we utilise it in training. As you said most ink pixels are beween 0-130 is it any usefull fore some sort of preprocessing? because i tried using only those pixels but didnt work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2247810,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "05/06/2023 09:38:21",
      "content": "<p>it is \"obvious\" that the pixel label are not correct.<br>\nbelow shows how the model predicts as training iterations proceed.</p>\n<p>[url=<a href=\"https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url\" target=\"_blank\">https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url</a>]<br>\n<img src=\"https://i.ibb.co/Ms3qv86/Selection-999-1945.png\" alt=\"https://i.ibb.co/Ms3qv86/Selection-999-1945.png\"></p>\n<p>you can see the predictions are flip flopping, i.e. the model \"sometimes learn some background as link\"</p>\n<p>you can relabel using the intensity histogram as discussed or some semi-supervised methods (i.e. selective pesudo label)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2269304,
          "author_name": "abdulkadirguner",
          "author_url": "",
          "post_date": "05/22/2023 11:17:36",
          "content": "<p>what if the private dataset is labelled in the same way (I guess it is) ?<br>\nwouldn't it be better to let the model learn labels 'as is', which in a sense 'fills the letters' in the same way as in the private set?<br>\nwouldn't it be harder to first find the 'corrected ink labels' and then to postprocess to fill the letters, rather then let the model learn all at once? <br>\nI am guessing,</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2272242,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/24/2023 11:11:14",
      "content": "<p>If the model can extract abstract features, does this kind of \"label noise\" successfully compliment by the model's expressiveness? I have a feeling of the difficulty in training with \"noisy\" label is simply coming from the lack of expressiveness of the model.</p>\n<p>In other viewpoint, we can consider this competition's label is \"weak\" label. The provided label only depicts corse segmentation map, instead of precise saliency map.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2230650": "We noticed that the boundary pixels behave abnormally and do not contribute well to the color analysis. Therefore, we excluded the boundary and instead focused on the inner part of the fragment.\n\n![Trimming process](https://user-images.githubusercontent.com/16557697/233792973-1c53dd47-2ee8-4dd4-b460-0248ab7e3fb1.png)\n\nAfter analyzing the ink's color distribution of the first fragment, it is clear that there are two overlapping normal distributions.\n\nThe first Gaussian-like distribution, ranging from 0-130, probably represents the genuine ink distribution, while the second one, ranging from 0-255, seems to account for noise.\n\nThe noise may originate from the annotation process of the letters, where the annotator likely focused on outlining the shape of each letter rather than capturing every individual pixel. Furthermore, ink degradation and sparsity over time could have also contributed to the noise.\n\n\n## Fragment 1\n![Fragment 1](https://user-images.githubusercontent.com/16557697/233792717-2e0b1fd0-692d-40c5-91da-c5be40fb61ba.png)\n\n## Fragment 2\n\nFragment 2 is more challenging and its histogram requires a bit of linear transformation (roughly 0.6* x + 42) processing but ultimately conforms to the same pattern.\n![Fragment 2](https://user-images.githubusercontent.com/16557697/233793079-9b331fc6-c956-4ad6-974e-ee71bc3abf6b.png)\n\n\nA few insights:\n\n1) It is reasonable to assume that pixels of **fragment 1** with a color value greater than 130 are unlikely to be ink pixels and can be ignored.\n\n2) We can generate a refined ground truth (GT) mask for **fragment 1** by eliminating the papyrus pixels (>130).\n\n![Refined mask](https://user-images.githubusercontent.com/16557697/233793741-c2c91ff1-27db-4643-96c5-e63ae3231fd7.png)\n\n3) A refined mask for **Fragment 2** needs to be generated considering the inverse transformation corresponding to the mentioned histogram transformation.",
    "2230833": "## Fragment 3\n\nThe same pattern applies to Fragment 3, where ink and noise can be distinguished. The spectrum of ink colors spans from 0 to 140.\n\n![Fragment 3](https://user-images.githubusercontent.com/16557697/233802053-e74b36ac-1497-49dd-803b-74d1762f3a38.png)\n\n\n![GIF](https://user-images.githubusercontent.com/16557697/233802776-1aab36bb-f912-448c-94d7-abdb9e28a27d.gif)",
    "2231499": "This is all very interesting. An obvious but important next step for anyone reading this is to see how this mask refinement affects model performance. This could probably be done with just about any working model, I would imagine.\n\nA few technical notes/questions:\n- What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.\n- The border shrinking seems fairly extreme. Did you try with thinner borders? I'm curious when the noise floor becomes too high. \n- It would also be interesting to see how this total effect is observed across characters. We have noticed that not all characters exhibit the same response in the baseline performance.\n- The data is 16bpc, but you're working with a 256 bin histogram. I'm curious if the trend is so cleanly visible at full dynamic range or if it needs the aggregation from being binned down.",
    "2235071": "We can reconstruct the \"true\" papyrus histogram by moving \"false\" ink pixel frequencies to papyrus pixel frequencies. As a result, the papyrus histogram appears as it should (roughly resembling a normal distribution).\n\n\n![Papyrus Reconstruction](https://user-images.githubusercontent.com/16557697/234257216-62c25fd2-6218-462d-8308-1cba7f268f63.png)\n\n\nThe distance `d` between the peak of the ink (64) and the peak of the refined papyrus (120) shows how recognizable the letters are. The larger this distance, the more easily the letters can be identified. A larger distance between these peaks signifies a more significant contrast between the two features.\n\nWhen the contrast between the objects is higher, it becomes easier for the human eye and/or neural network to differentiate between them. This is because the color or intensity values of the two objects are more distinct from each other, allowing for better visibility and easier separation.\n\nA more detailed analysis of all slices leads to an interesting observation. This pattern is not present in every slide. In fact, each fragment has its range of slices where the distance 'd' is sufficient. It is important to note that if both peaks are close, it indicates that the ink is so deeply buried in the noise that it becomes undetectable.\n\nIt seems reasonable that the detectable ink would appear in various locations across different scans. I managed to determine the following `good` slices. Some are significantly better than others, and I believe that selecting the optimal range for each slice is crucial in this competition.\n\n| Fragment    | Slices  | Ink Peak      |\n|-------------|---------|-----------|\n| Fragment 1  | 21-34   | 65 |\n| Fragment 2  | 25-38   | 88  |\n| Fragment 3  | 20-33   | 77  |",
    "2235077": "Thank you for your remarks. I truly appreciate them.\n\n>What are you calculating the histogram over? The full image stack for each fragment? Some subset of that? In the same way you shrank the boundary of the fragment, it would be interesting to shrink the image stack on the ends and see the effect on the signal.\n\nIt appears that every fragment has a unique range in which the inks are visible. A thorough analysis of all the slices may help identify the optimal range for each fragment. Kindly refer to my comment below for additional information",
    "2235081": "thanks a lot ! i was trying to find the best slices randomly !",
    "2237178": "You would see a thin edge when viewing a fragment from a lateral perspective. This animation demonstrates how fragment 1 appears from this viewpoint. It also shows the corresponding 1D ink mask. The ROI (Region of Interest) of the fragment (21-34) is highlighted.\"\n\n![fragment1](https://user-images.githubusercontent.com/16557697/234857771-00283d9a-af24-406d-98e5-f1175dd25a4f.gif)\n\nIt's interesting if the neural net can learn from this 1D input rather than using 2D patches. If so, we could use all 65 channels with coordY convolution to assist the neural network in distinguishing between the channels.",
    "2240436": "It is very interesting. However my analysis of the histogram doesn't show any menaingful difference in distribution. It would be great if you could provide some code and explanation of how did you receive these results.\nThere are also something that I would like you to elaborate:\n1. Whenever you say *channel* - do you assume slice number?\n2. The gray-levels are in range 0..65535 for fragment 1, did you normalize them to 0..255 or you are using some different source of data?\nBelow are two examples of histogram of fragment 1 for slices 26 and 30. Images are cropped to remove boundary pixels.\n**fragment 1, slice 26:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2F8854476a54dbd9445b426dad5ee1079a%2Fhistogram%20f%201%20s%2026.jpg?generation=1682867802742448&alt=media)\n**fragment 1, slice 30:**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5613794%2Fef39e1d7ea39ef3223e609ff9f1cb3bb%2Fhistogram%20f%201%20s%2030.jpg?generation=1682867821586031&alt=media)",
    "2240522": "pavelgonchar nice post and interesting insights! 👍 Thanks for sharing!🤝",
    "2240864": "I believe that 2 components of the signal that we see are scroll and air. There is no distinguishable ink signal.\nAlso there is a high chance that scrolls are scanned from the same direction from which infrared picture is taken, meaning layers 0-10 actually contain most ink information if any present at all.\nI think only features around bottom of the scroll that are useful for letter detection are rare geometrical deformations of scroll surface from excessive pressure applied by tool used to apply ink on it.\nThis deformation is probably what most notebooks focusing on layers 20-45 are learning to recognise.",
    "2247008": "The EDA is really good! But how do we utilise it in training. As you said most ink pixels are beween 0-130 is it any usefull fore some sort of preprocessing? because i tried using only those pixels but didnt work",
    "2247810": "it is \"obvious\" that the pixel label are not correct.\nbelow shows how the model predicts as training iterations proceed.\n\n[url=https://ibb.co/c17zBbL][img]https://i.ibb.co/XbvmQy8/Selection-999-1945.png[/img][/url]\n![https://i.ibb.co/Ms3qv86/Selection-999-1945.png](https://i.ibb.co/Ms3qv86/Selection-999-1945.png)\n\nyou can see the predictions are flip flopping, i.e. the model \"sometimes learn some background as link\"\n\nyou can relabel using the intensity histogram as discussed or some semi-supervised methods (i.e. selective pesudo label)",
    "2269304": "what if the private dataset is labelled in the same way (I guess it is) ?\nwouldn't it be better to let the model learn labels 'as is', which in a sense 'fills the letters' in the same way as in the private set?\nwouldn't it be harder to first find the 'corrected ink labels' and then to postprocess to fill the letters, rather then let the model learn all at once? \nI am guessing,",
    "2272242": "If the model can extract abstract features, does this kind of \"label noise\" successfully compliment by the model's expressiveness? I have a feeling of the difficulty in training with \"noisy\" label is simply coming from the lack of expressiveness of the model.\n\nIn other viewpoint, we can consider this competition's label is \"weak\" label. The provided label only depicts corse segmentation map, instead of precise saliency map."
  },
  "source": "meta"
}