{
  "id": 395676,
  "title": "Analogizing Ink detection to other problem domains",
  "url": "/competitions/vesuvius-challenge-ink-detection/discussion/395676",
  "author_name": "",
  "post_date": "2023-03-18T08:23:02.256945700Z",
  "votes": 50,
  "comment_count": 9,
  "views": 0,
  "content": "<p>There isn't much research about ink detection from papyrus CT images in the literature. So contesters in this competition will have a hard time simply using a state-of-the-art network for this task. But that doesn't mean we have to start from scratch. We can instead borrow ideas from similar problems that have received more attention. To find good analogies, we must first look at identifying characteristics of this problem and seek problems that have most of these characteristics in common, while adapting the solution where the problems differ. Characteristics of this problem include:</p>\n<p>Nature of the data:</p>\n<ul>\n<li>3d image data. More specifically, a CT image with a single channel (i.e. grayscale image)</li>\n<li>(extremely) Limited number of images. We only have 3 train images in the dataset.</li>\n<li>High resolution images</li>\n</ul>\n<p>Nature of the inference target:</p>\n<ul>\n<li>We are looking to classify each pixel as ink/no ink. In the literature, this would be called an image segmentation or pixel classification problem</li>\n<li>There is a single class.</li>\n<li>While our input is a 3d image, the target is a 2d mask. </li>\n</ul>\n<p>Nature of the relationship between the two:</p>\n<ul>\n<li>Not all spatial dimensions are created equal. While the x and y dimensions in this problem are translation invariant (it doesn't matter where the ink is on the papyrus), the depth dimension is not. Deeper voxels are less likely to have ink in them, while shallower voxels are more likely to have other materials or defects tarnishing them.</li>\n<li>The ink forms letters. Letters have spatial relationships, so a pixel with ink is likely to be close to other ink pixels. Meanwhile, there are many letters in the image. Compared to other analogs, letters are relatively small-scale structures, and taking them into account does not require analyzing the whole image.</li>\n</ul>\n<p>Looking at these characteristics, the closest analog that I have found is segmentation of medical images. Specifically, tumor segmentation, lung segmentation, bone segmentation, etc are all valid analogs. We have CT images as input, and a small amount of labeled images. The target is slightly different. While this challenge creates a 2d mask, medical image segmentation is typically a 3d one. However, this difference can be adapted by pooling or otherwise distilling the depth into a single value for each pixel. We will also need to keep in mind the differences in the relationship when seeking to import a solution. We don't need to care about large-scale structure as much as those problems do.</p>\n<p>Here is a list of references to get you started on seeking analogs:</p>\n<ul>\n<li>Section 2.2 of this <a href=\"https://www.sciencedirect.com/science/article/pii/B9780128181010000033#s0025\" target=\"_blank\">book</a> reviews different <br>\nneural networks used for tumor segmentation</li>\n<li>This <a href=\"https://bmcmedimaging.biomedcentral.com/articles/10.1186/s12880-020-00529-5#Sec9\" target=\"_blank\">paper</a> applies two neural network architectures to segment lung images as infected/not infected by COVID-19.</li>\n<li>This <a href=\"https://www.frontiersin.org/articles/10.3389/frobt.2020.00106/full#B41\" target=\"_blank\">paper</a> Applies deep supervision and attention to improve tumor segmentation.</li>\n</ul>\n<p>Here are some architectures mentioned in these papers:</p>\n<ul>\n<li><a href=\"https://www.cv-foundation.org/openaccess/content_cvpr_2015/papers/Long_Fully_Convolutional_Networks_2015_CVPR_paper.pdf\" target=\"_blank\">Fully-Connected CNNs</a></li>\n<li><a href=\"https://arxiv.org/abs/1505.04597\" target=\"_blank\">U-net</a>, a type of FC-CNN</li>\n<li><a href=\"https://arxiv.org/abs/1511.00561\" target=\"_blank\">Segnet</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/28688283/\" target=\"_blank\">Central-Focused CNN</a></li>\n<li><a href=\"https://arxiv.org/abs/1505.00880\" target=\"_blank\">MV-CNN</a></li>\n<li><a href=\"https://arxiv.org/abs/1409.5185\" target=\"_blank\">Deeply Supervised Nets</a></li>\n</ul>\n<p>Edit: I did some further research. Here are some more references:</p>\n<ul>\n<li><a href=\"https://link.springer.com/chapter/10.1007/978-3-319-46723-8_49\" target=\"_blank\">3D U-net</a> applied to medical scans with 2d label inputs. Note that as opposed to the problem here, this paper outputs 3d rather than 2d masks.</li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1361841516301839?via%3Dihub\" target=\"_blank\">3D FC-CNN with Conditional Random Fields, applied to brain lesion segmentation</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/abs/pii/S1361841517300725?via%3Dihub\" target=\"_blank\">Similar Approach to the last article, with Deep Supervision, applied to different tasks</a><br>\nThe following two papers seem like promising directions - They show significant improvements over U-net for medical image segmentation:</li>\n<li><a href=\"https://arxiv.org/pdf/1807.10165.pdf\" target=\"_blank\">U-net++</a></li>\n<li><a href=\"https://ieeexplore.ieee.org/abstract/document/9201310\" target=\"_blank\">MV-net</a></li>\n</ul>\n<p>I'm curious to hear other people's thoughts and which other analogs might be relevant. </p>",
  "messages": [
    {
      "id": "2186897",
      "postDate": "03/18/2023 08:23:02",
      "content": "<p>There isn't much research about ink detection from papyrus CT images in the literature. So contesters in this competition will have a hard time simply using a state-of-the-art network for this task. But that doesn't mean we have to start from scratch. We can instead borrow ideas from similar problems that have received more attention. To find good analogies, we must first look at identifying characteristics of this problem and seek problems that have most of these characteristics in common, while adapting the solution where the problems differ. Characteristics of this problem include:</p>\n<p>Nature of the data:</p>\n<ul>\n<li>3d image data. More specifically, a CT image with a single channel (i.e. grayscale image)</li>\n<li>(extremely) Limited number of images. We only have 3 train images in the dataset.</li>\n<li>High resolution images</li>\n</ul>\n<p>Nature of the inference target:</p>\n<ul>\n<li>We are looking to classify each pixel as ink/no ink. In the literature, this would be called an image segmentation or pixel classification problem</li>\n<li>There is a single class.</li>\n<li>While our input is a 3d image, the target is a 2d mask. </li>\n</ul>\n<p>Nature of the relationship between the two:</p>\n<ul>\n<li>Not all spatial dimensions are created equal. While the x and y dimensions in this problem are translation invariant (it doesn't matter where the ink is on the papyrus), the depth dimension is not. Deeper voxels are less likely to have ink in them, while shallower voxels are more likely to have other materials or defects tarnishing them.</li>\n<li>The ink forms letters. Letters have spatial relationships, so a pixel with ink is likely to be close to other ink pixels. Meanwhile, there are many letters in the image. Compared to other analogs, letters are relatively small-scale structures, and taking them into account does not require analyzing the whole image.</li>\n</ul>\n<p>Looking at these characteristics, the closest analog that I have found is segmentation of medical images. Specifically, tumor segmentation, lung segmentation, bone segmentation, etc are all valid analogs. We have CT images as input, and a small amount of labeled images. The target is slightly different. While this challenge creates a 2d mask, medical image segmentation is typically a 3d one. However, this difference can be adapted by pooling or otherwise distilling the depth into a single value for each pixel. We will also need to keep in mind the differences in the relationship when seeking to import a solution. We don't need to care about large-scale structure as much as those problems do.</p>\n<p>Here is a list of references to get you started on seeking analogs:</p>\n<ul>\n<li>Section 2.2 of this <a href=\"https://www.sciencedirect.com/science/article/pii/B9780128181010000033#s0025\" target=\"_blank\">book</a> reviews different <br>\nneural networks used for tumor segmentation</li>\n<li>This <a href=\"https://bmcmedimaging.biomedcentral.com/articles/10.1186/s12880-020-00529-5#Sec9\" target=\"_blank\">paper</a> applies two neural network architectures to segment lung images as infected/not infected by COVID-19.</li>\n<li>This <a href=\"https://www.frontiersin.org/articles/10.3389/frobt.2020.00106/full#B41\" target=\"_blank\">paper</a> Applies deep supervision and attention to improve tumor segmentation.</li>\n</ul>\n<p>Here are some architectures mentioned in these papers:</p>\n<ul>\n<li><a href=\"https://www.cv-foundation.org/openaccess/content_cvpr_2015/papers/Long_Fully_Convolutional_Networks_2015_CVPR_paper.pdf\" target=\"_blank\">Fully-Connected CNNs</a></li>\n<li><a href=\"https://arxiv.org/abs/1505.04597\" target=\"_blank\">U-net</a>, a type of FC-CNN</li>\n<li><a href=\"https://arxiv.org/abs/1511.00561\" target=\"_blank\">Segnet</a></li>\n<li><a href=\"https://pubmed.ncbi.nlm.nih.gov/28688283/\" target=\"_blank\">Central-Focused CNN</a></li>\n<li><a href=\"https://arxiv.org/abs/1505.00880\" target=\"_blank\">MV-CNN</a></li>\n<li><a href=\"https://arxiv.org/abs/1409.5185\" target=\"_blank\">Deeply Supervised Nets</a></li>\n</ul>\n<p>Edit: I did some further research. Here are some more references:</p>\n<ul>\n<li><a href=\"https://link.springer.com/chapter/10.1007/978-3-319-46723-8_49\" target=\"_blank\">3D U-net</a> applied to medical scans with 2d label inputs. Note that as opposed to the problem here, this paper outputs 3d rather than 2d masks.</li>\n<li><a href=\"https://www.sciencedirect.com/science/article/pii/S1361841516301839?via%3Dihub\" target=\"_blank\">3D FC-CNN with Conditional Random Fields, applied to brain lesion segmentation</a></li>\n<li><a href=\"https://www.sciencedirect.com/science/article/abs/pii/S1361841517300725?via%3Dihub\" target=\"_blank\">Similar Approach to the last article, with Deep Supervision, applied to different tasks</a><br>\nThe following two papers seem like promising directions - They show significant improvements over U-net for medical image segmentation:</li>\n<li><a href=\"https://arxiv.org/pdf/1807.10165.pdf\" target=\"_blank\">U-net++</a></li>\n<li><a href=\"https://ieeexplore.ieee.org/abstract/document/9201310\" target=\"_blank\">MV-net</a></li>\n</ul>\n<p>I'm curious to hear other people's thoughts and which other analogs might be relevant. </p>",
      "rawMarkdown": "There isn't much research about ink detection from papyrus CT images in the literature. So contesters in this competition will have a hard time simply using a state-of-the-art network for this task. But that doesn't mean we have to start from scratch. We can instead borrow ideas from similar problems that have received more attention. To find good analogies, we must first look at identifying characteristics of this problem and seek problems that have most of these characteristics in common, while adapting the solution where the problems differ. Characteristics of this problem include:\n\nNature of the data:\n- 3d image data. More specifically, a CT image with a single channel (i.e. grayscale image)\n- (extremely) Limited number of images. We only have 3 train images in the dataset.\n- High resolution images\n\nNature of the inference target:\n- We are looking to classify each pixel as ink/no ink. In the literature, this would be called an image segmentation or pixel classification problem\n- There is a single class.\n- While our input is a 3d image, the target is a 2d mask. \n\nNature of the relationship between the two:\n- Not all spatial dimensions are created equal. While the x and y dimensions in this problem are translation invariant (it doesn't matter where the ink is on the papyrus), the depth dimension is not. Deeper voxels are less likely to have ink in them, while shallower voxels are more likely to have other materials or defects tarnishing them.\n- The ink forms letters. Letters have spatial relationships, so a pixel with ink is likely to be close to other ink pixels. Meanwhile, there are many letters in the image. Compared to other analogs, letters are relatively small-scale structures, and taking them into account does not require analyzing the whole image.\n\nLooking at these characteristics, the closest analog that I have found is segmentation of medical images. Specifically, tumor segmentation, lung segmentation, bone segmentation, etc are all valid analogs. We have CT images as input, and a small amount of labeled images. The target is slightly different. While this challenge creates a 2d mask, medical image segmentation is typically a 3d one. However, this difference can be adapted by pooling or otherwise distilling the depth into a single value for each pixel. We will also need to keep in mind the differences in the relationship when seeking to import a solution. We don't need to care about large-scale structure as much as those problems do.\n\nHere is a list of references to get you started on seeking analogs:\n- Section 2.2 of this [book](https://www.sciencedirect.com/science/article/pii/B9780128181010000033#s0025) reviews different \nneural networks used for tumor segmentation\n- This [paper](https://bmcmedimaging.biomedcentral.com/articles/10.1186/s12880-020-00529-5#Sec9) applies two neural network architectures to segment lung images as infected/not infected by COVID-19.\n- This [paper](https://www.frontiersin.org/articles/10.3389/frobt.2020.00106/full#B41) Applies deep supervision and attention to improve tumor segmentation.\n\nHere are some architectures mentioned in these papers:\n- [Fully-Connected CNNs](https://www.cv-foundation.org/openaccess/content_cvpr_2015/papers/Long_Fully_Convolutional_Networks_2015_CVPR_paper.pdf)\n- [U-net](https://arxiv.org/abs/1505.04597), a type of FC-CNN\n- [Segnet](https://arxiv.org/abs/1511.00561)\n- [Central-Focused CNN](https://pubmed.ncbi.nlm.nih.gov/28688283/)\n- [MV-CNN](https://arxiv.org/abs/1505.00880)\n- [Deeply Supervised Nets](https://arxiv.org/abs/1409.5185)\n\nEdit: I did some further research. Here are some more references:\n\n- [3D U-net](https://link.springer.com/chapter/10.1007/978-3-319-46723-8_49) applied to medical scans with 2d label inputs. Note that as opposed to the problem here, this paper outputs 3d rather than 2d masks.\n- [3D FC-CNN with Conditional Random Fields, applied to brain lesion segmentation](https://www.sciencedirect.com/science/article/pii/S1361841516301839?via%3Dihub)\n- [Similar Approach to the last article, with Deep Supervision, applied to different tasks] (https://www.sciencedirect.com/science/article/abs/pii/S1361841517300725?via%3Dihub)\nThe following two papers seem like promising directions - They show significant improvements over U-net for medical image segmentation:\n- [U-net++](https://arxiv.org/pdf/1807.10165.pdf)\n- [MV-net](https://ieeexplore.ieee.org/abstract/document/9201310)\n \n\n\nI'm curious to hear other people's thoughts and which other analogs might be relevant.",
      "votes": null
    },
    {
      "id": "2187109",
      "postDate": "03/18/2023 12:33:02",
      "content": "<p>This provides an excellent breakdown of many of the properties and relationships of these datasets. I want to note that unlike organ segmentation tasks, where a reasonable shape prior might be incorporated, that's a more difficult task here. In the general case, you don't know the text, language, script, etc. of an unknown fragment or scroll.</p>\n<p>An additional thing I would like to note is resolution. These datasets have many times the spatial resolution of traditional medical CT data. So while you have 3 fragments of training data, the number of ink samples in each fragment can be significantly more than that. This leads to open questions about the area of support required to identify ink. </p>\n<p>I'm excited to see how this thread develops. Thanks!</p>",
      "rawMarkdown": "This provides an excellent breakdown of many of the properties and relationships of these datasets. I want to note that unlike organ segmentation tasks, where a reasonable shape prior might be incorporated, that's a more difficult task here. In the general case, you don't know the text, language, script, etc. of an unknown fragment or scroll.\n\nAn additional thing I would like to note is resolution. These datasets have many times the spatial resolution of traditional medical CT data. So while you have 3 fragments of training data, the number of ink samples in each fragment can be significantly more than that. This leads to open questions about the area of support required to identify ink. \n\nI'm excited to see how this thread develops. Thanks!",
      "votes": null
    },
    {
      "id": "2187287",
      "postDate": "03/18/2023 15:41:23",
      "content": "<p>Agreed with Seth, this is a great overview. I linked to it in Discord if you don't mind. :-)</p>",
      "rawMarkdown": "Agreed with Seth, this is a great overview. I linked to it in Discord if you don't mind. :-)",
      "votes": null
    },
    {
      "id": "2188769",
      "postDate": "03/20/2023 01:03:47",
      "content": "<p>I believe there is a lot of useful information near the surface of the papyrus, but those pixels are usually very close to the air, which is where relatively less information. Therefore, there should be an obvious information gap between the upper and lower ends of the surface, that is, it's important to find the surface of the papyrus which is in contact with the air.</p>\n<p>To find out the surface, I think a closer analogy may be AFM (Atomic Force Microscopy) probe, which will first lower the height to a certain point on the object surface, and then use its surrounding information to gradually scan through and then find out the information of the entire surface. Perhaps someone can write a similar scanning algorithm on these TIFF data.</p>",
      "rawMarkdown": "I believe there is a lot of useful information near the surface of the papyrus, but those pixels are usually very close to the air, which is where relatively less information. Therefore, there should be an obvious information gap between the upper and lower ends of the surface, that is, it's important to find the surface of the papyrus which is in contact with the air.\n\nTo find out the surface, I think a closer analogy may be AFM (Atomic Force Microscopy) probe, which will first lower the height to a certain point on the object surface, and then use its surrounding information to gradually scan through and then find out the information of the entire surface. Perhaps someone can write a similar scanning algorithm on these TIFF data.",
      "votes": null
    },
    {
      "id": "2196225",
      "postDate": "03/25/2023 08:32:13",
      "content": "<p>Good notes on the differences. I think there is still an important distinction when it comes to object scale between this and medical images. The objects that we are detecting still occupy a smaller region of the image than medical images, and I don't see a reason that an algorithm would need to know the full context of the image to identify letters. I haven't comprehensively tested but I think we can get away with classifying region by region in order to save computation.</p>",
      "rawMarkdown": "Good notes on the differences. I think there is still an important distinction when it comes to object scale between this and medical images. The objects that we are detecting still occupy a smaller region of the image than medical images, and I don't see a reason that an algorithm would need to know the full context of the image to identify letters. I haven't comprehensively tested but I think we can get away with classifying region by region in order to save computation.",
      "votes": null
    },
    {
      "id": "2196236",
      "postDate": "03/25/2023 08:37:23",
      "content": "<p>I agree. The paper that the team running this published also mentions how the ink is very shallow in these scrolls, about 6 µm thick. The dataset resolution is 3.2 µm, so the ink is only about two voxels deep. However, the papyrus is not perfectly flat. If you can somehow get an algorithm to identify the surface, you could get away with analyzing much less of the image and saving precious computation time. Regardless, the animations seem to indicate that some of the bottom layers are uninformative, so we can probably still discard them.</p>\n<p>Do you have some references for this probing technique?</p>",
      "rawMarkdown": "I agree. The paper that the team running this published also mentions how the ink is very shallow in these scrolls, about 6 µm thick. The dataset resolution is 3.2 µm, so the ink is only about two voxels deep. However, the papyrus is not perfectly flat. If you can somehow get an algorithm to identify the surface, you could get away with analyzing much less of the image and saving precious computation time. Regardless, the animations seem to indicate that some of the bottom layers are uninformative, so we can probably still discard them.\n\nDo you have some references for this probing technique?",
      "votes": null
    },
    {
      "id": "2198890",
      "postDate": "03/27/2023 11:07:05",
      "content": "<p>Two voxels is pretty thin for me. Is there a link to a paper that mentions this? Thanks. Regarding AFM, although having used it before, I don't know much about the mechanism behind it, so maybe someone who knows better can provide more information in this thread.</p>",
      "rawMarkdown": "Two voxels is pretty thin for me. Is there a link to a paper that mentions this? Thanks. Regarding AFM, although having used it before, I don't know much about the mechanism behind it, so maybe someone who knows better can provide more information in this thread.",
      "votes": null
    },
    {
      "id": "2199447",
      "postDate": "03/27/2023 17:50:58",
      "content": "<p>This paper mentions a total ink depth of 3-17µm: <a href=\"https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&amp;type=printable\" target=\"_blank\">https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&amp;type=printable</a></p>",
      "rawMarkdown": "This paper mentions a total ink depth of 3-17µm: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&type=printable",
      "votes": null
    },
    {
      "id": "2229476",
      "postDate": "04/21/2023 11:58:35",
      "content": "<p>Here are some other potential solutions:</p>\n<p>U-Net-based 3D architecture:<br>\nWe can adapt U-Net to a 3D setting by using 3D convolutional layers instead of 2D ones. After the final layer of the U-Net, you can use a global pooling layer and a fully connected layer to obtain a single output value for ink presence prediction.</p>\n<p>Residual Networks (ResNet) with 3D convolutions:<br>\nWe can adapt ResNet to work with 3D data by replacing 2D convolutions with 3D convolutions. Just like in the U-Net-based architecture, after the final layer, use a global pooling layer and a fully connected layer to output a single value for ink presence prediction.</p>\n<p>Volumetric CNN with Attention:<br>\nAttention mechanisms have been proven effective in various computer vision tasks. We can incorporate attention mechanisms, such as Squeeze-and-Excitation (SE) blocks or non-local blocks, into our 3D CNN architecture to allow the model to focus on important features. This approach can be combined with the other architectures mentioned above (e.g., U-Net or ResNet).</p>\n<p>Multi-view approach:<br>\nInstead of processing the entire 3D X-ray slide at once, we can maybe take 2D slices from different orientations (e.g., axial, coronal, and sagittal) and process them separately using 2D CNNs. The extracted features from these 2D CNNs can be concatenated and then passed through a fully connected layer to obtain the final prediction. This approach can reduce computational complexity and leverage pretrained 2D CNN architectures like ResNet or EfficientNet.</p>\n<p>Transfer learning:<br>\nWe can also take advantage of pre-trained 3D CNNs, such as 3D-ResNet or I3D, which have been trained on large-scale video classification datasets. Then fine-tune these models on your dataset by replacing the final classification layer to predict ink presence.</p>",
      "rawMarkdown": "Here are some other potential solutions:\n\nU-Net-based 3D architecture:\nWe can adapt U-Net to a 3D setting by using 3D convolutional layers instead of 2D ones. After the final layer of the U-Net, you can use a global pooling layer and a fully connected layer to obtain a single output value for ink presence prediction.\n\nResidual Networks (ResNet) with 3D convolutions:\nWe can adapt ResNet to work with 3D data by replacing 2D convolutions with 3D convolutions. Just like in the U-Net-based architecture, after the final layer, use a global pooling layer and a fully connected layer to output a single value for ink presence prediction.\n\nVolumetric CNN with Attention:\nAttention mechanisms have been proven effective in various computer vision tasks. We can incorporate attention mechanisms, such as Squeeze-and-Excitation (SE) blocks or non-local blocks, into our 3D CNN architecture to allow the model to focus on important features. This approach can be combined with the other architectures mentioned above (e.g., U-Net or ResNet).\n\nMulti-view approach:\nInstead of processing the entire 3D X-ray slide at once, we can maybe take 2D slices from different orientations (e.g., axial, coronal, and sagittal) and process them separately using 2D CNNs. The extracted features from these 2D CNNs can be concatenated and then passed through a fully connected layer to obtain the final prediction. This approach can reduce computational complexity and leverage pretrained 2D CNN architectures like ResNet or EfficientNet.\n\nTransfer learning:\nWe can also take advantage of pre-trained 3D CNNs, such as 3D-ResNet or I3D, which have been trained on large-scale video classification datasets. Then fine-tune these models on your dataset by replacing the final classification layer to predict ink presence.",
      "votes": null
    },
    {
      "id": "2251307",
      "postDate": "05/09/2023 08:47:38",
      "content": "<p>Thank you for the detailed notebook!</p>",
      "rawMarkdown": "Thank you for the detailed notebook!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2187109,
      "author_name": "csparker",
      "author_url": "",
      "post_date": "03/18/2023 12:33:02",
      "content": "<p>This provides an excellent breakdown of many of the properties and relationships of these datasets. I want to note that unlike organ segmentation tasks, where a reasonable shape prior might be incorporated, that's a more difficult task here. In the general case, you don't know the text, language, script, etc. of an unknown fragment or scroll.</p>\n<p>An additional thing I would like to note is resolution. These datasets have many times the spatial resolution of traditional medical CT data. So while you have 3 fragments of training data, the number of ink samples in each fragment can be significantly more than that. This leads to open questions about the area of support required to identify ink. </p>\n<p>I'm excited to see how this thread develops. Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2196225,
          "author_name": "pelegshilo",
          "author_url": "",
          "post_date": "03/25/2023 08:32:13",
          "content": "<p>Good notes on the differences. I think there is still an important distinction when it comes to object scale between this and medical images. The objects that we are detecting still occupy a smaller region of the image than medical images, and I don't see a reason that an algorithm would need to know the full context of the image to identify letters. I haven't comprehensively tested but I think we can get away with classifying region by region in order to save computation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2187287,
      "author_name": "jpposma",
      "author_url": "",
      "post_date": "03/18/2023 15:41:23",
      "content": "<p>Agreed with Seth, this is a great overview. I linked to it in Discord if you don't mind. :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2188769,
      "author_name": "yaohsiao123",
      "author_url": "",
      "post_date": "03/20/2023 01:03:47",
      "content": "<p>I believe there is a lot of useful information near the surface of the papyrus, but those pixels are usually very close to the air, which is where relatively less information. Therefore, there should be an obvious information gap between the upper and lower ends of the surface, that is, it's important to find the surface of the papyrus which is in contact with the air.</p>\n<p>To find out the surface, I think a closer analogy may be AFM (Atomic Force Microscopy) probe, which will first lower the height to a certain point on the object surface, and then use its surrounding information to gradually scan through and then find out the information of the entire surface. Perhaps someone can write a similar scanning algorithm on these TIFF data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2196236,
          "author_name": "pelegshilo",
          "author_url": "",
          "post_date": "03/25/2023 08:37:23",
          "content": "<p>I agree. The paper that the team running this published also mentions how the ink is very shallow in these scrolls, about 6 µm thick. The dataset resolution is 3.2 µm, so the ink is only about two voxels deep. However, the papyrus is not perfectly flat. If you can somehow get an algorithm to identify the surface, you could get away with analyzing much less of the image and saving precious computation time. Regardless, the animations seem to indicate that some of the bottom layers are uninformative, so we can probably still discard them.</p>\n<p>Do you have some references for this probing technique?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2198890,
              "author_name": "yaohsiao123",
              "author_url": "",
              "post_date": "03/27/2023 11:07:05",
              "content": "<p>Two voxels is pretty thin for me. Is there a link to a paper that mentions this? Thanks. Regarding AFM, although having used it before, I don't know much about the mechanism behind it, so maybe someone who knows better can provide more information in this thread.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2199447,
                  "author_name": "jpposma",
                  "author_url": "",
                  "post_date": "03/27/2023 17:50:58",
                  "content": "<p>This paper mentions a total ink depth of 3-17µm: <a href=\"https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&amp;type=printable\" target=\"_blank\">https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&amp;type=printable</a></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2229476,
      "author_name": "josipvrdoljak",
      "author_url": "",
      "post_date": "04/21/2023 11:58:35",
      "content": "<p>Here are some other potential solutions:</p>\n<p>U-Net-based 3D architecture:<br>\nWe can adapt U-Net to a 3D setting by using 3D convolutional layers instead of 2D ones. After the final layer of the U-Net, you can use a global pooling layer and a fully connected layer to obtain a single output value for ink presence prediction.</p>\n<p>Residual Networks (ResNet) with 3D convolutions:<br>\nWe can adapt ResNet to work with 3D data by replacing 2D convolutions with 3D convolutions. Just like in the U-Net-based architecture, after the final layer, use a global pooling layer and a fully connected layer to output a single value for ink presence prediction.</p>\n<p>Volumetric CNN with Attention:<br>\nAttention mechanisms have been proven effective in various computer vision tasks. We can incorporate attention mechanisms, such as Squeeze-and-Excitation (SE) blocks or non-local blocks, into our 3D CNN architecture to allow the model to focus on important features. This approach can be combined with the other architectures mentioned above (e.g., U-Net or ResNet).</p>\n<p>Multi-view approach:<br>\nInstead of processing the entire 3D X-ray slide at once, we can maybe take 2D slices from different orientations (e.g., axial, coronal, and sagittal) and process them separately using 2D CNNs. The extracted features from these 2D CNNs can be concatenated and then passed through a fully connected layer to obtain the final prediction. This approach can reduce computational complexity and leverage pretrained 2D CNN architectures like ResNet or EfficientNet.</p>\n<p>Transfer learning:<br>\nWe can also take advantage of pre-trained 3D CNNs, such as 3D-ResNet or I3D, which have been trained on large-scale video classification datasets. Then fine-tune these models on your dataset by replacing the final classification layer to predict ink presence.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2251307,
      "author_name": "namtranase",
      "author_url": "",
      "post_date": "05/09/2023 08:47:38",
      "content": "<p>Thank you for the detailed notebook!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2186897": "There isn't much research about ink detection from papyrus CT images in the literature. So contesters in this competition will have a hard time simply using a state-of-the-art network for this task. But that doesn't mean we have to start from scratch. We can instead borrow ideas from similar problems that have received more attention. To find good analogies, we must first look at identifying characteristics of this problem and seek problems that have most of these characteristics in common, while adapting the solution where the problems differ. Characteristics of this problem include:\n\nNature of the data:\n- 3d image data. More specifically, a CT image with a single channel (i.e. grayscale image)\n- (extremely) Limited number of images. We only have 3 train images in the dataset.\n- High resolution images\n\nNature of the inference target:\n- We are looking to classify each pixel as ink/no ink. In the literature, this would be called an image segmentation or pixel classification problem\n- There is a single class.\n- While our input is a 3d image, the target is a 2d mask. \n\nNature of the relationship between the two:\n- Not all spatial dimensions are created equal. While the x and y dimensions in this problem are translation invariant (it doesn't matter where the ink is on the papyrus), the depth dimension is not. Deeper voxels are less likely to have ink in them, while shallower voxels are more likely to have other materials or defects tarnishing them.\n- The ink forms letters. Letters have spatial relationships, so a pixel with ink is likely to be close to other ink pixels. Meanwhile, there are many letters in the image. Compared to other analogs, letters are relatively small-scale structures, and taking them into account does not require analyzing the whole image.\n\nLooking at these characteristics, the closest analog that I have found is segmentation of medical images. Specifically, tumor segmentation, lung segmentation, bone segmentation, etc are all valid analogs. We have CT images as input, and a small amount of labeled images. The target is slightly different. While this challenge creates a 2d mask, medical image segmentation is typically a 3d one. However, this difference can be adapted by pooling or otherwise distilling the depth into a single value for each pixel. We will also need to keep in mind the differences in the relationship when seeking to import a solution. We don't need to care about large-scale structure as much as those problems do.\n\nHere is a list of references to get you started on seeking analogs:\n- Section 2.2 of this [book](https://www.sciencedirect.com/science/article/pii/B9780128181010000033#s0025) reviews different \nneural networks used for tumor segmentation\n- This [paper](https://bmcmedimaging.biomedcentral.com/articles/10.1186/s12880-020-00529-5#Sec9) applies two neural network architectures to segment lung images as infected/not infected by COVID-19.\n- This [paper](https://www.frontiersin.org/articles/10.3389/frobt.2020.00106/full#B41) Applies deep supervision and attention to improve tumor segmentation.\n\nHere are some architectures mentioned in these papers:\n- [Fully-Connected CNNs](https://www.cv-foundation.org/openaccess/content_cvpr_2015/papers/Long_Fully_Convolutional_Networks_2015_CVPR_paper.pdf)\n- [U-net](https://arxiv.org/abs/1505.04597), a type of FC-CNN\n- [Segnet](https://arxiv.org/abs/1511.00561)\n- [Central-Focused CNN](https://pubmed.ncbi.nlm.nih.gov/28688283/)\n- [MV-CNN](https://arxiv.org/abs/1505.00880)\n- [Deeply Supervised Nets](https://arxiv.org/abs/1409.5185)\n\nEdit: I did some further research. Here are some more references:\n\n- [3D U-net](https://link.springer.com/chapter/10.1007/978-3-319-46723-8_49) applied to medical scans with 2d label inputs. Note that as opposed to the problem here, this paper outputs 3d rather than 2d masks.\n- [3D FC-CNN with Conditional Random Fields, applied to brain lesion segmentation](https://www.sciencedirect.com/science/article/pii/S1361841516301839?via%3Dihub)\n- [Similar Approach to the last article, with Deep Supervision, applied to different tasks] (https://www.sciencedirect.com/science/article/abs/pii/S1361841517300725?via%3Dihub)\nThe following two papers seem like promising directions - They show significant improvements over U-net for medical image segmentation:\n- [U-net++](https://arxiv.org/pdf/1807.10165.pdf)\n- [MV-net](https://ieeexplore.ieee.org/abstract/document/9201310)\n \n\n\nI'm curious to hear other people's thoughts and which other analogs might be relevant.",
    "2187109": "This provides an excellent breakdown of many of the properties and relationships of these datasets. I want to note that unlike organ segmentation tasks, where a reasonable shape prior might be incorporated, that's a more difficult task here. In the general case, you don't know the text, language, script, etc. of an unknown fragment or scroll.\n\nAn additional thing I would like to note is resolution. These datasets have many times the spatial resolution of traditional medical CT data. So while you have 3 fragments of training data, the number of ink samples in each fragment can be significantly more than that. This leads to open questions about the area of support required to identify ink. \n\nI'm excited to see how this thread develops. Thanks!",
    "2187287": "Agreed with Seth, this is a great overview. I linked to it in Discord if you don't mind. :-)",
    "2188769": "I believe there is a lot of useful information near the surface of the papyrus, but those pixels are usually very close to the air, which is where relatively less information. Therefore, there should be an obvious information gap between the upper and lower ends of the surface, that is, it's important to find the surface of the papyrus which is in contact with the air.\n\nTo find out the surface, I think a closer analogy may be AFM (Atomic Force Microscopy) probe, which will first lower the height to a certain point on the object surface, and then use its surrounding information to gradually scan through and then find out the information of the entire surface. Perhaps someone can write a similar scanning algorithm on these TIFF data.",
    "2196225": "Good notes on the differences. I think there is still an important distinction when it comes to object scale between this and medical images. The objects that we are detecting still occupy a smaller region of the image than medical images, and I don't see a reason that an algorithm would need to know the full context of the image to identify letters. I haven't comprehensively tested but I think we can get away with classifying region by region in order to save computation.",
    "2196236": "I agree. The paper that the team running this published also mentions how the ink is very shallow in these scrolls, about 6 µm thick. The dataset resolution is 3.2 µm, so the ink is only about two voxels deep. However, the papyrus is not perfectly flat. If you can somehow get an algorithm to identify the surface, you could get away with analyzing much less of the image and saving precious computation time. Regardless, the animations seem to indicate that some of the bottom layers are uninformative, so we can probably still discard them.\n\nDo you have some references for this probing technique?",
    "2198890": "Two voxels is pretty thin for me. Is there a link to a paper that mentions this? Thanks. Regarding AFM, although having used it before, I don't know much about the mechanism behind it, so maybe someone who knows better can provide more information in this thread.",
    "2199447": "This paper mentions a total ink depth of 3-17µm: https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0215775&type=printable",
    "2229476": "Here are some other potential solutions:\n\nU-Net-based 3D architecture:\nWe can adapt U-Net to a 3D setting by using 3D convolutional layers instead of 2D ones. After the final layer of the U-Net, you can use a global pooling layer and a fully connected layer to obtain a single output value for ink presence prediction.\n\nResidual Networks (ResNet) with 3D convolutions:\nWe can adapt ResNet to work with 3D data by replacing 2D convolutions with 3D convolutions. Just like in the U-Net-based architecture, after the final layer, use a global pooling layer and a fully connected layer to output a single value for ink presence prediction.\n\nVolumetric CNN with Attention:\nAttention mechanisms have been proven effective in various computer vision tasks. We can incorporate attention mechanisms, such as Squeeze-and-Excitation (SE) blocks or non-local blocks, into our 3D CNN architecture to allow the model to focus on important features. This approach can be combined with the other architectures mentioned above (e.g., U-Net or ResNet).\n\nMulti-view approach:\nInstead of processing the entire 3D X-ray slide at once, we can maybe take 2D slices from different orientations (e.g., axial, coronal, and sagittal) and process them separately using 2D CNNs. The extracted features from these 2D CNNs can be concatenated and then passed through a fully connected layer to obtain the final prediction. This approach can reduce computational complexity and leverage pretrained 2D CNN architectures like ResNet or EfficientNet.\n\nTransfer learning:\nWe can also take advantage of pre-trained 3D CNNs, such as 3D-ResNet or I3D, which have been trained on large-scale video classification datasets. Then fine-tune these models on your dataset by replacing the final classification layer to predict ink presence.",
    "2251307": "Thank you for the detailed notebook!"
  },
  "source": "meta"
}