{
  "id": 240882,
  "title": "Question about image bit depth",
  "url": "/competitions/siim-covid19-detection/discussion/240882",
  "author_name": "",
  "post_date": "2021-05-21T23:19:07.815780500Z",
  "votes": 4,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I am a newbie, and I don't yet understand the top tier level architectures in comps like this.</p>\n<p>My assumption is that this competition will include classification and object detection at the basic level.</p>\n<p>Are the final train and test images always in an 8 bit format? Or since DICOM CXRs are usually 10 bit or higher, do the models operate on tensors with a larger depth to avoid lossy issues?</p>\n<p>I understand there's not a significant loss in most cases when using a LUT to go from 10 to 8 bit, but it seems any loss is important in this context.</p>\n<p>I'm guessing most people use different forms of leveling (VOI, modality, histogram etc), and export multiple variations of each image?</p>",
  "messages": [
    {
      "id": "1318031",
      "postDate": "05/21/2021 23:19:07",
      "content": "<p>I am a newbie, and I don't yet understand the top tier level architectures in comps like this.</p>\n<p>My assumption is that this competition will include classification and object detection at the basic level.</p>\n<p>Are the final train and test images always in an 8 bit format? Or since DICOM CXRs are usually 10 bit or higher, do the models operate on tensors with a larger depth to avoid lossy issues?</p>\n<p>I understand there's not a significant loss in most cases when using a LUT to go from 10 to 8 bit, but it seems any loss is important in this context.</p>\n<p>I'm guessing most people use different forms of leveling (VOI, modality, histogram etc), and export multiple variations of each image?</p>",
      "rawMarkdown": "I am a newbie, and I don't yet understand the top tier level architectures in comps like this.\n\nMy assumption is that this competition will include classification and object detection at the basic level.\n\nAre the final train and test images always in an 8 bit format? Or since DICOM CXRs are usually 10 bit or higher, do the models operate on tensors with a larger depth to avoid lossy issues?\n\nI understand there's not a significant loss in most cases when using a LUT to go from 10 to 8 bit, but it seems any loss is important in this context.\n\nI'm guessing most people use different forms of leveling (VOI, modality, histogram etc), and export multiple variations of each image?",
      "votes": null
    },
    {
      "id": "1330269",
      "postDate": "05/31/2021 17:00:24",
      "content": "<p>Neural network uses float data, we get pixels from the picture and convert them from integer to float, so if you have access to more than 8bit data in DICOM it would be used as the float same way.</p>",
      "rawMarkdown": "Neural network uses float data, we get pixels from the picture and convert them from integer to float, so if you have access to more than 8bit data in DICOM it would be used as the float same way.",
      "votes": null
    },
    {
      "id": "1330513",
      "postDate": "05/31/2021 22:36:56",
      "content": "<p>I am also a complete newbie 😃, but you are clearly onto something here, and in your notebook posting, thanks. The general view might be: it’s float 32 bit, maybe 16 bit (either way exponent plus mantissa, 8-bit quantised tensors are more niche), the network will sort it out (linear/non-linear intensity distribution), nobody is looking at this so who cares about visualisation (that’s not the point!), it’s a CNN hence you just need to train it (less feature engineering), etc.</p>\n<p>Common sense dictates you use transfer learning so an existing pre-trained model, and normalise your data on the means of the original training set (so for each RGB channel, scale to a mean of zero, standard deviation of one). </p>\n<p>However what surely makes x-rays different is:</p>\n<ul>\n<li>images are monochrome, and the intensity scale is different, possibly non-linear</li>\n<li>abnormalities don’t have clear edges (unlike ribs!!!) so detection is different</li>\n<li>an opacity is a shape defined by fuzzy changes within a narrow intensity band</li>\n<li>unlike visible light x-rays penetrate depth-wise, there are no clear cuts where opaque objects occlude the deeper ones</li>\n</ul>\n<p>I definitely think there’s some magic image pre-processing that can be done here to improve detection, and unlike more general object detection problems, transfer learning may be of limited use beyond the earlier layers.</p>",
      "rawMarkdown": "I am also a complete newbie 😃, but you are clearly onto something here, and in your notebook posting, thanks. The general view might be: it’s float 32 bit, maybe 16 bit (either way exponent plus mantissa, 8-bit quantised tensors are more niche), the network will sort it out (linear/non-linear intensity distribution), nobody is looking at this so who cares about visualisation (that’s not the point!), it’s a CNN hence you just need to train it (less feature engineering), etc.\n\nCommon sense dictates you use transfer learning so an existing pre-trained model, and normalise your data on the means of the original training set (so for each RGB channel, scale to a mean of zero, standard deviation of one). \n\nHowever what surely makes x-rays different is:\n\n* images are monochrome, and the intensity scale is different, possibly non-linear\n* abnormalities don’t have clear edges (unlike ribs!!!) so detection is different\n* an opacity is a shape defined by fuzzy changes within a narrow intensity band\n* unlike visible light x-rays penetrate depth-wise, there are no clear cuts where opaque objects occlude the deeper ones\n\nI definitely think there’s some magic image pre-processing that can be done here to improve detection, and unlike more general object detection problems, transfer learning may be of limited use beyond the earlier layers.",
      "votes": null
    },
    {
      "id": "1330571",
      "postDate": "06/01/2021 00:58:52",
      "content": "<p>I made a quick notebook on how I think it should work. Rather than export images to JPG, why not extract the pixels and feed the raw values to the models? I assume pretrained CNNs scale inputs to match the depth of the images they're trained on, which defeats the purpose, but what if I build one from scratch.</p>\n<p>This notebook uses transfer learning with InceptionV3 and a Sequential model with raw pixel input. It didn't transform images or process them in any way. I just wanted to show my thought process.</p>\n<p><a href=\"https://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input\" target=\"_blank\">https://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input</a></p>\n<p>I'm new, so be gentle. And thanks for your input!</p>",
      "rawMarkdown": "I made a quick notebook on how I think it should work. Rather than export images to JPG, why not extract the pixels and feed the raw values to the models? I assume pretrained CNNs scale inputs to match the depth of the images they're trained on, which defeats the purpose, but what if I build one from scratch.\n\nThis notebook uses transfer learning with InceptionV3 and a Sequential model with raw pixel input. It didn't transform images or process them in any way. I just wanted to show my thought process.\n\nhttps://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input\n\nI'm new, so be gentle. And thanks for your input!",
      "votes": null
    },
    {
      "id": "1346073",
      "postDate": "06/12/2021 05:21:29",
      "content": "<p>Page not Found Error </p>",
      "rawMarkdown": "Page not Found Error",
      "votes": null
    },
    {
      "id": "1361481",
      "postDate": "06/22/2021 21:17:04",
      "content": "<p>Acck .. I realize I didn't share this notebook when I posted this. I get it now! Sorry for bumping an old post.</p>",
      "rawMarkdown": "Acck .. I realize I didn't share this notebook when I posted this. I get it now! Sorry for bumping an old post.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1330269,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "05/31/2021 17:00:24",
      "content": "<p>Neural network uses float data, we get pixels from the picture and convert them from integer to float, so if you have access to more than 8bit data in DICOM it would be used as the float same way.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1330513,
      "author_name": "benvenutto",
      "author_url": "",
      "post_date": "05/31/2021 22:36:56",
      "content": "<p>I am also a complete newbie 😃, but you are clearly onto something here, and in your notebook posting, thanks. The general view might be: it’s float 32 bit, maybe 16 bit (either way exponent plus mantissa, 8-bit quantised tensors are more niche), the network will sort it out (linear/non-linear intensity distribution), nobody is looking at this so who cares about visualisation (that’s not the point!), it’s a CNN hence you just need to train it (less feature engineering), etc.</p>\n<p>Common sense dictates you use transfer learning so an existing pre-trained model, and normalise your data on the means of the original training set (so for each RGB channel, scale to a mean of zero, standard deviation of one). </p>\n<p>However what surely makes x-rays different is:</p>\n<ul>\n<li>images are monochrome, and the intensity scale is different, possibly non-linear</li>\n<li>abnormalities don’t have clear edges (unlike ribs!!!) so detection is different</li>\n<li>an opacity is a shape defined by fuzzy changes within a narrow intensity band</li>\n<li>unlike visible light x-rays penetrate depth-wise, there are no clear cuts where opaque objects occlude the deeper ones</li>\n</ul>\n<p>I definitely think there’s some magic image pre-processing that can be done here to improve detection, and unlike more general object detection problems, transfer learning may be of limited use beyond the earlier layers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1330571,
      "author_name": "davidbroberts",
      "author_url": "",
      "post_date": "06/01/2021 00:58:52",
      "content": "<p>I made a quick notebook on how I think it should work. Rather than export images to JPG, why not extract the pixels and feed the raw values to the models? I assume pretrained CNNs scale inputs to match the depth of the images they're trained on, which defeats the purpose, but what if I build one from scratch.</p>\n<p>This notebook uses transfer learning with InceptionV3 and a Sequential model with raw pixel input. It didn't transform images or process them in any way. I just wanted to show my thought process.</p>\n<p><a href=\"https://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input\" target=\"_blank\">https://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input</a></p>\n<p>I'm new, so be gentle. And thanks for your input!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1346073,
          "author_name": "devanshchowdhury",
          "author_url": "",
          "post_date": "06/12/2021 05:21:29",
          "content": "<p>Page not Found Error </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1361481,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "06/22/2021 21:17:04",
          "content": "<p>Acck .. I realize I didn't share this notebook when I posted this. I get it now! Sorry for bumping an old post.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1318031": "I am a newbie, and I don't yet understand the top tier level architectures in comps like this.\n\nMy assumption is that this competition will include classification and object detection at the basic level.\n\nAre the final train and test images always in an 8 bit format? Or since DICOM CXRs are usually 10 bit or higher, do the models operate on tensors with a larger depth to avoid lossy issues?\n\nI understand there's not a significant loss in most cases when using a LUT to go from 10 to 8 bit, but it seems any loss is important in this context.\n\nI'm guessing most people use different forms of leveling (VOI, modality, histogram etc), and export multiple variations of each image?",
    "1330269": "Neural network uses float data, we get pixels from the picture and convert them from integer to float, so if you have access to more than 8bit data in DICOM it would be used as the float same way.",
    "1330513": "I am also a complete newbie 😃, but you are clearly onto something here, and in your notebook posting, thanks. The general view might be: it’s float 32 bit, maybe 16 bit (either way exponent plus mantissa, 8-bit quantised tensors are more niche), the network will sort it out (linear/non-linear intensity distribution), nobody is looking at this so who cares about visualisation (that’s not the point!), it’s a CNN hence you just need to train it (less feature engineering), etc.\n\nCommon sense dictates you use transfer learning so an existing pre-trained model, and normalise your data on the means of the original training set (so for each RGB channel, scale to a mean of zero, standard deviation of one). \n\nHowever what surely makes x-rays different is:\n\n* images are monochrome, and the intensity scale is different, possibly non-linear\n* abnormalities don’t have clear edges (unlike ribs!!!) so detection is different\n* an opacity is a shape defined by fuzzy changes within a narrow intensity band\n* unlike visible light x-rays penetrate depth-wise, there are no clear cuts where opaque objects occlude the deeper ones\n\nI definitely think there’s some magic image pre-processing that can be done here to improve detection, and unlike more general object detection problems, transfer learning may be of limited use beyond the earlier layers.",
    "1330571": "I made a quick notebook on how I think it should work. Rather than export images to JPG, why not extract the pixels and feed the raw values to the models? I assume pretrained CNNs scale inputs to match the depth of the images they're trained on, which defeats the purpose, but what if I build one from scratch.\n\nThis notebook uses transfer learning with InceptionV3 and a Sequential model with raw pixel input. It didn't transform images or process them in any way. I just wanted to show my thought process.\n\nhttps://www.kaggle.com/davidbroberts/dicom-full-range-pixels-as-cnn-input\n\nI'm new, so be gentle. And thanks for your input!",
    "1346073": "Page not Found Error",
    "1361481": "Acck .. I realize I didn't share this notebook when I posted this. I get it now! Sorry for bumping an old post."
  },
  "source": "meta"
}