{
  "id": 176926,
  "title": "Best way to use CT scans for better LB scores",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/176926",
  "author_name": "AnkurSingh",
  "post_date": "2020-08-24T06:00:21.463000",
  "votes": 40,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I have seen many high performing kernels using only the data in train.csv<br>\nThe 3 most widely use algorithms are Ridge regression, ElasticNet and LGBM.</p>\n<p>When it comes to using CT scans along side train.csv. I have found the following approaches:</p>\n<ul>\n<li>AutoEncoder to create encoded features</li>\n<li>CNNs to extract features (most commonly used is EfficientNet)</li>\n</ul>\n<p>Introducing encoded / extracted features along side train.csv seems to gives a small boost in performance. But I think we can get much more out of CT scans. </p>\n<p>Instead of using standard deep learning techniques, using domain specific techniques like the ones mentioned in discussion thread by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> can give much bigger boost. Here are some domain specific features that can be extracted from the CT scans: </p>\n<ul>\n<li>calculating lung volume</li>\n<li>chest circumference</li>\n<li>Histogram/kurtosis analysis</li>\n</ul>\n<p>I will be conducting an array of experiments to test these techniques and post the updated here in the thread. Also, all the contributions are welcome. If you try some of these features then please post the results here. It will help everyone.</p>",
  "messages": [
    {
      "id": 983236,
      "postDate": "2020-08-24T06:00:21.463Z",
      "content": "<p>I have seen many high performing kernels using only the data in train.csv<br>\nThe 3 most widely use algorithms are Ridge regression, ElasticNet and LGBM.</p>\n<p>When it comes to using CT scans along side train.csv. I have found the following approaches:</p>\n<ul>\n<li>AutoEncoder to create encoded features</li>\n<li>CNNs to extract features (most commonly used is EfficientNet)</li>\n</ul>\n<p>Introducing encoded / extracted features along side train.csv seems to gives a small boost in performance. But I think we can get much more out of CT scans. </p>\n<p>Instead of using standard deep learning techniques, using domain specific techniques like the ones mentioned in discussion thread by <a href=\"https://www.kaggle.com/sandorkonya\" target=\"_blank\">@sandorkonya</a> can give much bigger boost. Here are some domain specific features that can be extracted from the CT scans: </p>\n<ul>\n<li>calculating lung volume</li>\n<li>chest circumference</li>\n<li>Histogram/kurtosis analysis</li>\n</ul>\n<p>I will be conducting an array of experiments to test these techniques and post the updated here in the thread. Also, all the contributions are welcome. If you try some of these features then please post the results here. It will help everyone.</p>",
      "rawMarkdown": "I have seen many high performing kernels using only the data in train.csv\nThe 3 most widely use algorithms are Ridge regression, ElasticNet and LGBM.\n\nWhen it comes to using CT scans along side train.csv. I have found the following approaches:\n- AutoEncoder to create encoded features\n- CNNs to extract features (most commonly used is EfficientNet)\n\nIntroducing encoded / extracted features along side train.csv seems to gives a small boost in performance. But I think we can get much more out of CT scans. \n\nInstead of using standard deep learning techniques, using domain specific techniques like the ones mentioned in discussion thread by @sandorkonya can give much bigger boost. Here are some domain specific features that can be extracted from the CT scans: \n- calculating lung volume\n- chest circumference\n- Histogram/kurtosis analysis\n\nI will be conducting an array of experiments to test these techniques and post the updated here in the thread. Also, all the contributions are welcome. If you try some of these features then please post the results here. It will help everyone.",
      "votes": 39
    },
    {
      "id": 987063,
      "postDate": "2020-08-27T00:38:23.607Z",
      "content": "<p>I made my notebook public. I looked at whether small patches in the images are important. The results are interesting but inconclusive. My next step is to see if I can improve overall prediction score. </p>",
      "rawMarkdown": "I made my notebook public. I looked at whether small patches in the images are important. The results are interesting but inconclusive. My next step is to see if I can improve overall prediction score. ",
      "votes": 4
    },
    {
      "id": 1001018,
      "postDate": "2020-09-07T02:33:26.710Z",
      "content": "<p><a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a>, thanks for sharing, I will start playing tonight with the CT</p>",
      "rawMarkdown": "@ankursingh12, thanks for sharing, I will start playing tonight with the CT"
    },
    {
      "id": 999066,
      "postDate": "2020-09-05T10:07:53.110Z",
      "content": "<p>This seems good! Thank you for sharing. I will try the autoencoder approach .</p>",
      "rawMarkdown": "This seems good! Thank you for sharing. I will try the autoencoder approach ."
    },
    {
      "id": 984633,
      "postDate": "2020-08-25T07:40:34.380Z",
      "content": "<blockquote>\n  <p>calculating lung volume</p>\n</blockquote>\n<p>So, I tried to convert dicom 2d slice data to 3d voxel data (numpy array npy file), but the following error occurs <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/176879\" target=\"_blank\">here</a>.</p>\n<p>Somebody please help!</p>",
      "rawMarkdown": "> calculating lung volume\n\nSo, I tried to convert dicom 2d slice data to 3d voxel data (numpy array npy file), but the following error occurs [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/176879).\n\nSomebody please help!",
      "replies": [
        {
          "id": 985839,
          "postDate": "2020-08-26T03:43:13.853Z",
          "content": "<p>I have been through your code. The error says that there are missing slices. The <code>combine_slices()</code> that you are using assume the following conditions to true.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1537159%2F1923243632f764f03fdc7a0adb97eedd%2Fdicom.jpg?generation=1598424619155696&amp;alt=media\" alt=\"\"></p>\n<p>3rd point discusses about missing slices. If you have missing slices in the end, then its okay, else it will raise an error. Hope it helps.</p>\n<p>Even I found two corrupted images in the training set. But my ids are different yours. <br>\nAs of now, I am ignoring them. </p>",
          "rawMarkdown": "I have been through your code. The error says that there are missing slices. The `combine_slices()` that you are using assume the following conditions to true.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1537159%2F1923243632f764f03fdc7a0adb97eedd%2Fdicom.jpg?generation=1598424619155696&alt=media)\n\n3rd point discusses about missing slices. If you have missing slices in the end, then its okay, else it will raise an error. Hope it helps.\n\nEven I found two corrupted images in the training set. But my ids are different yours. \nAs of now, I am ignoring them. ",
          "votes": 2
        },
        {
          "id": 986043,
          "postDate": "2020-08-26T06:56:52.527Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 986636,
          "postDate": "2020-08-26T16:44:54.077Z",
          "content": "<p>Please sort them by slice location ..we tried for 1 image and it is working .. Below is the sample code .. <br>\n<a href=\"https://pydicom.github.io/pydicom/dev/auto_examples/image_processing/reslice.html\" target=\"_blank\">https://pydicom.github.io/pydicom/dev/auto_examples/image_processing/reslice.html</a></p>\n<p>code snippet</p>\n<h1>load the DICOM files</h1>\n<p>files = []<br>\nprint('glob: {}'.format(sys.argv[1]))<br>\nfor fname in glob.glob(sys.argv[1], recursive=False):<br>\n    print(\"loading: {}\".format(fname))<br>\n    files.append(pydicom.dcmread(fname))</p>\n<p>print(\"file count: {}\".format(len(files)))</p>\n<h1>skip files with no SliceLocation (eg scout views)</h1>\n<p>slices = []<br>\nskipcount = 0<br>\nfor f in files:<br>\n    if hasattr(f, 'SliceLocation'):<br>\n        slices.append(f)<br>\n    else:<br>\n        skipcount = skipcount + 1</p>\n<p>print(\"skipped, no SliceLocation: {}\".format(skipcount))</p>\n<h1>ensure they are in the correct order</h1>\n<p>slices = sorted(slices, key=lambda s: s.SliceLocation)</p>",
          "rawMarkdown": "Please sort them by slice location ..we tried for 1 image and it is working .. Below is the sample code .. \nhttps://pydicom.github.io/pydicom/dev/auto_examples/image_processing/reslice.html\n\n\ncode snippet\n# load the DICOM files\nfiles = []\nprint('glob: {}'.format(sys.argv[1]))\nfor fname in glob.glob(sys.argv[1], recursive=False):\n    print(\"loading: {}\".format(fname))\n    files.append(pydicom.dcmread(fname))\n\nprint(\"file count: {}\".format(len(files)))\n\n# skip files with no SliceLocation (eg scout views)\nslices = []\nskipcount = 0\nfor f in files:\n    if hasattr(f, 'SliceLocation'):\n        slices.append(f)\n    else:\n        skipcount = skipcount + 1\n\nprint(\"skipped, no SliceLocation: {}\".format(skipcount))\n\n# ensure they are in the correct order\nslices = sorted(slices, key=lambda s: s.SliceLocation)",
          "votes": 1
        },
        {
          "id": 987430,
          "postDate": "2020-08-27T08:50:39.397Z",
          "content": "<p>thanks for the info, i will give it a try. btw, found dataset someone upload a month ago:</p>\n<p><a href=\"https://www.kaggle.com/kjm1559/osic-dicom-3d-image\" target=\"_blank\">https://www.kaggle.com/kjm1559/osic-dicom-3d-image</a></p>",
          "rawMarkdown": "thanks for the info, i will give it a try. btw, found dataset someone upload a month ago:\n\nhttps://www.kaggle.com/kjm1559/osic-dicom-3d-image",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 987063,
      "author_name": "andy jennings",
      "author_url": "",
      "post_date": "2020-08-27T00:38:23.607000",
      "content": "<p>I made my notebook public. I looked at whether small patches in the images are important. The results are interesting but inconclusive. My next step is to see if I can improve overall prediction score. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1001018,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2020-09-07T02:33:26.710000",
      "content": "<p><a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a>, thanks for sharing, I will start playing tonight with the CT</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 999066,
      "author_name": "Aditya Baurai",
      "author_url": "",
      "post_date": "2020-09-05T10:07:53.110000",
      "content": "<p>This seems good! Thank you for sharing. I will try the autoencoder approach .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 984633,
      "author_name": "Ji Wong Park",
      "author_url": "",
      "post_date": "2020-08-25T07:40:34.380000",
      "content": "<blockquote>\n  <p>calculating lung volume</p>\n</blockquote>\n<p>So, I tried to convert dicom 2d slice data to 3d voxel data (numpy array npy file), but the following error occurs <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/176879\" target=\"_blank\">here</a>.</p>\n<p>Somebody please help!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 985839,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2020-08-26T03:43:13.853000",
          "content": "<p>I have been through your code. The error says that there are missing slices. The <code>combine_slices()</code> that you are using assume the following conditions to true.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1537159%2F1923243632f764f03fdc7a0adb97eedd%2Fdicom.jpg?generation=1598424619155696&amp;alt=media\" alt=\"\"></p>\n<p>3rd point discusses about missing slices. If you have missing slices in the end, then its okay, else it will raise an error. Hope it helps.</p>\n<p>Even I found two corrupted images in the training set. But my ids are different yours. <br>\nAs of now, I am ignoring them. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 986043,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-26T06:56:52.527000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 986636,
          "author_name": "Pulkit Mehta",
          "author_url": "",
          "post_date": "2020-08-26T16:44:54.077000",
          "content": "<p>Please sort them by slice location ..we tried for 1 image and it is working .. Below is the sample code .. <br>\n<a href=\"https://pydicom.github.io/pydicom/dev/auto_examples/image_processing/reslice.html\" target=\"_blank\">https://pydicom.github.io/pydicom/dev/auto_examples/image_processing/reslice.html</a></p>\n<p>code snippet</p>\n<h1>load the DICOM files</h1>\n<p>files = []<br>\nprint('glob: {}'.format(sys.argv[1]))<br>\nfor fname in glob.glob(sys.argv[1], recursive=False):<br>\n    print(\"loading: {}\".format(fname))<br>\n    files.append(pydicom.dcmread(fname))</p>\n<p>print(\"file count: {}\".format(len(files)))</p>\n<h1>skip files with no SliceLocation (eg scout views)</h1>\n<p>slices = []<br>\nskipcount = 0<br>\nfor f in files:<br>\n    if hasattr(f, 'SliceLocation'):<br>\n        slices.append(f)<br>\n    else:<br>\n        skipcount = skipcount + 1</p>\n<p>print(\"skipped, no SliceLocation: {}\".format(skipcount))</p>\n<h1>ensure they are in the correct order</h1>\n<p>slices = sorted(slices, key=lambda s: s.SliceLocation)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 987430,
          "author_name": "Ji Wong Park",
          "author_url": "",
          "post_date": "2020-08-27T08:50:39.397000",
          "content": "<p>thanks for the info, i will give it a try. btw, found dataset someone upload a month ago:</p>\n<p><a href=\"https://www.kaggle.com/kjm1559/osic-dicom-3d-image\" target=\"_blank\">https://www.kaggle.com/kjm1559/osic-dicom-3d-image</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "983236": "I have seen many high performing kernels using only the data in train.csv\nThe 3 most widely use algorithms are Ridge regression, ElasticNet and LGBM.\n\nWhen it comes to using CT scans along side train.csv. I have found the following approaches:\n- AutoEncoder to create encoded features\n- CNNs to extract features (most commonly used is EfficientNet)\n\nIntroducing encoded / extracted features along side train.csv seems to gives a small boost in performance. But I think we can get much more out of CT scans. \n\nInstead of using standard deep learning techniques, using domain specific techniques like the ones mentioned in discussion thread by @sandorkonya can give much bigger boost. Here are some domain specific features that can be extracted from the CT scans: \n- calculating lung volume\n- chest circumference\n- Histogram/kurtosis analysis\n\nI will be conducting an array of experiments to test these techniques and post the updated here in the thread. Also, all the contributions are welcome. If you try some of these features then please post the results here. It will help everyone.",
    "987063": "I made my notebook public. I looked at whether small patches in the images are important. The results are interesting but inconclusive. My next step is to see if I can improve overall prediction score. ",
    "1001018": "@ankursingh12, thanks for sharing, I will start playing tonight with the CT",
    "999066": "This seems good! Thank you for sharing. I will try the autoencoder approach .",
    "984633": "> calculating lung volume\n\nSo, I tried to convert dicom 2d slice data to 3d voxel data (numpy array npy file), but the following error occurs [here](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/176879).\n\nSomebody please help!"
  }
}