{
  "id": 170392,
  "title": "How image data is used for training?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/170392",
  "author_name": "Vijay Yadav",
  "post_date": "2020-07-27T13:30:48.539000",
  "votes": 11,
  "comment_count": 18,
  "views": 0,
  "content": "<p>I went through a few notebooks but didn't find out how they are using image data for training purposes. can anyone please explain what is going on?</p>",
  "messages": [
    {
      "id": 947775,
      "postDate": "2020-07-27T13:30:48.540Z",
      "content": "<p>I went through a few notebooks but didn't find out how they are using image data for training purposes. can anyone please explain what is going on?</p>",
      "rawMarkdown": "I went through a few notebooks but didn't find out how they are using image data for training purposes. can anyone please explain what is going on?",
      "votes": 10
    },
    {
      "id": 948672,
      "postDate": "2020-07-28T06:28:38.930Z",
      "content": "<p>The images are DICOM images of one single CT Scan. CT scans are not images rather videos, so the dataset contains frames of a single CT scan video.</p>",
      "rawMarkdown": "The images are DICOM images of one single CT Scan. CT scans are not images rather videos, so the dataset contains frames of a single CT scan video.",
      "votes": 1,
      "replies": [
        {
          "id": 948695,
          "postDate": "2020-07-28T06:43:07.073Z",
          "content": "<p>That's not correct, CT scans (as long as they are not sequential) are three-dimensional images when stacked.</p>",
          "rawMarkdown": "That's not correct, CT scans (as long as they are not sequential) are three-dimensional images when stacked.",
          "votes": 2,
          "replies": [
            {
              "id": 948723,
              "postDate": "2020-07-28T07:03:01.917Z",
              "content": "<p>Thanks for the reply <a href=\"/hojijoji\">@hojijoji</a> , i should correct my self. The images in the dataset are images of a single CT scan. CT scans are, as <a href=\"/hojijoji\">@hojijoji</a> said are 3D images when stacked, but in the dataset all the stacked images are individually present. \nFor Example, this GIF shows the CT scan with Patient Id ID00012637202177665765362 - <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2F6d456b529856448fbca7c4480f4a0bb1%2FID00012637202177665765362.gif?generation=1595919768770348&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thanks for the reply @hojijoji , i should correct my self. The images in the dataset are images of a single CT scan. CT scans are, as @hojijoji said are 3D images when stacked, but in the dataset all the stacked images are individually present. \nFor Example, this GIF shows the CT scan with Patient Id ID00012637202177665765362 -  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2F6d456b529856448fbca7c4480f4a0bb1%2FID00012637202177665765362.gif?generation=1595919768770348&amp;alt=media)\n"
            }
          ]
        },
        {
          "id": 948921,
          "postDate": "2020-07-28T10:20:22.997Z",
          "content": "<p>Thank you for replying <a href=\"/hojijoji\">@hojijoji</a> , I should correct myself, The dataset contains images of a single CT scan, and CT scans, as <a href=\"/hojijoji\">@hojijoji</a> said are 3 Dimensional images when stacked. In the given dataset we have individual images of a single CT scan of a Patient.</p>\n\n<p>For example, This is a GIF of CT scan of Patient ID00012637202177665765362,</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2Fda1ab4df2955500ba5c1f33541e4c739%2FID00012637202177665765362.gif?generation=1595931619408140&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you for replying @hojijoji , I should correct myself, The dataset contains images of a single CT scan, and CT scans, as @hojijoji said are 3 Dimensional images when stacked. In the given dataset we have individual images of a single CT scan of a Patient.\n\nFor example, This is a GIF of CT scan of Patient ID00012637202177665765362,\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2Fda1ab4df2955500ba5c1f33541e4c739%2FID00012637202177665765362.gif?generation=1595931619408140&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 948224,
      "postDate": "2020-07-27T18:44:01.677Z",
      "content": "<p><a href=\"/vijayyadav22\">@vijayyadav22</a> , \nas <a href=\"/hojijoji\">@hojijoji</a> mentioned in the comment section of <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">this discusion</a>, the image data of this competition has many flaws... heterogenous, totally different scan protocolls, some scans have even sequential images (meaning: not aquired within a single breathold maneuver).</p>\n\n<p>I am completely on the same opinion (especially the second one):</p>\n\n<blockquote>\n  <p>(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner).\n  (2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development\n  may go into anticipating data corruptions and artifacts rather than the real thing.</p>\n</blockquote>",
      "rawMarkdown": "@vijayyadav22 , \nas @hojijoji mentioned in the comment section of [this discusion](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044), the image data of this competition has many flaws... heterogenous, totally different scan protocolls, some scans have even sequential images (meaning: not aquired within a single breathold maneuver).\n\nI am completely on the same opinion (especially the second one):\n\n&gt; (1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner).\n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development\nmay go into anticipating data corruptions and artifacts rather than the real thing.\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 949301,
          "postDate": "2020-07-28T14:40:29.037Z",
          "content": "<p>Sir, can you please explain to me how to use these image data in training the model?</p>",
          "rawMarkdown": "Sir, can you please explain to me how to use these image data in training the model?"
        },
        {
          "id": 952196,
          "postDate": "2020-07-30T17:46:09.987Z",
          "content": "<p>if you know how to work with 3D Conv nets you can easily use these images for training.\nbut first, you have to pre-process these images as already mentioned above ~100x512x512 is a huge number and 3D convnet cannot handle that much data so first you have to reduce this number say 30x150x150 although it's still big.\nOnce you have these preprocessed data you can a model with two parallel branches\nBranch 1 -- will process and learn on tabular data (Dense layers)\nBranch 2 -- will process image data (3D Conv net)\nthen you can concatenate these branches together to form a single model use Keras functional API\nfollowing is my notebook which goes through preprocessing the Dicom data and saves the processed training images in npy format.\n<a href=\"https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data\">preprocessing dicom data</a>\n it has reduced each scan to a size of 30x150x150 which is still huge but you can use this approach to reduce the size as much as you want</p>",
          "rawMarkdown": "if you know how to work with 3D Conv nets you can easily use these images for training.\nbut first, you have to pre-process these images as already mentioned above ~100x512x512 is a huge number and 3D convnet cannot handle that much data so first you have to reduce this number say 30x150x150 although it's still big.\nOnce you have these preprocessed data you can a model with two parallel branches\nBranch 1 -- will process and learn on tabular data (Dense layers)\nBranch 2 -- will process image data (3D Conv net)\nthen you can concatenate these branches together to form a single model use Keras functional API\nfollowing is my notebook which goes through preprocessing the Dicom data and saves the processed training images in npy format.\n[preprocessing dicom data](https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data)\n it has reduced each scan to a size of 30x150x150 which is still huge but you can use this approach to reduce the size as much as you want",
          "votes": 3
        }
      ]
    },
    {
      "id": 949297,
      "postDate": "2020-07-28T14:39:02.287Z",
      "content": "<p>Thank you everyone, this discussion was very helpful in understanding the image but still, I do not understand whether image data was used in model training or not? </p>",
      "rawMarkdown": "Thank you everyone, this discussion was very helpful in understanding the image but still, I do not understand whether image data was used in model training or not? ",
      "replies": [
        {
          "id": 949302,
          "postDate": "2020-07-28T14:42:19.833Z",
          "content": "<p>The only public one with lots of upvotes is <a href=\"/carlossouza\">@carlossouza</a>: <a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\">https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular</a></p>\n\n<p>Though I suspect lots of us (myself including) are in the background trying to build models that use image data and improve on the score.</p>",
          "rawMarkdown": "The only public one with lots of upvotes is @carlossouza: https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\n\nThough I suspect lots of us (myself including) are in the background trying to build models that use image data and improve on the score.",
          "votes": 3
        },
        {
          "id": 949315,
          "postDate": "2020-07-28T14:54:22.743Z",
          "content": "<p>The submission you made does include image data for training?</p>",
          "rawMarkdown": "The submission you made does include image data for training?"
        },
        {
          "id": 949365,
          "postDate": "2020-07-28T15:33:40.513Z",
          "content": "<p>not got one that improves on the tabular data. The problem is that the image data has such high dimensionality (~100x512x512=25,000,000 features!) that using the image data increases the likelihood of simply overfitting the training data. That's why approaches like the one above try to reduce the dimensionality first.</p>",
          "rawMarkdown": "not got one that improves on the tabular data. The problem is that the image data has such high dimensionality (~100x512x512=25,000,000 features!) that using the image data increases the likelihood of simply overfitting the training data. That's why approaches like the one above try to reduce the dimensionality first.",
          "votes": 1
        },
        {
          "id": 949376,
          "postDate": "2020-07-28T15:43:06.997Z",
          "content": "<p>Thank you James for helping me out.</p>",
          "rawMarkdown": "Thank you James for helping me out."
        }
      ]
    },
    {
      "id": 986456,
      "postDate": "2020-08-26T14:12:29.413Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 980132,
      "postDate": "2020-08-21T10:52:20.680Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 980200,
          "postDate": "2020-08-21T11:48:33.137Z",
          "content": "<p>Firstly, The image data is just a single 3D scans. <br>\nSecondly, The CT Scans were taken on Week 0.<br>\nSince some training samples do not have FVC value for Week 0 then in that case you have to apply some measures - either you use the data and use the nearest week FVC as the target value or any other method. Also you can opt to not train over those samples which does not contain Week0 FVC value.</p>",
          "rawMarkdown": "Firstly, The image data is just a single 3D scans. \nSecondly, The CT Scans were taken on Week 0.\nSince some training samples do not have FVC value for Week 0 then in that case you have to apply some measures - either you use the data and use the nearest week FVC as the target value or any other method. Also you can opt to not train over those samples which does not contain Week0 FVC value."
        },
        {
          "id": 984879,
          "postDate": "2020-08-25T10:53:46.923Z",
          "content": "<p>Yes. Help needed! </p>\n<p>I can not understand the same point. How can you use a single 3D scan from the very first day with all the records of a particular patient in <strong>train.csv</strong>?</p>",
          "rawMarkdown": "Yes. Help needed! \n\nI can not understand the same point. How can you use a single 3D scan from the very first day with all the records of a particular patient in **train.csv**?"
        },
        {
          "id": 995298,
          "postDate": "2020-09-02T10:32:46.550Z",
          "content": "<p>Here the task is to predict the FVC values for the final 3 measurements of the patient. The challenge is to find a pattern in the first CT Scan in order to predict the FVC values. for example, you get your first CT scan of your lungs, the doctor uses it to predict your final three FVC measurements to be prepared in advance for best possible results and to provide the best possible treatment.</p>",
          "rawMarkdown": "Here the task is to predict the FVC values for the final 3 measurements of the patient. The challenge is to find a pattern in the first CT Scan in order to predict the FVC values. for example, you get your first CT scan of your lungs, the doctor uses it to predict your final three FVC measurements to be prepared in advance for best possible results and to provide the best possible treatment.\n\n"
        }
      ]
    },
    {
      "id": 953613,
      "postDate": "2020-07-31T23:15:09.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 948672,
      "author_name": "Sambhav Garg",
      "author_url": "",
      "post_date": "2020-07-28T06:28:38.930000",
      "content": "<p>The images are DICOM images of one single CT Scan. CT scans are not images rather videos, so the dataset contains frames of a single CT scan video.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 948695,
          "author_name": "Johannes Hofmanninger",
          "author_url": "",
          "post_date": "2020-07-28T06:43:07.073000",
          "content": "<p>That's not correct, CT scans (as long as they are not sequential) are three-dimensional images when stacked.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 948723,
              "author_name": "Sambhav Garg",
              "author_url": "",
              "post_date": "2020-07-28T07:03:01.917000",
              "content": "<p>Thanks for the reply <a href=\"/hojijoji\">@hojijoji</a> , i should correct my self. The images in the dataset are images of a single CT scan. CT scans are, as <a href=\"/hojijoji\">@hojijoji</a> said are 3D images when stacked, but in the dataset all the stacked images are individually present. \nFor Example, this GIF shows the CT scan with Patient Id ID00012637202177665765362 - <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2F6d456b529856448fbca7c4480f4a0bb1%2FID00012637202177665765362.gif?generation=1595919768770348&amp;alt=media\" alt=\"\"></p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 948921,
          "author_name": "Sambhav Garg",
          "author_url": "",
          "post_date": "2020-07-28T10:20:22.997000",
          "content": "<p>Thank you for replying <a href=\"/hojijoji\">@hojijoji</a> , I should correct myself, The dataset contains images of a single CT scan, and CT scans, as <a href=\"/hojijoji\">@hojijoji</a> said are 3 Dimensional images when stacked. In the given dataset we have individual images of a single CT scan of a Patient.</p>\n\n<p>For example, This is a GIF of CT scan of Patient ID00012637202177665765362,</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2841366%2Fda1ab4df2955500ba5c1f33541e4c739%2FID00012637202177665765362.gif?generation=1595931619408140&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 948224,
      "author_name": "dr. Konya",
      "author_url": "",
      "post_date": "2020-07-27T18:44:01.677000",
      "content": "<p><a href=\"/vijayyadav22\">@vijayyadav22</a> , \nas <a href=\"/hojijoji\">@hojijoji</a> mentioned in the comment section of <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044\">this discusion</a>, the image data of this competition has many flaws... heterogenous, totally different scan protocolls, some scans have even sequential images (meaning: not aquired within a single breathold maneuver).</p>\n\n<p>I am completely on the same opinion (especially the second one):</p>\n\n<blockquote>\n  <p>(1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner).\n  (2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development\n  may go into anticipating data corruptions and artifacts rather than the real thing.</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 949301,
          "author_name": "Vijay Yadav",
          "author_url": "",
          "post_date": "2020-07-28T14:40:29.037000",
          "content": "<p>Sir, can you please explain to me how to use these image data in training the model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 952196,
          "author_name": "Zain Ahmad",
          "author_url": "",
          "post_date": "2020-07-30T17:46:09.987000",
          "content": "<p>if you know how to work with 3D Conv nets you can easily use these images for training.\nbut first, you have to pre-process these images as already mentioned above ~100x512x512 is a huge number and 3D convnet cannot handle that much data so first you have to reduce this number say 30x150x150 although it's still big.\nOnce you have these preprocessed data you can a model with two parallel branches\nBranch 1 -- will process and learn on tabular data (Dense layers)\nBranch 2 -- will process image data (3D Conv net)\nthen you can concatenate these branches together to form a single model use Keras functional API\nfollowing is my notebook which goes through preprocessing the Dicom data and saves the processed training images in npy format.\n<a href=\"https://www.kaggle.com/zainahmad/preprocessing-the-dicom-data\">preprocessing dicom data</a>\n it has reduced each scan to a size of 30x150x150 which is still huge but you can use this approach to reduce the size as much as you want</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 949297,
      "author_name": "Vijay Yadav",
      "author_url": "",
      "post_date": "2020-07-28T14:39:02.287000",
      "content": "<p>Thank you everyone, this discussion was very helpful in understanding the image but still, I do not understand whether image data was used in model training or not? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 949302,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-07-28T14:42:19.833000",
          "content": "<p>The only public one with lots of upvotes is <a href=\"/carlossouza\">@carlossouza</a>: <a href=\"https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular\">https://www.kaggle.com/carlossouza/end-to-end-model-ct-scans-tabular</a></p>\n\n<p>Though I suspect lots of us (myself including) are in the background trying to build models that use image data and improve on the score.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 949315,
          "author_name": "Vijay Yadav",
          "author_url": "",
          "post_date": "2020-07-28T14:54:22.743000",
          "content": "<p>The submission you made does include image data for training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949365,
          "author_name": "jameschapman19",
          "author_url": "",
          "post_date": "2020-07-28T15:33:40.513000",
          "content": "<p>not got one that improves on the tabular data. The problem is that the image data has such high dimensionality (~100x512x512=25,000,000 features!) that using the image data increases the likelihood of simply overfitting the training data. That's why approaches like the one above try to reduce the dimensionality first.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 949376,
          "author_name": "Vijay Yadav",
          "author_url": "",
          "post_date": "2020-07-28T15:43:06.997000",
          "content": "<p>Thank you James for helping me out.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 986456,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-26T14:12:29.413000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 980132,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-21T10:52:20.680000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 980200,
          "author_name": "Sambhav Garg",
          "author_url": "",
          "post_date": "2020-08-21T11:48:33.137000",
          "content": "<p>Firstly, The image data is just a single 3D scans. <br>\nSecondly, The CT Scans were taken on Week 0.<br>\nSince some training samples do not have FVC value for Week 0 then in that case you have to apply some measures - either you use the data and use the nearest week FVC as the target value or any other method. Also you can opt to not train over those samples which does not contain Week0 FVC value.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 984879,
          "author_name": "AnkurSingh",
          "author_url": "",
          "post_date": "2020-08-25T10:53:46.923000",
          "content": "<p>Yes. Help needed! </p>\n<p>I can not understand the same point. How can you use a single 3D scan from the very first day with all the records of a particular patient in <strong>train.csv</strong>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 995298,
          "author_name": "Sambhav Garg",
          "author_url": "",
          "post_date": "2020-09-02T10:32:46.550000",
          "content": "<p>Here the task is to predict the FVC values for the final 3 measurements of the patient. The challenge is to find a pattern in the first CT Scan in order to predict the FVC values. for example, you get your first CT scan of your lungs, the doctor uses it to predict your final three FVC measurements to be prepared in advance for best possible results and to provide the best possible treatment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 953613,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-31T23:15:09.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "947775": "I went through a few notebooks but didn't find out how they are using image data for training purposes. can anyone please explain what is going on?",
    "948672": "The images are DICOM images of one single CT Scan. CT scans are not images rather videos, so the dataset contains frames of a single CT scan video.",
    "948224": "@vijayyadav22 , \nas @hojijoji mentioned in the comment section of [this discusion](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165044), the image data of this competition has many flaws... heterogenous, totally different scan protocolls, some scans have even sequential images (meaning: not aquired within a single breathold maneuver).\n\nI am completely on the same opinion (especially the second one):\n\n&gt; (1) The best performing algorithm will perform worse compared to what would be possible if clean unmodified data would have been used (data from the scanner).\n(2) The best performing algorithm wins not because it is the best in predicting FVC but because it is good in spotting flawed scans. E.g. there is a corrupted scan in the hidden test set and one algorithm wins because it detects the corruption and predicts the population average. An algorithm detects the wrong orientation etc. A lot of development\nmay go into anticipating data corruptions and artifacts rather than the real thing.\n\n",
    "949297": "Thank you everyone, this discussion was very helpful in understanding the image but still, I do not understand whether image data was used in model training or not? ",
    "986456": "",
    "980132": "",
    "953613": ""
  }
}