{
  "id": 180161,
  "title": "gdcm issue when submitting the notebook",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/180161",
  "author_name": "",
  "post_date": "2020-09-04T05:17:40.685849500Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I noticed that a common way to resolve the gdcm error is </p>\n<p><code>!conda install -c conda-forge gdcm -y</code></p>\n<p>However, we are not allowed to have internet when submitting the notebook, so this solution should not work when submitting for the score, I saw many discussions talked about it but seems no one has a solution, so I would like to bring it again to see if anyone figured it out recently.</p>\n<p>OR: although a different test set is used to calculate the score, actually the Patient ID of the hidden data is included in the train data (like the current Patient ID in submission.csv and test folder), and only the week information is different or maybe there are more patients but all the patient IDs are still in the train data? If this is the case, we actually can upload the processed results of the images and use them directly for the test data as long as the IDs match, since the image only contains the information of Week 0 so any data with the same Patient ID should share the same images, though the data have different FVC measurement in different weeks?</p>\n<p>Please correct me if I am wrong, thanks!</p>",
  "messages": [
    {
      "id": "997525",
      "postDate": "09/04/2020 05:17:40",
      "content": "<p>I noticed that a common way to resolve the gdcm error is </p>\n<p><code>!conda install -c conda-forge gdcm -y</code></p>\n<p>However, we are not allowed to have internet when submitting the notebook, so this solution should not work when submitting for the score, I saw many discussions talked about it but seems no one has a solution, so I would like to bring it again to see if anyone figured it out recently.</p>\n<p>OR: although a different test set is used to calculate the score, actually the Patient ID of the hidden data is included in the train data (like the current Patient ID in submission.csv and test folder), and only the week information is different or maybe there are more patients but all the patient IDs are still in the train data? If this is the case, we actually can upload the processed results of the images and use them directly for the test data as long as the IDs match, since the image only contains the information of Week 0 so any data with the same Patient ID should share the same images, though the data have different FVC measurement in different weeks?</p>\n<p>Please correct me if I am wrong, thanks!</p>",
      "rawMarkdown": "I noticed that a common way to resolve the gdcm error is \n\n```!conda install -c conda-forge gdcm -y```\n\nHowever, we are not allowed to have internet when submitting the notebook, so this solution should not work when submitting for the score, I saw many discussions talked about it but seems no one has a solution, so I would like to bring it again to see if anyone figured it out recently.\n\nOR: although a different test set is used to calculate the score, actually the Patient ID of the hidden data is included in the train data (like the current Patient ID in submission.csv and test folder), and only the week information is different or maybe there are more patients but all the patient IDs are still in the train data? If this is the case, we actually can upload the processed results of the images and use them directly for the test data as long as the IDs match, since the image only contains the information of Week 0 so any data with the same Patient ID should share the same images, though the data have different FVC measurement in different weeks?\n\nPlease correct me if I am wrong, thanks!",
      "votes": null
    },
    {
      "id": "997807",
      "postDate": "09/04/2020 08:56:22",
      "content": "<p><a href=\"https://www.kaggle.com/skyleov\" target=\"_blank\">@skyleov</a> What i did was that i split my training and inference code since the competition organizers mentioned GDCM is not needed for the hidden test set. Even in the train data only a few dicom files cause this issue. Therefore, I pre-process the train dicom files and save them as numpy arrays in a new dataset which i use in my train notebook to speed up training. In the inference notebook i include the entire pre-processing pipeline which i submit to the competition.</p>\n<p>P.S. The hidden test set Patient ID's and train data Patient ID's won't match because if some of the hidden test data was available to participants it would allow people to probe the public LB..</p>",
      "rawMarkdown": "skyleov What i did was that i split my training and inference code since the competition organizers mentioned GDCM is not needed for the hidden test set. Even in the train data only a few dicom files cause this issue. Therefore, I pre-process the train dicom files and save them as numpy arrays in a new dataset which i use in my train notebook to speed up training. In the inference notebook i include the entire pre-processing pipeline which i submit to the competition.\n\nP.S. The hidden test set Patient ID's and train data Patient ID's won't match because if some of the hidden test data was available to participants it would allow people to probe the public LB..",
      "votes": null
    },
    {
      "id": "998086",
      "postDate": "09/04/2020 13:40:07",
      "content": "<p>Thanks for the suggestion! If there is no damaged image in the hidden test data, then GDCM is not needed</p>",
      "rawMarkdown": "Thanks for the suggestion! If there is no damaged image in the hidden test data, then GDCM is not needed",
      "votes": null
    },
    {
      "id": "998098",
      "postDate": "09/04/2020 13:55:09",
      "content": "<p>Indeed, hidden set images are readable without GDCM. It is needed for only two scans in the training set.</p>",
      "rawMarkdown": "Indeed, hidden set images are readable without GDCM. It is needed for only two scans in the training set.",
      "votes": null
    },
    {
      "id": "998736",
      "postDate": "09/05/2020 02:56:15",
      "content": "<p>I have another question: if the hidden set Patient ID is not a subset of the train data Patient ID, how can load the images from the hidden set to the Kaggle notebook when submission?<br>\nFor example, in the local machine, we know that the Patient ID in sample_submission.csv is included in train data, and let's say a sample ID is ID00419637202311204720264, we can scan the image from<br>\n<code>/PATH/TO/train/ID00419637202311204720264/*.dcm</code></p>\n<p>but if the hidden test set is actually hidden, how to load the images?</p>",
      "rawMarkdown": "I have another question: if the hidden set Patient ID is not a subset of the train data Patient ID, how can load the images from the hidden set to the Kaggle notebook when submission?\nFor example, in the local machine, we know that the Patient ID in sample_submission.csv is included in train data, and let's say a sample ID is ID00419637202311204720264, we can scan the image from\n```/PATH/TO/train/ID00419637202311204720264/*.dcm```\n\nbut if the hidden test set is actually hidden, how to load the images?",
      "votes": null
    },
    {
      "id": "998740",
      "postDate": "09/05/2020 03:00:58",
      "content": "<p>uh, I may get it, it is in <code>/PATH/TO/test/xxx/*.dcm</code>, and the unique Patient ID in sample_submission in hidden data set is the same as the Patient ID in test.csv. so my code needs to handle this dynamically </p>\n<p>Please correct me if I am wrong</p>",
      "rawMarkdown": "uh, I may get it, it is in ```/PATH/TO/test/xxx/*.dcm```, and the unique Patient ID in sample_submission in hidden data set is the same as the Patient ID in test.csv. so my code needs to handle this dynamically \n\nPlease correct me if I am wrong",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 997807,
      "author_name": "yovinyahathugoda",
      "author_url": "",
      "post_date": "09/04/2020 08:56:22",
      "content": "<p><a href=\"https://www.kaggle.com/skyleov\" target=\"_blank\">@skyleov</a> What i did was that i split my training and inference code since the competition organizers mentioned GDCM is not needed for the hidden test set. Even in the train data only a few dicom files cause this issue. Therefore, I pre-process the train dicom files and save them as numpy arrays in a new dataset which i use in my train notebook to speed up training. In the inference notebook i include the entire pre-processing pipeline which i submit to the competition.</p>\n<p>P.S. The hidden test set Patient ID's and train data Patient ID's won't match because if some of the hidden test data was available to participants it would allow people to probe the public LB..</p>",
      "votes": null,
      "replies": [
        {
          "id": 998086,
          "author_name": "skyleov",
          "author_url": "",
          "post_date": "09/04/2020 13:40:07",
          "content": "<p>Thanks for the suggestion! If there is no damaged image in the hidden test data, then GDCM is not needed</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998098,
          "author_name": "ahmedhshahin",
          "author_url": "",
          "post_date": "09/04/2020 13:55:09",
          "content": "<p>Indeed, hidden set images are readable without GDCM. It is needed for only two scans in the training set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998736,
          "author_name": "skyleov",
          "author_url": "",
          "post_date": "09/05/2020 02:56:15",
          "content": "<p>I have another question: if the hidden set Patient ID is not a subset of the train data Patient ID, how can load the images from the hidden set to the Kaggle notebook when submission?<br>\nFor example, in the local machine, we know that the Patient ID in sample_submission.csv is included in train data, and let's say a sample ID is ID00419637202311204720264, we can scan the image from<br>\n<code>/PATH/TO/train/ID00419637202311204720264/*.dcm</code></p>\n<p>but if the hidden test set is actually hidden, how to load the images?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 998740,
          "author_name": "skyleov",
          "author_url": "",
          "post_date": "09/05/2020 03:00:58",
          "content": "<p>uh, I may get it, it is in <code>/PATH/TO/test/xxx/*.dcm</code>, and the unique Patient ID in sample_submission in hidden data set is the same as the Patient ID in test.csv. so my code needs to handle this dynamically </p>\n<p>Please correct me if I am wrong</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "997525": "I noticed that a common way to resolve the gdcm error is \n\n```!conda install -c conda-forge gdcm -y```\n\nHowever, we are not allowed to have internet when submitting the notebook, so this solution should not work when submitting for the score, I saw many discussions talked about it but seems no one has a solution, so I would like to bring it again to see if anyone figured it out recently.\n\nOR: although a different test set is used to calculate the score, actually the Patient ID of the hidden data is included in the train data (like the current Patient ID in submission.csv and test folder), and only the week information is different or maybe there are more patients but all the patient IDs are still in the train data? If this is the case, we actually can upload the processed results of the images and use them directly for the test data as long as the IDs match, since the image only contains the information of Week 0 so any data with the same Patient ID should share the same images, though the data have different FVC measurement in different weeks?\n\nPlease correct me if I am wrong, thanks!",
    "997807": "skyleov What i did was that i split my training and inference code since the competition organizers mentioned GDCM is not needed for the hidden test set. Even in the train data only a few dicom files cause this issue. Therefore, I pre-process the train dicom files and save them as numpy arrays in a new dataset which i use in my train notebook to speed up training. In the inference notebook i include the entire pre-processing pipeline which i submit to the competition.\n\nP.S. The hidden test set Patient ID's and train data Patient ID's won't match because if some of the hidden test data was available to participants it would allow people to probe the public LB..",
    "998086": "Thanks for the suggestion! If there is no damaged image in the hidden test data, then GDCM is not needed",
    "998098": "Indeed, hidden set images are readable without GDCM. It is needed for only two scans in the training set.",
    "998736": "I have another question: if the hidden set Patient ID is not a subset of the train data Patient ID, how can load the images from the hidden set to the Kaggle notebook when submission?\nFor example, in the local machine, we know that the Patient ID in sample_submission.csv is included in train data, and let's say a sample ID is ID00419637202311204720264, we can scan the image from\n```/PATH/TO/train/ID00419637202311204720264/*.dcm```\n\nbut if the hidden test set is actually hidden, how to load the images?",
    "998740": "uh, I may get it, it is in ```/PATH/TO/test/xxx/*.dcm```, and the unique Patient ID in sample_submission in hidden data set is the same as the Patient ID in test.csv. so my code needs to handle this dynamically \n\nPlease correct me if I am wrong"
  },
  "source": "meta"
}