{
  "id": 165228,
  "title": "Submission CSV Not Found",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/165228",
  "author_name": "",
  "post_date": "2020-07-08T22:52:35.658848100Z",
  "votes": 15,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi i have a running notebook <a href=\"https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics\">https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics</a> here , but when i submit it i have the following message <strong>Submission CSV Not Found</strong></p>\n\n<p>Quickly that's what i do:</p>\n\n<ul>\n<li><p>for both train or test (from sample submission) i try to load an image idenfied by the key patient-weeks if possible</p></li>\n<li><p>For these images, i trained a simple CNN with tensorflow-keras with the competition metrics</p></li>\n<li><p>For submission i predicted the loaded images with Conv net and for the missing images i use the FVC known in the public test csv (5 lines)</p></li>\n</ul>\n\n<p>So i don't know why the final submission csv is not found. Maybe when the code is run behind the hood there is an issue. But i can't figure it out. So i need your help, help me fix it. Thanks in advance</p>",
  "messages": [
    {
      "id": "920908",
      "postDate": "07/08/2020 22:52:35",
      "content": "<p>Hi i have a running notebook <a href=\"https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics\">https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics</a> here , but when i submit it i have the following message <strong>Submission CSV Not Found</strong></p>\n\n<p>Quickly that's what i do:</p>\n\n<ul>\n<li><p>for both train or test (from sample submission) i try to load an image idenfied by the key patient-weeks if possible</p></li>\n<li><p>For these images, i trained a simple CNN with tensorflow-keras with the competition metrics</p></li>\n<li><p>For submission i predicted the loaded images with Conv net and for the missing images i use the FVC known in the public test csv (5 lines)</p></li>\n</ul>\n\n<p>So i don't know why the final submission csv is not found. Maybe when the code is run behind the hood there is an issue. But i can't figure it out. So i need your help, help me fix it. Thanks in advance</p>",
      "rawMarkdown": "Hi i have a running notebook https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics here , but when i submit it i have the following message **Submission CSV Not Found**\n\nQuickly that's what i do:\n\n- for both train or test (from sample submission) i try to load an image idenfied by the key patient-weeks if possible\n\n- For these images, i trained a simple CNN with tensorflow-keras with the competition metrics\n\n- For submission i predicted the loaded images with Conv net and for the missing images i use the FVC known in the public test csv (5 lines)\n\nSo i don't know why the final submission csv is not found. Maybe when the code is run behind the hood there is an issue. But i can't figure it out. So i need your help, help me fix it. Thanks in advance",
      "votes": null
    },
    {
      "id": "920935",
      "postDate": "07/09/2020 00:03:34",
      "content": "<p>While trying to help you with your problem, I noted a few things worth mentioning.</p>\n\n<p>If I understand your code, you are mis-interpreting the Dicom file names.</p>\n\n<p>The dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.</p>\n\n<p>Each CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.</p>\n\n<p>There is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.</p>\n\n<p>I think your get_images function has \"train\" hardwired:</p>\n\n<p>img_path = f\"{ROOT}/train/{patient}/{week}.dcm\"</p>\n\n<p>This needs to be train or test, depending on your \"how\" variable.</p>\n\n<p>img_path = f\"{ROOT}/{how}/{patient}/{week}.dcm\"</p>\n\n<p>Once I made that change, however, I got an \"exceeded compute error\". You are reading in the dicom images. I don't know if trying to keep the test images in memory all at once is running out of memory. For debugging, decrease DESIRED_SIZE and see if that improves things.</p>\n\n<p>numpy.append makes a copy of the array each time, so that might be consuming your memory (not an expert in this area).</p>\n\n<p>Also, your log file says:</p>\n\n<p>Failed validating 'additionalProperties' in markdown_cell:</p>\n\n<p>I removed all the markdown cells and blank code cells.</p>\n\n<p>I note you have floating point numbers with many significant digits. Maybe round to 5 digits or so. Usually the submission scoring can handle them, but easier to read, and takes away one more unknown.</p>\n\n<p>Good luck,</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "While trying to help you with your problem, I noted a few things worth mentioning.\n\nIf I understand your code, you are mis-interpreting the Dicom file names.\n\nThe dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.\n\nEach CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.\n\nThere is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.\n\nI think your get_images function has \"train\" hardwired:\n\n  img_path = f\"{ROOT}/train/{patient}/{week}.dcm\"\n\nThis needs to be train or test, depending on your \"how\" variable.\n\n  img_path = f\"{ROOT}/{how}/{patient}/{week}.dcm\"\n\n\nOnce I made that change, however, I got an \"exceeded compute error\". You are reading in the dicom images. I don't know if trying to keep the test images in memory all at once is running out of memory. For debugging, decrease DESIRED_SIZE and see if that improves things.\n\nnumpy.append makes a copy of the array each time, so that might be consuming your memory (not an expert in this area).\n\nAlso, your log file says:\n\nFailed validating 'additionalProperties' in markdown_cell:\n\nI removed all the markdown cells and blank code cells.\n\nI note you have floating point numbers with many significant digits. Maybe round to 5 digits or so. Usually the submission scoring can handle them, but easier to read, and takes away one more unknown.\n\n\nGood luck,\n\n-Rich",
      "votes": null
    },
    {
      "id": "920947",
      "postDate": "07/09/2020 00:35:40",
      "content": "<p>Many thanks to you <a href=\"/richardepstein\">@richardepstein</a>  🙏 . By the way do you have an idea on how to make up the whole chest image per patient. </p>",
      "rawMarkdown": "Many thanks to you @richardepstein  🙏 . By the way do you have an idea on how to make up the whole chest image per patient.",
      "votes": null
    },
    {
      "id": "920963",
      "postDate": "07/09/2020 01:38:07",
      "content": "<p>I haven't built any 3-D models. I don't know of public pre-trained models for 3D (although I am sure they exist).</p>\n\n<p>One way to do a 2 1/2 D model would be to combine three slices into a single image, using each slice as a channel. So instead of (512 x 512 x 3) where you have 3 color channels, use those channels for slice1-slice2-slice3.</p>\n\n<p>The image names (1.dcm, 2.dcm, 3.dcm) are usually in the correct order for the images. (although some might be top to bottom and some bottom to top (and I know some are upside down)). You can use the Dicom tag \"Image Position\" to put them in correct order.</p>\n\n<p>-Rich</p>",
      "rawMarkdown": "I haven't built any 3-D models. I don't know of public pre-trained models for 3D (although I am sure they exist).\n\nOne way to do a 2 1/2 D model would be to combine three slices into a single image, using each slice as a channel. So instead of (512 x 512 x 3) where you have 3 color channels, use those channels for slice1-slice2-slice3.\n\nThe image names (1.dcm, 2.dcm, 3.dcm) are usually in the correct order for the images. (although some might be top to bottom and some bottom to top (and I know some are upside down)). You can use the Dicom tag \"Image Position\" to put them in correct order.\n\n-Rich",
      "votes": null
    },
    {
      "id": "920973",
      "postDate": "07/09/2020 01:49:41",
      "content": "<p>Check to see if this helps \n<a href=\"https://www.kaggle.com/vanausloos/full-preprocessing-tutorial\">https://www.kaggle.com/vanausloos/full-preprocessing-tutorial</a></p>\n\n<p>They load slices and stack them. I got this from some of the other threads for this competition </p>",
      "rawMarkdown": "Check to see if this helps \nhttps://www.kaggle.com/vanausloos/full-preprocessing-tutorial\n\nThey load slices and stack them. I got this from some of the other threads for this competition",
      "votes": null
    },
    {
      "id": "921313",
      "postDate": "07/09/2020 07:45:41",
      "content": "<p>I will have a look</p>",
      "rawMarkdown": "I will have a look",
      "votes": null
    },
    {
      "id": "921322",
      "postDate": "07/09/2020 07:59:03",
      "content": "<p>The erro is fixed as you said. Now i will try to correct my image processing blunder.</p>",
      "rawMarkdown": "The erro is fixed as you said. Now i will try to correct my image processing blunder.",
      "votes": null
    },
    {
      "id": "990959",
      "postDate": "08/30/2020 02:04:52",
      "content": "<p>I have the same problem and removed all blank cells and markdowns, but still get this error.. </p>",
      "rawMarkdown": "I have the same problem and removed all blank cells and markdowns, but still get this error..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 920935,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "07/09/2020 00:03:34",
      "content": "<p>While trying to help you with your problem, I noted a few things worth mentioning.</p>\n\n<p>If I understand your code, you are mis-interpreting the Dicom file names.</p>\n\n<p>The dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.</p>\n\n<p>Each CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.</p>\n\n<p>There is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.</p>\n\n<p>I think your get_images function has \"train\" hardwired:</p>\n\n<p>img_path = f\"{ROOT}/train/{patient}/{week}.dcm\"</p>\n\n<p>This needs to be train or test, depending on your \"how\" variable.</p>\n\n<p>img_path = f\"{ROOT}/{how}/{patient}/{week}.dcm\"</p>\n\n<p>Once I made that change, however, I got an \"exceeded compute error\". You are reading in the dicom images. I don't know if trying to keep the test images in memory all at once is running out of memory. For debugging, decrease DESIRED_SIZE and see if that improves things.</p>\n\n<p>numpy.append makes a copy of the array each time, so that might be consuming your memory (not an expert in this area).</p>\n\n<p>Also, your log file says:</p>\n\n<p>Failed validating 'additionalProperties' in markdown_cell:</p>\n\n<p>I removed all the markdown cells and blank code cells.</p>\n\n<p>I note you have floating point numbers with many significant digits. Maybe round to 5 digits or so. Usually the submission scoring can handle them, but easier to read, and takes away one more unknown.</p>\n\n<p>Good luck,</p>\n\n<p>-Rich</p>",
      "votes": null,
      "replies": [
        {
          "id": 920947,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "07/09/2020 00:35:40",
          "content": "<p>Many thanks to you <a href=\"/richardepstein\">@richardepstein</a>  🙏 . By the way do you have an idea on how to make up the whole chest image per patient. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920963,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "07/09/2020 01:38:07",
          "content": "<p>I haven't built any 3-D models. I don't know of public pre-trained models for 3D (although I am sure they exist).</p>\n\n<p>One way to do a 2 1/2 D model would be to combine three slices into a single image, using each slice as a channel. So instead of (512 x 512 x 3) where you have 3 color channels, use those channels for slice1-slice2-slice3.</p>\n\n<p>The image names (1.dcm, 2.dcm, 3.dcm) are usually in the correct order for the images. (although some might be top to bottom and some bottom to top (and I know some are upside down)). You can use the Dicom tag \"Image Position\" to put them in correct order.</p>\n\n<p>-Rich</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 920973,
          "author_name": "venkat555",
          "author_url": "",
          "post_date": "07/09/2020 01:49:41",
          "content": "<p>Check to see if this helps \n<a href=\"https://www.kaggle.com/vanausloos/full-preprocessing-tutorial\">https://www.kaggle.com/vanausloos/full-preprocessing-tutorial</a></p>\n\n<p>They load slices and stack them. I got this from some of the other threads for this competition </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 921313,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "07/09/2020 07:45:41",
          "content": "<p>I will have a look</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 921322,
          "author_name": "ulrich07",
          "author_url": "",
          "post_date": "07/09/2020 07:59:03",
          "content": "<p>The erro is fixed as you said. Now i will try to correct my image processing blunder.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 990959,
          "author_name": "timotta",
          "author_url": "",
          "post_date": "08/30/2020 02:04:52",
          "content": "<p>I have the same problem and removed all blank cells and markdowns, but still get this error.. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920908": "Hi i have a running notebook https://www.kaggle.com/ulrich07/osic-keras-starter-with-custom-metrics here , but when i submit it i have the following message **Submission CSV Not Found**\n\nQuickly that's what i do:\n\n- for both train or test (from sample submission) i try to load an image idenfied by the key patient-weeks if possible\n\n- For these images, i trained a simple CNN with tensorflow-keras with the competition metrics\n\n- For submission i predicted the loaded images with Conv net and for the missing images i use the FVC known in the public test csv (5 lines)\n\nSo i don't know why the final submission csv is not found. Maybe when the code is run behind the hood there is an issue. But i can't figure it out. So i need your help, help me fix it. Thanks in advance",
    "920935": "While trying to help you with your problem, I noted a few things worth mentioning.\n\nIf I understand your code, you are mis-interpreting the Dicom file names.\n\nThe dicom file names (1.dcm, 2. dcm, etc) are slice numbers, not weeks.\n\nEach CT scan was taken in a single day. If you look at the images, you'll see that they make up a whole chest, typically top to bottom.\n\nThere is only one CT scan per patient. The \"weeks\" variable in the train.csv file are the weeks (before or after the CT) that we have a FVC measurement.\n\nI think your get_images function has \"train\" hardwired:\n\n  img_path = f\"{ROOT}/train/{patient}/{week}.dcm\"\n\nThis needs to be train or test, depending on your \"how\" variable.\n\n  img_path = f\"{ROOT}/{how}/{patient}/{week}.dcm\"\n\n\nOnce I made that change, however, I got an \"exceeded compute error\". You are reading in the dicom images. I don't know if trying to keep the test images in memory all at once is running out of memory. For debugging, decrease DESIRED_SIZE and see if that improves things.\n\nnumpy.append makes a copy of the array each time, so that might be consuming your memory (not an expert in this area).\n\nAlso, your log file says:\n\nFailed validating 'additionalProperties' in markdown_cell:\n\nI removed all the markdown cells and blank code cells.\n\nI note you have floating point numbers with many significant digits. Maybe round to 5 digits or so. Usually the submission scoring can handle them, but easier to read, and takes away one more unknown.\n\n\nGood luck,\n\n-Rich",
    "920947": "Many thanks to you @richardepstein  🙏 . By the way do you have an idea on how to make up the whole chest image per patient.",
    "920963": "I haven't built any 3-D models. I don't know of public pre-trained models for 3D (although I am sure they exist).\n\nOne way to do a 2 1/2 D model would be to combine three slices into a single image, using each slice as a channel. So instead of (512 x 512 x 3) where you have 3 color channels, use those channels for slice1-slice2-slice3.\n\nThe image names (1.dcm, 2.dcm, 3.dcm) are usually in the correct order for the images. (although some might be top to bottom and some bottom to top (and I know some are upside down)). You can use the Dicom tag \"Image Position\" to put them in correct order.\n\n-Rich",
    "920973": "Check to see if this helps \nhttps://www.kaggle.com/vanausloos/full-preprocessing-tutorial\n\nThey load slices and stack them. I got this from some of the other threads for this competition",
    "921313": "I will have a look",
    "921322": "The erro is fixed as you said. Now i will try to correct my image processing blunder.",
    "990959": "I have the same problem and removed all blank cells and markdowns, but still get this error.."
  },
  "source": "meta"
}