{
  "id": 276502,
  "title": "Notebook Threw Exception followed by Submission Scoring Error",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/276502",
  "author_name": "",
  "post_date": "2021-10-05T00:41:29.124256700Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I got into this trouble and can not figure it out for a long time. First, I submitted a notebook that will load a trainned model and predict with the preprocessed testing data, a <em>submission.csv</em> is output at last. Then I got into the error of <code>Notebook Threw Exception</code>. That is strange because I have <em>save and run all</em> the notebook before submission and everything is ok.</p>\n<p>Then I tried with a notebook that directly loads a <em>submission.csv</em> file and then output it. However, I kept getting the error <code>Submission Scoring Error</code> this time. Here is the link of this notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/formtyan/submit-only-0-728-valid-auc/notebook</a> </p>\n<p>Accoring to kaggle's website<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a>, it seems that it will first substitute the original data set with its testing data set and then give a score. Should I include the whole work flow(data preprocessing, model initialization, predicting)?  </p>",
  "messages": [
    {
      "id": "1534481",
      "postDate": "10/05/2021 00:41:29",
      "content": "<p>I got into this trouble and can not figure it out for a long time. First, I submitted a notebook that will load a trainned model and predict with the preprocessed testing data, a <em>submission.csv</em> is output at last. Then I got into the error of <code>Notebook Threw Exception</code>. That is strange because I have <em>save and run all</em> the notebook before submission and everything is ok.</p>\n<p>Then I tried with a notebook that directly loads a <em>submission.csv</em> file and then output it. However, I kept getting the error <code>Submission Scoring Error</code> this time. Here is the link of this notebook: <a href=\"url\" target=\"_blank\">https://www.kaggle.com/formtyan/submit-only-0-728-valid-auc/notebook</a> </p>\n<p>Accoring to kaggle's website<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a>, it seems that it will first substitute the original data set with its testing data set and then give a score. Should I include the whole work flow(data preprocessing, model initialization, predicting)?  </p>",
      "rawMarkdown": "I got into this trouble and can not figure it out for a long time. First, I submitted a notebook that will load a trainned model and predict with the preprocessed testing data, a *submission.csv* is output at last. Then I got into the error of `Notebook Threw Exception`. That is strange because I have *save and run all* the notebook before submission and everything is ok.\n\nThen I tried with a notebook that directly loads a *submission.csv* file and then output it. However, I kept getting the error `Submission Scoring Error` this time. Here is the link of this notebook: [https://www.kaggle.com/formtyan/submit-only-0-728-valid-auc/notebook](url) \n\nAccoring to kaggle's website[https://www.kaggle.com/code-competition-debugging](url), it seems that it will first substitute the original data set with its testing data set and then give a score. Should I include the whole work flow(data preprocessing, model initialization, predicting)?",
      "votes": null
    },
    {
      "id": "1534518",
      "postDate": "10/05/2021 02:06:50",
      "content": "<p>Yes, since hidden test set will change during save and submit you need to assume there will be new data and you need to implement all the inference pipeline in the kernel including preprocessing, model loading and predicting. Checking other inference kernels might help.</p>",
      "rawMarkdown": "Yes, since hidden test set will change during save and submit you need to assume there will be new data and you need to implement all the inference pipeline in the kernel including preprocessing, model loading and predicting. Checking other inference kernels might help.",
      "votes": null
    },
    {
      "id": "1534593",
      "postDate": "10/05/2021 04:41:17",
      "content": "<p>Thanks a lot for your reply. I went through some others' notebooks and am still some confused about what should I include in my notbook for submission. </p>\n<p>Are they just substitute the test folder and the corresponding sample_submission.csv file(so the image id will also get changed as the new test folder) and then run the whole script again? </p>",
      "rawMarkdown": "Thanks a lot for your reply. I went through some others' notebooks and am still some confused about what should I include in my notbook for submission. \n\nAre they just substitute the test folder and the corresponding sample_submission.csv file(so the image id will also get changed as the new test folder) and then run the whole script again?",
      "votes": null
    },
    {
      "id": "1535495",
      "postDate": "10/05/2021 21:44:49",
      "content": "<p>Can we have 2 kernels: 1 CPU kernel for preprocessing data, 1GPU kernel for inference/prediction?</p>",
      "rawMarkdown": "Can we have 2 kernels: 1 CPU kernel for preprocessing data, 1GPU kernel for inference/prediction?",
      "votes": null
    },
    {
      "id": "1537770",
      "postDate": "10/07/2021 18:25:09",
      "content": "<p>Have anyone found something unexpected with the test data set? After continuous submission these days trying to find where the error is, it seems that my assumption about the <code>ImageOrientationPatient</code> data in some unseen test is actually wrong.</p>\n<p>My preprocessing step is under the assumption that the <code>ImageOrientationPatient</code> has the following form  in <code>x1, y1, x2, y2</code>: <code>[+-1, 0, 0, 0]</code>, <code>[0, +-1, 0, 0]</code>, '[+-1, 0, 0, +-1]', which holds on all the available data set. This step will raise an error when the <code>x1, y1, x2, y2</code> in not one of the three cases. My submission will always get into the state <code>Notebook Threw Exception</code> after running for several minutes. The submission will success when I comment out the code about raising exception.</p>\n<p>I am not sure about the above conclusion since there is no extral information on the error. If you also got into a similar error, then you can have a try. </p>\n<p>I am goint to include the code for the unexpected case in my preprocessing code and try again.</p>",
      "rawMarkdown": "Have anyone found something unexpected with the test data set? After continuous submission these days trying to find where the error is, it seems that my assumption about the `ImageOrientationPatient` data in some unseen test is actually wrong.\n\nMy preprocessing step is under the assumption that the `ImageOrientationPatient` has the following form  in `x1, y1, x2, y2`: `[+-1, 0, 0, 0]`, `[0, +-1, 0, 0]`, '[+-1, 0, 0, +-1]', which holds on all the available data set. This step will raise an error when the `x1, y1, x2, y2` in not one of the three cases. My submission will always get into the state `Notebook Threw Exception` after running for several minutes. The submission will success when I comment out the code about raising exception.\n\nI am not sure about the above conclusion since there is no extral information on the error. If you also got into a similar error, then you can have a try. \n\nI am goint to include the code for the unexpected case in my preprocessing code and try again.",
      "votes": null
    },
    {
      "id": "1537800",
      "postDate": "10/07/2021 19:26:06",
      "content": "<p>I may not get what you mean in your question. I think you may be facing the problem that the <code>save &amp; run all</code> process together with the submission process takes too much time. My solution for this is to use multi-workers in dataloader(P.S I am using pytorch for this competition). Since in mycode, prediction does not take much time while data preprocessing is very slow.</p>\n<p>According to Kaggle's website(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu</a>),  <em>GPU backed instances have less CPU power and RAM</em> . I only used CPU in the script for submission, since the preprocessing code uses CPU. (However, I find the available CPU memory is almost unchanged no matter if I turned the GPU on)</p>",
      "rawMarkdown": "I may not get what you mean in your question. I think you may be facing the problem that the `save & run all` process together with the submission process takes too much time. My solution for this is to use multi-workers in dataloader(P.S I am using pytorch for this competition). Since in mycode, prediction does not take much time while data preprocessing is very slow.\n\nAccording to Kaggle's website([https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu](url)),  *GPU backed instances have less CPU power and RAM* . I only used CPU in the script for submission, since the preprocessing code uses CPU. (However, I find the available CPU memory is almost unchanged no matter if I turned the GPU on)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1534518,
      "author_name": "keremt",
      "author_url": "",
      "post_date": "10/05/2021 02:06:50",
      "content": "<p>Yes, since hidden test set will change during save and submit you need to assume there will be new data and you need to implement all the inference pipeline in the kernel including preprocessing, model loading and predicting. Checking other inference kernels might help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1534593,
          "author_name": "formtyan",
          "author_url": "",
          "post_date": "10/05/2021 04:41:17",
          "content": "<p>Thanks a lot for your reply. I went through some others' notebooks and am still some confused about what should I include in my notbook for submission. </p>\n<p>Are they just substitute the test folder and the corresponding sample_submission.csv file(so the image id will also get changed as the new test folder) and then run the whole script again? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1535495,
      "author_name": "saruaralam",
      "author_url": "",
      "post_date": "10/05/2021 21:44:49",
      "content": "<p>Can we have 2 kernels: 1 CPU kernel for preprocessing data, 1GPU kernel for inference/prediction?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1537800,
          "author_name": "formtyan",
          "author_url": "",
          "post_date": "10/07/2021 19:26:06",
          "content": "<p>I may not get what you mean in your question. I think you may be facing the problem that the <code>save &amp; run all</code> process together with the submission process takes too much time. My solution for this is to use multi-workers in dataloader(P.S I am using pytorch for this competition). Since in mycode, prediction does not take much time while data preprocessing is very slow.</p>\n<p>According to Kaggle's website(<a href=\"url\" target=\"_blank\">https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu</a>),  <em>GPU backed instances have less CPU power and RAM</em> . I only used CPU in the script for submission, since the preprocessing code uses CPU. (However, I find the available CPU memory is almost unchanged no matter if I turned the GPU on)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1537770,
      "author_name": "formtyan",
      "author_url": "",
      "post_date": "10/07/2021 18:25:09",
      "content": "<p>Have anyone found something unexpected with the test data set? After continuous submission these days trying to find where the error is, it seems that my assumption about the <code>ImageOrientationPatient</code> data in some unseen test is actually wrong.</p>\n<p>My preprocessing step is under the assumption that the <code>ImageOrientationPatient</code> has the following form  in <code>x1, y1, x2, y2</code>: <code>[+-1, 0, 0, 0]</code>, <code>[0, +-1, 0, 0]</code>, '[+-1, 0, 0, +-1]', which holds on all the available data set. This step will raise an error when the <code>x1, y1, x2, y2</code> in not one of the three cases. My submission will always get into the state <code>Notebook Threw Exception</code> after running for several minutes. The submission will success when I comment out the code about raising exception.</p>\n<p>I am not sure about the above conclusion since there is no extral information on the error. If you also got into a similar error, then you can have a try. </p>\n<p>I am goint to include the code for the unexpected case in my preprocessing code and try again.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1534481": "I got into this trouble and can not figure it out for a long time. First, I submitted a notebook that will load a trainned model and predict with the preprocessed testing data, a *submission.csv* is output at last. Then I got into the error of `Notebook Threw Exception`. That is strange because I have *save and run all* the notebook before submission and everything is ok.\n\nThen I tried with a notebook that directly loads a *submission.csv* file and then output it. However, I kept getting the error `Submission Scoring Error` this time. Here is the link of this notebook: [https://www.kaggle.com/formtyan/submit-only-0-728-valid-auc/notebook](url) \n\nAccoring to kaggle's website[https://www.kaggle.com/code-competition-debugging](url), it seems that it will first substitute the original data set with its testing data set and then give a score. Should I include the whole work flow(data preprocessing, model initialization, predicting)?",
    "1534518": "Yes, since hidden test set will change during save and submit you need to assume there will be new data and you need to implement all the inference pipeline in the kernel including preprocessing, model loading and predicting. Checking other inference kernels might help.",
    "1534593": "Thanks a lot for your reply. I went through some others' notebooks and am still some confused about what should I include in my notbook for submission. \n\nAre they just substitute the test folder and the corresponding sample_submission.csv file(so the image id will also get changed as the new test folder) and then run the whole script again?",
    "1535495": "Can we have 2 kernels: 1 CPU kernel for preprocessing data, 1GPU kernel for inference/prediction?",
    "1537770": "Have anyone found something unexpected with the test data set? After continuous submission these days trying to find where the error is, it seems that my assumption about the `ImageOrientationPatient` data in some unseen test is actually wrong.\n\nMy preprocessing step is under the assumption that the `ImageOrientationPatient` has the following form  in `x1, y1, x2, y2`: `[+-1, 0, 0, 0]`, `[0, +-1, 0, 0]`, '[+-1, 0, 0, +-1]', which holds on all the available data set. This step will raise an error when the `x1, y1, x2, y2` in not one of the three cases. My submission will always get into the state `Notebook Threw Exception` after running for several minutes. The submission will success when I comment out the code about raising exception.\n\nI am not sure about the above conclusion since there is no extral information on the error. If you also got into a similar error, then you can have a try. \n\nI am goint to include the code for the unexpected case in my preprocessing code and try again.",
    "1537800": "I may not get what you mean in your question. I think you may be facing the problem that the `save & run all` process together with the submission process takes too much time. My solution for this is to use multi-workers in dataloader(P.S I am using pytorch for this competition). Since in mycode, prediction does not take much time while data preprocessing is very slow.\n\nAccording to Kaggle's website([https://www.kaggle.com/dansbecker/running-kaggle-kernels-with-a-gpu](url)),  *GPU backed instances have less CPU power and RAM* . I only used CPU in the script for submission, since the preprocessing code uses CPU. (However, I find the available CPU memory is almost unchanged no matter if I turned the GPU on)"
  },
  "source": "meta"
}