{
  "id": 483379,
  "title": "Notebook Threw Exception Your notebook hit an unhandled error while rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/483379",
  "author_name": "Jamie",
  "post_date": "2024-03-12T06:03:44.923000",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hello, as you may notice, I am quite new to Kaggle competitions. <br>\nI am facing the Notebook Threw Exception Error. The notebook does provide a submission.csv output, so I assume there is no problem with my work. But I don't really understand where the error is occurring… </p>",
  "messages": [
    {
      "id": 2692908,
      "postDate": "2024-03-12T06:03:44.923Z",
      "content": "<p>Hello, as you may notice, I am quite new to Kaggle competitions. <br>\nI am facing the Notebook Threw Exception Error. The notebook does provide a submission.csv output, so I assume there is no problem with my work. But I don't really understand where the error is occurring… </p>",
      "rawMarkdown": "Hello, as you may notice, I am quite new to Kaggle competitions. \nI am facing the Notebook Threw Exception Error. The notebook does provide a submission.csv output, so I assume there is no problem with my work. But I don't really understand where the error is occurring... ",
      "votes": 1
    },
    {
      "id": 2693250,
      "postDate": "2024-03-12T10:05:51.400Z",
      "content": "<p>I have been playing this competition since the beginning of the competition. (now two weeks after halftime)</p>\n<p>I have encountered similar problems more than ten times in this competition. My suggestion is to add one file at a time to my own code so that I can clearly identify which file is causing the error.</p>\n<p>In the process of model inference, predictions should be made in batches, such as predicting 10000 data at once.</p>\n<p>Finally, you can take a look at my baseline:<a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break</a></p>",
      "rawMarkdown": "I have been playing this competition since the beginning of the competition. (now two weeks after halftime)\n\nI have encountered similar problems more than ten times in this competition. My suggestion is to add one file at a time to my own code so that I can clearly identify which file is causing the error.\n\nIn the process of model inference, predictions should be made in batches, such as predicting 10000 data at once.\n\nFinally, you can take a look at my baseline:https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break",
      "votes": 2,
      "replies": [
        {
          "id": 2697602,
          "postDate": "2024-03-15T02:26:59.693Z",
          "content": "<p>Hello yunsuxiaozi!</p>\n<p>After trying to elaborate more on why the error occured, I had assumed the private data set was quite different from the public data, so I had to review my original code when handling the missing or the NaN data in the features, I believe now I have been able to solve prob.</p>\n<p>Thanks for your advice!</p>",
          "rawMarkdown": "Hello yunsuxiaozi!\n\nAfter trying to elaborate more on why the error occured, I had assumed the private data set was quite different from the public data, so I had to review my original code when handling the missing or the NaN data in the features, I believe now I have been able to solve prob.\n\nThanks for your advice!",
          "replies": [
            {
              "id": 2723171,
              "postDate": "2024-03-30T02:10:07.103Z",
              "content": "<p>HI Jamie, <br>\nCan you give more details/examples of how you dealt with that error? Or what you did to avoid the error? </p>\n<p>My previous runs went OK. Then I tried imputing missings in numeric variables; also, introducing \"unknown\" and \"missing\" as categories to categorical variables and encoding them…and got that error.</p>\n<p>Thanks! </p>",
              "rawMarkdown": "HI Jamie, \nCan you give more details/examples of how you dealt with that error? Or what you did to avoid the error? \n\nMy previous runs went OK. Then I tried imputing missings in numeric variables; also, introducing \"unknown\" and \"missing\" as categories to categorical variables and encoding them...and got that error.\n\nThanks! "
            },
            {
              "id": 2724692,
              "postDate": "2024-03-31T03:40:12.227Z",
              "content": "<p>I am facing the same issue. Can you please elaborate on how to solve this issue.</p>",
              "rawMarkdown": "I am facing the same issue. Can you please elaborate on how to solve this issue."
            },
            {
              "id": 2724766,
              "postDate": "2024-03-31T04:56:30.973Z",
              "content": "<p>Hello, Varuni!<br>\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.<br>\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).</p>\n<p>Another reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run. <br>\nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories.</p>",
              "rawMarkdown": "Hello, Varuni!\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).\n\nAnother reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run. \nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories."
            },
            {
              "id": 2724767,
              "postDate": "2024-03-31T04:56:44.170Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2724768,
              "postDate": "2024-03-31T04:57:07.767Z",
              "content": "<p>Hello, Rodrigo!<br>\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.<br>\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).</p>\n<p>Another reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run.<br>\nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories.</p>",
              "rawMarkdown": "Hello, Rodrigo!\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).\n\nAnother reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run.\nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories."
            },
            {
              "id": 2725593,
              "postDate": "2024-03-31T17:16:59.927Z",
              "content": "<p>Thank you very much Jamie. I had already taken care of the extra categorical data that may come up in the test data as well as memory issues. I am treating the submission notebook solely with test data and hence making sure the RAM is going overboard. But, still got the error and hence wanted to check with you.</p>\n<p>Your response however did help. I looked closely at the categoricals and found one boolean feature that was different for the train and test sets. I corrected that and now it is working fine. Thank you very much for your response.</p>",
              "rawMarkdown": "Thank you very much Jamie. I had already taken care of the extra categorical data that may come up in the test data as well as memory issues. I am treating the submission notebook solely with test data and hence making sure the RAM is going overboard. But, still got the error and hence wanted to check with you.\n\nYour response however did help. I looked closely at the categoricals and found one boolean feature that was different for the train and test sets. I corrected that and now it is working fine. Thank you very much for your response.",
              "votes": 1
            },
            {
              "id": 2727888,
              "postDate": "2024-04-02T01:37:50.063Z",
              "content": "<p>Hi Varuni! Suppose the handling process of the extra categorical data and the RAM usage didn't solve the process. In that case, the only problem may exist in the feature engineering process for the train and test data set. I suggest you to take a look at some codes that were posted by other co to find out their attempt on the competition. </p>\n<p>Oh, its a small thing to mention and this might not be a problem, but if you are assigning a name for the prediction column in the submission file, make sure that it is the same as the requirement. I believe it was \"case_id\" and \"score\"</p>",
              "rawMarkdown": "Hi Varuni! Suppose the handling process of the extra categorical data and the RAM usage didn't solve the process. In that case, the only problem may exist in the feature engineering process for the train and test data set. I suggest you to take a look at some codes that were posted by other co to find out their attempt on the competition. \n\nOh, its a small thing to mention and this might not be a problem, but if you are assigning a name for the prediction column in the submission file, make sure that it is the same as the requirement. I believe it was \"case_id\" and \"score\""
            },
            {
              "id": 2727940,
              "postDate": "2024-04-02T02:39:04.490Z",
              "content": "<p>Your response is indeed precious. Thank you again for getting back. I reused the code from the starter notebook to make sure that I don't mess up with my submission files. </p>\n<p>I could figure out the feature that was causing the error (flag feature that was being read as categorical in train and int in test). Besides it was not adding any value to the model and hence just dropped the feature and that corrected the error. </p>",
              "rawMarkdown": "Your response is indeed precious. Thank you again for getting back. I reused the code from the starter notebook to make sure that I don't mess up with my submission files. \n\nI could figure out the feature that was causing the error (flag feature that was being read as categorical in train and int in test). Besides it was not adding any value to the model and hence just dropped the feature and that corrected the error. "
            },
            {
              "id": 2728562,
              "postDate": "2024-04-02T10:09:08.263Z",
              "content": "<p>No problem! Im happy that you were able to find out the reason that was causing the error. Good luck with your progression!</p>",
              "rawMarkdown": "No problem! Im happy that you were able to find out the reason that was causing the error. Good luck with your progression!",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2693250,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "2024-03-12T10:05:51.400000",
      "content": "<p>I have been playing this competition since the beginning of the competition. (now two weeks after halftime)</p>\n<p>I have encountered similar problems more than ten times in this competition. My suggestion is to add one file at a time to my own code so that I can clearly identify which file is causing the error.</p>\n<p>In the process of model inference, predictions should be made in batches, such as predicting 10000 data at once.</p>\n<p>Finally, you can take a look at my baseline:<a href=\"https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break\" target=\"_blank\">https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2697602,
          "author_name": "Jamie",
          "author_url": "",
          "post_date": "2024-03-15T02:26:59.693000",
          "content": "<p>Hello yunsuxiaozi!</p>\n<p>After trying to elaborate more on why the error occured, I had assumed the private data set was quite different from the public data, so I had to review my original code when handling the missing or the NaN data in the features, I believe now I have been able to solve prob.</p>\n<p>Thanks for your advice!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2723171,
              "author_name": "Rodrigo M Carrillo Larco",
              "author_url": "",
              "post_date": "2024-03-30T02:10:07.103000",
              "content": "<p>HI Jamie, <br>\nCan you give more details/examples of how you dealt with that error? Or what you did to avoid the error? </p>\n<p>My previous runs went OK. Then I tried imputing missings in numeric variables; also, introducing \"unknown\" and \"missing\" as categories to categorical variables and encoding them…and got that error.</p>\n<p>Thanks! </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2724692,
              "author_name": "Varuni Rao",
              "author_url": "",
              "post_date": "2024-03-31T03:40:12.227000",
              "content": "<p>I am facing the same issue. Can you please elaborate on how to solve this issue.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2724766,
              "author_name": "Jamie",
              "author_url": "",
              "post_date": "2024-03-31T04:56:30.973000",
              "content": "<p>Hello, Varuni!<br>\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.<br>\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).</p>\n<p>Another reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run. <br>\nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2724767,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-03-31T04:56:44.170000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2724768,
              "author_name": "Jamie",
              "author_url": "",
              "post_date": "2024-03-31T04:57:07.767000",
              "content": "<p>Hello, Rodrigo!<br>\nWhen I faced this problem, I realized that there are mainly two reasons why this error occurs in Kaggle.<br>\n1 reason may lie because Kaggle only allows 16 GB RAM usage per notebook. Due to this, I don't think that we are able to perform large iterations during our model training (like cross-validation with 2000 iterations).</p>\n<p>Another reason (which was my problem) may be because of the way you handle the features and the way you optimize those NaN and unknown values. Unlike the test data provided to us prior, the submission data contains external categorical data which may throw an error during the run.<br>\nFor example, the pre-trained model with the unknown values adjusted for training might not be able to handle the test data. So, I think for my case, you need to adjust the way you handle the unknown categories.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2725593,
              "author_name": "Varuni Rao",
              "author_url": "",
              "post_date": "2024-03-31T17:16:59.927000",
              "content": "<p>Thank you very much Jamie. I had already taken care of the extra categorical data that may come up in the test data as well as memory issues. I am treating the submission notebook solely with test data and hence making sure the RAM is going overboard. But, still got the error and hence wanted to check with you.</p>\n<p>Your response however did help. I looked closely at the categoricals and found one boolean feature that was different for the train and test sets. I corrected that and now it is working fine. Thank you very much for your response.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2727888,
              "author_name": "Jamie",
              "author_url": "",
              "post_date": "2024-04-02T01:37:50.063000",
              "content": "<p>Hi Varuni! Suppose the handling process of the extra categorical data and the RAM usage didn't solve the process. In that case, the only problem may exist in the feature engineering process for the train and test data set. I suggest you to take a look at some codes that were posted by other co to find out their attempt on the competition. </p>\n<p>Oh, its a small thing to mention and this might not be a problem, but if you are assigning a name for the prediction column in the submission file, make sure that it is the same as the requirement. I believe it was \"case_id\" and \"score\"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2727940,
              "author_name": "Varuni Rao",
              "author_url": "",
              "post_date": "2024-04-02T02:39:04.490000",
              "content": "<p>Your response is indeed precious. Thank you again for getting back. I reused the code from the starter notebook to make sure that I don't mess up with my submission files. </p>\n<p>I could figure out the feature that was causing the error (flag feature that was being read as categorical in train and int in test). Besides it was not adding any value to the model and hence just dropped the feature and that corrected the error. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2728562,
              "author_name": "Jamie",
              "author_url": "",
              "post_date": "2024-04-02T10:09:08.263000",
              "content": "<p>No problem! Im happy that you were able to find out the reason that was causing the error. Good luck with your progression!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2692908": "Hello, as you may notice, I am quite new to Kaggle competitions. \nI am facing the Notebook Threw Exception Error. The notebook does provide a submission.csv output, so I assume there is no problem with my work. But I don't really understand where the error is occurring... ",
    "2693250": "I have been playing this competition since the beginning of the competition. (now two weeks after halftime)\n\nI have encountered similar problems more than ten times in this competition. My suggestion is to add one file at a time to my own code so that I can clearly identify which file is causing the error.\n\nIn the process of model inference, predictions should be made in batches, such as predicting 10000 data at once.\n\nFinally, you can take a look at my baseline:https://www.kaggle.com/code/yunsuxiaozi/home-credit-baseline-after-break"
  }
}