{
  "id": 190590,
  "title": "Kaggle Submission Scoring Error",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190590",
  "author_name": "",
  "post_date": "2020-10-12T13:30:02.664463100Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hello All,</p>\n<p>I'm getting a submission error while submitting my notebook - <a href=\"https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927\" target=\"_blank\">https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927</a></p>\n<p>I initially thought it was because of some of the joins &amp; missing value treatment performed by me on the test dataset but then I also hardcoded my actual predictions to 0.5 and still getting the error.</p>\n<p>I saw a different topic created on this but wasn't able to figure out a solution from there. Any recommendations on how to debug and avoid this?</p>",
  "messages": [
    {
      "id": "1047321",
      "postDate": "10/12/2020 13:30:02",
      "content": "<p>Hello All,</p>\n<p>I'm getting a submission error while submitting my notebook - <a href=\"https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927\" target=\"_blank\">https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927</a></p>\n<p>I initially thought it was because of some of the joins &amp; missing value treatment performed by me on the test dataset but then I also hardcoded my actual predictions to 0.5 and still getting the error.</p>\n<p>I saw a different topic created on this but wasn't able to figure out a solution from there. Any recommendations on how to debug and avoid this?</p>",
      "rawMarkdown": "Hello All,\n\nI'm getting a submission error while submitting my notebook - https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927\n\nI initially thought it was because of some of the joins & missing value treatment performed by me on the test dataset but then I also hardcoded my actual predictions to 0.5 and still getting the error.\n\nI saw a different topic created on this but wasn't able to figure out a solution from there. Any recommendations on how to debug and avoid this?",
      "votes": null
    },
    {
      "id": "1047434",
      "postDate": "10/12/2020 15:54:22",
      "content": "<p>Hi Arpit,<br>\nI ran across similar issues and ended up expending my submissions for the day. At that point I decided to test my code using 1000 rows from train set as the test_df. Turns out there were many label encoding issues due to lecture introduced nans. You could check if you have a similar problem or try this approach yourself.</p>",
      "rawMarkdown": "Hi Arpit,\nI ran across similar issues and ended up expending my submissions for the day. At that point I decided to test my code using 1000 rows from train set as the test_df. Turns out there were many label encoding issues due to lecture introduced nans. You could check if you have a similar problem or try this approach yourself.",
      "votes": null
    },
    {
      "id": "1047442",
      "postDate": "10/12/2020 15:57:15",
      "content": "<p>Hi Abhimanyu,</p>\n<p>I thought that joins maybe an issue too therefore I removed all joins and just tried submitting with 0.5 as probability for all test_df records but still getting an error.</p>",
      "rawMarkdown": "Hi Abhimanyu,\n\nI thought that joins maybe an issue too therefore I removed all joins and just tried submitting with 0.5 as probability for all test_df records but still getting an error.",
      "votes": null
    },
    {
      "id": "1047448",
      "postDate": "10/12/2020 16:03:45",
      "content": "<p>That is disconcerting! Can you share your notebook?</p>",
      "rawMarkdown": "That is disconcerting! Can you share your notebook?",
      "votes": null
    },
    {
      "id": "1047455",
      "postDate": "10/12/2020 16:10:59",
      "content": "<p>Yes it's a part of my topic description above</p>",
      "rawMarkdown": "Yes it's a part of my topic description above",
      "votes": null
    },
    {
      "id": "1047468",
      "postDate": "10/12/2020 16:21:11",
      "content": "<p>You seem to be merging questions.csv here. That is the most likely source of nans considering some content_ids will not be in it. Consider using the first 1000 rows of train.csv as test_df:</p>\n<p>iter_test = [train_df].</p>\n<p>You'll at least get some errors messages </p>",
      "rawMarkdown": "You seem to be merging questions.csv here. That is the most likely source of nans considering some content_ids will not be in it. Consider using the first 1000 rows of train.csv as test_df:\n\niter_test = [train_df].\n\nYou'll at least get some errors messages",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1047434,
      "author_name": "abhimanyud",
      "author_url": "",
      "post_date": "10/12/2020 15:54:22",
      "content": "<p>Hi Arpit,<br>\nI ran across similar issues and ended up expending my submissions for the day. At that point I decided to test my code using 1000 rows from train set as the test_df. Turns out there were many label encoding issues due to lecture introduced nans. You could check if you have a similar problem or try this approach yourself.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1047442,
          "author_name": "arpitsolanki14",
          "author_url": "",
          "post_date": "10/12/2020 15:57:15",
          "content": "<p>Hi Abhimanyu,</p>\n<p>I thought that joins maybe an issue too therefore I removed all joins and just tried submitting with 0.5 as probability for all test_df records but still getting an error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047448,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/12/2020 16:03:45",
          "content": "<p>That is disconcerting! Can you share your notebook?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047455,
          "author_name": "arpitsolanki14",
          "author_url": "",
          "post_date": "10/12/2020 16:10:59",
          "content": "<p>Yes it's a part of my topic description above</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1047468,
          "author_name": "abhimanyud",
          "author_url": "",
          "post_date": "10/12/2020 16:21:11",
          "content": "<p>You seem to be merging questions.csv here. That is the most likely source of nans considering some content_ids will not be in it. Consider using the first 1000 rows of train.csv as test_df:</p>\n<p>iter_test = [train_df].</p>\n<p>You'll at least get some errors messages </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1047321": "Hello All,\n\nI'm getting a submission error while submitting my notebook - https://www.kaggle.com/arpitsolanki14/riiid-eda-basic-rf?scriptVersionId=44553927\n\nI initially thought it was because of some of the joins & missing value treatment performed by me on the test dataset but then I also hardcoded my actual predictions to 0.5 and still getting the error.\n\nI saw a different topic created on this but wasn't able to figure out a solution from there. Any recommendations on how to debug and avoid this?",
    "1047434": "Hi Arpit,\nI ran across similar issues and ended up expending my submissions for the day. At that point I decided to test my code using 1000 rows from train set as the test_df. Turns out there were many label encoding issues due to lecture introduced nans. You could check if you have a similar problem or try this approach yourself.",
    "1047442": "Hi Abhimanyu,\n\nI thought that joins maybe an issue too therefore I removed all joins and just tried submitting with 0.5 as probability for all test_df records but still getting an error.",
    "1047448": "That is disconcerting! Can you share your notebook?",
    "1047455": "Yes it's a part of my topic description above",
    "1047468": "You seem to be merging questions.csv here. That is the most likely source of nans considering some content_ids will not be in it. Consider using the first 1000 rows of train.csv as test_df:\n\niter_test = [train_df].\n\nYou'll at least get some errors messages"
  },
  "source": "meta"
}