{
  "id": 189987,
  "title": "Submission Scoring Error",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189987",
  "author_name": "David M",
  "post_date": "2020-10-09T15:53:42.008000",
  "votes": 11,
  "comment_count": 17,
  "views": 0,
  "content": "<p>When I try to submit, it says Notebook Running for hours and then I get submission scoring error. My whole notebook runs in 6 minutes without error but then when I click upload it runs indefinitely. I'm using R based on <a href=\"https://www.kaggle.com/rdboyes/r-starter-notebook\" target=\"_blank\">this notebook</a> so maybe that has something to do with it? Has anyone used R for this competition?</p>",
  "messages": [
    {
      "id": 1044222,
      "postDate": "2020-10-09T15:53:42.010Z",
      "content": "<p>When I try to submit, it says Notebook Running for hours and then I get submission scoring error. My whole notebook runs in 6 minutes without error but then when I click upload it runs indefinitely. I'm using R based on <a href=\"https://www.kaggle.com/rdboyes/r-starter-notebook\" target=\"_blank\">this notebook</a> so maybe that has something to do with it? Has anyone used R for this competition?</p>",
      "rawMarkdown": "When I try to submit, it says Notebook Running for hours and then I get submission scoring error. My whole notebook runs in 6 minutes without error but then when I click upload it runs indefinitely. I'm using R based on [this notebook](https://www.kaggle.com/rdboyes/r-starter-notebook) so maybe that has something to do with it? Has anyone used R for this competition?",
      "votes": 11
    },
    {
      "id": 1045731,
      "postDate": "2020-10-11T01:16:42.020Z",
      "content": "<p>Hi! Don't forget that the submission will rerun your code and generate a new submission on a different test dataset.</p>\n<p>Therefore committing and after submitting runs your notebook 2 times.<br>\nTo reduce the submission time one trick I use is to save the notebook(don't commit simple save the notebook) with a dummy submission file in the output(make sure there is a submission.csv file in the output. You can create an empty one , when saving choose to keep the current notebook output in the advance options). <br>\nAfter you save your notebook, go to the output and submit the submission file like that. </p>",
      "rawMarkdown": "Hi! Don't forget that the submission will rerun your code and generate a new submission on a different test dataset.\n\nTherefore committing and after submitting runs your notebook 2 times.\nTo reduce the submission time one trick I use is to save the notebook(don't commit simple save the notebook) with a dummy submission file in the output(make sure there is a submission.csv file in the output. You can create an empty one , when saving choose to keep the current notebook output in the advance options). \nAfter you save your notebook, go to the output and submit the submission file like that. ",
      "votes": 3
    },
    {
      "id": 1044619,
      "postDate": "2020-10-10T01:13:22.603Z",
      "content": "<pre><code>def do_predict_df(test_df, sample_prediction_df):\n    if len(sample_prediction_df)==0: return sample_prediction_df\n\n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = d['user_id'].map(user_model).fillna(0.5)\n    s1 = d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df['answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n</code></pre>\n<p>i have the same issue.  simple kernel gives \"Submission Scoring Error\" ?</p>",
      "rawMarkdown": "```\n\ndef do_predict_df(test_df, sample_prediction_df):\n    if len(sample_prediction_df)==0: return sample_prediction_df\n    \n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = d['user_id'].map(user_model).fillna(0.5)\n    s1 = d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df['answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n\n```\n\ni have the same issue.  simple kernel gives \"Submission Scoring Error\" ?",
      "votes": 2,
      "replies": [
        {
          "id": 1044630,
          "postDate": "2020-10-10T01:57:05.200Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1044667,
          "postDate": "2020-10-10T02:58:00.647Z",
          "content": "<pre><code>def do_predict_df(test_df, sample_prediction_df):\n    if sample_prediction_df.empty():\n        return sample_prediction_df\n\n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df.loc[d.index, 'answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n</code></pre>\n<p>This works though (&lt;10 mins) i got the LB score of .5, will try to replicate what you did Heng!</p>",
          "rawMarkdown": "```\ndef do_predict_df(test_df, sample_prediction_df):\n    if sample_prediction_df.empty():\n        return sample_prediction_df\n\n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df.loc[d.index, 'answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n```\n\nThis works though (<10 mins) i got the LB score of .5, will try to replicate what you did Heng!"
        },
        {
          "id": 1046025,
          "postDate": "2020-10-11T09:21:54.867Z",
          "content": "<pre><code>    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n</code></pre>\n<p>uncomment these will not run, although I don't know why and has been stuck for days</p>",
          "rawMarkdown": "```\n    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n```\n\nuncomment these will not run, although I don't know why and has been stuck for days"
        },
        {
          "id": 1046034,
          "postDate": "2020-10-11T09:38:19.697Z",
          "content": "<p>I see, swapped them with <code>defaultdict(lambda :0.6)</code> to mimic; And it ran in less than 5 mins now.</p>",
          "rawMarkdown": "I see, swapped them with ``` defaultdict(lambda :0.6)``` to mimic; And it ran in less than 5 mins now."
        },
        {
          "id": 1046056,
          "postDate": "2020-10-11T10:10:11.747Z",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<p>thanks for the hint. defaultdict works. i think the assumption .fillna(0.5) is wrong</p>\n<p><a href=\"https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365</a><br>\naverage prediction baseline : LB 0.741</p>",
          "rawMarkdown": "@adityaecdrid \n\nthanks for the hint. defaultdict works. i think the assumption .fillna(0.5) is wrong\n\nhttps://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\naverage prediction baseline : LB 0.741",
          "votes": 2
        },
        {
          "id": 1046060,
          "postDate": "2020-10-11T10:15:59.530Z",
          "content": "<p>Very happy to help Heng!</p>",
          "rawMarkdown": "Very happy to help Heng!"
        },
        {
          "id": 1046160,
          "postDate": "2020-10-11T12:02:46.123Z",
          "content": "<p>How could .fillna() cause trouble though? I think I'm having the same problem, but when using the real data it does inevitably happen that I divide by zero and get NaN's which need to be replaced…</p>",
          "rawMarkdown": "How could .fillna() cause trouble though? I think I'm having the same problem, but when using the real data it does inevitably happen that I divide by zero and get NaN's which need to be replaced..."
        }
      ]
    },
    {
      "id": 1044533,
      "postDate": "2020-10-09T21:10:52.003Z",
      "content": "<p>Same problem here. My notebook runs in 5 minutes and uses about 5-6 Gb memory. But when submitted for scoring, it runs many hours and finishes with \"Submission Scoring Error\". </p>\n<p>It's in Python. And I believe I've handled possible cases with unseen users_id, content_id. Also don't want to believe it could be memory issue, as I'm using less data/memory compared to public kernels, which are passing scoring successfully. </p>\n<p>Probably there's some trivial explanation, but until it's found, looks strange. Will have to dig more.</p>",
      "rawMarkdown": "Same problem here. My notebook runs in 5 minutes and uses about 5-6 Gb memory. But when submitted for scoring, it runs many hours and finishes with \"Submission Scoring Error\". \n\nIt's in Python. And I believe I've handled possible cases with unseen users_id, content_id. Also don't want to believe it could be memory issue, as I'm using less data/memory compared to public kernels, which are passing scoring successfully. \n\nProbably there's some trivial explanation, but until it's found, looks strange. Will have to dig more.",
      "votes": 2,
      "replies": [
        {
          "id": 1046401,
          "postDate": "2020-10-11T16:08:24.793Z",
          "content": "<p>Did you figure out your problem?</p>",
          "rawMarkdown": "Did you figure out your problem?"
        },
        {
          "id": 1046437,
          "postDate": "2020-10-11T16:52:33.293Z",
          "content": "<p>Partly. I figured out, which group of features caused the issue, and for now simply removed them. But didn't find the exact reason for failure yet - will return to this part later when won't have other ideas for features.</p>",
          "rawMarkdown": "Partly. I figured out, which group of features caused the issue, and for now simply removed them. But didn't find the exact reason for failure yet - will return to this part later when won't have other ideas for features.",
          "votes": 2
        },
        {
          "id": 1049179,
          "postDate": "2020-10-14T07:18:39.607Z",
          "content": "<p><a href=\"https://www.kaggle.com/alijs1\" target=\"_blank\">@alijs1</a>  I got same problem with you. Did you do the feature engineer and train the model on Kaggle Kernel? I generate the features and models offline then upload them for online prediction, but finnally \"submission score error\" :(</p>",
          "rawMarkdown": "@alijs1  I got same problem with you. Did you do the feature engineer and train the model on Kaggle Kernel? I generate the features and models offline then upload them for online prediction, but finnally \"submission score error\" :("
        }
      ]
    },
    {
      "id": 1044487,
      "postDate": "2020-10-09T20:07:43.463Z",
      "content": "<p>Happens with me as well, I am using python notebook</p>",
      "rawMarkdown": "Happens with me as well, I am using python notebook",
      "votes": 2
    },
    {
      "id": 1050519,
      "postDate": "2020-10-15T13:48:55.617Z",
      "content": "<p>I also get this problem.<br>\nI already run and save my notebook completely.<br>\nMy output is updated too.<br>\nWhen I submit, I got error.<br>\nBut description said \"No additional details provided for this error\".<br>\nPlease help</p>",
      "rawMarkdown": "I also get this problem.\nI already run and save my notebook completely.\nMy output is updated too.\nWhen I submit, I got error.\nBut description said \"No additional details provided for this error\".\nPlease help"
    },
    {
      "id": 1046272,
      "postDate": "2020-10-11T14:12:49.483Z",
      "content": "<p>May be there is some code error which is not giving output in desired format.</p>",
      "rawMarkdown": "May be there is some code error which is not giving output in desired format."
    },
    {
      "id": 1044501,
      "postDate": "2020-10-09T20:26:49.207Z",
      "content": "<p>I'm also experiencing this. My code runs in about an hour but it's been six hours so far without a score returned. I assume that that it has something to do w/ memory/RAM issues.</p>",
      "rawMarkdown": "I'm also experiencing this. My code runs in about an hour but it's been six hours so far without a score returned. I assume that that it has something to do w/ memory/RAM issues."
    }
  ],
  "comments": [
    {
      "id": 1045731,
      "author_name": "Jude TCHAYE",
      "author_url": "",
      "post_date": "2020-10-11T01:16:42.020000",
      "content": "<p>Hi! Don't forget that the submission will rerun your code and generate a new submission on a different test dataset.</p>\n<p>Therefore committing and after submitting runs your notebook 2 times.<br>\nTo reduce the submission time one trick I use is to save the notebook(don't commit simple save the notebook) with a dummy submission file in the output(make sure there is a submission.csv file in the output. You can create an empty one , when saving choose to keep the current notebook output in the advance options). <br>\nAfter you save your notebook, go to the output and submit the submission file like that. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1044619,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-10-10T01:13:22.603000",
      "content": "<pre><code>def do_predict_df(test_df, sample_prediction_df):\n    if len(sample_prediction_df)==0: return sample_prediction_df\n\n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = d['user_id'].map(user_model).fillna(0.5)\n    s1 = d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df['answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n</code></pre>\n<p>i have the same issue.  simple kernel gives \"Submission Scoring Error\" ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1044630,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-10T01:57:05.200000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1044667,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-10T02:58:00.647000",
          "content": "<pre><code>def do_predict_df(test_df, sample_prediction_df):\n    if sample_prediction_df.empty():\n        return sample_prediction_df\n\n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df.loc[d.index, 'answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n</code></pre>\n<p>This works though (&lt;10 mins) i got the LB score of .5, will try to replicate what you did Heng!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046025,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-11T09:21:54.867000",
          "content": "<pre><code>    s0 = 0.5 # d['user_id'].map(user_model).fillna(0.5)\n    s1 = 0.5 # d['content_id'].map(content_model).fillna(0.5) \n</code></pre>\n<p>uncomment these will not run, although I don't know why and has been stuck for days</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046034,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-11T09:38:19.697000",
          "content": "<p>I see, swapped them with <code>defaultdict(lambda :0.6)</code> to mimic; And it ran in less than 5 mins now.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046056,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-10-11T10:10:11.747000",
          "content": "<p><a href=\"https://www.kaggle.com/adityaecdrid\" target=\"_blank\">@adityaecdrid</a> </p>\n<p>thanks for the hint. defaultdict works. i think the assumption .fillna(0.5) is wrong</p>\n<p><a href=\"https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365\" target=\"_blank\">https://www.kaggle.com/hengck23/notebookcdb764afc6?scriptVersionId=44475365</a><br>\naverage prediction baseline : LB 0.741</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1046060,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-11T10:15:59.530000",
          "content": "<p>Very happy to help Heng!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046160,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-11T12:02:46.123000",
          "content": "<p>How could .fillna() cause trouble though? I think I'm having the same problem, but when using the real data it does inevitably happen that I divide by zero and get NaN's which need to be replaced…</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044533,
      "author_name": "alijs",
      "author_url": "",
      "post_date": "2020-10-09T21:10:52.003000",
      "content": "<p>Same problem here. My notebook runs in 5 minutes and uses about 5-6 Gb memory. But when submitted for scoring, it runs many hours and finishes with \"Submission Scoring Error\". </p>\n<p>It's in Python. And I believe I've handled possible cases with unseen users_id, content_id. Also don't want to believe it could be memory issue, as I'm using less data/memory compared to public kernels, which are passing scoring successfully. </p>\n<p>Probably there's some trivial explanation, but until it's found, looks strange. Will have to dig more.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1046401,
          "author_name": "David M",
          "author_url": "",
          "post_date": "2020-10-11T16:08:24.793000",
          "content": "<p>Did you figure out your problem?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1046437,
          "author_name": "alijs",
          "author_url": "",
          "post_date": "2020-10-11T16:52:33.293000",
          "content": "<p>Partly. I figured out, which group of features caused the issue, and for now simply removed them. But didn't find the exact reason for failure yet - will return to this part later when won't have other ideas for features.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1049179,
          "author_name": "Ethan",
          "author_url": "",
          "post_date": "2020-10-14T07:18:39.607000",
          "content": "<p><a href=\"https://www.kaggle.com/alijs1\" target=\"_blank\">@alijs1</a>  I got same problem with you. Did you do the feature engineer and train the model on Kaggle Kernel? I generate the features and models offline then upload them for online prediction, but finnally \"submission score error\" :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1044487,
      "author_name": "Sudeep Shouche",
      "author_url": "",
      "post_date": "2020-10-09T20:07:43.463000",
      "content": "<p>Happens with me as well, I am using python notebook</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1050519,
      "author_name": "KevinnWu",
      "author_url": "",
      "post_date": "2020-10-15T13:48:55.617000",
      "content": "<p>I also get this problem.<br>\nI already run and save my notebook completely.<br>\nMy output is updated too.<br>\nWhen I submit, I got error.<br>\nBut description said \"No additional details provided for this error\".<br>\nPlease help</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1046272,
      "author_name": "Harsh Verma",
      "author_url": "",
      "post_date": "2020-10-11T14:12:49.483000",
      "content": "<p>May be there is some code error which is not giving output in desired format.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1044501,
      "author_name": "Nick Sarris",
      "author_url": "",
      "post_date": "2020-10-09T20:26:49.207000",
      "content": "<p>I'm also experiencing this. My code runs in about an hour but it's been six hours so far without a score returned. I assume that that it has something to do w/ memory/RAM issues.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1044222": "When I try to submit, it says Notebook Running for hours and then I get submission scoring error. My whole notebook runs in 6 minutes without error but then when I click upload it runs indefinitely. I'm using R based on [this notebook](https://www.kaggle.com/rdboyes/r-starter-notebook) so maybe that has something to do with it? Has anyone used R for this competition?",
    "1045731": "Hi! Don't forget that the submission will rerun your code and generate a new submission on a different test dataset.\n\nTherefore committing and after submitting runs your notebook 2 times.\nTo reduce the submission time one trick I use is to save the notebook(don't commit simple save the notebook) with a dummy submission file in the output(make sure there is a submission.csv file in the output. You can create an empty one , when saving choose to keep the current notebook output in the advance options). \nAfter you save your notebook, go to the output and submit the submission file like that. ",
    "1044619": "```\n\ndef do_predict_df(test_df, sample_prediction_df):\n    if len(sample_prediction_df)==0: return sample_prediction_df\n    \n    d = test_df[test_df['content_type_id'] == 0]\n    s0 = d['user_id'].map(user_model).fillna(0.5)\n    s1 = d['content_id'].map(content_model).fillna(0.5) \n    sample_prediction_df['answered_correctly'] = (s0+s1)/2\n    return sample_prediction_df\n\n\niter_test = env.iter_test()\nfor t, (test_df, sample_prediction_df) in enumerate(iter_test):\n\n    predict_df = do_predict_df(test_df, sample_prediction_df)\n    env.predict(predict_df)\n    print('loop at %d :'%(t), test_df.shape, sample_prediction_df.shape, predict_df.shape)\n\n```\n\ni have the same issue.  simple kernel gives \"Submission Scoring Error\" ?",
    "1044533": "Same problem here. My notebook runs in 5 minutes and uses about 5-6 Gb memory. But when submitted for scoring, it runs many hours and finishes with \"Submission Scoring Error\". \n\nIt's in Python. And I believe I've handled possible cases with unseen users_id, content_id. Also don't want to believe it could be memory issue, as I'm using less data/memory compared to public kernels, which are passing scoring successfully. \n\nProbably there's some trivial explanation, but until it's found, looks strange. Will have to dig more.",
    "1044487": "Happens with me as well, I am using python notebook",
    "1050519": "I also get this problem.\nI already run and save my notebook completely.\nMy output is updated too.\nWhen I submit, I got error.\nBut description said \"No additional details provided for this error\".\nPlease help",
    "1046272": "May be there is some code error which is not giving output in desired format.",
    "1044501": "I'm also experiencing this. My code runs in about an hour but it's been six hours so far without a score returned. I assume that that it has something to do w/ memory/RAM issues."
  }
}