{
  "id": 197065,
  "title": "submission stops after 4 iterations",
  "url": "/competitions/riiid-test-answer-prediction/discussion/197065",
  "author_name": "",
  "post_date": "2020-11-14T05:14:50.659005200Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey, i tried to submit my prediction and i have turned off the internet of my notebook.<br>\nwhen  i click  <strong>save and run all</strong> , the process finishes in like 20 seconds!<br>\nThen i check the log and output file, i do have submmisons.csv file, but very short one, with only 108 rows, i first thought that maybe the env.iter_test() didn't work, so i add log for each iteration, the log file shows that it stops at iteration 4, and i have no other information.</p>\n<p>here is my code.</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\n\n############\n# all the files needed and pre-train model are loaded there.\n############\n\n# MAKE PREDICTION\niter_test = env.iter_test()\ntest_prev = pd.DataFrame()\nn_iter = 1\nfor (df_test, sample_prediction_df) in iter_test: \n\n    logging.info('\\n*CURRENT ITERATION*:-{}-'.format(n_iter))\n    test, user_summary, lecture_summary = processing_test(df_test, test_prev,  questions lectures, user_summary, q_stats_train,  lecture_summary,  features,  question_tags)\n\n    test_prev = df_test.copy()\n    test['answered_correctly'] = model.predict(test.drop(['row_id','group_num'], axis=1))\n    test.set_index('group_num', inplace=True)\n    env.predict(test[['row_id', 'answered_correctly']])\n    n_iter += 1\n</code></pre>\n<p>here is the log that i created:</p>\n<pre><code>Loading all summaries:\n\nShape of user summary:(393656, 11),\nShape of lecture summary:(393656, 23),\nShape of question stats:(13523, 5) \n- Successful\n\n# of features to add(extract):48\n\nQuestions info:(13523, 5)\nLectures info:(418, 4)\n\nLoading model...\nmodel BernoulliNB() loaded -- Successful!...\n\n*CURRENT ITERATION*:-1-\nNumExpr defaulting to 4 threads.\n\nsplit into :new_test (18, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (18, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (18, 55)\n\n*CURRENT ITERATION*:-2-\n\nsplit into :new_test (27, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (27, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (27, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (18, 11)\n\nUpdating user_summary:\nDone!\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-3-\n\nsplit into :new_test (26, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (26, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (26, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (27, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-4-\n\nsplit into :new_test (33, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (33, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (33, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (26, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n</code></pre>",
  "messages": [
    {
      "id": "1077928",
      "postDate": "11/14/2020 05:14:50",
      "content": "<p>Hey, i tried to submit my prediction and i have turned off the internet of my notebook.<br>\nwhen  i click  <strong>save and run all</strong> , the process finishes in like 20 seconds!<br>\nThen i check the log and output file, i do have submmisons.csv file, but very short one, with only 108 rows, i first thought that maybe the env.iter_test() didn't work, so i add log for each iteration, the log file shows that it stops at iteration 4, and i have no other information.</p>\n<p>here is my code.</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\n\n############\n# all the files needed and pre-train model are loaded there.\n############\n\n# MAKE PREDICTION\niter_test = env.iter_test()\ntest_prev = pd.DataFrame()\nn_iter = 1\nfor (df_test, sample_prediction_df) in iter_test: \n\n    logging.info('\\n*CURRENT ITERATION*:-{}-'.format(n_iter))\n    test, user_summary, lecture_summary = processing_test(df_test, test_prev,  questions lectures, user_summary, q_stats_train,  lecture_summary,  features,  question_tags)\n\n    test_prev = df_test.copy()\n    test['answered_correctly'] = model.predict(test.drop(['row_id','group_num'], axis=1))\n    test.set_index('group_num', inplace=True)\n    env.predict(test[['row_id', 'answered_correctly']])\n    n_iter += 1\n</code></pre>\n<p>here is the log that i created:</p>\n<pre><code>Loading all summaries:\n\nShape of user summary:(393656, 11),\nShape of lecture summary:(393656, 23),\nShape of question stats:(13523, 5) \n- Successful\n\n# of features to add(extract):48\n\nQuestions info:(13523, 5)\nLectures info:(418, 4)\n\nLoading model...\nmodel BernoulliNB() loaded -- Successful!...\n\n*CURRENT ITERATION*:-1-\nNumExpr defaulting to 4 threads.\n\nsplit into :new_test (18, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (18, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (18, 55)\n\n*CURRENT ITERATION*:-2-\n\nsplit into :new_test (27, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (27, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (27, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (18, 11)\n\nUpdating user_summary:\nDone!\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-3-\n\nsplit into :new_test (26, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (26, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (26, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (27, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-4-\n\nsplit into :new_test (33, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user &amp; question summary by userID...\ndf returned by [add_questions_summary] has shape of (33, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (33, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (26, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n</code></pre>",
      "rawMarkdown": "Hey, i tried to submit my prediction and i have turned off the internet of my notebook.\nwhen  i click  **save and run all** , the process finishes in like 20 seconds!\nThen i check the log and output file, i do have submmisons.csv file, but very short one, with only 108 rows, i first thought that maybe the env.iter_test() didn't work, so i add log for each iteration, the log file shows that it stops at iteration 4, and i have no other information.\n\nhere is my code.\n\n```\n\nimport riiideducation\nenv = riiideducation.make_env()\n\n############\n# all the files needed and pre-train model are loaded there.\n############\n\n# MAKE PREDICTION\niter_test = env.iter_test()\ntest_prev = pd.DataFrame()\nn_iter = 1\nfor (df_test, sample_prediction_df) in iter_test: \n\n    logging.info('\\n*CURRENT ITERATION*:-{}-'.format(n_iter))\n    test, user_summary, lecture_summary = processing_test(df_test, test_prev,  questions lectures, user_summary, q_stats_train,  lecture_summary,  features,  question_tags)\n\n    test_prev = df_test.copy()\n    test['answered_correctly'] = model.predict(test.drop(['row_id','group_num'], axis=1))\n    test.set_index('group_num', inplace=True)\n    env.predict(test[['row_id', 'answered_correctly']])\n    n_iter += 1\n```\nhere is the log that i created:\n\n```\nLoading all summaries:\n\nShape of user summary:(393656, 11),\nShape of lecture summary:(393656, 23),\nShape of question stats:(13523, 5) \n- Successful\n\n# of features to add(extract):48\n\nQuestions info:(13523, 5)\nLectures info:(418, 4)\n\nLoading model...\nmodel BernoulliNB() loaded -- Successful!...\n\n*CURRENT ITERATION*:-1-\nNumExpr defaulting to 4 threads.\n\nsplit into :new_test (18, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (18, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (18, 55)\n\n*CURRENT ITERATION*:-2-\n\nsplit into :new_test (27, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (27, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (27, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (18, 11)\n\nUpdating user_summary:\nDone!\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-3-\n\nsplit into :new_test (26, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (26, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (26, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (27, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-4-\n\nsplit into :new_test (33, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (33, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (33, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (26, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n```",
      "votes": null
    },
    {
      "id": "1077949",
      "postDate": "11/14/2020 05:56:38",
      "content": "<p>I'm having the same issue</p>",
      "rawMarkdown": "I'm having the same issue",
      "votes": null
    },
    {
      "id": "1077981",
      "postDate": "11/14/2020 07:17:34",
      "content": "<p>save and run all is not the submission. When you submit, a unseen hidden dataset is used, which is much larger than the example_test.csv.</p>",
      "rawMarkdown": "save and run all is not the submission. When you submit, a unseen hidden dataset is used, which is much larger than the example_test.csv.",
      "votes": null
    },
    {
      "id": "1077995",
      "postDate": "11/14/2020 07:41:54",
      "content": "<p>so it is ok that the notebook finishes produce a small submission.csv, which i can select from submit predictions and use this one?</p>\n<p>i tried to submit with this file, but there is a Submission Scoring Error!!</p>",
      "rawMarkdown": "so it is ok that the notebook finishes produce a small submission.csv, which i can select from submit predictions and use this one?\n\ni tried to submit with this file, but there is a Submission Scoring Error!!",
      "votes": null
    },
    {
      "id": "1077998",
      "postDate": "11/14/2020 07:45:21",
      "content": "<p>It means you have bug in your code. I suggest to use the official start notebook , and build on it step by step</p>",
      "rawMarkdown": "It means you have bug in your code. I suggest to use the official start notebook , and build on it step by step",
      "votes": null
    },
    {
      "id": "1078228",
      "postDate": "11/14/2020 14:22:15",
      "content": "<p>Yes, i think you are right, i am able to submit after i comment out some of the code…</p>",
      "rawMarkdown": "Yes, i think you are right, i am able to submit after i comment out some of the code...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077949,
      "author_name": "ahensley",
      "author_url": "",
      "post_date": "11/14/2020 05:56:38",
      "content": "<p>I'm having the same issue</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1077981,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "11/14/2020 07:17:34",
      "content": "<p>save and run all is not the submission. When you submit, a unseen hidden dataset is used, which is much larger than the example_test.csv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077995,
          "author_name": "toddgm",
          "author_url": "",
          "post_date": "11/14/2020 07:41:54",
          "content": "<p>so it is ok that the notebook finishes produce a small submission.csv, which i can select from submit predictions and use this one?</p>\n<p>i tried to submit with this file, but there is a Submission Scoring Error!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1077998,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "11/14/2020 07:45:21",
          "content": "<p>It means you have bug in your code. I suggest to use the official start notebook , and build on it step by step</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1078228,
          "author_name": "toddgm",
          "author_url": "",
          "post_date": "11/14/2020 14:22:15",
          "content": "<p>Yes, i think you are right, i am able to submit after i comment out some of the code…</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1077928": "Hey, i tried to submit my prediction and i have turned off the internet of my notebook.\nwhen  i click  **save and run all** , the process finishes in like 20 seconds!\nThen i check the log and output file, i do have submmisons.csv file, but very short one, with only 108 rows, i first thought that maybe the env.iter_test() didn't work, so i add log for each iteration, the log file shows that it stops at iteration 4, and i have no other information.\n\nhere is my code.\n\n```\n\nimport riiideducation\nenv = riiideducation.make_env()\n\n############\n# all the files needed and pre-train model are loaded there.\n############\n\n# MAKE PREDICTION\niter_test = env.iter_test()\ntest_prev = pd.DataFrame()\nn_iter = 1\nfor (df_test, sample_prediction_df) in iter_test: \n\n    logging.info('\\n*CURRENT ITERATION*:-{}-'.format(n_iter))\n    test, user_summary, lecture_summary = processing_test(df_test, test_prev,  questions lectures, user_summary, q_stats_train,  lecture_summary,  features,  question_tags)\n\n    test_prev = df_test.copy()\n    test['answered_correctly'] = model.predict(test.drop(['row_id','group_num'], axis=1))\n    test.set_index('group_num', inplace=True)\n    env.predict(test[['row_id', 'answered_correctly']])\n    n_iter += 1\n```\nhere is the log that i created:\n\n```\nLoading all summaries:\n\nShape of user summary:(393656, 11),\nShape of lecture summary:(393656, 23),\nShape of question stats:(13523, 5) \n- Successful\n\n# of features to add(extract):48\n\nQuestions info:(13523, 5)\nLectures info:(418, 4)\n\nLoading model...\nmodel BernoulliNB() loaded -- Successful!...\n\n*CURRENT ITERATION*:-1-\nNumExpr defaulting to 4 threads.\n\nsplit into :new_test (18, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (18, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (18, 55)\n\n*CURRENT ITERATION*:-2-\n\nsplit into :new_test (27, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (27, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (27, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (18, 11)\n\nUpdating user_summary:\nDone!\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-3-\n\nsplit into :new_test (26, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (26, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (26, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (27, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n\n*CURRENT ITERATION*:-4-\n\nsplit into :new_test (33, 11) and ltest(0, 15)\n\nAdding top 10 question tags...\nAdding  question parts...\n\nLooking up statistics from user & question summary by userID...\ndf returned by [add_questions_summary] has shape of (33, 34)\n\nLooking up statistics lecture summary by userID...\ndf returned by [add_lecture_summary] has shape of (33, 55)\n\nUpdating test_prev:\nDone!\ndf returned by [update_test_prev] has shape of (26, 11)\n\nUpdating user_summary:\ndf returned by [update_user_summary] has shape of (393657, 11)\n```",
    "1077949": "I'm having the same issue",
    "1077981": "save and run all is not the submission. When you submit, a unseen hidden dataset is used, which is much larger than the example_test.csv.",
    "1077995": "so it is ok that the notebook finishes produce a small submission.csv, which i can select from submit predictions and use this one?\n\ni tried to submit with this file, but there is a Submission Scoring Error!!",
    "1077998": "It means you have bug in your code. I suggest to use the official start notebook , and build on it step by step",
    "1078228": "Yes, i think you are right, i am able to submit after i comment out some of the code..."
  },
  "source": "meta"
}