{
  "id": 201156,
  "title": "Submission failing",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201156",
  "author_name": "",
  "post_date": "2020-12-03T12:23:35.818826200Z",
  "votes": null,
  "comment_count": 6,
  "views": 0,
  "content": "<p>My notebook is generating submission.csv. when I am submitting the file, the notebook takes forever to run and after 9 hours it shows \"submission scoring error\". Need suggestions.</p>",
  "messages": [
    {
      "id": "1100868",
      "postDate": "12/03/2020 12:23:35",
      "content": "<p>My notebook is generating submission.csv. when I am submitting the file, the notebook takes forever to run and after 9 hours it shows \"submission scoring error\". Need suggestions.</p>",
      "rawMarkdown": "My notebook is generating submission.csv. when I am submitting the file, the notebook takes forever to run and after 9 hours it shows \"submission scoring error\". Need suggestions.",
      "votes": null
    },
    {
      "id": "1100947",
      "postDate": "12/03/2020 13:52:29",
      "content": "<p>It's probably taking too long to process the test set. It's about 50k iterations with the median batch size somewhere in the range of 15-20 I think</p>",
      "rawMarkdown": "It's probably taking too long to process the test set. It's about 50k iterations with the median batch size somewhere in the range of 15-20 I think",
      "votes": null
    },
    {
      "id": "1100959",
      "postDate": "12/03/2020 14:09:40",
      "content": "<p>I trained the model for 50 iterations and commit time of the notebook is 20 mins. How come the model cant predict 50k*(15 or 20) rows in 8 hours?</p>",
      "rawMarkdown": "I trained the model for 50 iterations and commit time of the notebook is 20 mins. How come the model cant predict 50k*(15 or 20) rows in 8 hours?",
      "votes": null
    },
    {
      "id": "1101012",
      "postDate": "12/03/2020 14:50:58",
      "content": "<p>That you'll have to check yourself. Depends a lot on the type of pipeline that you have made.</p>\n<p>Try to simulate the test data and see if your pipeline is able to handle it in less than 8 hours or not.</p>",
      "rawMarkdown": "That you'll have to check yourself. Depends a lot on the type of pipeline that you have made.\n\nTry to simulate the test data and see if your pipeline is able to handle it in less than 8 hours or not.",
      "votes": null
    },
    {
      "id": "1101122",
      "postDate": "12/03/2020 16:38:59",
      "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Thanks. I am working on that.</p>",
      "rawMarkdown": "abdurrafae Thanks. I am working on that.",
      "votes": null
    },
    {
      "id": "1101164",
      "postDate": "12/03/2020 17:23:24",
      "content": "<p>You could try <a href=\"https://www.kaggle.com/calebeverett/riiid-mock-test-iterator\" target=\"_blank\">this</a> - same api as test and runs 2.5 million records without state update or model prediction in ~15 minutes - similar to test.</p>\n<p><a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Here</a> also is another emulator by an experienced kaggler.</p>",
      "rawMarkdown": "You could try [this](https://www.kaggle.com/calebeverett/riiid-mock-test-iterator) - same api as test and runs 2.5 million records without state update or model prediction in ~15 minutes - similar to test.\n\n[Here](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) also is another emulator by an experienced kaggler.",
      "votes": null
    },
    {
      "id": "1105426",
      "postDate": "12/07/2020 21:20:10",
      "content": "<p>Expect to see roughly 2.5 million questions in the hidden test set. This tells us that the given code. start_time= start_time= time.time()<br>\nfor (test_df, sample_prediction_df) in iter_test:<br>\n    test_df['answered_correctly'] =  model.predict_proba(test_df[col_fit2])[:,1]<br>\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])<br>\nprint(\"--- %s seconds ---\" % (time.time() - start_time))<br>\nMust be completed in less than 1.2 seconds<br>\nEvidence:<br>\n2500000/108 = 23148,1481481<br>\n23148,1481481*1.2=27777,7777777<br>\n(27777,7777777/60)/60 = 7,71604938269 hours<br>\n+- 1 hours in store</p>",
      "rawMarkdown": "Expect to see roughly 2.5 million questions in the hidden test set. This tells us that the given code. start_time= start_time= time.time()\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df['answered_correctly'] =  model.predict_proba(test_df[col_fit2])[:,1]\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\nprint(\"--- %s seconds ---\" % (time.time() - start_time))\nMust be completed in less than 1.2 seconds\nEvidence:\n2500000/108 = 23148,1481481\n23148,1481481*1.2=27777,7777777\n(27777,7777777/60)/60 = 7,71604938269 hours\n+- 1 hours in store",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1100947,
      "author_name": "abdurrafae",
      "author_url": "",
      "post_date": "12/03/2020 13:52:29",
      "content": "<p>It's probably taking too long to process the test set. It's about 50k iterations with the median batch size somewhere in the range of 15-20 I think</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100959,
          "author_name": "debojit23",
          "author_url": "",
          "post_date": "12/03/2020 14:09:40",
          "content": "<p>I trained the model for 50 iterations and commit time of the notebook is 20 mins. How come the model cant predict 50k*(15 or 20) rows in 8 hours?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101012,
          "author_name": "abdurrafae",
          "author_url": "",
          "post_date": "12/03/2020 14:50:58",
          "content": "<p>That you'll have to check yourself. Depends a lot on the type of pipeline that you have made.</p>\n<p>Try to simulate the test data and see if your pipeline is able to handle it in less than 8 hours or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101122,
          "author_name": "debojit23",
          "author_url": "",
          "post_date": "12/03/2020 16:38:59",
          "content": "<p><a href=\"https://www.kaggle.com/abdurrafae\" target=\"_blank\">@abdurrafae</a> Thanks. I am working on that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1101164,
          "author_name": "calebeverett",
          "author_url": "",
          "post_date": "12/03/2020 17:23:24",
          "content": "<p>You could try <a href=\"https://www.kaggle.com/calebeverett/riiid-mock-test-iterator\" target=\"_blank\">this</a> - same api as test and runs 2.5 million records without state update or model prediction in ~15 minutes - similar to test.</p>\n<p><a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">Here</a> also is another emulator by an experienced kaggler.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1105426,
      "author_name": "zefirchik",
      "author_url": "",
      "post_date": "12/07/2020 21:20:10",
      "content": "<p>Expect to see roughly 2.5 million questions in the hidden test set. This tells us that the given code. start_time= start_time= time.time()<br>\nfor (test_df, sample_prediction_df) in iter_test:<br>\n    test_df['answered_correctly'] =  model.predict_proba(test_df[col_fit2])[:,1]<br>\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])<br>\nprint(\"--- %s seconds ---\" % (time.time() - start_time))<br>\nMust be completed in less than 1.2 seconds<br>\nEvidence:<br>\n2500000/108 = 23148,1481481<br>\n23148,1481481*1.2=27777,7777777<br>\n(27777,7777777/60)/60 = 7,71604938269 hours<br>\n+- 1 hours in store</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1100868": "My notebook is generating submission.csv. when I am submitting the file, the notebook takes forever to run and after 9 hours it shows \"submission scoring error\". Need suggestions.",
    "1100947": "It's probably taking too long to process the test set. It's about 50k iterations with the median batch size somewhere in the range of 15-20 I think",
    "1100959": "I trained the model for 50 iterations and commit time of the notebook is 20 mins. How come the model cant predict 50k*(15 or 20) rows in 8 hours?",
    "1101012": "That you'll have to check yourself. Depends a lot on the type of pipeline that you have made.\n\nTry to simulate the test data and see if your pipeline is able to handle it in less than 8 hours or not.",
    "1101122": "abdurrafae Thanks. I am working on that.",
    "1101164": "You could try [this](https://www.kaggle.com/calebeverett/riiid-mock-test-iterator) - same api as test and runs 2.5 million records without state update or model prediction in ~15 minutes - similar to test.\n\n[Here](https://www.kaggle.com/its7171/time-series-api-iter-test-emulator) also is another emulator by an experienced kaggler.",
    "1105426": "Expect to see roughly 2.5 million questions in the hidden test set. This tells us that the given code. start_time= start_time= time.time()\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df['answered_correctly'] =  model.predict_proba(test_df[col_fit2])[:,1]\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\nprint(\"--- %s seconds ---\" % (time.time() - start_time))\nMust be completed in less than 1.2 seconds\nEvidence:\n2500000/108 = 23148,1481481\n23148,1481481*1.2=27777,7777777\n(27777,7777777/60)/60 = 7,71604938269 hours\n+- 1 hours in store"
  },
  "source": "meta"
}