{
  "id": 206396,
  "title": "Notebook gets timeout on submission",
  "url": "/competitions/riiid-test-answer-prediction/discussion/206396",
  "author_name": "",
  "post_date": "2020-12-24T12:02:44.202038400Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi, </p>\n<p>I'm now only a month in machine learning and new to pandas.<br>\nAs I don't know anything about gradient boost I started to solve this competition via an <br>\npytorch neural net. My approach does certainly not make much sense, but for now my<br>\nonly goal is to get a running submission.<br>\nI had quite some problems with memory, so I reduced my training data to 2.5 million.<br>\nMy notebook now finishes at around one hour.<br>\nBut when I try to submit my submission.csv and it is reran,<br>\nthe notebook runs forever and after a night I get a Timeout error for the submission.</p>\n<p>As it is the same notebook and thus the same training data size I think on the submission<br>\nthere will be the same time needed for training my model as on the previous run I made myself.</p>\n<p>And then there is the test data. As far as I read there will be 2.5 million questions, but is this <br>\nequivalent to the test data rows, as the test data consists of answers not questions, isn't it?</p>\n<p>On my notebook run I measured the time for a test data iteration. It outputs 108 rows and has 4 iterations. This took overall 0.72 seconds while spending around  0.03 on one iteration. As this does not sum up, the most time must be somehow spend in the line \"for (test_df, sample_prediction_df) in iter_test:\" what I cant explain. But also if I take 2.5 mio rows (which I dont know is the correct number for test data) it should take around six hours.<br>\nAdditionally one hour for training the model it should be finished after 7 hours, which it isn't.</p>\n<p>But if the time is spent on \"for (test_df, sample_prediction_df) in iter_test:\" line, it means I cannot improve anything on this.</p>\n<p>So what am I doing wrong, and how can I get the submission done in time?</p>\n<p>If needed I could share my notebook, but always with the hint that my approach is presumably nonsense and I'm doing weird stuff because of really not knowing  pandas or machine learning very good. :-)</p>\n<p>Greetings and merry Xmas</p>",
  "messages": [
    {
      "id": "1125131",
      "postDate": "12/24/2020 12:02:44",
      "content": "<p>Hi, </p>\n<p>I'm now only a month in machine learning and new to pandas.<br>\nAs I don't know anything about gradient boost I started to solve this competition via an <br>\npytorch neural net. My approach does certainly not make much sense, but for now my<br>\nonly goal is to get a running submission.<br>\nI had quite some problems with memory, so I reduced my training data to 2.5 million.<br>\nMy notebook now finishes at around one hour.<br>\nBut when I try to submit my submission.csv and it is reran,<br>\nthe notebook runs forever and after a night I get a Timeout error for the submission.</p>\n<p>As it is the same notebook and thus the same training data size I think on the submission<br>\nthere will be the same time needed for training my model as on the previous run I made myself.</p>\n<p>And then there is the test data. As far as I read there will be 2.5 million questions, but is this <br>\nequivalent to the test data rows, as the test data consists of answers not questions, isn't it?</p>\n<p>On my notebook run I measured the time for a test data iteration. It outputs 108 rows and has 4 iterations. This took overall 0.72 seconds while spending around  0.03 on one iteration. As this does not sum up, the most time must be somehow spend in the line \"for (test_df, sample_prediction_df) in iter_test:\" what I cant explain. But also if I take 2.5 mio rows (which I dont know is the correct number for test data) it should take around six hours.<br>\nAdditionally one hour for training the model it should be finished after 7 hours, which it isn't.</p>\n<p>But if the time is spent on \"for (test_df, sample_prediction_df) in iter_test:\" line, it means I cannot improve anything on this.</p>\n<p>So what am I doing wrong, and how can I get the submission done in time?</p>\n<p>If needed I could share my notebook, but always with the hint that my approach is presumably nonsense and I'm doing weird stuff because of really not knowing  pandas or machine learning very good. :-)</p>\n<p>Greetings and merry Xmas</p>",
      "rawMarkdown": "Hi, \n\nI'm now only a month in machine learning and new to pandas.\nAs I don't know anything about gradient boost I started to solve this competition via an \npytorch neural net. My approach does certainly not make much sense, but for now my\nonly goal is to get a running submission.\nI had quite some problems with memory, so I reduced my training data to 2.5 million.\nMy notebook now finishes at around one hour.\nBut when I try to submit my submission.csv and it is reran,\nthe notebook runs forever and after a night I get a Timeout error for the submission.\n\nAs it is the same notebook and thus the same training data size I think on the submission\nthere will be the same time needed for training my model as on the previous run I made myself.\n\nAnd then there is the test data. As far as I read there will be 2.5 million questions, but is this \nequivalent to the test data rows, as the test data consists of answers not questions, isn't it?\n\nOn my notebook run I measured the time for a test data iteration. It outputs 108 rows and has 4 iterations. This took overall 0.72 seconds while spending around  0.03 on one iteration. As this does not sum up, the most time must be somehow spend in the line \"for (test_df, sample_prediction_df) in iter_test:\" what I cant explain. But also if I take 2.5 mio rows (which I dont know is the correct number for test data) it should take around six hours.\nAdditionally one hour for training the model it should be finished after 7 hours, which it isn't.\n\nBut if the time is spent on \"for (test_df, sample_prediction_df) in iter_test:\" line, it means I cannot improve anything on this.\n\nSo what am I doing wrong, and how can I get the submission done in time?\n\nIf needed I could share my notebook, but always with the hint that my approach is presumably nonsense and I'm doing weird stuff because of really not knowing  pandas or machine learning very good. :-)\n\nGreetings and merry Xmas",
      "votes": null
    },
    {
      "id": "1127771",
      "postDate": "12/26/2020 20:58:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/logarith2020\" target=\"_blank\">@logarith2020</a> ,<br>\n  I am new to ML. But I have learned that you could save the trained model which takes 2 hrs and just load it in a code within seconds and do predictions. So by doing so you can reduce the training time.<br>\n  But even doing so my code is not fast enough. If yours could do better that would help. I think.<br>\n  I have submitted my code with a small tweek from this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198414\" target=\"_blank\">discussion</a>. I am not sure about the result, If I had any I will update. </p>",
      "rawMarkdown": "Hi @logarith2020 ,\n  I am new to ML. But I have learned that you could save the trained model which takes 2 hrs and just load it in a code within seconds and do predictions. So by doing so you can reduce the training time.\n  But even doing so my code is not fast enough. If yours could do better that would help. I think.\n  I have submitted my code with a small tweek from this [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198414). I am not sure about the result, If I had any I will update.",
      "votes": null
    },
    {
      "id": "1128179",
      "postDate": "12/27/2020 08:24:38",
      "content": "<p><a href=\"https://www.kaggle.com/abishekrobin\" target=\"_blank\">@abishekrobin</a> Thanks for your hint. I tried the small tweak from the discussion with the 8 hours, but I still run in the same timeout problem. <br>\nNow I give it a try with a trained model, but I don't expect any difference, as now for testing purposes I trained my model only on 20000 rows, which is normally done in some seconds.</p>\n<p>It's making me nuts, that I don't have a clue where it could take so long. The example test set works well.</p>",
      "rawMarkdown": "abishekrobin Thanks for your hint. I tried the small tweak from the discussion with the 8 hours, but I still run in the same timeout problem. \nNow I give it a try with a trained model, but I don't expect any difference, as now for testing purposes I trained my model only on 20000 rows, which is normally done in some seconds.\n\nIt's making me nuts, that I don't have a clue where it could take so long. The example test set works well.",
      "votes": null
    },
    {
      "id": "1128500",
      "postDate": "12/27/2020 13:41:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/logarith2020\" target=\"_blank\">@logarith2020</a> Don't compromise your training. What I do is train my model in my local system for two hours and save the model and upload it to the competition notebook in server. And then the 8hr trick worked for me. make sure you are  predicting only for 'test_df = test_df[test_df['content_type_id'] == 0]', not all the test data given to you. I think this line did the trick for me after 8 failure notebook submissions.</p>",
      "rawMarkdown": "Hi @logarith2020 Don't compromise your training. What I do is train my model in my local system for two hours and save the model and upload it to the competition notebook in server. And then the 8hr trick worked for me. make sure you are  predicting only for 'test_df = test_df[test_df['content_type_id'] == 0]', not all the test data given to you. I think this line did the trick for me after 8 failure notebook submissions.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1127771,
      "author_name": "abishekrobin",
      "author_url": "",
      "post_date": "12/26/2020 20:58:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/logarith2020\" target=\"_blank\">@logarith2020</a> ,<br>\n  I am new to ML. But I have learned that you could save the trained model which takes 2 hrs and just load it in a code within seconds and do predictions. So by doing so you can reduce the training time.<br>\n  But even doing so my code is not fast enough. If yours could do better that would help. I think.<br>\n  I have submitted my code with a small tweek from this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198414\" target=\"_blank\">discussion</a>. I am not sure about the result, If I had any I will update. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128179,
      "author_name": "logarith2020",
      "author_url": "",
      "post_date": "12/27/2020 08:24:38",
      "content": "<p><a href=\"https://www.kaggle.com/abishekrobin\" target=\"_blank\">@abishekrobin</a> Thanks for your hint. I tried the small tweak from the discussion with the 8 hours, but I still run in the same timeout problem. <br>\nNow I give it a try with a trained model, but I don't expect any difference, as now for testing purposes I trained my model only on 20000 rows, which is normally done in some seconds.</p>\n<p>It's making me nuts, that I don't have a clue where it could take so long. The example test set works well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128500,
      "author_name": "abishekrobin",
      "author_url": "",
      "post_date": "12/27/2020 13:41:02",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/logarith2020\" target=\"_blank\">@logarith2020</a> Don't compromise your training. What I do is train my model in my local system for two hours and save the model and upload it to the competition notebook in server. And then the 8hr trick worked for me. make sure you are  predicting only for 'test_df = test_df[test_df['content_type_id'] == 0]', not all the test data given to you. I think this line did the trick for me after 8 failure notebook submissions.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125131": "Hi, \n\nI'm now only a month in machine learning and new to pandas.\nAs I don't know anything about gradient boost I started to solve this competition via an \npytorch neural net. My approach does certainly not make much sense, but for now my\nonly goal is to get a running submission.\nI had quite some problems with memory, so I reduced my training data to 2.5 million.\nMy notebook now finishes at around one hour.\nBut when I try to submit my submission.csv and it is reran,\nthe notebook runs forever and after a night I get a Timeout error for the submission.\n\nAs it is the same notebook and thus the same training data size I think on the submission\nthere will be the same time needed for training my model as on the previous run I made myself.\n\nAnd then there is the test data. As far as I read there will be 2.5 million questions, but is this \nequivalent to the test data rows, as the test data consists of answers not questions, isn't it?\n\nOn my notebook run I measured the time for a test data iteration. It outputs 108 rows and has 4 iterations. This took overall 0.72 seconds while spending around  0.03 on one iteration. As this does not sum up, the most time must be somehow spend in the line \"for (test_df, sample_prediction_df) in iter_test:\" what I cant explain. But also if I take 2.5 mio rows (which I dont know is the correct number for test data) it should take around six hours.\nAdditionally one hour for training the model it should be finished after 7 hours, which it isn't.\n\nBut if the time is spent on \"for (test_df, sample_prediction_df) in iter_test:\" line, it means I cannot improve anything on this.\n\nSo what am I doing wrong, and how can I get the submission done in time?\n\nIf needed I could share my notebook, but always with the hint that my approach is presumably nonsense and I'm doing weird stuff because of really not knowing  pandas or machine learning very good. :-)\n\nGreetings and merry Xmas",
    "1127771": "Hi @logarith2020 ,\n  I am new to ML. But I have learned that you could save the trained model which takes 2 hrs and just load it in a code within seconds and do predictions. So by doing so you can reduce the training time.\n  But even doing so my code is not fast enough. If yours could do better that would help. I think.\n  I have submitted my code with a small tweek from this [discussion](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/198414). I am not sure about the result, If I had any I will update.",
    "1128179": "abishekrobin Thanks for your hint. I tried the small tweak from the discussion with the 8 hours, but I still run in the same timeout problem. \nNow I give it a try with a trained model, but I don't expect any difference, as now for testing purposes I trained my model only on 20000 rows, which is normally done in some seconds.\n\nIt's making me nuts, that I don't have a clue where it could take so long. The example test set works well.",
    "1128500": "Hi @logarith2020 Don't compromise your training. What I do is train my model in my local system for two hours and save the model and upload it to the competition notebook in server. And then the 8hr trick worked for me. make sure you are  predicting only for 'test_df = test_df[test_df['content_type_id'] == 0]', not all the test data given to you. I think this line did the trick for me after 8 failure notebook submissions."
  },
  "source": "meta"
}