{
  "id": 208250,
  "title": "Notebook timeout",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208250",
  "author_name": "",
  "post_date": "2021-01-02T14:45:01.919363500Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I keep having notebook timeout errors. I don't understand why because, in local, I have a time by row of about 3.5ms (it needs to be below 12ms) and, on the sample rows on kaggle I'm about 7ms by row. I use a Transformers-like model. I don't use the GPU in inference as the transfer cpu -&gt; gpu takes time and the pipeline seems faster on every aspect on CPU (which is weird). The only part that the gpu should be faster is when applying the transformers but, strangely, the cpu version is faster (around 45ms by inference groups, using the iter_test emulator many uses).</p>\n<p>I'm running out of ideas on how to debug this.<br>\nAny idea ?<br>\nThanks</p>",
  "messages": [
    {
      "id": "1135833",
      "postDate": "01/02/2021 14:45:01",
      "content": "<p>I keep having notebook timeout errors. I don't understand why because, in local, I have a time by row of about 3.5ms (it needs to be below 12ms) and, on the sample rows on kaggle I'm about 7ms by row. I use a Transformers-like model. I don't use the GPU in inference as the transfer cpu -&gt; gpu takes time and the pipeline seems faster on every aspect on CPU (which is weird). The only part that the gpu should be faster is when applying the transformers but, strangely, the cpu version is faster (around 45ms by inference groups, using the iter_test emulator many uses).</p>\n<p>I'm running out of ideas on how to debug this.<br>\nAny idea ?<br>\nThanks</p>",
      "rawMarkdown": "I keep having notebook timeout errors. I don't understand why because, in local, I have a time by row of about 3.5ms (it needs to be below 12ms) and, on the sample rows on kaggle I'm about 7ms by row. I use a Transformers-like model. I don't use the GPU in inference as the transfer cpu -> gpu takes time and the pipeline seems faster on every aspect on CPU (which is weird). The only part that the gpu should be faster is when applying the transformers but, strangely, the cpu version is faster (around 45ms by inference groups, using the iter_test emulator many uses).\n\nI'm running out of ideas on how to debug this.\nAny idea ?\nThanks",
      "votes": null
    },
    {
      "id": "1135962",
      "postDate": "01/02/2021 16:18:06",
      "content": "<p>My code (feature engineering, pytorch model, state updates, etc) is in the \"model.predict\" line 7 below. I tried (below) to compute the predictions on the first 10% of the data (line 6) but always outputs the sample_prediction_df instead of the computed predictions. This code works within the limit (take about 1h30).<br>\nWhenever I comment line 8-9 (that erase the prediction_df), I got a timeout error.</p>\n<p>I guess it means that it's not my computations that are taking too much time but what happens after that. The env.predict call could take a lot of time depending on the input.<br>\nI checked the format of the prediction_df compared to the sample_prediction_df. The only difference is that prediction_df has the group_num as column instead of index (but at the last line I select only row_id and answered_correctly). I will move it to the index in a next submission but I would be surprised if it's the cause of the timeout.</p>\n<pre><code>rows_treated = 0\nfrom time import time\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated &lt; 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)    \n        prediction_df = sample_prediction_df # commenting these 2 lines create a timeout error\n        prediction_df['answered_correctly'] = 0.62 # commenting these 2 lines create a timeout error\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n</code></pre>",
      "rawMarkdown": "My code (feature engineering, pytorch model, state updates, etc) is in the \"model.predict\" line 7 below. I tried (below) to compute the predictions on the first 10% of the data (line 6) but always outputs the sample_prediction_df instead of the computed predictions. This code works within the limit (take about 1h30).\nWhenever I comment line 8-9 (that erase the prediction_df), I got a timeout error.\n\nI guess it means that it's not my computations that are taking too much time but what happens after that. The env.predict call could take a lot of time depending on the input.\nI checked the format of the prediction_df compared to the sample_prediction_df. The only difference is that prediction_df has the group_num as column instead of index (but at the last line I select only row_id and answered_correctly). I will move it to the index in a next submission but I would be surprised if it's the cause of the timeout.\n\n```\nrows_treated = 0\nfrom time import time\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated < 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)    \n        prediction_df = sample_prediction_df # commenting these 2 lines create a timeout error\n        prediction_df['answered_correctly'] = 0.62 # commenting these 2 lines create a timeout error\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n```",
      "votes": null
    },
    {
      "id": "1136052",
      "postDate": "01/02/2021 17:32:45",
      "content": "<p>Make sure you remove lecture rows in your predictions.</p>",
      "rawMarkdown": "Make sure you remove lecture rows in your predictions.",
      "votes": null
    },
    {
      "id": "1136093",
      "postDate": "01/02/2021 18:11:26",
      "content": "<p>Using the iter_test emulator of <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a>, it should have caught this kind of problem. Here are the code I run for this emulator, having added  checks on the row_ids to be sure I don't submit lecture rows for example.</p>\n<pre><code>for iteration, (current_test, current_prediction_df) in enumerate(iter_test):\n    row_ids_before = current_test['row_id'][current_test['content_type_id']==0].values\n    current_test['answered_correctly'] = np.nan\n    current_test = preprocess_inference_batch(inference_batch=current_test)\n    current_test = model.predict(students=students,\n                                 inference_batch=current_test)\n    current_test = current_test.sort_values(\"row_id\")\n    row_ids_after = current_test['row_id'].values\n    assert len(row_ids_before) == len(row_ids_after)\n    assert set(row_ids_before) == set(row_ids_after)\n    set_predict(current_test.loc[:,['row_id', 'answered_correctly']])\n</code></pre>",
      "rawMarkdown": "Using the iter_test emulator of https://www.kaggle.com/its7171/time-series-api-iter-test-emulator, it should have caught this kind of problem. Here are the code I run for this emulator, having added  checks on the row_ids to be sure I don't submit lecture rows for example.\n\n```\nfor iteration, (current_test, current_prediction_df) in enumerate(iter_test):\n    row_ids_before = current_test['row_id'][current_test['content_type_id']==0].values\n    current_test['answered_correctly'] = np.nan\n    current_test = preprocess_inference_batch(inference_batch=current_test)\n    current_test = model.predict(students=students,\n                                 inference_batch=current_test)\n    current_test = current_test.sort_values(\"row_id\")\n    row_ids_after = current_test['row_id'].values\n    assert len(row_ids_before) == len(row_ids_after)\n    assert set(row_ids_before) == set(row_ids_after)\n    set_predict(current_test.loc[:,['row_id', 'answered_correctly']])\n```",
      "votes": null
    },
    {
      "id": "1136102",
      "postDate": "01/02/2021 18:17:14",
      "content": "<p>My last submission of the day raised an error in 20 minutes. I tried to copy my predictions to the sample_prediction_df so that I avoid problems with the group_num index or anything else from the format of the dfs. I added an assert to check row_ids are aligned and equal and I guess the error likely comes from this assert. I don't understand how my code can run fine on the iter_test emulator (which checks lecture rows are not predicted) while failing here. I'll check again that I remove lecture rows.</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated &lt; 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)\n        prediction_df.sort_values('row_id', inplace=True)\n        sample_prediction_df.sort_values('row_id', inplace=True)\n        assert (prediction_df.row_id.values == sample_prediction_df.row_id.values).all()\n        sample_prediction_df['answered_correctly'] = prediction_df.answered_correctly.values\n        prediction_df = sample_prediction_df\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    #prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n</code></pre>",
      "rawMarkdown": "My last submission of the day raised an error in 20 minutes. I tried to copy my predictions to the sample_prediction_df so that I avoid problems with the group_num index or anything else from the format of the dfs. I added an assert to check row_ids are aligned and equal and I guess the error likely comes from this assert. I don't understand how my code can run fine on the iter_test emulator (which checks lecture rows are not predicted) while failing here. I'll check again that I remove lecture rows.\n\n```\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated < 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)\n        prediction_df.sort_values('row_id', inplace=True)\n        sample_prediction_df.sort_values('row_id', inplace=True)\n        assert (prediction_df.row_id.values == sample_prediction_df.row_id.values).all()\n        sample_prediction_df['answered_correctly'] = prediction_df.answered_correctly.values\n        prediction_df = sample_prediction_df\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    #prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n```",
      "votes": null
    },
    {
      "id": "1138243",
      "postDate": "01/04/2021 14:21:22",
      "content": "<p>I found the solution of my problem. I was actually producing too many rows (duplicates) and, somehow, kaggle doesn't detect that kind of errors and doesn't produce a submission error but a timeout one. </p>",
      "rawMarkdown": "I found the solution of my problem. I was actually producing too many rows (duplicates) and, somehow, kaggle doesn't detect that kind of errors and doesn't produce a submission error but a timeout one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1135962,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "01/02/2021 16:18:06",
      "content": "<p>My code (feature engineering, pytorch model, state updates, etc) is in the \"model.predict\" line 7 below. I tried (below) to compute the predictions on the first 10% of the data (line 6) but always outputs the sample_prediction_df instead of the computed predictions. This code works within the limit (take about 1h30).<br>\nWhenever I comment line 8-9 (that erase the prediction_df), I got a timeout error.</p>\n<p>I guess it means that it's not my computations that are taking too much time but what happens after that. The env.predict call could take a lot of time depending on the input.<br>\nI checked the format of the prediction_df compared to the sample_prediction_df. The only difference is that prediction_df has the group_num as column instead of index (but at the last line I select only row_id and answered_correctly). I will move it to the index in a next submission but I would be surprised if it's the cause of the timeout.</p>\n<pre><code>rows_treated = 0\nfrom time import time\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated &lt; 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)    \n        prediction_df = sample_prediction_df # commenting these 2 lines create a timeout error\n        prediction_df['answered_correctly'] = 0.62 # commenting these 2 lines create a timeout error\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1136052,
          "author_name": "mingpan07",
          "author_url": "",
          "post_date": "01/02/2021 17:32:45",
          "content": "<p>Make sure you remove lecture rows in your predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136093,
          "author_name": "rodolphelampe",
          "author_url": "",
          "post_date": "01/02/2021 18:11:26",
          "content": "<p>Using the iter_test emulator of <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a>, it should have caught this kind of problem. Here are the code I run for this emulator, having added  checks on the row_ids to be sure I don't submit lecture rows for example.</p>\n<pre><code>for iteration, (current_test, current_prediction_df) in enumerate(iter_test):\n    row_ids_before = current_test['row_id'][current_test['content_type_id']==0].values\n    current_test['answered_correctly'] = np.nan\n    current_test = preprocess_inference_batch(inference_batch=current_test)\n    current_test = model.predict(students=students,\n                                 inference_batch=current_test)\n    current_test = current_test.sort_values(\"row_id\")\n    row_ids_after = current_test['row_id'].values\n    assert len(row_ids_before) == len(row_ids_after)\n    assert set(row_ids_before) == set(row_ids_after)\n    set_predict(current_test.loc[:,['row_id', 'answered_correctly']])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1136102,
          "author_name": "rodolphelampe",
          "author_url": "",
          "post_date": "01/02/2021 18:17:14",
          "content": "<p>My last submission of the day raised an error in 20 minutes. I tried to copy my predictions to the sample_prediction_df so that I avoid problems with the group_num index or anything else from the format of the dfs. I added an assert to check row_ids are aligned and equal and I guess the error likely comes from this assert. I don't understand how my code can run fine on the iter_test emulator (which checks lecture rows are not predicted) while failing here. I'll check again that I remove lecture rows.</p>\n<pre><code>for (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated &lt; 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)\n        prediction_df.sort_values('row_id', inplace=True)\n        sample_prediction_df.sort_values('row_id', inplace=True)\n        assert (prediction_df.row_id.values == sample_prediction_df.row_id.values).all()\n        sample_prediction_df['answered_correctly'] = prediction_df.answered_correctly.values\n        prediction_df = sample_prediction_df\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    #prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1138243,
      "author_name": "rodolphelampe",
      "author_url": "",
      "post_date": "01/04/2021 14:21:22",
      "content": "<p>I found the solution of my problem. I was actually producing too many rows (duplicates) and, somehow, kaggle doesn't detect that kind of errors and doesn't produce a submission error but a timeout one. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1135833": "I keep having notebook timeout errors. I don't understand why because, in local, I have a time by row of about 3.5ms (it needs to be below 12ms) and, on the sample rows on kaggle I'm about 7ms by row. I use a Transformers-like model. I don't use the GPU in inference as the transfer cpu -> gpu takes time and the pipeline seems faster on every aspect on CPU (which is weird). The only part that the gpu should be faster is when applying the transformers but, strangely, the cpu version is faster (around 45ms by inference groups, using the iter_test emulator many uses).\n\nI'm running out of ideas on how to debug this.\nAny idea ?\nThanks",
    "1135962": "My code (feature engineering, pytorch model, state updates, etc) is in the \"model.predict\" line 7 below. I tried (below) to compute the predictions on the first 10% of the data (line 6) but always outputs the sample_prediction_df instead of the computed predictions. This code works within the limit (take about 1h30).\nWhenever I comment line 8-9 (that erase the prediction_df), I got a timeout error.\n\nI guess it means that it's not my computations that are taking too much time but what happens after that. The env.predict call could take a lot of time depending on the input.\nI checked the format of the prediction_df compared to the sample_prediction_df. The only difference is that prediction_df has the group_num as column instead of index (but at the last line I select only row_id and answered_correctly). I will move it to the index in a next submission but I would be surprised if it's the cause of the timeout.\n\n```\nrows_treated = 0\nfrom time import time\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated < 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)    \n        prediction_df = sample_prediction_df # commenting these 2 lines create a timeout error\n        prediction_df['answered_correctly'] = 0.62 # commenting these 2 lines create a timeout error\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n```",
    "1136052": "Make sure you remove lecture rows in your predictions.",
    "1136093": "Using the iter_test emulator of https://www.kaggle.com/its7171/time-series-api-iter-test-emulator, it should have caught this kind of problem. Here are the code I run for this emulator, having added  checks on the row_ids to be sure I don't submit lecture rows for example.\n\n```\nfor iteration, (current_test, current_prediction_df) in enumerate(iter_test):\n    row_ids_before = current_test['row_id'][current_test['content_type_id']==0].values\n    current_test['answered_correctly'] = np.nan\n    current_test = preprocess_inference_batch(inference_batch=current_test)\n    current_test = model.predict(students=students,\n                                 inference_batch=current_test)\n    current_test = current_test.sort_values(\"row_id\")\n    row_ids_after = current_test['row_id'].values\n    assert len(row_ids_before) == len(row_ids_after)\n    assert set(row_ids_before) == set(row_ids_after)\n    set_predict(current_test.loc[:,['row_id', 'answered_correctly']])\n```",
    "1136102": "My last submission of the day raised an error in 20 minutes. I tried to copy my predictions to the sample_prediction_df so that I avoid problems with the group_num index or anything else from the format of the dfs. I added an assert to check row_ids are aligned and equal and I guess the error likely comes from this assert. I don't understand how my code can run fine on the iter_test emulator (which checks lecture rows are not predicted) while failing here. I'll check again that I remove lecture rows.\n\n```\nfor (test_df, sample_prediction_df) in iter_test:\n    test_df = preprocess_inference_batch(inference_batch=test_df)\n    test_df.loc[:, \"answered_correctly\"] = np.nan    \n    if rows_treated < 0.1 * 2.5e6:\n        prediction_df = model.predict(students=students, inference_batch=test_df)\n        prediction_df.sort_values('row_id', inplace=True)\n        sample_prediction_df.sort_values('row_id', inplace=True)\n        assert (prediction_df.row_id.values == sample_prediction_df.row_id.values).all()\n        sample_prediction_df['answered_correctly'] = prediction_df.answered_correctly.values\n        prediction_df = sample_prediction_df\n    else:\n        prediction_df = sample_prediction_df\n        prediction_df['answered_correctly'] = 0.62\n    #prediction_df = prediction_df.sort_values(\"row_id\")\n    env.predict(prediction_df.loc[:, ['row_id', 'answered_correctly']])\n    rows_treated += len(prediction_df)\n```",
    "1138243": "I found the solution of my problem. I was actually producing too many rows (duplicates) and, somehow, kaggle doesn't detect that kind of errors and doesn't produce a submission error but a timeout one."
  },
  "source": "meta"
}