{
  "id": 202090,
  "title": "over fitting or bug?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/202090",
  "author_name": "",
  "post_date": "2020-12-08T09:24:54.701537400Z",
  "votes": 1,
  "comment_count": 14,
  "views": 0,
  "content": "<p>cv score:0.77<br>\nlb score:0.71<br>\n😂😂😂<br>\nI used cv strategy from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>\n<p>Feature engineering and training data：<br>\n`train = features_(train)<br>\nvalidation = features_(validation)<br>\n……………..<br>\ny_train = train['answered_correctly']<br>\ntrain = train[features]</p>\n<p>y_val = validation['answered_correctly']<br>\nvalidation = validation[features]`</p>\n<p>Prediction:<br>\n<code>for test_df, _ in iter_test:\n    test_df = features_(test_df)\n    test_df['answered_correctly'] = model.predict(test_df[features])\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])</code></p>\n<p>Where is the problem?<br>\nOverfitting or some bug?</p>\n<p>I am really appreciate if someone can help me.😆</p>",
  "messages": [
    {
      "id": "1105867",
      "postDate": "12/08/2020 09:24:54",
      "content": "<p>cv score:0.77<br>\nlb score:0.71<br>\n😂😂😂<br>\nI used cv strategy from <a href=\"https://www.kaggle.com/its7171/cv-strategy\" target=\"_blank\">https://www.kaggle.com/its7171/cv-strategy</a></p>\n<p>Feature engineering and training data：<br>\n`train = features_(train)<br>\nvalidation = features_(validation)<br>\n……………..<br>\ny_train = train['answered_correctly']<br>\ntrain = train[features]</p>\n<p>y_val = validation['answered_correctly']<br>\nvalidation = validation[features]`</p>\n<p>Prediction:<br>\n<code>for test_df, _ in iter_test:\n    test_df = features_(test_df)\n    test_df['answered_correctly'] = model.predict(test_df[features])\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])</code></p>\n<p>Where is the problem?<br>\nOverfitting or some bug?</p>\n<p>I am really appreciate if someone can help me.😆</p>",
      "rawMarkdown": "cv score:0.77\nlb score:0.71\n😂😂😂\nI used cv strategy from https://www.kaggle.com/its7171/cv-strategy\n\nFeature engineering and training data：\n`train = features_(train)\nvalidation = features_(validation)\n.................\ny_train = train['answered_correctly']\ntrain = train[features]\n\ny_val = validation['answered_correctly']\nvalidation = validation[features]`\n\nPrediction:\n`for test_df, _ in iter_test:\n    test_df = features_(test_df)\n    test_df['answered_correctly'] = model.predict(test_df[features])\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])`\n\nWhere is the problem?\nOverfitting or some bug?\n\nI am really appreciate if someone can help me.😆",
      "votes": null
    },
    {
      "id": "1105975",
      "postDate": "12/08/2020 11:34:32",
      "content": "<p>the same here used all the train from train_cv_1 and validate on valid_cv_1 : <br>\ncv score : 0.79<br>\nLb : 0.711<br>\ni think overfitting … no ? any other similar cases ?  </p>",
      "rawMarkdown": "the same here used all the train from train_cv_1 and validate on valid_cv_1 : \ncv score : 0.79\nLb : 0.711\ni think overfitting ... no ? any other similar cases ?",
      "votes": null
    },
    {
      "id": "1106068",
      "postDate": "12/08/2020 13:32:34",
      "content": "<p>Could be bug. Got such issue by wrong computation of some features during inference. Same when I uploaded the wrong dump (pandas) containing some stats (i.e. average correct answer per question), trained with 90M but uploaded only with 10M.<br>\nCould be data leak too.  </p>",
      "rawMarkdown": "Could be bug. Got such issue by wrong computation of some features during inference. Same when I uploaded the wrong dump (pandas) containing some stats (i.e. average correct answer per question), trained with 90M but uploaded only with 10M.\nCould be data leak too.",
      "votes": null
    },
    {
      "id": "1108431",
      "postDate": "12/10/2020 16:31:20",
      "content": "<p>Aggregations with target is costly …. i am getting .82 something with 70 million dataset with xgb gpu ..<br>\nbut in leaderboard its .74 </p>\n<p>Has to be data leak with aggregating targets….. have  to find some other aggragations… may be a shake up in leader board</p>",
      "rawMarkdown": "Aggregations with target is costly .... i am getting .82 something with 70 million dataset with xgb gpu ..\nbut in leaderboard its .74 \n\nHas to be data leak with aggregating targets..... have  to find some other aggragations... may be a shake up in leader board",
      "votes": null
    },
    {
      "id": "1124527",
      "postDate": "12/24/2020 02:52:38",
      "content": "<p>I am also encountering the same issue, and using the same CV strategy. I made many changes at once to my solution and now I'm not sure of what's causing so big of a difference between my CV (~80) and LB (~70).</p>\n<p>But I'm positive that there is information leakage in the concept or execution in some features I created, although train and validation sets were split before feature engineering.</p>",
      "rawMarkdown": "I am also encountering the same issue, and using the same CV strategy. I made many changes at once to my solution and now I'm not sure of what's causing so big of a difference between my CV (~80) and LB (~70).\n\nBut I'm positive that there is information leakage in the concept or execution in some features I created, although train and validation sets were split before feature engineering.",
      "votes": null
    },
    {
      "id": "1124737",
      "postDate": "12/24/2020 07:08:41",
      "content": "<p>Hello all , <br>\nI was experimenting user with part and target aggregations to see the overall results.<br>\nThe code successfully runs but after i submit i get this error<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Ffa374fc612b9c6902a8c66edee4d462d%2FScreenshot%20(2).png?generation=1608793686884347&amp;alt=media\" alt=\"\"></p>\n<p>I don't know how to debug this?<br>\nI have used all the submission trying to make this work for two days now.<br>\nBut i have no clue on how to do it .<br>\nI seem to be handling new user 275030867 with part 5 correctly as you will notice in the output.<br>\nThe first element of user_part_count, user_part_sum, user_part_mean seems to be correct according to me.</p>\n<p>Could you guys please look into it ? <br>\nHere is the notebook i was referring to  <a href=\"https://www.kaggle.com/ptrikp/part-taregt-aggregations\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-taregt-aggregations</a></p>\n<p>Thanks</p>",
      "rawMarkdown": "Hello all , \nI was experimenting user with part and target aggregations to see the overall results.\nThe code successfully runs but after i submit i get this error\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Ffa374fc612b9c6902a8c66edee4d462d%2FScreenshot%20(2).png?generation=1608793686884347&alt=media)\n\nI don't know how to debug this?\nI have used all the submission trying to make this work for two days now.\nBut i have no clue on how to do it .\nI seem to be handling new user 275030867 with part 5 correctly as you will notice in the output.\nThe first element of user_part_count, user_part_sum, user_part_mean seems to be correct according to me.\n\n\nCould you guys please look into it ? \nHere is the notebook i was referring to  https://www.kaggle.com/ptrikp/part-taregt-aggregations\n\nThanks",
      "votes": null
    },
    {
      "id": "1124958",
      "postDate": "12/24/2020 09:45:28",
      "content": "<p>And after modifying the code, and after running for 1 hour i get this error now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fc8e89ee1a86ff261068d6f29cd6c1a97%2FScreenshot%20(3).png?generation=1608803120127593&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "And after modifying the code, and after running for 1 hour i get this error now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fc8e89ee1a86ff261068d6f29cd6c1a97%2FScreenshot%20(3).png?generation=1608803120127593&alt=media)",
      "votes": null
    },
    {
      "id": "1125135",
      "postDate": "12/24/2020 12:05:22",
      "content": "<p>To me, it seems you are missing lecture content in your <code>prior_test_df</code><br>\nI would put <code>prior_test_df = test_df.copy()</code> right after where the first if statement ends. </p>\n<p>example_test does not have any content_type with lecture. However this statement <code>prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])</code> will include lecture content also.  </p>",
      "rawMarkdown": "To me, it seems you are missing lecture content in your `prior_test_df`\nI would put `prior_test_df = test_df.copy()` right after where the first if statement ends. \n\nexample_test does not have any content_type with lecture. However this statement `prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])` will include lecture content also.",
      "votes": null
    },
    {
      "id": "1125139",
      "postDate": "12/24/2020 12:10:44",
      "content": "<p>If i do that how will i get part which comes after merging questions_df ?<br>\nThis code gives me the part<br>\ntest_df =  pd.merge(test_df, questions_df, left_on='content_id',right_on='question_id', how='left')</p>\n<p>Did you mean after this i do a copy --&gt; prior_test_df = test_df.copy()</p>\n<p>What does it matter if i do prior_test_df = test_df.copy() at the very end ?</p>\n<p>You said </p>\n<blockquote>\n  <p>To me, it seems you are missing lecture content in your prior_test_df</p>\n</blockquote>\n<p>I can't understand this ? what is this ?</p>",
      "rawMarkdown": "If i do that how will i get part which comes after merging questions_df ?\nThis code gives me the part\ntest_df =  pd.merge(test_df, questions_df, left_on='content_id',right_on='question_id', how='left')\n\nDid you mean after this i do a copy --> prior_test_df = test_df.copy()\n\nWhat does it matter if i do prior_test_df = test_df.copy() at the very end ?\n\nYou said \n> To me, it seems you are missing lecture content in your prior_test_df\n\nI can't understand this ? what is this ?",
      "votes": null
    },
    {
      "id": "1125172",
      "postDate": "12/24/2020 12:52:07",
      "content": "<p>For predictions we only care about <code>content_type_id==0</code> however this statement <code>prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])</code> will include information about  <code>content_type_id==1</code> also. </p>\n<p><code>content_type_id==1</code> =&gt; Lecture </p>\n<p>So <code>prior_test_df</code> should include both <code>content_type_ids</code>. If it helps my code structure looks like this - </p>\n<pre><code>previous_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n  if previous_test_df is not None:\n      previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n      # update here\n\n  previous_test_df = test_df.copy()\n  test_df = pd.merge(test_df, questions_df, left_on = 'content_id', right_on = 'question_id', how = 'left')\n  test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True)\n....\n  test_df[TARGET] =  model.predict(test_df[FEATURES])\n</code></pre>",
      "rawMarkdown": "For predictions we only care about `content_type_id==0` however this statement `prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])` will include information about  `content_type_id==1` also. \n\n`content_type_id==1` => Lecture \n\nSo `prior_test_df` should include both `content_type_ids`. If it helps my code structure looks like this - \n\n```\nprevious_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n  if previous_test_df is not None:\n      previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n      # update here\n  \n  previous_test_df = test_df.copy()\n  test_df = pd.merge(test_df, questions_df, left_on = 'content_id', right_on = 'question_id', how = 'left')\n  test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True)\n....\n  test_df[TARGET] =  model.predict(test_df[FEATURES])\n\n```",
      "votes": null
    },
    {
      "id": "1125175",
      "postDate": "12/24/2020 13:05:10",
      "content": "<p>If i do this i will not get the part that i am working with so this update will not run<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fcc18a5d5485e86a45d673ab6d6d74c62%2FScreenshot%20(4).png?generation=1608814903576037&amp;alt=media\" alt=\"\"></p>\n<p>So i copied after merging question_df. Here is the notebook but the same error i also modified the code for getting user part data<br>\n<a href=\"https://www.kaggle.com/ptrikp/part-and-target-agg\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-and-target-agg</a></p>\n<p>I have already exhausted the limit .so i will try tomorrow ant let you know .</p>",
      "rawMarkdown": "If i do this i will not get the part that i am working with so this update will not run\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fcc18a5d5485e86a45d673ab6d6d74c62%2FScreenshot%20(4).png?generation=1608814903576037&alt=media)\n\nSo i copied after merging question_df. Here is the notebook but the same error i also modified the code for getting user part data\nhttps://www.kaggle.com/ptrikp/part-and-target-agg\n\nI have already exhausted the limit .so i will try tomorrow ant let you know .",
      "votes": null
    },
    {
      "id": "1125212",
      "postDate": "12/24/2020 13:51:32",
      "content": "<p>From my experience, if you use something like a cumulative avg score than don't forget to shift the value</p>",
      "rawMarkdown": "From my experience, if you use something like a cumulative avg score than don't forget to shift the value",
      "votes": null
    },
    {
      "id": "1125622",
      "postDate": "12/24/2020 21:50:55",
      "content": "<p>I looked at your notebook and it seems to create the sample submission correctly, probably because there are no lectures rows in this small test set (~100 rows). So I would say you are probably getting this error due to some issue in the pipeline that involves lectures df.</p>\n<p>I noticed that when using the API you are saying:</p>\n<p><code>env.predict(test_df[['row_id', target]])</code></p>\n<p>Maybe in this cell you should do your preprocessing on the test_df with the lectures rows, and then when submitting you'd say:</p>\n<p><code>env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'target']])</code></p>\n<p>It is hard to debug errors in this API part. Sometimes I just print a few things in between the preprocessing steps.</p>",
      "rawMarkdown": "I looked at your notebook and it seems to create the sample submission correctly, probably because there are no lectures rows in this small test set (~100 rows). So I would say you are probably getting this error due to some issue in the pipeline that involves lectures df.\n\nI noticed that when using the API you are saying:\n\n`env.predict(test_df[['row_id', target]])`\n\nMaybe in this cell you should do your preprocessing on the test_df with the lectures rows, and then when submitting you'd say:\n\n`env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'target']])`\n\n\nIt is hard to debug errors in this API part. Sometimes I just print a few things in between the preprocessing steps.",
      "votes": null
    },
    {
      "id": "1125766",
      "postDate": "12/25/2020 03:29:33",
      "content": "<p>I changed the ordering of the copy statement and added some redundant code like this </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fb02c953a2bf3aa8c8ecb6443c75ffbfe%2FScreenshot%20(5).png?generation=1608866731069236&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F11e250085a09632eef9a0f67d9e578db%2FScreenshot%20(6).png?generation=1608866742513185&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fa6a8f66fac85865981d5a9625c5121af%2FScreenshot%20(7).png?generation=1608866775123337&amp;alt=media\" alt=\"\"></p>\n<p>Its been half hour and its running may be it will complete successfully.<br>\nI hope so <br>\nI will let you know if something happens in between as in some notebook after 1 hour i got Notebook Exceeded Allowed Compute</p>\n<p>Thanks all for the suggestions</p>",
      "rawMarkdown": "I changed the ordering of the copy statement and added some redundant code like this \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fb02c953a2bf3aa8c8ecb6443c75ffbfe%2FScreenshot%20(5).png?generation=1608866731069236&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F11e250085a09632eef9a0f67d9e578db%2FScreenshot%20(6).png?generation=1608866742513185&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fa6a8f66fac85865981d5a9625c5121af%2FScreenshot%20(7).png?generation=1608866775123337&alt=media)\n\nIts been half hour and its running may be it will complete successfully.\nI hope so \nI will let you know if something happens in between as in some notebook after 1 hour i got Notebook Exceeded Allowed Compute\n\nThanks all for the suggestions",
      "votes": null
    },
    {
      "id": "1125799",
      "postDate": "12/25/2020 04:45:00",
      "content": "<p>And now the error is Notebook Exceeded Allowed Compute after running for 2 hours<br>\nIts getting frustrating now</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fecaa69d7788e0fadf06a0212b7b10e56%2FScreenshot%20(8).png?generation=1608871441716385&amp;alt=media\" alt=\"\"></p>\n<p>By the way the notebook is this <br>\n<a href=\"https://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937</a><br>\nVersion 2</p>",
      "rawMarkdown": "And now the error is Notebook Exceeded Allowed Compute after running for 2 hours\nIts getting frustrating now\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fecaa69d7788e0fadf06a0212b7b10e56%2FScreenshot%20(8).png?generation=1608871441716385&alt=media)\n\nBy the way the notebook is this \nhttps://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937\nVersion 2",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1105975,
      "author_name": "rpygamer",
      "author_url": "",
      "post_date": "12/08/2020 11:34:32",
      "content": "<p>the same here used all the train from train_cv_1 and validate on valid_cv_1 : <br>\ncv score : 0.79<br>\nLb : 0.711<br>\ni think overfitting … no ? any other similar cases ?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1106068,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/08/2020 13:32:34",
          "content": "<p>Could be bug. Got such issue by wrong computation of some features during inference. Same when I uploaded the wrong dump (pandas) containing some stats (i.e. average correct answer per question), trained with 90M but uploaded only with 10M.<br>\nCould be data leak too.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108431,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/10/2020 16:31:20",
          "content": "<p>Aggregations with target is costly …. i am getting .82 something with 70 million dataset with xgb gpu ..<br>\nbut in leaderboard its .74 </p>\n<p>Has to be data leak with aggregating targets….. have  to find some other aggragations… may be a shake up in leader board</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1124527,
          "author_name": "hinepo",
          "author_url": "",
          "post_date": "12/24/2020 02:52:38",
          "content": "<p>I am also encountering the same issue, and using the same CV strategy. I made many changes at once to my solution and now I'm not sure of what's causing so big of a difference between my CV (~80) and LB (~70).</p>\n<p>But I'm positive that there is information leakage in the concept or execution in some features I created, although train and validation sets were split before feature engineering.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125212,
          "author_name": "marcellosusanto",
          "author_url": "",
          "post_date": "12/24/2020 13:51:32",
          "content": "<p>From my experience, if you use something like a cumulative avg score than don't forget to shift the value</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124737,
      "author_name": "ptrikp",
      "author_url": "",
      "post_date": "12/24/2020 07:08:41",
      "content": "<p>Hello all , <br>\nI was experimenting user with part and target aggregations to see the overall results.<br>\nThe code successfully runs but after i submit i get this error<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Ffa374fc612b9c6902a8c66edee4d462d%2FScreenshot%20(2).png?generation=1608793686884347&amp;alt=media\" alt=\"\"></p>\n<p>I don't know how to debug this?<br>\nI have used all the submission trying to make this work for two days now.<br>\nBut i have no clue on how to do it .<br>\nI seem to be handling new user 275030867 with part 5 correctly as you will notice in the output.<br>\nThe first element of user_part_count, user_part_sum, user_part_mean seems to be correct according to me.</p>\n<p>Could you guys please look into it ? <br>\nHere is the notebook i was referring to  <a href=\"https://www.kaggle.com/ptrikp/part-taregt-aggregations\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-taregt-aggregations</a></p>\n<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 1124958,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/24/2020 09:45:28",
          "content": "<p>And after modifying the code, and after running for 1 hour i get this error now.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fc8e89ee1a86ff261068d6f29cd6c1a97%2FScreenshot%20(3).png?generation=1608803120127593&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125135,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "12/24/2020 12:05:22",
          "content": "<p>To me, it seems you are missing lecture content in your <code>prior_test_df</code><br>\nI would put <code>prior_test_df = test_df.copy()</code> right after where the first if statement ends. </p>\n<p>example_test does not have any content_type with lecture. However this statement <code>prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])</code> will include lecture content also.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125139,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/24/2020 12:10:44",
          "content": "<p>If i do that how will i get part which comes after merging questions_df ?<br>\nThis code gives me the part<br>\ntest_df =  pd.merge(test_df, questions_df, left_on='content_id',right_on='question_id', how='left')</p>\n<p>Did you mean after this i do a copy --&gt; prior_test_df = test_df.copy()</p>\n<p>What does it matter if i do prior_test_df = test_df.copy() at the very end ?</p>\n<p>You said </p>\n<blockquote>\n  <p>To me, it seems you are missing lecture content in your prior_test_df</p>\n</blockquote>\n<p>I can't understand this ? what is this ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125172,
          "author_name": "rashmibanthia",
          "author_url": "",
          "post_date": "12/24/2020 12:52:07",
          "content": "<p>For predictions we only care about <code>content_type_id==0</code> however this statement <code>prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])</code> will include information about  <code>content_type_id==1</code> also. </p>\n<p><code>content_type_id==1</code> =&gt; Lecture </p>\n<p>So <code>prior_test_df</code> should include both <code>content_type_ids</code>. If it helps my code structure looks like this - </p>\n<pre><code>previous_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n  if previous_test_df is not None:\n      previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n      # update here\n\n  previous_test_df = test_df.copy()\n  test_df = pd.merge(test_df, questions_df, left_on = 'content_id', right_on = 'question_id', how = 'left')\n  test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True)\n....\n  test_df[TARGET] =  model.predict(test_df[FEATURES])\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125175,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/24/2020 13:05:10",
          "content": "<p>If i do this i will not get the part that i am working with so this update will not run<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fcc18a5d5485e86a45d673ab6d6d74c62%2FScreenshot%20(4).png?generation=1608814903576037&amp;alt=media\" alt=\"\"></p>\n<p>So i copied after merging question_df. Here is the notebook but the same error i also modified the code for getting user part data<br>\n<a href=\"https://www.kaggle.com/ptrikp/part-and-target-agg\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-and-target-agg</a></p>\n<p>I have already exhausted the limit .so i will try tomorrow ant let you know .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125622,
          "author_name": "hinepo",
          "author_url": "",
          "post_date": "12/24/2020 21:50:55",
          "content": "<p>I looked at your notebook and it seems to create the sample submission correctly, probably because there are no lectures rows in this small test set (~100 rows). So I would say you are probably getting this error due to some issue in the pipeline that involves lectures df.</p>\n<p>I noticed that when using the API you are saying:</p>\n<p><code>env.predict(test_df[['row_id', target]])</code></p>\n<p>Maybe in this cell you should do your preprocessing on the test_df with the lectures rows, and then when submitting you'd say:</p>\n<p><code>env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'target']])</code></p>\n<p>It is hard to debug errors in this API part. Sometimes I just print a few things in between the preprocessing steps.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125766,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/25/2020 03:29:33",
          "content": "<p>I changed the ordering of the copy statement and added some redundant code like this </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fb02c953a2bf3aa8c8ecb6443c75ffbfe%2FScreenshot%20(5).png?generation=1608866731069236&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F11e250085a09632eef9a0f67d9e578db%2FScreenshot%20(6).png?generation=1608866742513185&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fa6a8f66fac85865981d5a9625c5121af%2FScreenshot%20(7).png?generation=1608866775123337&amp;alt=media\" alt=\"\"></p>\n<p>Its been half hour and its running may be it will complete successfully.<br>\nI hope so <br>\nI will let you know if something happens in between as in some notebook after 1 hour i got Notebook Exceeded Allowed Compute</p>\n<p>Thanks all for the suggestions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1125799,
          "author_name": "ptrikp",
          "author_url": "",
          "post_date": "12/25/2020 04:45:00",
          "content": "<p>And now the error is Notebook Exceeded Allowed Compute after running for 2 hours<br>\nIts getting frustrating now</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fecaa69d7788e0fadf06a0212b7b10e56%2FScreenshot%20(8).png?generation=1608871441716385&amp;alt=media\" alt=\"\"></p>\n<p>By the way the notebook is this <br>\n<a href=\"https://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937\" target=\"_blank\">https://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937</a><br>\nVersion 2</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1105867": "cv score:0.77\nlb score:0.71\n😂😂😂\nI used cv strategy from https://www.kaggle.com/its7171/cv-strategy\n\nFeature engineering and training data：\n`train = features_(train)\nvalidation = features_(validation)\n.................\ny_train = train['answered_correctly']\ntrain = train[features]\n\ny_val = validation['answered_correctly']\nvalidation = validation[features]`\n\nPrediction:\n`for test_df, _ in iter_test:\n    test_df = features_(test_df)\n    test_df['answered_correctly'] = model.predict(test_df[features])\n    env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])`\n\nWhere is the problem?\nOverfitting or some bug?\n\nI am really appreciate if someone can help me.😆",
    "1105975": "the same here used all the train from train_cv_1 and validate on valid_cv_1 : \ncv score : 0.79\nLb : 0.711\ni think overfitting ... no ? any other similar cases ?",
    "1106068": "Could be bug. Got such issue by wrong computation of some features during inference. Same when I uploaded the wrong dump (pandas) containing some stats (i.e. average correct answer per question), trained with 90M but uploaded only with 10M.\nCould be data leak too.",
    "1108431": "Aggregations with target is costly .... i am getting .82 something with 70 million dataset with xgb gpu ..\nbut in leaderboard its .74 \n\nHas to be data leak with aggregating targets..... have  to find some other aggragations... may be a shake up in leader board",
    "1124527": "I am also encountering the same issue, and using the same CV strategy. I made many changes at once to my solution and now I'm not sure of what's causing so big of a difference between my CV (~80) and LB (~70).\n\nBut I'm positive that there is information leakage in the concept or execution in some features I created, although train and validation sets were split before feature engineering.",
    "1124737": "Hello all , \nI was experimenting user with part and target aggregations to see the overall results.\nThe code successfully runs but after i submit i get this error\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Ffa374fc612b9c6902a8c66edee4d462d%2FScreenshot%20(2).png?generation=1608793686884347&alt=media)\n\nI don't know how to debug this?\nI have used all the submission trying to make this work for two days now.\nBut i have no clue on how to do it .\nI seem to be handling new user 275030867 with part 5 correctly as you will notice in the output.\nThe first element of user_part_count, user_part_sum, user_part_mean seems to be correct according to me.\n\n\nCould you guys please look into it ? \nHere is the notebook i was referring to  https://www.kaggle.com/ptrikp/part-taregt-aggregations\n\nThanks",
    "1124958": "And after modifying the code, and after running for 1 hour i get this error now.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fc8e89ee1a86ff261068d6f29cd6c1a97%2FScreenshot%20(3).png?generation=1608803120127593&alt=media)",
    "1125135": "To me, it seems you are missing lecture content in your `prior_test_df`\nI would put `prior_test_df = test_df.copy()` right after where the first if statement ends. \n\nexample_test does not have any content_type with lecture. However this statement `prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])` will include lecture content also.",
    "1125139": "If i do that how will i get part which comes after merging questions_df ?\nThis code gives me the part\ntest_df =  pd.merge(test_df, questions_df, left_on='content_id',right_on='question_id', how='left')\n\nDid you mean after this i do a copy --> prior_test_df = test_df.copy()\n\nWhat does it matter if i do prior_test_df = test_df.copy() at the very end ?\n\nYou said \n> To me, it seems you are missing lecture content in your prior_test_df\n\nI can't understand this ? what is this ?",
    "1125172": "For predictions we only care about `content_type_id==0` however this statement `prior_test_df[target] = eval(test_df['prior_group_answers_correct'].iloc[0])` will include information about  `content_type_id==1` also. \n\n`content_type_id==1` => Lecture \n\nSo `prior_test_df` should include both `content_type_ids`. If it helps my code structure looks like this - \n\n```\nprevious_test_df = None\nfor (test_df, sample_prediction_df) in iter_test:\n  if previous_test_df is not None:\n      previous_test_df[TARGET] = eval(test_df[\"prior_group_answers_correct\"].iloc[0])\n      # update here\n  \n  previous_test_df = test_df.copy()\n  test_df = pd.merge(test_df, questions_df, left_on = 'content_id', right_on = 'question_id', how = 'left')\n  test_df = test_df[test_df['content_type_id'] == 0].reset_index(drop=True)\n....\n  test_df[TARGET] =  model.predict(test_df[FEATURES])\n\n```",
    "1125175": "If i do this i will not get the part that i am working with so this update will not run\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fcc18a5d5485e86a45d673ab6d6d74c62%2FScreenshot%20(4).png?generation=1608814903576037&alt=media)\n\nSo i copied after merging question_df. Here is the notebook but the same error i also modified the code for getting user part data\nhttps://www.kaggle.com/ptrikp/part-and-target-agg\n\nI have already exhausted the limit .so i will try tomorrow ant let you know .",
    "1125212": "From my experience, if you use something like a cumulative avg score than don't forget to shift the value",
    "1125622": "I looked at your notebook and it seems to create the sample submission correctly, probably because there are no lectures rows in this small test set (~100 rows). So I would say you are probably getting this error due to some issue in the pipeline that involves lectures df.\n\nI noticed that when using the API you are saying:\n\n`env.predict(test_df[['row_id', target]])`\n\nMaybe in this cell you should do your preprocessing on the test_df with the lectures rows, and then when submitting you'd say:\n\n`env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'target']])`\n\n\nIt is hard to debug errors in this API part. Sometimes I just print a few things in between the preprocessing steps.",
    "1125766": "I changed the ordering of the copy statement and added some redundant code like this \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fb02c953a2bf3aa8c8ecb6443c75ffbfe%2FScreenshot%20(5).png?generation=1608866731069236&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2F11e250085a09632eef9a0f67d9e578db%2FScreenshot%20(6).png?generation=1608866742513185&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fa6a8f66fac85865981d5a9625c5121af%2FScreenshot%20(7).png?generation=1608866775123337&alt=media)\n\nIts been half hour and its running may be it will complete successfully.\nI hope so \nI will let you know if something happens in between as in some notebook after 1 hour i got Notebook Exceeded Allowed Compute\n\nThanks all for the suggestions",
    "1125799": "And now the error is Notebook Exceeded Allowed Compute after running for 2 hours\nIts getting frustrating now\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4947407%2Fecaa69d7788e0fadf06a0212b7b10e56%2FScreenshot%20(8).png?generation=1608871441716385&alt=media)\n\nBy the way the notebook is this \nhttps://www.kaggle.com/ptrikp/part-taregt-aggregations?scriptVersionId=50201937\nVersion 2"
  },
  "source": "meta"
}