{
  "id": 57413,
  "title": "Score difference between local score and final score",
  "url": "/competitions/avito-demand-prediction/discussion/57413",
  "author_name": "",
  "post_date": "2018-05-23T13:11:08.158908600Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I split train data into two parts: one for local train and one for local test. My local score is 0.2290 in my computer.</p>\n\n<p>After that, I downloaded the test data for prediction and upload the answer to Kaggle. I just got 0.2676.  The difference is quite big.</p>\n\n<p>What is reason of this phenomenon?</p>",
  "messages": [
    {
      "id": "332609",
      "postDate": "05/23/2018 13:11:08",
      "content": "<p>I split train data into two parts: one for local train and one for local test. My local score is 0.2290 in my computer.</p>\n\n<p>After that, I downloaded the test data for prediction and upload the answer to Kaggle. I just got 0.2676.  The difference is quite big.</p>\n\n<p>What is reason of this phenomenon?</p>",
      "rawMarkdown": "I split train data into two parts: one for local train and one for local test. My local score is 0.2290 in my computer.\n\nAfter that, I downloaded the test data for prediction and upload the answer to Kaggle. I just got 0.2676.  The difference is quite big.\n\nWhat is reason of this phenomenon?",
      "votes": null
    },
    {
      "id": "332645",
      "postDate": "05/23/2018 14:00:37",
      "content": "<p>Try 5 Kfold split  ! </p>",
      "rawMarkdown": "Try 5 Kfold split  !",
      "votes": null
    },
    {
      "id": "332695",
      "postDate": "05/23/2018 15:29:10",
      "content": "<p>I tried 5 Kfold. The average score is 0.2298.</p>\n\n<p>I think I get stuck in other problem.</p>",
      "rawMarkdown": "I tried 5 Kfold. The average score is 0.2298.\n\nI think I get stuck in other problem.",
      "votes": null
    },
    {
      "id": "332702",
      "postDate": "05/23/2018 15:35:31",
      "content": "<p>Maybe you evaluate with fit instead of val ? (even though ideally you should evaluate on both to see if you overfit)</p>",
      "rawMarkdown": "Maybe you evaluate with fit instead of val ? (even though ideally you should evaluate on both to see if you overfit)",
      "votes": null
    },
    {
      "id": "332793",
      "postDate": "05/23/2018 18:53:29",
      "content": "<p>@ Kwok Kang Chuen </p>\n\n<p>The logical answer is that you have Overfit your model to the local validation data. ( Since local CV is 0.2290 )\nHowever, the difference seems to be very big for an overfit. You could have made an error ERROR in calculating TEST VECTOR  - Like switching price and item_seq_number columns in the test vector.  </p>\n\n<p>Also, listing out other typical overfitting pitfalls. </p>\n\n<p>Typical Overfitting pit falls </p>\n\n<ul>\n<li>You could have fitted your word vectors on training and test data. Try tokenizing only for training data. </li>\n<li>A suboptimal learning rate value. You can change this to see local CV increase but your </li>\n<li>It could also be model specific \nFor ex:  LGBM - Using max_depth -  15 os opposed to 5/7  and in a neural network too many dense layers with little or \nno Drop out value </li>\n</ul>\n\n<p>Hope this helps. </p>",
      "rawMarkdown": "Kwok Kang Chuen \n\nThe logical answer is that you have Overfit your model to the local validation data. ( Since local CV is 0.2290 )\nHowever, the difference seems to be very big for an overfit. You could have made an error ERROR in calculating TEST VECTOR  - Like switching price and item_seq_number columns in the test vector.  \n\nAlso, listing out other typical overfitting pitfalls. \n\nTypical Overfitting pit falls \n\n- You could have fitted your word vectors on training and test data. Try tokenizing only for training data. \n- A suboptimal learning rate value. You can change this to see local CV increase but your \n- It could also be model specific \n  For ex:  LGBM - Using max_depth -  15 os opposed to 5/7  and in a neural network too many dense layers with little or \n no Drop out value \n\nHope this helps.",
      "votes": null
    },
    {
      "id": "332856",
      "postDate": "05/23/2018 22:57:58",
      "content": "<p>I don't believe there is so big difference in data distributions so I guess that might happen due to data leak. Check twice for exampe target encoding or the whole feature processing pipeline</p>",
      "rawMarkdown": "I don't believe there is so big difference in data distributions so I guess that might happen due to data leak. Check twice for exampe target encoding or the whole feature processing pipeline",
      "votes": null
    },
    {
      "id": "332868",
      "postDate": "05/24/2018 00:11:57",
      "content": "<p>You overfit the training set. Likely because you have some sort of data leakage.</p>",
      "rawMarkdown": "You overfit the training set. Likely because you have some sort of data leakage.",
      "votes": null
    },
    {
      "id": "332909",
      "postDate": "05/24/2018 03:02:01",
      "content": "<p>0.2676 is a pretty far distance away from simple baseline models, there's almost certainly something you're failing to do with respect to the training data. Are you sure that your pre processing methods are the same on the test data? </p>\n\n<p>For instance, if you categorically embed, you want to re-use the same embedding from training. If you normalize a data column, you also want to use the coefficients from the train set, not test set. </p>",
      "rawMarkdown": "0.2676 is a pretty far distance away from simple baseline models, there's almost certainly something you're failing to do with respect to the training data. Are you sure that your pre processing methods are the same on the test data? \n\nFor instance, if you categorically embed, you want to re-use the same embedding from training. If you normalize a data column, you also want to use the coefficients from the train set, not test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 332645,
      "author_name": "adilztn",
      "author_url": "",
      "post_date": "05/23/2018 14:00:37",
      "content": "<p>Try 5 Kfold split  ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 332695,
          "author_name": "ivankwok",
          "author_url": "",
          "post_date": "05/23/2018 15:29:10",
          "content": "<p>I tried 5 Kfold. The average score is 0.2298.</p>\n\n<p>I think I get stuck in other problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 332702,
          "author_name": "adilztn",
          "author_url": "",
          "post_date": "05/23/2018 15:35:31",
          "content": "<p>Maybe you evaluate with fit instead of val ? (even though ideally you should evaluate on both to see if you overfit)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 332793,
      "author_name": "shanth84",
      "author_url": "",
      "post_date": "05/23/2018 18:53:29",
      "content": "<p>@ Kwok Kang Chuen </p>\n\n<p>The logical answer is that you have Overfit your model to the local validation data. ( Since local CV is 0.2290 )\nHowever, the difference seems to be very big for an overfit. You could have made an error ERROR in calculating TEST VECTOR  - Like switching price and item_seq_number columns in the test vector.  </p>\n\n<p>Also, listing out other typical overfitting pitfalls. </p>\n\n<p>Typical Overfitting pit falls </p>\n\n<ul>\n<li>You could have fitted your word vectors on training and test data. Try tokenizing only for training data. </li>\n<li>A suboptimal learning rate value. You can change this to see local CV increase but your </li>\n<li>It could also be model specific \nFor ex:  LGBM - Using max_depth -  15 os opposed to 5/7  and in a neural network too many dense layers with little or \nno Drop out value </li>\n</ul>\n\n<p>Hope this helps. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332856,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "05/23/2018 22:57:58",
      "content": "<p>I don't believe there is so big difference in data distributions so I guess that might happen due to data leak. Check twice for exampe target encoding or the whole feature processing pipeline</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332868,
      "author_name": "arroqc",
      "author_url": "",
      "post_date": "05/24/2018 00:11:57",
      "content": "<p>You overfit the training set. Likely because you have some sort of data leakage.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 332909,
      "author_name": "vannak",
      "author_url": "",
      "post_date": "05/24/2018 03:02:01",
      "content": "<p>0.2676 is a pretty far distance away from simple baseline models, there's almost certainly something you're failing to do with respect to the training data. Are you sure that your pre processing methods are the same on the test data? </p>\n\n<p>For instance, if you categorically embed, you want to re-use the same embedding from training. If you normalize a data column, you also want to use the coefficients from the train set, not test set. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "332609": "I split train data into two parts: one for local train and one for local test. My local score is 0.2290 in my computer.\n\nAfter that, I downloaded the test data for prediction and upload the answer to Kaggle. I just got 0.2676.  The difference is quite big.\n\nWhat is reason of this phenomenon?",
    "332645": "Try 5 Kfold split  !",
    "332695": "I tried 5 Kfold. The average score is 0.2298.\n\nI think I get stuck in other problem.",
    "332702": "Maybe you evaluate with fit instead of val ? (even though ideally you should evaluate on both to see if you overfit)",
    "332793": "Kwok Kang Chuen \n\nThe logical answer is that you have Overfit your model to the local validation data. ( Since local CV is 0.2290 )\nHowever, the difference seems to be very big for an overfit. You could have made an error ERROR in calculating TEST VECTOR  - Like switching price and item_seq_number columns in the test vector.  \n\nAlso, listing out other typical overfitting pitfalls. \n\nTypical Overfitting pit falls \n\n- You could have fitted your word vectors on training and test data. Try tokenizing only for training data. \n- A suboptimal learning rate value. You can change this to see local CV increase but your \n- It could also be model specific \n  For ex:  LGBM - Using max_depth -  15 os opposed to 5/7  and in a neural network too many dense layers with little or \n no Drop out value \n\nHope this helps.",
    "332856": "I don't believe there is so big difference in data distributions so I guess that might happen due to data leak. Check twice for exampe target encoding or the whole feature processing pipeline",
    "332868": "You overfit the training set. Likely because you have some sort of data leakage.",
    "332909": "0.2676 is a pretty far distance away from simple baseline models, there's almost certainly something you're failing to do with respect to the training data. Are you sure that your pre processing methods are the same on the test data? \n\nFor instance, if you categorically embed, you want to re-use the same embedding from training. If you normalize a data column, you also want to use the coefficients from the train set, not test set."
  },
  "source": "meta"
}