{
  "id": 412220,
  "title": "How to avoid leakage if using previous predictions as features for multi-fold model",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/412220",
  "author_name": "Ya Xu",
  "post_date": "2023-05-22T18:50:19.028000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>It's the first time I tried this strategy and found it problematic to use.</p>\n<p>Whenever I tried to apply this strategy to my model, CV got boosted greatly but LB dropped sharply. I figure out it could be the way how I implement it. During the training stage, I filled previous predictions with the result of each validation set from groupKfold. So in my case, I fill 1/5 prediction of the full dataset for each 5-fold and got a full filling of predictions after doing it 5 times. But during the submission stage, the predictions were calculated by the average of 5 folds, thus causing a mismatch of behavior between models. (or maybe I was wrong, it doesn't cause a mismatch?)</p>\n<p>If I want to keep my model synchronized, I should change the behavior from either side. I can't find a way to change how the submission part work. If I'm going to synchronize my training part, I have to predict the full dataset for each fold and averaged it at the end. But it also means I have to predict the training split part with the model just trained from it, this seems like a leakage for me. (or it actually isn't?) </p>\n<p>Using a single-fold model seems like a solution to this problem. But is it still possible to use this strategy for a multi-fold model?</p>\n<h2>update:</h2>\n<p>I did some quick research and experiment. Single-fold model is <strong>not</strong> a solution since it will certainly cause leakage.</p>\n<p>It seems people usually use another different model to make previous prediction as features, this can avoid leakage but require additional resources for another model.</p>",
  "messages": [
    {
      "id": 2269864,
      "postDate": "2023-05-22T18:50:19.030Z",
      "content": "<p>It's the first time I tried this strategy and found it problematic to use.</p>\n<p>Whenever I tried to apply this strategy to my model, CV got boosted greatly but LB dropped sharply. I figure out it could be the way how I implement it. During the training stage, I filled previous predictions with the result of each validation set from groupKfold. So in my case, I fill 1/5 prediction of the full dataset for each 5-fold and got a full filling of predictions after doing it 5 times. But during the submission stage, the predictions were calculated by the average of 5 folds, thus causing a mismatch of behavior between models. (or maybe I was wrong, it doesn't cause a mismatch?)</p>\n<p>If I want to keep my model synchronized, I should change the behavior from either side. I can't find a way to change how the submission part work. If I'm going to synchronize my training part, I have to predict the full dataset for each fold and averaged it at the end. But it also means I have to predict the training split part with the model just trained from it, this seems like a leakage for me. (or it actually isn't?) </p>\n<p>Using a single-fold model seems like a solution to this problem. But is it still possible to use this strategy for a multi-fold model?</p>\n<h2>update:</h2>\n<p>I did some quick research and experiment. Single-fold model is <strong>not</strong> a solution since it will certainly cause leakage.</p>\n<p>It seems people usually use another different model to make previous prediction as features, this can avoid leakage but require additional resources for another model.</p>",
      "rawMarkdown": "It's the first time I tried this strategy and found it problematic to use.\n\nWhenever I tried to apply this strategy to my model, CV got boosted greatly but LB dropped sharply. I figure out it could be the way how I implement it. During the training stage, I filled previous predictions with the result of each validation set from groupKfold. So in my case, I fill 1/5 prediction of the full dataset for each 5-fold and got a full filling of predictions after doing it 5 times. But during the submission stage, the predictions were calculated by the average of 5 folds, thus causing a mismatch of behavior between models. (or maybe I was wrong, it doesn't cause a mismatch?)\n\nIf I want to keep my model synchronized, I should change the behavior from either side. I can't find a way to change how the submission part work. If I'm going to synchronize my training part, I have to predict the full dataset for each fold and averaged it at the end. But it also means I have to predict the training split part with the model just trained from it, this seems like a leakage for me. (or it actually isn't?) \n\nUsing a single-fold model seems like a solution to this problem. But is it still possible to use this strategy for a multi-fold model?\n\n\n\n## update:\n\nI did some quick research and experiment. Single-fold model is **not** a solution since it will certainly cause leakage.\n \nIt seems people usually use another different model to make previous prediction as features, this can avoid leakage but require additional resources for another model.\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2269864": "It's the first time I tried this strategy and found it problematic to use.\n\nWhenever I tried to apply this strategy to my model, CV got boosted greatly but LB dropped sharply. I figure out it could be the way how I implement it. During the training stage, I filled previous predictions with the result of each validation set from groupKfold. So in my case, I fill 1/5 prediction of the full dataset for each 5-fold and got a full filling of predictions after doing it 5 times. But during the submission stage, the predictions were calculated by the average of 5 folds, thus causing a mismatch of behavior between models. (or maybe I was wrong, it doesn't cause a mismatch?)\n\nIf I want to keep my model synchronized, I should change the behavior from either side. I can't find a way to change how the submission part work. If I'm going to synchronize my training part, I have to predict the full dataset for each fold and averaged it at the end. But it also means I have to predict the training split part with the model just trained from it, this seems like a leakage for me. (or it actually isn't?) \n\nUsing a single-fold model seems like a solution to this problem. But is it still possible to use this strategy for a multi-fold model?\n\n\n\n## update:\n\nI did some quick research and experiment. Single-fold model is **not** a solution since it will certainly cause leakage.\n \nIt seems people usually use another different model to make previous prediction as features, this can avoid leakage but require additional resources for another model.\n"
  }
}