{
  "id": 396620,
  "title": "How to deal with the brand new train set",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/396620",
  "author_name": "",
  "post_date": "2023-03-22T10:35:43.143864800Z",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Test data for this competition was unintentionally made available for a period of time. <br>\n(see <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202</a> for detail).</p>\n<p>Host released  all of the leaked data, which contains the entire previous test set.<br>\nNow train set contains old train sessions (11779) and old test sessions (11783) which  increases the size of the training data by 100%.<br>\nOld train sessions and old test sessions  are games played between Oct 2020 and Nov 2022.</p>\n<p>So how I could use this new train set ?</p>\n<p>Here are some basic ideas:</p>\n<h2>Is my model overfitting the LB ?</h2>\n<p>Use one of your submitred model (trained with old train set) to predict old test sessions for which now we have labels.</p>\n<p>If score on old test sessions is (much) worse the LB then your model is probably overfitting LB.</p>\n<h2>How my model performs with 100% data augmentation</h2>\n<p>Retrain one of your model on old train sessions + old test sessions and compare validation score between model trained on old train sessions and new train set.<br>\nIf validation improved with new train set your model takes avantage of  data augmentation.   </p>\n<p><a href=\"https://www.kaggle.com/datasets/steubk/psp-sessions-in-new-trainset\" target=\"_blank\">here</a> you can find new train sessions separated between old train and old test.</p>",
  "messages": [
    {
      "id": "2191957",
      "postDate": "03/22/2023 10:35:43",
      "content": "<p>Test data for this competition was unintentionally made available for a period of time. <br>\n(see <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202</a> for detail).</p>\n<p>Host released  all of the leaked data, which contains the entire previous test set.<br>\nNow train set contains old train sessions (11779) and old test sessions (11783) which  increases the size of the training data by 100%.<br>\nOld train sessions and old test sessions  are games played between Oct 2020 and Nov 2022.</p>\n<p>So how I could use this new train set ?</p>\n<p>Here are some basic ideas:</p>\n<h2>Is my model overfitting the LB ?</h2>\n<p>Use one of your submitred model (trained with old train set) to predict old test sessions for which now we have labels.</p>\n<p>If score on old test sessions is (much) worse the LB then your model is probably overfitting LB.</p>\n<h2>How my model performs with 100% data augmentation</h2>\n<p>Retrain one of your model on old train sessions + old test sessions and compare validation score between model trained on old train sessions and new train set.<br>\nIf validation improved with new train set your model takes avantage of  data augmentation.   </p>\n<p><a href=\"https://www.kaggle.com/datasets/steubk/psp-sessions-in-new-trainset\" target=\"_blank\">here</a> you can find new train sessions separated between old train and old test.</p>",
      "rawMarkdown": "Test data for this competition was unintentionally made available for a period of time. \n(see https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202 for detail).\n\nHost released  all of the leaked data, which contains the entire previous test set.\nNow train set contains old train sessions (11779) and old test sessions (11783) which  increases the size of the training data by 100%.\nOld train sessions and old test sessions  are games played between Oct 2020 and Nov 2022.\n\nSo how I could use this new train set ?\n\nHere are some basic ideas:\n\n## Is my model overfitting the LB ?\n\nUse one of your submitred model (trained with old train set) to predict old test sessions for which now we have labels.\n\nIf score on old test sessions is (much) worse the LB then your model is probably overfitting LB.\n\n\n\n## How my model performs with 100% data augmentation\n\nRetrain one of your model on old train sessions + old test sessions and compare validation score between model trained on old train sessions and new train set.\nIf validation improved with new train set your model takes avantage of  data augmentation.   \n\n[here](https://www.kaggle.com/datasets/steubk/psp-sessions-in-new-trainset) you can find new train sessions separated between old train and old test.",
      "votes": null
    },
    {
      "id": "2191961",
      "postDate": "03/22/2023 10:37:01",
      "content": "<p>Hey. The last link doesn't work :(</p>",
      "rawMarkdown": "Hey. The last link doesn't work :(",
      "votes": null
    },
    {
      "id": "2191964",
      "postDate": "03/22/2023 10:41:43",
      "content": "<p>link updated ! thank you</p>",
      "rawMarkdown": "link updated ! thank you",
      "votes": null
    },
    {
      "id": "2191966",
      "postDate": "03/22/2023 10:43:34",
      "content": "<p>Thank you 🙏 , this is very helpful</p>",
      "rawMarkdown": "Thank you 🙏 , this is very helpful",
      "votes": null
    },
    {
      "id": "2192116",
      "postDate": "03/22/2023 12:40:40",
      "content": "<p>Thanks for the info, I wonder if you could share how much your model improved from the data augmentation? Only if you think that you can share this, thanks!</p>",
      "rawMarkdown": "Thanks for the info, I wonder if you could share how much your model improved from the data augmentation? Only if you think that you can share this, thanks!",
      "votes": null
    },
    {
      "id": "2192207",
      "postDate": "03/22/2023 14:06:02",
      "content": "<p>it's 0.003 for my best model</p>",
      "rawMarkdown": "it's 0.003 for my best model",
      "votes": null
    },
    {
      "id": "2192239",
      "postDate": "03/22/2023 14:33:07",
      "content": "<p>Thanks, it gave me a bit more, but probably due the fact that I submit only with the 1st fold model and the double size of the data set.</p>",
      "rawMarkdown": "Thanks, it gave me a bit more, but probably due the fact that I submit only with the 1st fold model and the double size of the data set.",
      "votes": null
    },
    {
      "id": "2192971",
      "postDate": "03/23/2023 02:37:37",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> , 0.003 for CV or LB? In my case, adding more data improved CV by only 0.001, but LB improved by 0.003</p>",
      "rawMarkdown": "Hi @steubk , 0.003 for CV or LB? In my case, adding more data improved CV by only 0.001, but LB improved by 0.003",
      "votes": null
    },
    {
      "id": "2193279",
      "postDate": "03/23/2023 07:31:42",
      "content": "<p>in my case in  +0.03 for both cv and lb</p>",
      "rawMarkdown": "in my case in  +0.03 for both cv and lb",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2191961,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "03/22/2023 10:37:01",
      "content": "<p>Hey. The last link doesn't work :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 2191964,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "03/22/2023 10:41:43",
          "content": "<p>link updated ! thank you</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2191966,
      "author_name": "ihebch",
      "author_url": "",
      "post_date": "03/22/2023 10:43:34",
      "content": "<p>Thank you 🙏 , this is very helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2192116,
      "author_name": "dimitrislev",
      "author_url": "",
      "post_date": "03/22/2023 12:40:40",
      "content": "<p>Thanks for the info, I wonder if you could share how much your model improved from the data augmentation? Only if you think that you can share this, thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2192207,
          "author_name": "steubk",
          "author_url": "",
          "post_date": "03/22/2023 14:06:02",
          "content": "<p>it's 0.003 for my best model</p>",
          "votes": null,
          "replies": [
            {
              "id": 2192239,
              "author_name": "dimitrislev",
              "author_url": "",
              "post_date": "03/22/2023 14:33:07",
              "content": "<p>Thanks, it gave me a bit more, but probably due the fact that I submit only with the 1st fold model and the double size of the data set.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2192971,
              "author_name": "mathormad",
              "author_url": "",
              "post_date": "03/23/2023 02:37:37",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> , 0.003 for CV or LB? In my case, adding more data improved CV by only 0.001, but LB improved by 0.003</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2193279,
                  "author_name": "steubk",
                  "author_url": "",
                  "post_date": "03/23/2023 07:31:42",
                  "content": "<p>in my case in  +0.03 for both cv and lb</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2191957": "Test data for this competition was unintentionally made available for a period of time. \n(see https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202 for detail).\n\nHost released  all of the leaked data, which contains the entire previous test set.\nNow train set contains old train sessions (11779) and old test sessions (11783) which  increases the size of the training data by 100%.\nOld train sessions and old test sessions  are games played between Oct 2020 and Nov 2022.\n\nSo how I could use this new train set ?\n\nHere are some basic ideas:\n\n## Is my model overfitting the LB ?\n\nUse one of your submitred model (trained with old train set) to predict old test sessions for which now we have labels.\n\nIf score on old test sessions is (much) worse the LB then your model is probably overfitting LB.\n\n\n\n## How my model performs with 100% data augmentation\n\nRetrain one of your model on old train sessions + old test sessions and compare validation score between model trained on old train sessions and new train set.\nIf validation improved with new train set your model takes avantage of  data augmentation.   \n\n[here](https://www.kaggle.com/datasets/steubk/psp-sessions-in-new-trainset) you can find new train sessions separated between old train and old test.",
    "2191961": "Hey. The last link doesn't work :(",
    "2191964": "link updated ! thank you",
    "2191966": "Thank you 🙏 , this is very helpful",
    "2192116": "Thanks for the info, I wonder if you could share how much your model improved from the data augmentation? Only if you think that you can share this, thanks!",
    "2192207": "it's 0.003 for my best model",
    "2192239": "Thanks, it gave me a bit more, but probably due the fact that I submit only with the 1st fold model and the double size of the data set.",
    "2192971": "Hi @steubk , 0.003 for CV or LB? In my case, adding more data improved CV by only 0.001, but LB improved by 0.003",
    "2193279": "in my case in  +0.03 for both cv and lb"
  },
  "source": "meta"
}