{
  "id": 420315,
  "title": "8th Position: (External Dataset) Good Preprocessing = +0.002 for Single Model",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420315",
  "author_name": "",
  "post_date": "2023-06-30T08:32:55.761898100Z",
  "votes": 16,
  "comment_count": 3,
  "views": 0,
  "content": "<p>…and we sadly missed it!</p>\n<p>In this discussion, I’m sharing about my data leak preprocessing in the competition. The raw data can be downloaded from <a href=\"https://opengamedata.fielddaylab.wisc.edu/gamedata.php?game=JOWILDER\" target=\"_blank\">Field Day Lab's Open Game Data website</a>. Our <strong>full solution</strong> has also been detailed out <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420528\" target=\"_blank\">here</a>.</p>\n<p>Before beginning, I would like to thank the host and Kaggle for this competition. Despite its issues and challenges, I'm sure all parties have given their best and this has definitely been a memorable learning opportunity.</p>\n<p>Second, I'm grateful to be part of my team with <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> , <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> and <a href=\"https://www.kaggle.com/martasprg\" target=\"_blank\">@martasprg</a>. Your knowledge and hardwork inspired me a lot. It was fortunate that even when we missed our best shots, we still managed to achieve a shake-up and snatch a gold medal, all thanks to our talented and persisting team members.</p>\n<p>Finally, thank you all who participated and shared your experience/insights. This community has been and will always be the best part of Kaggle!</p>\n<p>Now let's begin!</p>\n<h1>Leaked Data Collection</h1>\n<p>I found out about this leaked dataset quite late (about 3 weeks before the deadline) when seeing the thread by <a href=\"https://www.kaggle.com/glipko\" target=\"_blank\">@glipko</a> - thank you! I then came across <a href=\"https://www.kaggle.com/datasets/rsakata/psp-raw-dataset\" target=\"_blank\">this dataset</a> by <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a>. Thank you both so much for sharing!</p>\n<p>I realized that, between these two datasets, there were some fils that exist in one but not the other. When combined together, their total distribution is below</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total sessions</td>\n<td>263,881</td>\n</tr>\n<tr>\n<td>Overlapping sessions with original train set</td>\n<td>14,505</td>\n</tr>\n<tr>\n<td>Sessions completing only level 0-4</td>\n<td>57,091</td>\n</tr>\n<tr>\n<td>Sessions completing only level 0-12</td>\n<td>21,623</td>\n</tr>\n<tr>\n<td>Sessions completing all levels (0-22)</td>\n<td>11,607</td>\n</tr>\n</tbody>\n</table>\n<p><em>(the rest are unusable &amp; non-overlapping sessions)</em></p>\n<h1>Leaked Data Preprocessing</h1>\n<p>I spent the last 3 weeks (until the very last hours) of the competition trying to preprocess the raw leaked data. Assuming that processing stays the same between from the original training to the hidden test set, I used the 14k overlapping sessions as a benchmark to compare my processed version with the original data. Due to some discrepancies from within the raw data itself, we could not obtain complete alignment. We got the similarity to 99.99% for labels and 99.5% for input, and stopped there.</p>\n<p>Here is the final version of our preprocessing notebook: <a href=\"https://www.kaggle.com/code/hoangnguyen719/pspfgp-additional-data-processed-v3-2\" target=\"_blank\">link to notebook</a></p>\n<h1>Leaked Data Usage</h1>\n<p>Due to the limited time left in the competition, we could utilize only the 11k sessions that complete all levels.</p>\n<p>However, one big problem remains: using our privately processed leaked data (hereby called <em>private extra data</em>) brings in much lower <strong>public LB</strong> improvement than using the processed data provided by <a href=\"https://www.kaggle.com/glipko\" target=\"_blank\">@glipko</a> <a href=\"https://www.kaggle.com/datasets/glipko/additional-data-predict-students-performance\" target=\"_blank\">here</a> (hereby called <em>public extra data</em>). With <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a>'s best FE and XGB, the public extra data improves LB by 0.001-0.003, but private extra data <strong>decreased</strong> our best pipeline's public LB score by 0.002. I did several comparisons and concluded that my private version is more similar to the original training data than the public version, and couldn’t understand why its publib LB's performance was worse.</p>\n<p>With hindsight, we now all know it was because the public LB's and the private LB's test sets are from different distributions.</p>\n<p>In addition, our ensembles (which uses only the public extra data) returned better LB results overall. At the end, we decided not to use the submission with private extra data. Similar to other teams, we missed our best submissions and learned a lesson about not overfitting public LB.</p>\n<h1>Performance Summary</h1>\n<p>Below are results of some of our models when using the leaked data:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Submission selected</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGB 1 (original training data + public extra data)</td>\n<td>0.70127</td>\n<td>0.704</td>\n<td>0.701</td>\n<td>No</td>\n</tr>\n<tr>\n<td>XGB 1 (original training data + private extra data)</td>\n<td>0.70162</td>\n<td>0.702</td>\n<td><strong>0.704</strong></td>\n<td>No</td>\n</tr>\n<tr>\n<td>XGB 2 (original training data + public extra data)</td>\n<td>0.70148</td>\n<td><strong>0.705</strong></td>\n<td>0.700</td>\n<td><strong>Yes</strong></td>\n</tr>\n<tr>\n<td>XGB 2 (original training data + private extra data)</td>\n<td><strong>0.7019</strong></td>\n<td>0.703</td>\n<td><strong>0.704</strong></td>\n<td>No</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": "2323893",
      "postDate": "06/30/2023 08:32:55",
      "content": "<p>…and we sadly missed it!</p>\n<p>In this discussion, I’m sharing about my data leak preprocessing in the competition. The raw data can be downloaded from <a href=\"https://opengamedata.fielddaylab.wisc.edu/gamedata.php?game=JOWILDER\" target=\"_blank\">Field Day Lab's Open Game Data website</a>. Our <strong>full solution</strong> has also been detailed out <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420528\" target=\"_blank\">here</a>.</p>\n<p>Before beginning, I would like to thank the host and Kaggle for this competition. Despite its issues and challenges, I'm sure all parties have given their best and this has definitely been a memorable learning opportunity.</p>\n<p>Second, I'm grateful to be part of my team with <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> , <a href=\"https://www.kaggle.com/shinomoriaoshi\" target=\"_blank\">@shinomoriaoshi</a> and <a href=\"https://www.kaggle.com/martasprg\" target=\"_blank\">@martasprg</a>. Your knowledge and hardwork inspired me a lot. It was fortunate that even when we missed our best shots, we still managed to achieve a shake-up and snatch a gold medal, all thanks to our talented and persisting team members.</p>\n<p>Finally, thank you all who participated and shared your experience/insights. This community has been and will always be the best part of Kaggle!</p>\n<p>Now let's begin!</p>\n<h1>Leaked Data Collection</h1>\n<p>I found out about this leaked dataset quite late (about 3 weeks before the deadline) when seeing the thread by <a href=\"https://www.kaggle.com/glipko\" target=\"_blank\">@glipko</a> - thank you! I then came across <a href=\"https://www.kaggle.com/datasets/rsakata/psp-raw-dataset\" target=\"_blank\">this dataset</a> by <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a>. Thank you both so much for sharing!</p>\n<p>I realized that, between these two datasets, there were some fils that exist in one but not the other. When combined together, their total distribution is below</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Total sessions</td>\n<td>263,881</td>\n</tr>\n<tr>\n<td>Overlapping sessions with original train set</td>\n<td>14,505</td>\n</tr>\n<tr>\n<td>Sessions completing only level 0-4</td>\n<td>57,091</td>\n</tr>\n<tr>\n<td>Sessions completing only level 0-12</td>\n<td>21,623</td>\n</tr>\n<tr>\n<td>Sessions completing all levels (0-22)</td>\n<td>11,607</td>\n</tr>\n</tbody>\n</table>\n<p><em>(the rest are unusable &amp; non-overlapping sessions)</em></p>\n<h1>Leaked Data Preprocessing</h1>\n<p>I spent the last 3 weeks (until the very last hours) of the competition trying to preprocess the raw leaked data. Assuming that processing stays the same between from the original training to the hidden test set, I used the 14k overlapping sessions as a benchmark to compare my processed version with the original data. Due to some discrepancies from within the raw data itself, we could not obtain complete alignment. We got the similarity to 99.99% for labels and 99.5% for input, and stopped there.</p>\n<p>Here is the final version of our preprocessing notebook: <a href=\"https://www.kaggle.com/code/hoangnguyen719/pspfgp-additional-data-processed-v3-2\" target=\"_blank\">link to notebook</a></p>\n<h1>Leaked Data Usage</h1>\n<p>Due to the limited time left in the competition, we could utilize only the 11k sessions that complete all levels.</p>\n<p>However, one big problem remains: using our privately processed leaked data (hereby called <em>private extra data</em>) brings in much lower <strong>public LB</strong> improvement than using the processed data provided by <a href=\"https://www.kaggle.com/glipko\" target=\"_blank\">@glipko</a> <a href=\"https://www.kaggle.com/datasets/glipko/additional-data-predict-students-performance\" target=\"_blank\">here</a> (hereby called <em>public extra data</em>). With <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a>'s best FE and XGB, the public extra data improves LB by 0.001-0.003, but private extra data <strong>decreased</strong> our best pipeline's public LB score by 0.002. I did several comparisons and concluded that my private version is more similar to the original training data than the public version, and couldn’t understand why its publib LB's performance was worse.</p>\n<p>With hindsight, we now all know it was because the public LB's and the private LB's test sets are from different distributions.</p>\n<p>In addition, our ensembles (which uses only the public extra data) returned better LB results overall. At the end, we decided not to use the submission with private extra data. Similar to other teams, we missed our best submissions and learned a lesson about not overfitting public LB.</p>\n<h1>Performance Summary</h1>\n<p>Below are results of some of our models when using the leaked data:</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n<th>Submission selected</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>XGB 1 (original training data + public extra data)</td>\n<td>0.70127</td>\n<td>0.704</td>\n<td>0.701</td>\n<td>No</td>\n</tr>\n<tr>\n<td>XGB 1 (original training data + private extra data)</td>\n<td>0.70162</td>\n<td>0.702</td>\n<td><strong>0.704</strong></td>\n<td>No</td>\n</tr>\n<tr>\n<td>XGB 2 (original training data + public extra data)</td>\n<td>0.70148</td>\n<td><strong>0.705</strong></td>\n<td>0.700</td>\n<td><strong>Yes</strong></td>\n</tr>\n<tr>\n<td>XGB 2 (original training data + private extra data)</td>\n<td><strong>0.7019</strong></td>\n<td>0.703</td>\n<td><strong>0.704</strong></td>\n<td>No</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "…and we sadly missed it!\n\nIn this discussion, I’m sharing about my data leak preprocessing in the competition. The raw data can be downloaded from [Field Day Lab's Open Game Data website](https://opengamedata.fielddaylab.wisc.edu/gamedata.php?game=JOWILDER). Our **full solution** has also been detailed out [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420528).\n\nBefore beginning, I would like to thank the host and Kaggle for this competition. Despite its issues and challenges, I'm sure all parties have given their best and this has definitely been a memorable learning opportunity.\n\nSecond, I'm grateful to be part of my team with @chaudharypriyanshu , @shinomoriaoshi and @martasprg. Your knowledge and hardwork inspired me a lot. It was fortunate that even when we missed our best shots, we still managed to achieve a shake-up and snatch a gold medal, all thanks to our talented and persisting team members.\n\nFinally, thank you all who participated and shared your experience/insights. This community has been and will always be the best part of Kaggle!\n\nNow let's begin!\n\n# Leaked Data Collection\nI found out about this leaked dataset quite late (about 3 weeks before the deadline) when seeing the thread by @glipko - thank you! I then came across [this dataset](https://www.kaggle.com/datasets/rsakata/psp-raw-dataset) by @rsakata. Thank you both so much for sharing!\n\nI realized that, between these two datasets, there were some fils that exist in one but not the other. When combined together, their total distribution is below\n| | |\n| --- | --- |\n| Total sessions | 263,881  |\n| Overlapping sessions with original train set | 14,505  |\n| Sessions completing only level 0-4 | 57,091  |\n| Sessions completing only level 0-12 | 21,623  |\n| Sessions completing all levels (0-22) | 11,607  |\n\n*(the rest are unusable & non-overlapping sessions)*\n\n# Leaked Data Preprocessing\nI spent the last 3 weeks (until the very last hours) of the competition trying to preprocess the raw leaked data. Assuming that processing stays the same between from the original training to the hidden test set, I used the 14k overlapping sessions as a benchmark to compare my processed version with the original data. Due to some discrepancies from within the raw data itself, we could not obtain complete alignment. We got the similarity to 99.99% for labels and 99.5% for input, and stopped there.\n\nHere is the final version of our preprocessing notebook: [link to notebook](https://www.kaggle.com/code/hoangnguyen719/pspfgp-additional-data-processed-v3-2)\n\n# Leaked Data Usage\nDue to the limited time left in the competition, we could utilize only the 11k sessions that complete all levels.\n\nHowever, one big problem remains: using our privately processed leaked data (hereby called *private extra data*) brings in much lower **public LB** improvement than using the processed data provided by @glipko [here](https://www.kaggle.com/datasets/glipko/additional-data-predict-students-performance) (hereby called *public extra data*). With @chaudharypriyanshu's best FE and XGB, the public extra data improves LB by 0.001-0.003, but private extra data **decreased** our best pipeline's public LB score by 0.002. I did several comparisons and concluded that my private version is more similar to the original training data than the public version, and couldn’t understand why its publib LB's performance was worse.\n\nWith hindsight, we now all know it was because the public LB's and the private LB's test sets are from different distributions.\n\nIn addition, our ensembles (which uses only the public extra data) returned better LB results overall. At the end, we decided not to use the submission with private extra data. Similar to other teams, we missed our best submissions and learned a lesson about not overfitting public LB.\n\n# Performance Summary\nBelow are results of some of our models when using the leaked data:\n \n| Model | CV | Public LB | Private LB | Submission selected |\n| --- | --- |\n| XGB 1 (original training data + public extra data) | 0.70127 | 0.704 | 0.701 | No |\n| XGB 1 (original training data + private extra data) | 0.70162 | 0.702 | **0.704** | No |\n| XGB 2 (original training data + public extra data) | 0.70148 | **0.705** | 0.700 | **Yes** |\n| XGB 2 (original training data + private extra data) | **0.7019** | 0.703 | **0.704** | No |",
      "votes": null
    },
    {
      "id": "2325933",
      "postDate": "07/01/2023 17:33:49",
      "content": "<p>Interesting, thanks for sharing!</p>\n<p>Did you compare CV <em>excluding</em> private data? From the chart above, I suspect the CV comparisons include CV score of the extra data itself?</p>\n<p>In other words, if you trained on either A or A+B, then comparing CV of A to CV of A+B might be less helpful than comparing CV of A to CV only on A (exclude CV on B from result). </p>",
      "rawMarkdown": "Interesting, thanks for sharing!\n\nDid you compare CV *excluding* private data? From the chart above, I suspect the CV comparisons include CV score of the extra data itself?\n\nIn other words, if you trained on either A or A+B, then comparing CV of A to CV of A+B might be less helpful than comparing CV of A to CV only on A (exclude CV on B from result).",
      "votes": null
    },
    {
      "id": "2326343",
      "postDate": "07/02/2023 04:34:47",
      "content": "<p>We compared CVs excluding External data, we only used external data for training. Since we needed previous questions predictions as meta-features, we created CV like this</p>\n<pre><code>\ntraining_data = pd.concat(train[fold != i],external[fold != i],axis=)\nvalidation_data = train[fold == i]  \nexternal_data_for_saving meta_feature = external[fold==i] \n</code></pre>\n<p>Hence, <strong>we validated our models only using the original training data provided by Kaggle</strong>, however, we did split the external data for training and saving meta-features that can be used across different questions.</p>\n<p>Our CV was reliable till we reached .703. Above 0.703, the scores became quite random for us (i.e we did not observe the difference &gt;=+0.002, b/w LB and CV that we were observing before, which lead to taking .705 submission due to LB &amp; CV difference&gt;= +0.002 difference).</p>\n<p>More about our Validation strategy can be found here:<a href=\"https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129</a></p>",
      "rawMarkdown": "We compared CVs excluding External data, we only used external data for training. Since we needed previous questions predictions as meta-features, we created CV like this\n\n```python\n## for  fold=i\ntraining_data = pd.concat(train[fold != i],external[fold != i],axis=0)\nvalidation_data = train[fold == i]  #for validation\nexternal_data_for_saving meta_feature = external[fold==i] #for meta feature generation\n```\nHence, **we validated our models only using the original training data provided by Kaggle**, however, we did split the external data for training and saving meta-features that can be used across different questions.\n\nOur CV was reliable till we reached .703. Above 0.703, the scores became quite random for us (i.e we did not observe the difference >=+0.002, b/w LB and CV that we were observing before, which lead to taking .705 submission due to LB & CV difference>= +0.002 difference).\n\nMore about our Validation strategy can be found here:<https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129>",
      "votes": null
    },
    {
      "id": "2329401",
      "postDate": "07/04/2023 08:55:46",
      "content": "<p>Sorry for the late response.</p>\n<p>I don't quite understand your last question. Are you saying that we should have done CV on only-A, on only-B, and on (A+B)?</p>\n<p>If yes, then I agree what you said would have been more thorough. However, there were two reasons why we could not do so:</p>\n<ol>\n<li>We did not have enough time - we started working with the extra data 3 weeks before deadline and had to spend half the time left to build &amp; debug the processing, and eventually did not have enough time &amp; submissions to test all possible cases.</li>\n<li>Both <em>private extra data</em> and <em>public extra data</em> differ from the original training data more or less, which we concluded was due to errors in preprocessing &amp; sources of the data. Given this assumption, we believed models trained on <strong>extra data alone</strong> wouldn't be safe to submit, and, going back to point (1) above, decided to spare our submissions for other more hopeful things (FE explorations, ensemble, etc.).</li>\n</ol>\n<p>Hope this makes sense, but if it doesn't/is not what you asked about, feel free to comment to continue the discussion!</p>",
      "rawMarkdown": "Sorry for the late response.\n\nI don't quite understand your last question. Are you saying that we should have done CV on only-A, on only-B, and on (A+B)?\n\nIf yes, then I agree what you said would have been more thorough. However, there were two reasons why we could not do so:\n1. We did not have enough time - we started working with the extra data 3 weeks before deadline and had to spend half the time left to build & debug the processing, and eventually did not have enough time & submissions to test all possible cases.\n2. Both *private extra data* and *public extra data* differ from the original training data more or less, which we concluded was due to errors in preprocessing & sources of the data. Given this assumption, we believed models trained on **extra data alone** wouldn't be safe to submit, and, going back to point (1) above, decided to spare our submissions for other more hopeful things (FE explorations, ensemble, etc.).\n\nHope this makes sense, but if it doesn't/is not what you asked about, feel free to comment to continue the discussion!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2325933,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/01/2023 17:33:49",
      "content": "<p>Interesting, thanks for sharing!</p>\n<p>Did you compare CV <em>excluding</em> private data? From the chart above, I suspect the CV comparisons include CV score of the extra data itself?</p>\n<p>In other words, if you trained on either A or A+B, then comparing CV of A to CV of A+B might be less helpful than comparing CV of A to CV only on A (exclude CV on B from result). </p>",
      "votes": null,
      "replies": [
        {
          "id": 2326343,
          "author_name": "chaudharypriyanshu",
          "author_url": "",
          "post_date": "07/02/2023 04:34:47",
          "content": "<p>We compared CVs excluding External data, we only used external data for training. Since we needed previous questions predictions as meta-features, we created CV like this</p>\n<pre><code>\ntraining_data = pd.concat(train[fold != i],external[fold != i],axis=)\nvalidation_data = train[fold == i]  \nexternal_data_for_saving meta_feature = external[fold==i] \n</code></pre>\n<p>Hence, <strong>we validated our models only using the original training data provided by Kaggle</strong>, however, we did split the external data for training and saving meta-features that can be used across different questions.</p>\n<p>Our CV was reliable till we reached .703. Above 0.703, the scores became quite random for us (i.e we did not observe the difference &gt;=+0.002, b/w LB and CV that we were observing before, which lead to taking .705 submission due to LB &amp; CV difference&gt;= +0.002 difference).</p>\n<p>More about our Validation strategy can be found here:<a href=\"https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129\" target=\"_blank\">https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2329401,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "07/04/2023 08:55:46",
          "content": "<p>Sorry for the late response.</p>\n<p>I don't quite understand your last question. Are you saying that we should have done CV on only-A, on only-B, and on (A+B)?</p>\n<p>If yes, then I agree what you said would have been more thorough. However, there were two reasons why we could not do so:</p>\n<ol>\n<li>We did not have enough time - we started working with the extra data 3 weeks before deadline and had to spend half the time left to build &amp; debug the processing, and eventually did not have enough time &amp; submissions to test all possible cases.</li>\n<li>Both <em>private extra data</em> and <em>public extra data</em> differ from the original training data more or less, which we concluded was due to errors in preprocessing &amp; sources of the data. Given this assumption, we believed models trained on <strong>extra data alone</strong> wouldn't be safe to submit, and, going back to point (1) above, decided to spare our submissions for other more hopeful things (FE explorations, ensemble, etc.).</li>\n</ol>\n<p>Hope this makes sense, but if it doesn't/is not what you asked about, feel free to comment to continue the discussion!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2323893": "…and we sadly missed it!\n\nIn this discussion, I’m sharing about my data leak preprocessing in the competition. The raw data can be downloaded from [Field Day Lab's Open Game Data website](https://opengamedata.fielddaylab.wisc.edu/gamedata.php?game=JOWILDER). Our **full solution** has also been detailed out [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420528).\n\nBefore beginning, I would like to thank the host and Kaggle for this competition. Despite its issues and challenges, I'm sure all parties have given their best and this has definitely been a memorable learning opportunity.\n\nSecond, I'm grateful to be part of my team with @chaudharypriyanshu , @shinomoriaoshi and @martasprg. Your knowledge and hardwork inspired me a lot. It was fortunate that even when we missed our best shots, we still managed to achieve a shake-up and snatch a gold medal, all thanks to our talented and persisting team members.\n\nFinally, thank you all who participated and shared your experience/insights. This community has been and will always be the best part of Kaggle!\n\nNow let's begin!\n\n# Leaked Data Collection\nI found out about this leaked dataset quite late (about 3 weeks before the deadline) when seeing the thread by @glipko - thank you! I then came across [this dataset](https://www.kaggle.com/datasets/rsakata/psp-raw-dataset) by @rsakata. Thank you both so much for sharing!\n\nI realized that, between these two datasets, there were some fils that exist in one but not the other. When combined together, their total distribution is below\n| | |\n| --- | --- |\n| Total sessions | 263,881  |\n| Overlapping sessions with original train set | 14,505  |\n| Sessions completing only level 0-4 | 57,091  |\n| Sessions completing only level 0-12 | 21,623  |\n| Sessions completing all levels (0-22) | 11,607  |\n\n*(the rest are unusable & non-overlapping sessions)*\n\n# Leaked Data Preprocessing\nI spent the last 3 weeks (until the very last hours) of the competition trying to preprocess the raw leaked data. Assuming that processing stays the same between from the original training to the hidden test set, I used the 14k overlapping sessions as a benchmark to compare my processed version with the original data. Due to some discrepancies from within the raw data itself, we could not obtain complete alignment. We got the similarity to 99.99% for labels and 99.5% for input, and stopped there.\n\nHere is the final version of our preprocessing notebook: [link to notebook](https://www.kaggle.com/code/hoangnguyen719/pspfgp-additional-data-processed-v3-2)\n\n# Leaked Data Usage\nDue to the limited time left in the competition, we could utilize only the 11k sessions that complete all levels.\n\nHowever, one big problem remains: using our privately processed leaked data (hereby called *private extra data*) brings in much lower **public LB** improvement than using the processed data provided by @glipko [here](https://www.kaggle.com/datasets/glipko/additional-data-predict-students-performance) (hereby called *public extra data*). With @chaudharypriyanshu's best FE and XGB, the public extra data improves LB by 0.001-0.003, but private extra data **decreased** our best pipeline's public LB score by 0.002. I did several comparisons and concluded that my private version is more similar to the original training data than the public version, and couldn’t understand why its publib LB's performance was worse.\n\nWith hindsight, we now all know it was because the public LB's and the private LB's test sets are from different distributions.\n\nIn addition, our ensembles (which uses only the public extra data) returned better LB results overall. At the end, we decided not to use the submission with private extra data. Similar to other teams, we missed our best submissions and learned a lesson about not overfitting public LB.\n\n# Performance Summary\nBelow are results of some of our models when using the leaked data:\n \n| Model | CV | Public LB | Private LB | Submission selected |\n| --- | --- |\n| XGB 1 (original training data + public extra data) | 0.70127 | 0.704 | 0.701 | No |\n| XGB 1 (original training data + private extra data) | 0.70162 | 0.702 | **0.704** | No |\n| XGB 2 (original training data + public extra data) | 0.70148 | **0.705** | 0.700 | **Yes** |\n| XGB 2 (original training data + private extra data) | **0.7019** | 0.703 | **0.704** | No |",
    "2325933": "Interesting, thanks for sharing!\n\nDid you compare CV *excluding* private data? From the chart above, I suspect the CV comparisons include CV score of the extra data itself?\n\nIn other words, if you trained on either A or A+B, then comparing CV of A to CV of A+B might be less helpful than comparing CV of A to CV only on A (exclude CV on B from result).",
    "2326343": "We compared CVs excluding External data, we only used external data for training. Since we needed previous questions predictions as meta-features, we created CV like this\n\n```python\n## for  fold=i\ntraining_data = pd.concat(train[fold != i],external[fold != i],axis=0)\nvalidation_data = train[fold == i]  #for validation\nexternal_data_for_saving meta_feature = external[fold==i] #for meta feature generation\n```\nHence, **we validated our models only using the original training data provided by Kaggle**, however, we did split the external data for training and saving meta-features that can be used across different questions.\n\nOur CV was reliable till we reached .703. Above 0.703, the scores became quite random for us (i.e we did not observe the difference >=+0.002, b/w LB and CV that we were observing before, which lead to taking .705 submission due to LB & CV difference>= +0.002 difference).\n\nMore about our Validation strategy can be found here:<https://www.kaggle.com/code/chaudharypriyanshu/8th-place-solution-training?scriptVersionId=135015129>",
    "2329401": "Sorry for the late response.\n\nI don't quite understand your last question. Are you saying that we should have done CV on only-A, on only-B, and on (A+B)?\n\nIf yes, then I agree what you said would have been more thorough. However, there were two reasons why we could not do so:\n1. We did not have enough time - we started working with the extra data 3 weeks before deadline and had to spend half the time left to build & debug the processing, and eventually did not have enough time & submissions to test all possible cases.\n2. Both *private extra data* and *public extra data* differ from the original training data more or less, which we concluded was due to errors in preprocessing & sources of the data. Given this assumption, we believed models trained on **extra data alone** wouldn't be safe to submit, and, going back to point (1) above, decided to spare our submissions for other more hopeful things (FE explorations, ensemble, etc.).\n\nHope this makes sense, but if it doesn't/is not what you asked about, feel free to comment to continue the discussion!"
  },
  "source": "meta"
}