{
  "id": 420274,
  "title": "3rd place solution(yyykrk part)",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420274",
  "author_name": "",
  "post_date": "2023-06-30T04:04:58.316289300Z",
  "votes": 21,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thank you very much to the team members( <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a>). It was my first time working in a team. I learned a lot, so I would like to participate in teams more actively in the future.</p>\n<p>The overall solution is explained here: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235</a>, so I would like to provide additional explanation about the part I worked on.</p>\n<p>My source codes: <a href=\"https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play\" target=\"_blank\">https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play</a></p>\n<h2>The main topics are as follows</h2>\n<ul>\n<li>Model summary</li>\n<li>Created features</li>\n<li>Feature selection by out-of-fold to avoid leakage</li>\n<li>Generate additional dataset from raw log data</li>\n</ul>\n<h2>Model summary</h2>\n<ul>\n<li><p>Individual model for each question (18 binary classifiers)</p></li>\n<li><p>Single Best Model (5folds XGBoost)</p>\n<p>CV: 0.702, Public LB: 0.700, Private LB: 0.701</p></li>\n</ul>\n<h2>Created features</h2>\n<h3>Features in our final submissions</h3>\n<ul>\n<li>Elapsed time diff between the previous level group end and the current level group start<ul>\n<li>This works because it almost represents the time the user takes to answer questions.</li></ul></li>\n<li>Elapsed time and index count between flag events<ul>\n<li>What I call flag events are events that must be passed during game progression. Although some events can occur in any order, most of them have a specific order.</li>\n<li>These events can be extracted from users who have answered all 18 questions correctly. Some of them are users who take the shortest action as if they already knew the answer and finish the game with a small number of records. In addition, the next action to be taken is appended to the notebook when the flag event is passed, so I actually checked it while playing the game.</li></ul></li>\n<li>Prediction probabilities for previous questions<ul>\n<li>All prediction values from the questions before the target question<ul>\n<li>For example, to predict q3, then use q1 and q2 prediction probabilities</li></ul></li>\n<li>Sum of the most recent M (M=1,2,…) prediction probabilities</li></ul></li>\n<li>Aggregate features<ul>\n<li>elapsed time diff</li>\n<li>screen_coor_x diff, screen_coor_y diff</li>\n<li>index per elapsed time</li>\n<li>Some boolean<ul>\n<li>elapsed time diff &lt; 0</li>\n<li>elapsed time diff &gt; th</li></ul></li>\n<li>index counts and elapsed time for each category as a percentage of the total</li>\n<li>hover_duration, page, …</li></ul></li>\n<li>Datetime</li>\n<li>Screen aspect (based on max value and min value of screen_coor_x and screen_coor_y)</li>\n</ul>\n<h3>Features not in our final submissions</h3>\n<p>These are the features removed because no correlation between cv and public scores.</p>\n<ul>\n<li>Some aggregate features<ul>\n<li>elapsed time diff forward (shift: -1)</li>\n<li>elapsed time max - elapsed time current</li>\n<li>room_fqid in the current record == room_fqid in the previous record</li></ul></li>\n<li>Some prediction probabilities features<ul>\n<li>Sum of percentages relative to the average correct rate</li>\n<li>Sum of probabilities or percentages binned</li></ul></li>\n<li>Meta features<ul>\n<li>Aggregate features for oof prediction probabilities (not previous question)</li>\n<li>The 2019 Data Science Bowl's 3rd place solution will be helpful.<ul>\n<li><a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a></li></ul></li></ul></li>\n<li>Script type (normal, nohumor, dry, nosnark)<ul>\n<li>This can be predicted from the unique text that appears</li></ul></li>\n<li>Sub fqids/room_fqids broken down comma by comma</li>\n</ul>\n<h2>Feature selection by out-of-fold to avoid leakage</h2>\n<p>When feature selection, I did it by out of fold (per fold×question). Otherwise, it would cause leakage, unfairly boost the CV score, and affect negatively ensemble weights, etc.</p>\n<p>Advantages:</p>\n<ul>\n<li>Able to evaluate fairly the CV score of a model</li>\n<li>When ensembles, prevent calculating mistakenly higher weights for feature-selected models</li>\n</ul>\n<p>Disadvantages:</p>\n<ul>\n<li>More difficult to manage features</li>\n<li>Increased computational cost due to the total number of features required<ul>\n<li>For example, when using the top 300 features, the total number of features is as follows:<ul>\n<li>NOT OOF: 300</li>\n<li>OOF: (fold1 300) | (fold2 300) | ・・・|(fold K 300) ≥ 300</li></ul></li>\n<li>Of course, in the most cases, this difference is small.</li></ul></li>\n</ul>\n<p>In my case, even though the CV score did not change with feature selection, the public score became worse, so I switched to a policy of no feature selection. However, when I saw the top solutions creating thousands of features and implementing feature selection, I now wish I had created a larger number of features and done feature selection.</p>\n<h2>Additional dataset</h2>\n<p>I created almost perfectly reproduces train.csv from raw log data. Since the format slightly differs between old and recent data, encoding was troublesome, but the following data was obtained in the end.</p>\n<p>Additional dataset:</p>\n<ul>\n<li>Complete sessions: 10,000+</li>\n<li>Incomplete sessions (retire players): 200,000+</li>\n</ul>\n<p>I used the above incomplete sessions, and while the CV score increased by about 0.002, there was no positive impact on the private score. As in the excellent solution of 7th place: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119</a>, adding the maximum level of each session might have provided good results for me.</p>",
  "messages": [
    {
      "id": "2323587",
      "postDate": "06/30/2023 04:04:58",
      "content": "<p>Thank you very much to the team members( <a href=\"https://www.kaggle.com/wimwim\" target=\"_blank\">@wimwim</a> <a href=\"https://www.kaggle.com/kingychiu\" target=\"_blank\">@kingychiu</a> <a href=\"https://www.kaggle.com/tangtunyu\" target=\"_blank\">@tangtunyu</a>). It was my first time working in a team. I learned a lot, so I would like to participate in teams more actively in the future.</p>\n<p>The overall solution is explained here: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235</a>, so I would like to provide additional explanation about the part I worked on.</p>\n<p>My source codes: <a href=\"https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play\" target=\"_blank\">https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play</a></p>\n<h2>The main topics are as follows</h2>\n<ul>\n<li>Model summary</li>\n<li>Created features</li>\n<li>Feature selection by out-of-fold to avoid leakage</li>\n<li>Generate additional dataset from raw log data</li>\n</ul>\n<h2>Model summary</h2>\n<ul>\n<li><p>Individual model for each question (18 binary classifiers)</p></li>\n<li><p>Single Best Model (5folds XGBoost)</p>\n<p>CV: 0.702, Public LB: 0.700, Private LB: 0.701</p></li>\n</ul>\n<h2>Created features</h2>\n<h3>Features in our final submissions</h3>\n<ul>\n<li>Elapsed time diff between the previous level group end and the current level group start<ul>\n<li>This works because it almost represents the time the user takes to answer questions.</li></ul></li>\n<li>Elapsed time and index count between flag events<ul>\n<li>What I call flag events are events that must be passed during game progression. Although some events can occur in any order, most of them have a specific order.</li>\n<li>These events can be extracted from users who have answered all 18 questions correctly. Some of them are users who take the shortest action as if they already knew the answer and finish the game with a small number of records. In addition, the next action to be taken is appended to the notebook when the flag event is passed, so I actually checked it while playing the game.</li></ul></li>\n<li>Prediction probabilities for previous questions<ul>\n<li>All prediction values from the questions before the target question<ul>\n<li>For example, to predict q3, then use q1 and q2 prediction probabilities</li></ul></li>\n<li>Sum of the most recent M (M=1,2,…) prediction probabilities</li></ul></li>\n<li>Aggregate features<ul>\n<li>elapsed time diff</li>\n<li>screen_coor_x diff, screen_coor_y diff</li>\n<li>index per elapsed time</li>\n<li>Some boolean<ul>\n<li>elapsed time diff &lt; 0</li>\n<li>elapsed time diff &gt; th</li></ul></li>\n<li>index counts and elapsed time for each category as a percentage of the total</li>\n<li>hover_duration, page, …</li></ul></li>\n<li>Datetime</li>\n<li>Screen aspect (based on max value and min value of screen_coor_x and screen_coor_y)</li>\n</ul>\n<h3>Features not in our final submissions</h3>\n<p>These are the features removed because no correlation between cv and public scores.</p>\n<ul>\n<li>Some aggregate features<ul>\n<li>elapsed time diff forward (shift: -1)</li>\n<li>elapsed time max - elapsed time current</li>\n<li>room_fqid in the current record == room_fqid in the previous record</li></ul></li>\n<li>Some prediction probabilities features<ul>\n<li>Sum of percentages relative to the average correct rate</li>\n<li>Sum of probabilities or percentages binned</li></ul></li>\n<li>Meta features<ul>\n<li>Aggregate features for oof prediction probabilities (not previous question)</li>\n<li>The 2019 Data Science Bowl's 3rd place solution will be helpful.<ul>\n<li><a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\" target=\"_blank\">https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388</a></li></ul></li></ul></li>\n<li>Script type (normal, nohumor, dry, nosnark)<ul>\n<li>This can be predicted from the unique text that appears</li></ul></li>\n<li>Sub fqids/room_fqids broken down comma by comma</li>\n</ul>\n<h2>Feature selection by out-of-fold to avoid leakage</h2>\n<p>When feature selection, I did it by out of fold (per fold×question). Otherwise, it would cause leakage, unfairly boost the CV score, and affect negatively ensemble weights, etc.</p>\n<p>Advantages:</p>\n<ul>\n<li>Able to evaluate fairly the CV score of a model</li>\n<li>When ensembles, prevent calculating mistakenly higher weights for feature-selected models</li>\n</ul>\n<p>Disadvantages:</p>\n<ul>\n<li>More difficult to manage features</li>\n<li>Increased computational cost due to the total number of features required<ul>\n<li>For example, when using the top 300 features, the total number of features is as follows:<ul>\n<li>NOT OOF: 300</li>\n<li>OOF: (fold1 300) | (fold2 300) | ・・・|(fold K 300) ≥ 300</li></ul></li>\n<li>Of course, in the most cases, this difference is small.</li></ul></li>\n</ul>\n<p>In my case, even though the CV score did not change with feature selection, the public score became worse, so I switched to a policy of no feature selection. However, when I saw the top solutions creating thousands of features and implementing feature selection, I now wish I had created a larger number of features and done feature selection.</p>\n<h2>Additional dataset</h2>\n<p>I created almost perfectly reproduces train.csv from raw log data. Since the format slightly differs between old and recent data, encoding was troublesome, but the following data was obtained in the end.</p>\n<p>Additional dataset:</p>\n<ul>\n<li>Complete sessions: 10,000+</li>\n<li>Incomplete sessions (retire players): 200,000+</li>\n</ul>\n<p>I used the above incomplete sessions, and while the CV score increased by about 0.002, there was no positive impact on the private score. As in the excellent solution of 7th place: <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119</a>, adding the maximum level of each session might have provided good results for me.</p>",
      "rawMarkdown": "Thank you very much to the team members( @wimwim @kingychiu @tangtunyu). It was my first time working in a team. I learned a lot, so I would like to participate in teams more actively in the future.\n\nThe overall solution is explained here: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235, so I would like to provide additional explanation about the part I worked on.\n\nMy source codes: https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play\n\n## The main topics are as follows\n\n- Model summary\n- Created features\n- Feature selection by out-of-fold to avoid leakage\n- Generate additional dataset from raw log data\n\n## Model summary\n\n- Individual model for each question (18 binary classifiers)\n- Single Best Model (5folds XGBoost)\n    \n    CV: 0.702, Public LB: 0.700, Private LB: 0.701\n    \n\n## Created features\n\n### Features in our final submissions\n\n- Elapsed time diff between the previous level group end and the current level group start\n    - This works because it almost represents the time the user takes to answer questions.\n- Elapsed time and index count between flag events\n    - What I call flag events are events that must be passed during game progression. Although some events can occur in any order, most of them have a specific order.\n    - These events can be extracted from users who have answered all 18 questions correctly. Some of them are users who take the shortest action as if they already knew the answer and finish the game with a small number of records. In addition, the next action to be taken is appended to the notebook when the flag event is passed, so I actually checked it while playing the game.\n- Prediction probabilities for previous questions\n    - All prediction values from the questions before the target question\n        - For example, to predict q3, then use q1 and q2 prediction probabilities\n    - Sum of the most recent M (M=1,2,…) prediction probabilities\n- Aggregate features\n    - elapsed time diff\n    - screen_coor_x diff, screen_coor_y diff\n    - index per elapsed time\n    - Some boolean\n        - elapsed time diff < 0\n        - elapsed time diff > th\n    - index counts and elapsed time for each category as a percentage of the total\n    - hover_duration, page, …\n- Datetime\n- Screen aspect (based on max value and min value of screen_coor_x and screen_coor_y)\n\n### Features not in our final submissions\n\nThese are the features removed because no correlation between cv and public scores.\n\n- Some aggregate features\n    - elapsed time diff forward (shift: -1)\n    - elapsed time max - elapsed time current\n    - room_fqid in the current record == room_fqid in the previous record\n- Some prediction probabilities features\n    - Sum of percentages relative to the average correct rate\n    - Sum of probabilities or percentages binned\n- Meta features\n    - Aggregate features for oof prediction probabilities (not previous question)\n    - The 2019 Data Science Bowl's 3rd place solution will be helpful.\n        - https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\n- Script type (normal, nohumor, dry, nosnark)\n    - This can be predicted from the unique text that appears\n- Sub fqids/room_fqids broken down comma by comma\n\n## Feature selection by out-of-fold to avoid leakage\n\nWhen feature selection, I did it by out of fold (per fold×question). Otherwise, it would cause leakage, unfairly boost the CV score, and affect negatively ensemble weights, etc.\n\nAdvantages:\n\n- Able to evaluate fairly the CV score of a model\n- When ensembles, prevent calculating mistakenly higher weights for feature-selected models\n\nDisadvantages:\n\n- More difficult to manage features\n- Increased computational cost due to the total number of features required\n    - For example, when using the top 300 features, the total number of features is as follows:\n        - NOT OOF: 300\n        - OOF: (fold1 300) | (fold2 300) | ・・・|(fold K 300) ≥ 300\n    - Of course, in the most cases, this difference is small.\n\nIn my case, even though the CV score did not change with feature selection, the public score became worse, so I switched to a policy of no feature selection. However, when I saw the top solutions creating thousands of features and implementing feature selection, I now wish I had created a larger number of features and done feature selection.\n\n## Additional dataset\n\nI created almost perfectly reproduces train.csv from raw log data. Since the format slightly differs between old and recent data, encoding was troublesome, but the following data was obtained in the end.\n\nAdditional dataset:\n\n- Complete sessions: 10,000+\n- Incomplete sessions (retire players): 200,000+\n\nI used the above incomplete sessions, and while the CV score increased by about 0.002, there was no positive impact on the private score. As in the excellent solution of 7th place: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119, adding the maximum level of each session might have provided good results for me.",
      "votes": null
    },
    {
      "id": "2323621",
      "postDate": "06/30/2023 04:53:59",
      "content": "<p><a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> congrats with 3 place and gold medal! Very interesting solution!</p>",
      "rawMarkdown": "yyykrk congrats with 3 place and gold medal! Very interesting solution!",
      "votes": null
    },
    {
      "id": "2323648",
      "postDate": "06/30/2023 05:18:08",
      "content": "<p>Your experimentation speed brings a lot of value to the team as well! </p>",
      "rawMarkdown": "Your experimentation speed brings a lot of value to the team as well!",
      "votes": null
    },
    {
      "id": "2323690",
      "postDate": "06/30/2023 05:55:47",
      "content": "<p>Nice! Congrats on the win!</p>",
      "rawMarkdown": "Nice! Congrats on the win!",
      "votes": null
    },
    {
      "id": "2323846",
      "postDate": "06/30/2023 07:38:15",
      "content": "<p>Congratulations on coming 3rd <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> 🎉🎉🎉</p>",
      "rawMarkdown": "Congratulations on coming 3rd @yyykrk 🎉🎉🎉",
      "votes": null
    },
    {
      "id": "2324090",
      "postDate": "06/30/2023 11:28:44",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "2324099",
      "postDate": "06/30/2023 11:33:52",
      "content": "<p>I am very grateful to you because without your approach and models we would not have won the Gold Medal!</p>",
      "rawMarkdown": "I am very grateful to you because without your approach and models we would not have won the Gold Medal!",
      "votes": null
    },
    {
      "id": "2343922",
      "postDate": "07/14/2023 05:17:48",
      "content": "<p>Congratulations. Thanks for sharing the solution. Its a very useful learning resource. <br>\nDid you try any advanced programming techniques?</p>",
      "rawMarkdown": "Congratulations. Thanks for sharing the solution. Its a very useful learning resource. \nDid you try any advanced programming techniques?",
      "votes": null
    },
    {
      "id": "2344197",
      "postDate": "07/14/2023 10:39:34",
      "content": "<p>Thanks. I haven't tried anything too special.</p>",
      "rawMarkdown": "Thanks. I haven't tried anything too special.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323621,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "06/30/2023 04:53:59",
      "content": "<p><a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> congrats with 3 place and gold medal! Very interesting solution!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2324090,
          "author_name": "yyykrk",
          "author_url": "",
          "post_date": "06/30/2023 11:28:44",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2323648,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "06/30/2023 05:18:08",
      "content": "<p>Your experimentation speed brings a lot of value to the team as well! </p>",
      "votes": null,
      "replies": [
        {
          "id": 2324099,
          "author_name": "yyykrk",
          "author_url": "",
          "post_date": "06/30/2023 11:33:52",
          "content": "<p>I am very grateful to you because without your approach and models we would not have won the Gold Medal!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2323690,
      "author_name": "orereisrael",
      "author_url": "",
      "post_date": "06/30/2023 05:55:47",
      "content": "<p>Nice! Congrats on the win!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323846,
      "author_name": "swapnilchowdhury",
      "author_url": "",
      "post_date": "06/30/2023 07:38:15",
      "content": "<p>Congratulations on coming 3rd <a href=\"https://www.kaggle.com/yyykrk\" target=\"_blank\">@yyykrk</a> 🎉🎉🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2343922,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/14/2023 05:17:48",
      "content": "<p>Congratulations. Thanks for sharing the solution. Its a very useful learning resource. <br>\nDid you try any advanced programming techniques?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2344197,
          "author_name": "yyykrk",
          "author_url": "",
          "post_date": "07/14/2023 10:39:34",
          "content": "<p>Thanks. I haven't tried anything too special.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2323587": "Thank you very much to the team members( @wimwim @kingychiu @tangtunyu). It was my first time working in a team. I learned a lot, so I would like to participate in teams more actively in the future.\n\nThe overall solution is explained here: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420235, so I would like to provide additional explanation about the part I worked on.\n\nMy source codes: https://github.com/krk-san/3rd-place-solution-Predict-Student-Performance-from-Game-Play\n\n## The main topics are as follows\n\n- Model summary\n- Created features\n- Feature selection by out-of-fold to avoid leakage\n- Generate additional dataset from raw log data\n\n## Model summary\n\n- Individual model for each question (18 binary classifiers)\n- Single Best Model (5folds XGBoost)\n    \n    CV: 0.702, Public LB: 0.700, Private LB: 0.701\n    \n\n## Created features\n\n### Features in our final submissions\n\n- Elapsed time diff between the previous level group end and the current level group start\n    - This works because it almost represents the time the user takes to answer questions.\n- Elapsed time and index count between flag events\n    - What I call flag events are events that must be passed during game progression. Although some events can occur in any order, most of them have a specific order.\n    - These events can be extracted from users who have answered all 18 questions correctly. Some of them are users who take the shortest action as if they already knew the answer and finish the game with a small number of records. In addition, the next action to be taken is appended to the notebook when the flag event is passed, so I actually checked it while playing the game.\n- Prediction probabilities for previous questions\n    - All prediction values from the questions before the target question\n        - For example, to predict q3, then use q1 and q2 prediction probabilities\n    - Sum of the most recent M (M=1,2,…) prediction probabilities\n- Aggregate features\n    - elapsed time diff\n    - screen_coor_x diff, screen_coor_y diff\n    - index per elapsed time\n    - Some boolean\n        - elapsed time diff < 0\n        - elapsed time diff > th\n    - index counts and elapsed time for each category as a percentage of the total\n    - hover_duration, page, …\n- Datetime\n- Screen aspect (based on max value and min value of screen_coor_x and screen_coor_y)\n\n### Features not in our final submissions\n\nThese are the features removed because no correlation between cv and public scores.\n\n- Some aggregate features\n    - elapsed time diff forward (shift: -1)\n    - elapsed time max - elapsed time current\n    - room_fqid in the current record == room_fqid in the previous record\n- Some prediction probabilities features\n    - Sum of percentages relative to the average correct rate\n    - Sum of probabilities or percentages binned\n- Meta features\n    - Aggregate features for oof prediction probabilities (not previous question)\n    - The 2019 Data Science Bowl's 3rd place solution will be helpful.\n        - https://www.kaggle.com/c/data-science-bowl-2019/discussion/127388\n- Script type (normal, nohumor, dry, nosnark)\n    - This can be predicted from the unique text that appears\n- Sub fqids/room_fqids broken down comma by comma\n\n## Feature selection by out-of-fold to avoid leakage\n\nWhen feature selection, I did it by out of fold (per fold×question). Otherwise, it would cause leakage, unfairly boost the CV score, and affect negatively ensemble weights, etc.\n\nAdvantages:\n\n- Able to evaluate fairly the CV score of a model\n- When ensembles, prevent calculating mistakenly higher weights for feature-selected models\n\nDisadvantages:\n\n- More difficult to manage features\n- Increased computational cost due to the total number of features required\n    - For example, when using the top 300 features, the total number of features is as follows:\n        - NOT OOF: 300\n        - OOF: (fold1 300) | (fold2 300) | ・・・|(fold K 300) ≥ 300\n    - Of course, in the most cases, this difference is small.\n\nIn my case, even though the CV score did not change with feature selection, the public score became worse, so I switched to a policy of no feature selection. However, when I saw the top solutions creating thousands of features and implementing feature selection, I now wish I had created a larger number of features and done feature selection.\n\n## Additional dataset\n\nI created almost perfectly reproduces train.csv from raw log data. Since the format slightly differs between old and recent data, encoding was troublesome, but the following data was obtained in the end.\n\nAdditional dataset:\n\n- Complete sessions: 10,000+\n- Incomplete sessions (retire players): 200,000+\n\nI used the above incomplete sessions, and while the CV score increased by about 0.002, there was no positive impact on the private score. As in the excellent solution of 7th place: https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420119, adding the maximum level of each session might have provided good results for me.",
    "2323621": "yyykrk congrats with 3 place and gold medal! Very interesting solution!",
    "2323648": "Your experimentation speed brings a lot of value to the team as well!",
    "2323690": "Nice! Congrats on the win!",
    "2323846": "Congratulations on coming 3rd @yyykrk 🎉🎉🎉",
    "2324090": "Thank you!",
    "2324099": "I am very grateful to you because without your approach and models we would not have won the Gold Medal!",
    "2343922": "Congratulations. Thanks for sharing the solution. Its a very useful learning resource. \nDid you try any advanced programming techniques?",
    "2344197": "Thanks. I haven't tried anything too special."
  },
  "source": "meta"
}