{
  "id": 420046,
  "title": "9th Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/makotu-9th-place-solution",
  "author_name": "",
  "post_date": "2023-06-29T03:44:49Z",
  "votes": 81,
  "comment_count": 29,
  "views": 0,
  "content": "<p>First of all, I would first like to thank all the participants who dedicated so much time and effort to this competition, as well as the hosts and the management. While I still believe there are areas in which the administration of the competition could be improved, in this post, I will focus solely on discussing my solution.</p>\n<h2>Overview</h2>\n<p>I didn't do anything particularly special.</p>\n<p>Mainly, I just kept adding features to improve the accuracy of the single model. For each question, I built models using LightGBM and Catboost, and took the simple average of the two models. I used the models with the highest CV scores as the final candidates. The high CV submit also got almost the best score for private.</p>\n<ul>\n<li>LightGBM CV：0.7018 LB(Public)：0.703 LB(Private)：0.702</li>\n<li>Catboost CV:0.7011 LB(Public)：0.7 LB(Private)：0.701</li>\n<li>Merge(final submit) CV:0.7024 LB(Public)：0.7 LB(Private)：0.702</li>\n</ul>\n<h2>Features</h2>\n<p>Rather than introducing each of the numerous features I created, I'll share my overall approach and discuss a few specific features that particularly contributed to the accuracy.</p>\n<p>As already demonstrated in public notebooks, an important element in this competition was \"how much time one spends playing.\" To delve deeper, I felt that \"how much time it took from a certain point to another\" was crucial, so I created many features related to this.</p>\n<ul>\n<li>Checkpoint feature (as I named it)<br>\nIn this game, there are events that almost every player will inevitably experience. For instance, every user will find a notebook and see the message \"found it!\" I identified these \"events that almost every user goes through,\" and used the time taken between these events (i.e., the elapsed time from event A to event B) and the number of clicks as features. This seemed to significantly contribute to the accuracy.</li>\n</ul>\n<p>I created such features in various patterns, like the elapsed time from viewing text A to text B, the elapsed time from one fqid to the next, the elapsed time from one room to the next, and so on.</p>\n<ul>\n<li>Other than this, I obviously included features like the time elapsed for each level and the average coordinates, as introduced in the public notebooks.</li>\n</ul>\n<p>The number of features increases with the level group. Ultimately, the feature counts were as follows:<br>\nLevel group 0-4: 3009 features Level group 5-12: 9747 features Level group 13-22: 18610 features</p>\n<h2>Modeling Approach</h2>\n<ul>\n<li><p>I chose to use 10-fold rather than 5-fold as it gave slightly higher CV scores (around +0.0005).</p></li>\n<li><p>I used straified Kfold with the number of correct answers for 18 questions of the user. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)</p></li>\n<li><p>For feature selection, I simply used the top 500 features based on their importance. To prevent leakage, I selected the feature importance for each fold, and retrained the model for each fold. For example, when training fold1, I first train with all features, then select the features using the fold1 model, and retrain the fold1 model with the top 500 features.</p></li>\n<li><p>I used the prediction probabilities of previous questions as features. For example, when predicting question 3, I used the prediction probabilities for questions 1 and 2. When predicting question 15, I used the prediction probabilities for questions 1 through 14, etc. (this improved CV by around +0.001)</p></li>\n<li><p>Inference took the following amounts of time:<br>\nLightGBM: 90min Catboost: 120min Merge: 150min<br>\nLightGBM's inference became significantly faster by compiling the model with a library called \"lleaves.\"<br>\n<a href=\"https://github.com/siboehm/lleaves\" target=\"_blank\">https://github.com/siboehm/lleaves</a><br>\nI'm sharing my inference code. Features not mentioned here can be somewhat understood by looking at it.<br>\n<a href=\"https://www.kaggle.com/code/mhyodo/restart-model-merge-v1\" target=\"_blank\">https://www.kaggle.com/code/mhyodo/restart-model-merge-v1</a></p></li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li>NN models (I tried several types, such as LSTM and MLP, but none contributed to the CV)</li>\n</ul>\n<p>Lastly, I've seen posts from others where the LB score was higher than the CV score, but in my case, they were pretty much the same. I was expecting some sort of shakeup, but I didn't anticipate making it into the top 10. I'm curious as to how those with higher LB scores achieved this, as I was unable to significantly increase my LB score. <br>\nAnyway, thank you! If the ranking is confirmed, I can become a new GrandMaster!</p>",
  "messages": [
    {
      "id": "2321980",
      "postDate": "06/29/2023 03:00:43",
      "content": "<p>First of all, I would first like to thank all the participants who dedicated so much time and effort to this competition, as well as the hosts and the management. While I still believe there are areas in which the administration of the competition could be improved, in this post, I will focus solely on discussing my solution.</p>\n<h2>Overview</h2>\n<p>I didn't do anything particularly special.</p>\n<p>Mainly, I just kept adding features to improve the accuracy of the single model. For each question, I built models using LightGBM and Catboost, and took the simple average of the two models. I used the models with the highest CV scores as the final candidates. The high CV submit also got almost the best score for private.</p>\n<ul>\n<li>LightGBM CV：0.7018 LB(Public)：0.703 LB(Private)：0.702</li>\n<li>Catboost CV:0.7011 LB(Public)：0.7 LB(Private)：0.701</li>\n<li>Merge(final submit) CV:0.7024 LB(Public)：0.7 LB(Private)：0.702</li>\n</ul>\n<h2>Features</h2>\n<p>Rather than introducing each of the numerous features I created, I'll share my overall approach and discuss a few specific features that particularly contributed to the accuracy.</p>\n<p>As already demonstrated in public notebooks, an important element in this competition was \"how much time one spends playing.\" To delve deeper, I felt that \"how much time it took from a certain point to another\" was crucial, so I created many features related to this.</p>\n<ul>\n<li>Checkpoint feature (as I named it)<br>\nIn this game, there are events that almost every player will inevitably experience. For instance, every user will find a notebook and see the message \"found it!\" I identified these \"events that almost every user goes through,\" and used the time taken between these events (i.e., the elapsed time from event A to event B) and the number of clicks as features. This seemed to significantly contribute to the accuracy.</li>\n</ul>\n<p>I created such features in various patterns, like the elapsed time from viewing text A to text B, the elapsed time from one fqid to the next, the elapsed time from one room to the next, and so on.</p>\n<ul>\n<li>Other than this, I obviously included features like the time elapsed for each level and the average coordinates, as introduced in the public notebooks.</li>\n</ul>\n<p>The number of features increases with the level group. Ultimately, the feature counts were as follows:<br>\nLevel group 0-4: 3009 features Level group 5-12: 9747 features Level group 13-22: 18610 features</p>\n<h2>Modeling Approach</h2>\n<ul>\n<li><p>I chose to use 10-fold rather than 5-fold as it gave slightly higher CV scores (around +0.0005).</p></li>\n<li><p>I used straified Kfold with the number of correct answers for 18 questions of the user. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)</p></li>\n<li><p>For feature selection, I simply used the top 500 features based on their importance. To prevent leakage, I selected the feature importance for each fold, and retrained the model for each fold. For example, when training fold1, I first train with all features, then select the features using the fold1 model, and retrain the fold1 model with the top 500 features.</p></li>\n<li><p>I used the prediction probabilities of previous questions as features. For example, when predicting question 3, I used the prediction probabilities for questions 1 and 2. When predicting question 15, I used the prediction probabilities for questions 1 through 14, etc. (this improved CV by around +0.001)</p></li>\n<li><p>Inference took the following amounts of time:<br>\nLightGBM: 90min Catboost: 120min Merge: 150min<br>\nLightGBM's inference became significantly faster by compiling the model with a library called \"lleaves.\"<br>\n<a href=\"https://github.com/siboehm/lleaves\" target=\"_blank\">https://github.com/siboehm/lleaves</a><br>\nI'm sharing my inference code. Features not mentioned here can be somewhat understood by looking at it.<br>\n<a href=\"https://www.kaggle.com/code/mhyodo/restart-model-merge-v1\" target=\"_blank\">https://www.kaggle.com/code/mhyodo/restart-model-merge-v1</a></p></li>\n</ul>\n<h2>What Didn't Work</h2>\n<ul>\n<li>NN models (I tried several types, such as LSTM and MLP, but none contributed to the CV)</li>\n</ul>\n<p>Lastly, I've seen posts from others where the LB score was higher than the CV score, but in my case, they were pretty much the same. I was expecting some sort of shakeup, but I didn't anticipate making it into the top 10. I'm curious as to how those with higher LB scores achieved this, as I was unable to significantly increase my LB score. <br>\nAnyway, thank you! If the ranking is confirmed, I can become a new GrandMaster!</p>",
      "rawMarkdown": "First of all, I would first like to thank all the participants who dedicated so much time and effort to this competition, as well as the hosts and the management. While I still believe there are areas in which the administration of the competition could be improved, in this post, I will focus solely on discussing my solution.\n\n\n## Overview\nI didn't do anything particularly special.\n\nMainly, I just kept adding features to improve the accuracy of the single model. For each question, I built models using LightGBM and Catboost, and took the simple average of the two models. I used the models with the highest CV scores as the final candidates. The high CV submit also got almost the best score for private.\n\n- LightGBM CV：0.7018 LB(Public)：0.703 LB(Private)：0.702\n- Catboost CV:0.7011 LB(Public)：0.7 LB(Private)：0.701\n- Merge(final submit) CV:0.7024 LB(Public)：0.7 LB(Private)：0.702\n\n\n## Features\nRather than introducing each of the numerous features I created, I'll share my overall approach and discuss a few specific features that particularly contributed to the accuracy.\n\nAs already demonstrated in public notebooks, an important element in this competition was \"how much time one spends playing.\" To delve deeper, I felt that \"how much time it took from a certain point to another\" was crucial, so I created many features related to this.\n\n- Checkpoint feature (as I named it)\nIn this game, there are events that almost every player will inevitably experience. For instance, every user will find a notebook and see the message \"found it!\" I identified these \"events that almost every user goes through,\" and used the time taken between these events (i.e., the elapsed time from event A to event B) and the number of clicks as features. This seemed to significantly contribute to the accuracy.\n\nI created such features in various patterns, like the elapsed time from viewing text A to text B, the elapsed time from one fqid to the next, the elapsed time from one room to the next, and so on.\n\n- Other than this, I obviously included features like the time elapsed for each level and the average coordinates, as introduced in the public notebooks.\n\nThe number of features increases with the level group. Ultimately, the feature counts were as follows:\nLevel group 0-4: 3009 features Level group 5-12: 9747 features Level group 13-22: 18610 features\n\n\n## Modeling Approach\n- I chose to use 10-fold rather than 5-fold as it gave slightly higher CV scores (around +0.0005).\n- I used straified Kfold with the number of correct answers for 18 questions of the user. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)\n- For feature selection, I simply used the top 500 features based on their importance. To prevent leakage, I selected the feature importance for each fold, and retrained the model for each fold. For example, when training fold1, I first train with all features, then select the features using the fold1 model, and retrain the fold1 model with the top 500 features.\n- I used the prediction probabilities of previous questions as features. For example, when predicting question 3, I used the prediction probabilities for questions 1 and 2. When predicting question 15, I used the prediction probabilities for questions 1 through 14, etc. (this improved CV by around +0.001)\n\n- Inference took the following amounts of time:\nLightGBM: 90min Catboost: 120min Merge: 150min\nLightGBM's inference became significantly faster by compiling the model with a library called \"lleaves.\"\nhttps://github.com/siboehm/lleaves\nI'm sharing my inference code. Features not mentioned here can be somewhat understood by looking at it.\nhttps://www.kaggle.com/code/mhyodo/restart-model-merge-v1\n\n\n## What Didn't Work\n- NN models (I tried several types, such as LSTM and MLP, but none contributed to the CV)\n\n\n\nLastly, I've seen posts from others where the LB score was higher than the CV score, but in my case, they were pretty much the same. I was expecting some sort of shakeup, but I didn't anticipate making it into the top 10. I'm curious as to how those with higher LB scores achieved this, as I was unable to significantly increase my LB score. \nAnyway, thank you! If the ranking is confirmed, I can become a new GrandMaster!",
      "votes": null
    },
    {
      "id": "2321984",
      "postDate": "06/29/2023 03:03:32",
      "content": "<p>Congratulations on your soon-coming GM title! Great work!</p>",
      "rawMarkdown": "Congratulations on your soon-coming GM title! Great work!",
      "votes": null
    },
    {
      "id": "2321994",
      "postDate": "06/29/2023 03:15:22",
      "content": "<p>Congratulations Makotu achieving solo gold 9th and Congratulations for become Kaggle competition Grandmaster!</p>\n<p>I also observed that my best CV models had <code>CV score = Public LB score = Private LB score</code></p>",
      "rawMarkdown": "Congratulations Makotu achieving solo gold 9th and Congratulations for become Kaggle competition Grandmaster!\n\nI also observed that my best CV models had `CV score = Public LB score = Private LB score`",
      "votes": null
    },
    {
      "id": "2321998",
      "postDate": "06/29/2023 03:18:35",
      "content": "<p>Thanks for sharing and big congrats on your GM title!!</p>",
      "rawMarkdown": "Thanks for sharing and big congrats on your GM title!!",
      "votes": null
    },
    {
      "id": "2322032",
      "postDate": "06/29/2023 03:36:27",
      "content": "<p>Congrats on the gold medal and the incoming GM title!</p>\n<p>Can I ask, what was your CV strategy? (stratified multi-label/time-series window/something else?)</p>",
      "rawMarkdown": "Congrats on the gold medal and the incoming GM title!\n\nCan I ask, what was your CV strategy? (stratified multi-label/time-series window/something else?)",
      "votes": null
    },
    {
      "id": "2322038",
      "postDate": "06/29/2023 03:43:36",
      "content": "<p>Oh, yes, I forgot to mention that. Thank you very much. I will add it later.<br>\nI used straified Kfold with the number of correct answers for 18 questions of the USER. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)</p>",
      "rawMarkdown": "Oh, yes, I forgot to mention that. Thank you very much. I will add it later.\nI used straified Kfold with the number of correct answers for 18 questions of the USER. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)",
      "votes": null
    },
    {
      "id": "2322070",
      "postDate": "06/29/2023 04:24:31",
      "content": "<p>Did you use the extra 7000 sessions from the leaked data? I think many teams used it to get top public/private LB.</p>",
      "rawMarkdown": "Did you use the extra 7000 sessions from the leaked data? I think many teams used it to get top public/private LB.",
      "votes": null
    },
    {
      "id": "2322083",
      "postDate": "06/29/2023 04:30:55",
      "content": "<p>I didn't use any extra data.</p>",
      "rawMarkdown": "I didn't use any extra data.",
      "votes": null
    },
    {
      "id": "2322095",
      "postDate": "06/29/2023 04:38:02",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> congrats with 9 place and gold medal!</p>",
      "rawMarkdown": "mhyodo congrats with 9 place and gold medal!",
      "votes": null
    },
    {
      "id": "2322100",
      "postDate": "06/29/2023 04:41:30",
      "content": "<p>That leaked data seems to be the secret recipe for some top public LB positions. However I think some of us utilized it too much and caused overfitting/shake-downs.</p>\n<p>Kudos to you for building a gold-medal solution without the need of that data!</p>",
      "rawMarkdown": "That leaked data seems to be the secret recipe for some top public LB positions. However I think some of us utilized it too much and caused overfitting/shake-downs.\n\nKudos to you for building a gold-medal solution without the need of that data!",
      "votes": null
    },
    {
      "id": "2322108",
      "postDate": "06/29/2023 04:45:51",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations!",
      "votes": null
    },
    {
      "id": "2322116",
      "postDate": "06/29/2023 04:53:06",
      "content": "<p>Congratulations on becoming a GrandMaster! Thank you so much for sharing!👍👍👍</p>",
      "rawMarkdown": "Congratulations on becoming a GrandMaster! Thank you so much for sharing!👍👍👍",
      "votes": null
    },
    {
      "id": "2322128",
      "postDate": "06/29/2023 05:06:30",
      "content": "<p>Though we didn’t select unfortunately 🙁 (due to low public lb)but with extra data we saw better private score around 0.703 0.704 , will post soon</p>",
      "rawMarkdown": "Though we didn’t select unfortunately 🙁 (due to low public lb)but with extra data we saw better private score around 0.703 0.704 , will post soon",
      "votes": null
    },
    {
      "id": "2322476",
      "postDate": "06/29/2023 09:58:30",
      "content": "<p>congratulations. Great work! I must say I agree with you that don't try NN models. Based on my experience, it isn't easy to get a great score from NN models in a run, even if the task seems to be designed for NN models💥. </p>",
      "rawMarkdown": "congratulations. Great work! I must say I agree with you that don't try NN models. Based on my experience, it isn't easy to get a great score from NN models in a run, even if the task seems to be designed for NN models💥.",
      "votes": null
    },
    {
      "id": "2322700",
      "postDate": "06/29/2023 12:49:37",
      "content": "<p>Congratulations on achieving solo gold and GM! <br>\nIn my case, what significantly contributed to higher LB scores were the ensemble of LightGBM and NN models, as well as adjusting the threshold. The following are the scores when the threshold was changed:</p>\n<ul>\n<li>Threshold: 0.617<ul>\n<li>Public Score: 0.706</li>\n<li>Private Score: 0.702<br>\n<br></li></ul></li>\n<li>Threshold: 0.623<ul>\n<li>Public Score: 0.705</li>\n<li>Private Score: 0.702<br>\n<br></li></ul></li>\n<li>Threshold: 0.629<ul>\n<li>Public Score: 0.703</li>\n<li>Private Score: 0.701</li></ul></li>\n</ul>",
      "rawMarkdown": "Congratulations on achieving solo gold and GM! \nIn my case, what significantly contributed to higher LB scores were the ensemble of LightGBM and NN models, as well as adjusting the threshold. The following are the scores when the threshold was changed:\n\n\n- Threshold: 0.617\n    - Public Score: 0.706\n    - Private Score: 0.702\n</br>\n- Threshold: 0.623\n    - Public Score: 0.705\n    - Private Score: 0.702\n</br>\n- Threshold: 0.629\n    - Public Score: 0.703\n    - Private Score: 0.701",
      "votes": null
    },
    {
      "id": "2322747",
      "postDate": "06/29/2023 13:28:53",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> Congratulations on your achievement ! The checkpoint feature you developed, based on the time taken between specific events, seems particularly insightful.</p>",
      "rawMarkdown": "mhyodo Congratulations on your achievement ! The checkpoint feature you developed, based on the time taken between specific events, seems particularly insightful.",
      "votes": null
    },
    {
      "id": "2323463",
      "postDate": "06/30/2023 01:04:34",
      "content": "<p>I see, thank you for your info! I had not tried either.</p>",
      "rawMarkdown": "I see, thank you for your info! I had not tried either.",
      "votes": null
    },
    {
      "id": "2323493",
      "postDate": "06/30/2023 02:02:27",
      "content": "<p>We also observed threshold played a part, best for us was 0.605</p>",
      "rawMarkdown": "We also observed threshold played a part, best for us was 0.605",
      "votes": null
    },
    {
      "id": "2323654",
      "postDate": "06/30/2023 05:20:24",
      "content": "<p>Congratulations on your 9th place finish in the competition! Your approach of combining LightGBM and Catboost models through averaging was effective. The checkpoint feature and utilization of prediction probabilities of previous questions were smart additions. Well done!</p>",
      "rawMarkdown": "Congratulations on your 9th place finish in the competition! Your approach of combining LightGBM and Catboost models through averaging was effective. The checkpoint feature and utilization of prediction probabilities of previous questions were smart additions. Well done!",
      "votes": null
    },
    {
      "id": "2323769",
      "postDate": "06/30/2023 06:41:33",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>  Thank you for sharing and congratulations! Next time I wish to get the first place!</p>",
      "rawMarkdown": "mhyodo  Thank you for sharing and congratulations! Next time I wish to get the first place!",
      "votes": null
    },
    {
      "id": "2324139",
      "postDate": "06/30/2023 12:09:44",
      "content": "<p>thanks for sharing and congrats on gold</p>",
      "rawMarkdown": "thanks for sharing and congrats on gold",
      "votes": null
    },
    {
      "id": "2326959",
      "postDate": "07/02/2023 15:00:30",
      "content": "<p>👍👍Very glad to see your circle turn yellow</p>",
      "rawMarkdown": "👍👍Very glad to see your circle turn yellow",
      "votes": null
    },
    {
      "id": "2327239",
      "postDate": "07/02/2023 18:47:19",
      "content": "<p>I also came up with checkpoint features. (But didn't do all the great work to build a robust solution, of course :) ) Did you notice that there are four different text options, each session_id only getting one of them?</p>\n<p>If you spend enough effort, (I did NOT, lol) you can have more checkpoints through matching up equivalent texts. But not sure if having more checkpoints would really even improve score. </p>",
      "rawMarkdown": "I also came up with checkpoint features. (But didn't do all the great work to build a robust solution, of course :) ) Did you notice that there are four different text options, each session_id only getting one of them?\n\nIf you spend enough effort, (I did NOT, lol) you can have more checkpoints through matching up equivalent texts. But not sure if having more checkpoints would really even improve score.",
      "votes": null
    },
    {
      "id": "2327248",
      "postDate": "07/02/2023 18:57:35",
      "content": "<p>Great solution！Thanks for sharing and congratulations on becoming a GrandMaster! </p>",
      "rawMarkdown": "Great solution！Thanks for sharing and congratulations on becoming a GrandMaster!",
      "votes": null
    },
    {
      "id": "2328869",
      "postDate": "07/03/2023 23:47:01",
      "content": "<p>Congratulations on your solo gold and earning the title of Competitions Grandmaster!</p>\n<p>I've tried a similar approach, but I haven't been able to achieve a high score with LightGBM (always worse than Catboost.) Would you mind sharing how you set your hyperparameters for the models? And could you also explain your tuning process? </p>",
      "rawMarkdown": "Congratulations on your solo gold and earning the title of Competitions Grandmaster!\n\nI've tried a similar approach, but I haven't been able to achieve a high score with LightGBM (always worse than Catboost.) Would you mind sharing how you set your hyperparameters for the models? And could you also explain your tuning process?",
      "votes": null
    },
    {
      "id": "2328890",
      "postDate": "07/04/2023 00:32:11",
      "content": "<p>Here are my parameters.  <br>\nSince the features are very large and prone to overfitting, making min_data_in_leaf large enough and taking feature_fraction small enough improved the CV in my case.</p>\n<pre><code>param = {\n    : , \n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : , \n    : -,\n    :\n}\n</code></pre>\n<p>I didn't spend much time on tuning, I tried several patterns by hand, guessing the parameters that seemed important. (Ex. max_depth, min_data_in_leaf, feature_fraction</p>",
      "rawMarkdown": "Here are my parameters.  \nSince the features are very large and prone to overfitting, making min_data_in_leaf large enough and taking feature_fraction small enough improved the CV in my case.\n\n```python\nparam = {\n    'boosting': 'gbdt', \n    'objective': 'binary',\n    'metric': 'binary_logloss',\n    'learning_rate': 0.02,\n    'max_depth': 7,\n    'min_data_in_leaf': 300,\n    'bagging_fraction': 0.5,\n    'feature_fraction': 0.1, # Level group 0-4: 0.3  Level group 5-12: 0.1  Level group 13-22: 0.05\n    'verbose': -1,\n    'seed':1208\n}\n```\n\nI didn't spend much time on tuning, I tried several patterns by hand, guessing the parameters that seemed important. (Ex. max_depth, min_data_in_leaf, feature_fraction",
      "votes": null
    },
    {
      "id": "2328893",
      "postDate": "07/04/2023 00:35:38",
      "content": "<p>Thanks Sirius!<br>\nCongrats to you too on your high placement in the KDDcup!</p>",
      "rawMarkdown": "Thanks Sirius!\nCongrats to you too on your high placement in the KDDcup!",
      "votes": null
    },
    {
      "id": "2328895",
      "postDate": "07/04/2023 00:37:26",
      "content": "<p>Yes, I noticed.<br>\nHowever, in my case, even if I made the text consistent, the CV did not increase that much.</p>",
      "rawMarkdown": "Yes, I noticed.\nHowever, in my case, even if I made the text consistent, the CV did not increase that much.",
      "votes": null
    },
    {
      "id": "2328896",
      "postDate": "07/04/2023 00:37:51",
      "content": "<p>Thank you for your prompt response! Your advice is very helpful. Once again, congratulations on becoming a Grandmaster!</p>",
      "rawMarkdown": "Thank you for your prompt response! Your advice is very helpful. Once again, congratulations on becoming a Grandmaster!",
      "votes": null
    },
    {
      "id": "2342418",
      "postDate": "07/12/2023 19:27:51",
      "content": "<p>Congrats on being a GM!</p>",
      "rawMarkdown": "Congrats on being a GM!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2321984,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 03:03:32",
      "content": "<p>Congratulations on your soon-coming GM title! Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2321994,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/29/2023 03:15:22",
      "content": "<p>Congratulations Makotu achieving solo gold 9th and Congratulations for become Kaggle competition Grandmaster!</p>\n<p>I also observed that my best CV models had <code>CV score = Public LB score = Private LB score</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2321998,
      "author_name": "hookman",
      "author_url": "",
      "post_date": "06/29/2023 03:18:35",
      "content": "<p>Thanks for sharing and big congrats on your GM title!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322032,
      "author_name": "informhunter",
      "author_url": "",
      "post_date": "06/29/2023 03:36:27",
      "content": "<p>Congrats on the gold medal and the incoming GM title!</p>\n<p>Can I ask, what was your CV strategy? (stratified multi-label/time-series window/something else?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322038,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "06/29/2023 03:43:36",
          "content": "<p>Oh, yes, I forgot to mention that. Thank you very much. I will add it later.<br>\nI used straified Kfold with the number of correct answers for 18 questions of the USER. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2322070,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 04:24:31",
      "content": "<p>Did you use the extra 7000 sessions from the leaked data? I think many teams used it to get top public/private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322083,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "06/29/2023 04:30:55",
          "content": "<p>I didn't use any extra data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2322100,
              "author_name": "hoangnguyen719",
              "author_url": "",
              "post_date": "06/29/2023 04:41:30",
              "content": "<p>That leaked data seems to be the secret recipe for some top public LB positions. However I think some of us utilized it too much and caused overfitting/shake-downs.</p>\n<p>Kudos to you for building a gold-medal solution without the need of that data!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2322128,
                  "author_name": "gauravbrills",
                  "author_url": "",
                  "post_date": "06/29/2023 05:06:30",
                  "content": "<p>Though we didn’t select unfortunately 🙁 (due to low public lb)but with extra data we saw better private score around 0.703 0.704 , will post soon</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2322095,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "06/29/2023 04:38:02",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> congrats with 9 place and gold medal!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322108,
      "author_name": "joey1001",
      "author_url": "",
      "post_date": "06/29/2023 04:45:51",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322116,
      "author_name": "kaggleaau",
      "author_url": "",
      "post_date": "06/29/2023 04:53:06",
      "content": "<p>Congratulations on becoming a GrandMaster! Thank you so much for sharing!👍👍👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322476,
      "author_name": "carloszonetgmailcom",
      "author_url": "",
      "post_date": "06/29/2023 09:58:30",
      "content": "<p>congratulations. Great work! I must say I agree with you that don't try NN models. Based on my experience, it isn't easy to get a great score from NN models in a run, even if the task seems to be designed for NN models💥. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2322700,
      "author_name": "takoihiraokazu",
      "author_url": "",
      "post_date": "06/29/2023 12:49:37",
      "content": "<p>Congratulations on achieving solo gold and GM! <br>\nIn my case, what significantly contributed to higher LB scores were the ensemble of LightGBM and NN models, as well as adjusting the threshold. The following are the scores when the threshold was changed:</p>\n<ul>\n<li>Threshold: 0.617<ul>\n<li>Public Score: 0.706</li>\n<li>Private Score: 0.702<br>\n<br></li></ul></li>\n<li>Threshold: 0.623<ul>\n<li>Public Score: 0.705</li>\n<li>Private Score: 0.702<br>\n<br></li></ul></li>\n<li>Threshold: 0.629<ul>\n<li>Public Score: 0.703</li>\n<li>Private Score: 0.701</li></ul></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 2323463,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "06/30/2023 01:04:34",
          "content": "<p>I see, thank you for your info! I had not tried either.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2323493,
              "author_name": "gauravbrills",
              "author_url": "",
              "post_date": "06/30/2023 02:02:27",
              "content": "<p>We also observed threshold played a part, best for us was 0.605</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2322747,
      "author_name": "akshayvyas02",
      "author_url": "",
      "post_date": "06/29/2023 13:28:53",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a> Congratulations on your achievement ! The checkpoint feature you developed, based on the time taken between specific events, seems particularly insightful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323654,
      "author_name": "poojach7611",
      "author_url": "",
      "post_date": "06/30/2023 05:20:24",
      "content": "<p>Congratulations on your 9th place finish in the competition! Your approach of combining LightGBM and Catboost models through averaging was effective. The checkpoint feature and utilization of prediction probabilities of previous questions were smart additions. Well done!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323769,
      "author_name": "eugeniyosetrov",
      "author_url": "",
      "post_date": "06/30/2023 06:41:33",
      "content": "<p><a href=\"https://www.kaggle.com/mhyodo\" target=\"_blank\">@mhyodo</a>  Thank you for sharing and congratulations! Next time I wish to get the first place!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2324139,
      "author_name": "joek47",
      "author_url": "",
      "post_date": "06/30/2023 12:09:44",
      "content": "<p>thanks for sharing and congrats on gold</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2326959,
      "author_name": "sirius81",
      "author_url": "",
      "post_date": "07/02/2023 15:00:30",
      "content": "<p>👍👍Very glad to see your circle turn yellow</p>",
      "votes": null,
      "replies": [
        {
          "id": 2328893,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "07/04/2023 00:35:38",
          "content": "<p>Thanks Sirius!<br>\nCongrats to you too on your high placement in the KDDcup!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2327239,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/02/2023 18:47:19",
      "content": "<p>I also came up with checkpoint features. (But didn't do all the great work to build a robust solution, of course :) ) Did you notice that there are four different text options, each session_id only getting one of them?</p>\n<p>If you spend enough effort, (I did NOT, lol) you can have more checkpoints through matching up equivalent texts. But not sure if having more checkpoints would really even improve score. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2328895,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "07/04/2023 00:37:26",
          "content": "<p>Yes, I noticed.<br>\nHowever, in my case, even if I made the text consistent, the CV did not increase that much.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2327248,
      "author_name": "zhuyile",
      "author_url": "",
      "post_date": "07/02/2023 18:57:35",
      "content": "<p>Great solution！Thanks for sharing and congratulations on becoming a GrandMaster! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2328869,
      "author_name": "maruichi01",
      "author_url": "",
      "post_date": "07/03/2023 23:47:01",
      "content": "<p>Congratulations on your solo gold and earning the title of Competitions Grandmaster!</p>\n<p>I've tried a similar approach, but I haven't been able to achieve a high score with LightGBM (always worse than Catboost.) Would you mind sharing how you set your hyperparameters for the models? And could you also explain your tuning process? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2328890,
          "author_name": "mhyodo",
          "author_url": "",
          "post_date": "07/04/2023 00:32:11",
          "content": "<p>Here are my parameters.  <br>\nSince the features are very large and prone to overfitting, making min_data_in_leaf large enough and taking feature_fraction small enough improved the CV in my case.</p>\n<pre><code>param = {\n    : , \n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : ,\n    : , \n    : -,\n    :\n}\n</code></pre>\n<p>I didn't spend much time on tuning, I tried several patterns by hand, guessing the parameters that seemed important. (Ex. max_depth, min_data_in_leaf, feature_fraction</p>",
          "votes": null,
          "replies": [
            {
              "id": 2328896,
              "author_name": "maruichi01",
              "author_url": "",
              "post_date": "07/04/2023 00:37:51",
              "content": "<p>Thank you for your prompt response! Your advice is very helpful. Once again, congratulations on becoming a Grandmaster!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2342418,
      "author_name": "kingychiu",
      "author_url": "",
      "post_date": "07/12/2023 19:27:51",
      "content": "<p>Congrats on being a GM!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2321980": "First of all, I would first like to thank all the participants who dedicated so much time and effort to this competition, as well as the hosts and the management. While I still believe there are areas in which the administration of the competition could be improved, in this post, I will focus solely on discussing my solution.\n\n\n## Overview\nI didn't do anything particularly special.\n\nMainly, I just kept adding features to improve the accuracy of the single model. For each question, I built models using LightGBM and Catboost, and took the simple average of the two models. I used the models with the highest CV scores as the final candidates. The high CV submit also got almost the best score for private.\n\n- LightGBM CV：0.7018 LB(Public)：0.703 LB(Private)：0.702\n- Catboost CV:0.7011 LB(Public)：0.7 LB(Private)：0.701\n- Merge(final submit) CV:0.7024 LB(Public)：0.7 LB(Private)：0.702\n\n\n## Features\nRather than introducing each of the numerous features I created, I'll share my overall approach and discuss a few specific features that particularly contributed to the accuracy.\n\nAs already demonstrated in public notebooks, an important element in this competition was \"how much time one spends playing.\" To delve deeper, I felt that \"how much time it took from a certain point to another\" was crucial, so I created many features related to this.\n\n- Checkpoint feature (as I named it)\nIn this game, there are events that almost every player will inevitably experience. For instance, every user will find a notebook and see the message \"found it!\" I identified these \"events that almost every user goes through,\" and used the time taken between these events (i.e., the elapsed time from event A to event B) and the number of clicks as features. This seemed to significantly contribute to the accuracy.\n\nI created such features in various patterns, like the elapsed time from viewing text A to text B, the elapsed time from one fqid to the next, the elapsed time from one room to the next, and so on.\n\n- Other than this, I obviously included features like the time elapsed for each level and the average coordinates, as introduced in the public notebooks.\n\nThe number of features increases with the level group. Ultimately, the feature counts were as follows:\nLevel group 0-4: 3009 features Level group 5-12: 9747 features Level group 13-22: 18610 features\n\n\n## Modeling Approach\n- I chose to use 10-fold rather than 5-fold as it gave slightly higher CV scores (around +0.0005).\n- I used straified Kfold with the number of correct answers for 18 questions of the user. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)\n- For feature selection, I simply used the top 500 features based on their importance. To prevent leakage, I selected the feature importance for each fold, and retrained the model for each fold. For example, when training fold1, I first train with all features, then select the features using the fold1 model, and retrain the fold1 model with the top 500 features.\n- I used the prediction probabilities of previous questions as features. For example, when predicting question 3, I used the prediction probabilities for questions 1 and 2. When predicting question 15, I used the prediction probabilities for questions 1 through 14, etc. (this improved CV by around +0.001)\n\n- Inference took the following amounts of time:\nLightGBM: 90min Catboost: 120min Merge: 150min\nLightGBM's inference became significantly faster by compiling the model with a library called \"lleaves.\"\nhttps://github.com/siboehm/lleaves\nI'm sharing my inference code. Features not mentioned here can be somewhat understood by looking at it.\nhttps://www.kaggle.com/code/mhyodo/restart-model-merge-v1\n\n\n## What Didn't Work\n- NN models (I tried several types, such as LSTM and MLP, but none contributed to the CV)\n\n\n\nLastly, I've seen posts from others where the LB score was higher than the CV score, but in my case, they were pretty much the same. I was expecting some sort of shakeup, but I didn't anticipate making it into the top 10. I'm curious as to how those with higher LB scores achieved this, as I was unable to significantly increase my LB score. \nAnyway, thank you! If the ranking is confirmed, I can become a new GrandMaster!",
    "2321984": "Congratulations on your soon-coming GM title! Great work!",
    "2321994": "Congratulations Makotu achieving solo gold 9th and Congratulations for become Kaggle competition Grandmaster!\n\nI also observed that my best CV models had `CV score = Public LB score = Private LB score`",
    "2321998": "Thanks for sharing and big congrats on your GM title!!",
    "2322032": "Congrats on the gold medal and the incoming GM title!\n\nCan I ask, what was your CV strategy? (stratified multi-label/time-series window/something else?)",
    "2322038": "Oh, yes, I forgot to mention that. Thank you very much. I will add it later.\nI used straified Kfold with the number of correct answers for 18 questions of the USER. So, the model is made with the distribution of the total number of correct answers of the users almost aligned. (However, I don't think it would be much different with a simple K-fold)",
    "2322070": "Did you use the extra 7000 sessions from the leaked data? I think many teams used it to get top public/private LB.",
    "2322083": "I didn't use any extra data.",
    "2322095": "mhyodo congrats with 9 place and gold medal!",
    "2322100": "That leaked data seems to be the secret recipe for some top public LB positions. However I think some of us utilized it too much and caused overfitting/shake-downs.\n\nKudos to you for building a gold-medal solution without the need of that data!",
    "2322108": "Thanks for sharing and congratulations!",
    "2322116": "Congratulations on becoming a GrandMaster! Thank you so much for sharing!👍👍👍",
    "2322128": "Though we didn’t select unfortunately 🙁 (due to low public lb)but with extra data we saw better private score around 0.703 0.704 , will post soon",
    "2322476": "congratulations. Great work! I must say I agree with you that don't try NN models. Based on my experience, it isn't easy to get a great score from NN models in a run, even if the task seems to be designed for NN models💥.",
    "2322700": "Congratulations on achieving solo gold and GM! \nIn my case, what significantly contributed to higher LB scores were the ensemble of LightGBM and NN models, as well as adjusting the threshold. The following are the scores when the threshold was changed:\n\n\n- Threshold: 0.617\n    - Public Score: 0.706\n    - Private Score: 0.702\n</br>\n- Threshold: 0.623\n    - Public Score: 0.705\n    - Private Score: 0.702\n</br>\n- Threshold: 0.629\n    - Public Score: 0.703\n    - Private Score: 0.701",
    "2322747": "mhyodo Congratulations on your achievement ! The checkpoint feature you developed, based on the time taken between specific events, seems particularly insightful.",
    "2323463": "I see, thank you for your info! I had not tried either.",
    "2323493": "We also observed threshold played a part, best for us was 0.605",
    "2323654": "Congratulations on your 9th place finish in the competition! Your approach of combining LightGBM and Catboost models through averaging was effective. The checkpoint feature and utilization of prediction probabilities of previous questions were smart additions. Well done!",
    "2323769": "mhyodo  Thank you for sharing and congratulations! Next time I wish to get the first place!",
    "2324139": "thanks for sharing and congrats on gold",
    "2326959": "👍👍Very glad to see your circle turn yellow",
    "2327239": "I also came up with checkpoint features. (But didn't do all the great work to build a robust solution, of course :) ) Did you notice that there are four different text options, each session_id only getting one of them?\n\nIf you spend enough effort, (I did NOT, lol) you can have more checkpoints through matching up equivalent texts. But not sure if having more checkpoints would really even improve score.",
    "2327248": "Great solution！Thanks for sharing and congratulations on becoming a GrandMaster!",
    "2328869": "Congratulations on your solo gold and earning the title of Competitions Grandmaster!\n\nI've tried a similar approach, but I haven't been able to achieve a high score with LightGBM (always worse than Catboost.) Would you mind sharing how you set your hyperparameters for the models? And could you also explain your tuning process?",
    "2328890": "Here are my parameters.  \nSince the features are very large and prone to overfitting, making min_data_in_leaf large enough and taking feature_fraction small enough improved the CV in my case.\n\n```python\nparam = {\n    'boosting': 'gbdt', \n    'objective': 'binary',\n    'metric': 'binary_logloss',\n    'learning_rate': 0.02,\n    'max_depth': 7,\n    'min_data_in_leaf': 300,\n    'bagging_fraction': 0.5,\n    'feature_fraction': 0.1, # Level group 0-4: 0.3  Level group 5-12: 0.1  Level group 13-22: 0.05\n    'verbose': -1,\n    'seed':1208\n}\n```\n\nI didn't spend much time on tuning, I tried several patterns by hand, guessing the parameters that seemed important. (Ex. max_depth, min_data_in_leaf, feature_fraction",
    "2328893": "Thanks Sirius!\nCongrats to you too on your high placement in the KDDcup!",
    "2328895": "Yes, I noticed.\nHowever, in my case, even if I made the text consistent, the CV did not increase that much.",
    "2328896": "Thank you for your prompt response! Your advice is very helpful. Once again, congratulations on becoming a Grandmaster!",
    "2342418": "Congrats on being a GM!"
  },
  "source": "meta"
}