{
  "id": 420119,
  "title": "7th Place Solution (Efficiency 1st)",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420119",
  "author_name": "Jack (Japan)",
  "post_date": "2023-06-29T09:35:17.036000",
  "votes": 102,
  "comment_count": 32,
  "views": 0,
  "content": "<p>I am pleased to have fought the long and hard competition with all of you here.<br>\nHere I would like to outline my solution.</p>\n<h2>Overview</h2>\n<ul>\n<li>To make predictions for 18 questions, I trained 3 LightGBM models, one for each level_group. The reason I did not build a model for each question was primarily to reduce inference time.</li>\n<li>Most of the features I have created are features based on the time difference between two consecutive actions. (More on this later.)</li>\n<li>The CV score was improved by about 0.002 by adding raw data published by the competition host.</li>\n<li>Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.</li>\n</ul>\n<p>The notebooks reproducing my submission are as follows:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-1-save-data\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-1-save-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-2-process-raw-data\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-2-process-raw-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-4-test-inference\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-4-test-inference</a></li>\n</ul>\n<h2>Feature Engineering</h2>\n<p>The six variables (level, name, event_name, room_fqid, fqid, and text) were concatenated as aggregation keys, and the time difference from the previous or following record was summed for each key and used as the feature. If written in pandas-like code, <br>\n<code>df.groupby(['level', 'name', 'event_name', 'room_fqid', 'fqid', 'text'])['elapsed_time_diff'].sum()</code></p>\n<p>In addition to the time difference from the previous or following records, the number of occurrences of each key is also added as a feature. Since these features can be calculated by sequentially reading the user's session, they can be calculated very efficiently by treating the data as the Python list instead of using Pandas.</p>\n<p>Furthermore, the record whose event_name is 'notification_click' is considered as a important event, and the time difference between the two events is added to the feature.</p>\n<p>The procedures for calculating these features can be found by reading the third published notebook.</p>\n<h2>Modeling</h2>\n<p>Since the variety of keys (combinations of six variables) is very large, I reduced features before training by excluding in advance rare combinations that appear only in a small number of sessions. However, since the number of features still amounted to several thousand, I first trained LightGBM with a large learning rate (0.1) and performed feature selection based on gain feature importance. The training was then performed again with a smaller learning rate (0.02) using 500 to 700 features.</p>\n<p>In the second modeling, raw data published by the host (<a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>) was included in the training. Although I was unable to reproduce the host's train.csv file completely, but I was able to reproduce it approximately using the second published notebook.</p>\n<p>Many of the sessions included in this data were different in nature from the competition data because they did not complete the game until the end. In fact, users who left the game midway through tended to have lower percentages of correct responses. To reflect this difference, the maximum level of each session was added as a feature.</p>\n<p>When training the model for the last level_group, I augmented the label of the second level_group, which contributed to the improvement in accuracy. I believe that the reason for this is that overfitting was suppressed by using more data to determine the split point when splitting nodes of decision trees. However, for the first and second level_groups, this data augmentation method did not contribute to improve accuracy in local validation.</p>\n<p>The CV/LB scores of my best submission is:</p>\n<ul>\n<li>CV: 0.7034</li>\n<li>Public LB: 0.703</li>\n<li>Private LB: 0.703</li>\n</ul>\n<h2>Other Remarks</h2>\n<ul>\n<li>For stability of evaluation, 4-fold CV was repeated three times with different seeds.</li>\n<li>Based on the validation results, the threshold was set at 0.625. No adjustment was made for each question.</li>\n<li>To reduce inference time, models trained in CV were not used, but retrained models using all data were used for inference.</li>\n</ul>",
  "messages": [
    {
      "id": 2322444,
      "postDate": "2023-06-29T09:35:17.037Z",
      "content": "<p>I am pleased to have fought the long and hard competition with all of you here.<br>\nHere I would like to outline my solution.</p>\n<h2>Overview</h2>\n<ul>\n<li>To make predictions for 18 questions, I trained 3 LightGBM models, one for each level_group. The reason I did not build a model for each question was primarily to reduce inference time.</li>\n<li>Most of the features I have created are features based on the time difference between two consecutive actions. (More on this later.)</li>\n<li>The CV score was improved by about 0.002 by adding raw data published by the competition host.</li>\n<li>Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.</li>\n</ul>\n<p>The notebooks reproducing my submission are as follows:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-1-save-data\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-1-save-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-2-process-raw-data\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-2-process-raw-data</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb</a></li>\n<li><a href=\"https://www.kaggle.com/code/rsakata/psp-4-test-inference\" target=\"_blank\">https://www.kaggle.com/code/rsakata/psp-4-test-inference</a></li>\n</ul>\n<h2>Feature Engineering</h2>\n<p>The six variables (level, name, event_name, room_fqid, fqid, and text) were concatenated as aggregation keys, and the time difference from the previous or following record was summed for each key and used as the feature. If written in pandas-like code, <br>\n<code>df.groupby(['level', 'name', 'event_name', 'room_fqid', 'fqid', 'text'])['elapsed_time_diff'].sum()</code></p>\n<p>In addition to the time difference from the previous or following records, the number of occurrences of each key is also added as a feature. Since these features can be calculated by sequentially reading the user's session, they can be calculated very efficiently by treating the data as the Python list instead of using Pandas.</p>\n<p>Furthermore, the record whose event_name is 'notification_click' is considered as a important event, and the time difference between the two events is added to the feature.</p>\n<p>The procedures for calculating these features can be found by reading the third published notebook.</p>\n<h2>Modeling</h2>\n<p>Since the variety of keys (combinations of six variables) is very large, I reduced features before training by excluding in advance rare combinations that appear only in a small number of sessions. However, since the number of features still amounted to several thousand, I first trained LightGBM with a large learning rate (0.1) and performed feature selection based on gain feature importance. The training was then performed again with a smaller learning rate (0.02) using 500 to 700 features.</p>\n<p>In the second modeling, raw data published by the host (<a href=\"https://fielddaylab.wisc.edu/opengamedata/\" target=\"_blank\">https://fielddaylab.wisc.edu/opengamedata/</a>) was included in the training. Although I was unable to reproduce the host's train.csv file completely, but I was able to reproduce it approximately using the second published notebook.</p>\n<p>Many of the sessions included in this data were different in nature from the competition data because they did not complete the game until the end. In fact, users who left the game midway through tended to have lower percentages of correct responses. To reflect this difference, the maximum level of each session was added as a feature.</p>\n<p>When training the model for the last level_group, I augmented the label of the second level_group, which contributed to the improvement in accuracy. I believe that the reason for this is that overfitting was suppressed by using more data to determine the split point when splitting nodes of decision trees. However, for the first and second level_groups, this data augmentation method did not contribute to improve accuracy in local validation.</p>\n<p>The CV/LB scores of my best submission is:</p>\n<ul>\n<li>CV: 0.7034</li>\n<li>Public LB: 0.703</li>\n<li>Private LB: 0.703</li>\n</ul>\n<h2>Other Remarks</h2>\n<ul>\n<li>For stability of evaluation, 4-fold CV was repeated three times with different seeds.</li>\n<li>Based on the validation results, the threshold was set at 0.625. No adjustment was made for each question.</li>\n<li>To reduce inference time, models trained in CV were not used, but retrained models using all data were used for inference.</li>\n</ul>",
      "rawMarkdown": "I am pleased to have fought the long and hard competition with all of you here.\nHere I would like to outline my solution.\n\n## Overview\n\n- To make predictions for 18 questions, I trained 3 LightGBM models, one for each level_group. The reason I did not build a model for each question was primarily to reduce inference time.\n- Most of the features I have created are features based on the time difference between two consecutive actions. (More on this later.)\n- The CV score was improved by about 0.002 by adding raw data published by the competition host.\n- Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.\n\nThe notebooks reproducing my submission are as follows:\n- https://www.kaggle.com/code/rsakata/psp-1-save-data\n- https://www.kaggle.com/code/rsakata/psp-2-process-raw-data\n- https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb\n- https://www.kaggle.com/code/rsakata/psp-4-test-inference\n\n## Feature Engineering\n\nThe six variables (level, name, event_name, room_fqid, fqid, and text) were concatenated as aggregation keys, and the time difference from the previous or following record was summed for each key and used as the feature. If written in pandas-like code, \n`df.groupby(['level', 'name', 'event_name', 'room_fqid', 'fqid', 'text'])['elapsed_time_diff'].sum()`\n\nIn addition to the time difference from the previous or following records, the number of occurrences of each key is also added as a feature. Since these features can be calculated by sequentially reading the user's session, they can be calculated very efficiently by treating the data as the Python list instead of using Pandas.\n\nFurthermore, the record whose event_name is 'notification_click' is considered as a important event, and the time difference between the two events is added to the feature.\n\nThe procedures for calculating these features can be found by reading the third published notebook.\n\n## Modeling\n\nSince the variety of keys (combinations of six variables) is very large, I reduced features before training by excluding in advance rare combinations that appear only in a small number of sessions. However, since the number of features still amounted to several thousand, I first trained LightGBM with a large learning rate (0.1) and performed feature selection based on gain feature importance. The training was then performed again with a smaller learning rate (0.02) using 500 to 700 features.\n\nIn the second modeling, raw data published by the host (https://fielddaylab.wisc.edu/opengamedata/) was included in the training. Although I was unable to reproduce the host's train.csv file completely, but I was able to reproduce it approximately using the second published notebook.\n\nMany of the sessions included in this data were different in nature from the competition data because they did not complete the game until the end. In fact, users who left the game midway through tended to have lower percentages of correct responses. To reflect this difference, the maximum level of each session was added as a feature.\n\nWhen training the model for the last level_group, I augmented the label of the second level_group, which contributed to the improvement in accuracy. I believe that the reason for this is that overfitting was suppressed by using more data to determine the split point when splitting nodes of decision trees. However, for the first and second level_groups, this data augmentation method did not contribute to improve accuracy in local validation.\n\nThe CV/LB scores of my best submission is:\n- CV: 0.7034\n- Public LB: 0.703\n- Private LB: 0.703\n\n## Other Remarks\n\n- For stability of evaluation, 4-fold CV was repeated three times with different seeds.\n- Based on the validation results, the threshold was set at 0.625. No adjustment was made for each question.\n- To reduce inference time, models trained in CV were not used, but retrained models using all data were used for inference.\n",
      "votes": 102
    },
    {
      "id": 2322540,
      "postDate": "2023-06-29T10:36:16.857Z",
      "content": "<blockquote>\n  <p>Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.</p>\n</blockquote>\n<p>Wow, it's shocking to know you can build a model with such level of efficiency and competitiveness at the same time.</p>\n<p>I observed during my daily work that the list structure is far more efficient than dataframe when possible to use, but the grammar is very tricky and very difficult to master for large programs. It's very generous of you to make those notebooks public, I will learn a lot from them for sure.</p>",
      "rawMarkdown": ">Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.\n\nWow, it's shocking to know you can build a model with such level of efficiency and competitiveness at the same time.\n\nI observed during my daily work that the list structure is far more efficient than dataframe when possible to use, but the grammar is very tricky and very difficult to master for large programs. It's very generous of you to make those notebooks public, I will learn a lot from them for sure.",
      "votes": 3
    },
    {
      "id": 2337344,
      "postDate": "2023-07-10T04:24:47.510Z",
      "content": "<p>Congratulations on your 7th place solution in the competition! <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> I notice in the second notebook when processing raw datasets, different data corresponding to different file_date is processed differently. For example: file_date == \"20220701\"; file_date &lt;= \"20200901\";and else are processed in a different way. May I know how you realize the split point and the corresponding appropriate processing method for datasets in different stages? </p>",
      "rawMarkdown": "Congratulations on your 7th place solution in the competition! @rsakata I notice in the second notebook when processing raw datasets, different data corresponding to different file_date is processed differently. For example: file_date == \"20220701\"; file_date <= \"20200901\";and else are processed in a different way. May I know how you realize the split point and the corresponding appropriate processing method for datasets in different stages? ",
      "votes": 1,
      "replies": [
        {
          "id": 2337901,
          "postDate": "2023-07-10T12:16:34.930Z",
          "content": "<p>Another question is that, in your 3rd notebook, in the Feature Engineering cell, I found several codes as follows:</p>\n<pre><code>     i  ((df_session)):\n         keys[i]  self.valid_keys:   \n            \n            feature_idx = self.map_key[keys[i]]\n             level_group &lt;= :\n                features[feature_idx] += \n            \n            feature_idx += (self.map_key)\n             level_group &lt;=   i &gt; :\n                features[feature_idx] += values_time[i] - values_time[i-]\n            \n            feature_idx += (self.map_key)\n             i &lt; (df_session) - :\n                features[feature_idx] += values_time[i+] - values_time[i]\n</code></pre>\n<p>Does these code mean that the third level_group(level_group=2) is excluded during calculating the #cound,#bdiff,and #fdiff,  and why is that?   <br>\nAgain thank you for your great job and share, hope you could answer.</p>",
          "rawMarkdown": "Another question is that, in your 3rd notebook, in the Feature Engineering cell, I found several codes as follows:\n\n        for i in range(len(df_session)):\n            if keys[i] in self.valid_keys:   \n                # count\n                feature_idx = self.map_key[keys[i]]\n                if level_group <= 1:\n                    features[feature_idx] += 1\n                # bdiff\n                feature_idx += len(self.map_key)\n                if level_group <= 1 and i > 0:\n                    features[feature_idx] += values_time[i] - values_time[i-1]\n                # fdiff\n                feature_idx += len(self.map_key)\n                if i < len(df_session) - 1:\n                    features[feature_idx] += values_time[i+1] - values_time[i]\nDoes these code mean that the third level_group(level_group=2) is excluded during calculating the #cound,#bdiff,and #fdiff,  and why is that?   \nAgain thank you for your great job and share, hope you could answer.",
          "replies": [
            {
              "id": 2338131,
              "postDate": "2023-07-10T14:52:30.093Z",
              "content": "<p>Regarding the first question, it is simply the result of a series of modifications to the program to eliminate errors that were output due to changes in the data format.</p>\n<p>Regarding the second question, note that for the third level_group, bdiff and count are skipped, but fdiff is calculated. The reason for this is that the validation results showed that the contribution was small for bdiff and count, and the validation score for the third level_group was almost the same without including it.</p>",
              "rawMarkdown": "Regarding the first question, it is simply the result of a series of modifications to the program to eliminate errors that were output due to changes in the data format.\n\nRegarding the second question, note that for the third level_group, bdiff and count are skipped, but fdiff is calculated. The reason for this is that the validation results showed that the contribution was small for bdiff and count, and the validation score for the third level_group was almost the same without including it."
            },
            {
              "id": 2338190,
              "postDate": "2023-07-10T15:47:15.593Z",
              "content": "<p>Thanks  a lot! That really helps me!</p>",
              "rawMarkdown": "Thanks  a lot! That really helps me!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2333948,
      "postDate": "2023-07-07T10:03:32.580Z",
      "content": "<p>Congratulations!!<br>\ni want to be like you,just study hard!<br>\nThank you for sharing tips.</p>",
      "rawMarkdown": "Congratulations!!\ni want to be like you,just study hard!\nThank you for sharing tips.",
      "votes": 1
    },
    {
      "id": 2331492,
      "postDate": "2023-07-05T15:29:00.553Z",
      "content": "<p>It's amazing just to be in the top 10! You'll most probably ace the next one 👍</p>",
      "rawMarkdown": "It's amazing just to be in the top 10! You'll most probably ace the next one 👍",
      "votes": 1
    },
    {
      "id": 2329664,
      "postDate": "2023-07-04T12:24:06.603Z",
      "content": "<p>Congratulations great job solo cash gold! Great work hope to learn from your excellent work.</p>",
      "rawMarkdown": "Congratulations great job solo cash gold! Great work hope to learn from your excellent work.",
      "votes": 1
    },
    {
      "id": 2324717,
      "postDate": "2023-06-30T20:24:20.727Z",
      "content": "<p>Good experience</p>",
      "rawMarkdown": "Good experience",
      "votes": 1
    },
    {
      "id": 2324256,
      "postDate": "2023-06-30T13:51:29.937Z",
      "content": "<p>Thank you for sharing your solution.<br>\nThis competition was my first experience participating in Kaggle.<br>\nSeeing your solution has further motivated me to continue participating in Kaggle.</p>",
      "rawMarkdown": "Thank you for sharing your solution.\nThis competition was my first experience participating in Kaggle.\nSeeing your solution has further motivated me to continue participating in Kaggle.",
      "votes": 1
    },
    {
      "id": 2323013,
      "postDate": "2023-06-29T16:15:06.210Z",
      "content": "<p>Congratulations, Jack! And I really thank you so much for sharing your great value solution. I am amazed at the model's score, efficiency, and stability all in one. I will make good use of it for my improvements.</p>",
      "rawMarkdown": "Congratulations, Jack! And I really thank you so much for sharing your great value solution. I am amazed at the model's score, efficiency, and stability all in one. I will make good use of it for my improvements.",
      "votes": 1
    },
    {
      "id": 2322804,
      "postDate": "2023-06-29T14:01:49.643Z",
      "content": "<p>Thank you for publishing the solution.</p>\n<p>I have one question. <br>\nI am wondering what process you used to identify \"notification_click\" as important.</p>",
      "rawMarkdown": "Thank you for publishing the solution.\n\nI have one question. \nI am wondering what process you used to identify \"notification_click\" as important.",
      "votes": 1,
      "replies": [
        {
          "id": 2323402,
          "postDate": "2023-06-29T23:28:26.723Z",
          "content": "<p>Thanks for your comment.<br>\nIn looking at the data with my own eyes, \"notification_click\" seemed to be some sort of checkpoint event.</p>",
          "rawMarkdown": "Thanks for your comment.\nIn looking at the data with my own eyes, \"notification_click\" seemed to be some sort of checkpoint event.",
          "votes": 3,
          "replies": [
            {
              "id": 2323467,
              "postDate": "2023-06-30T01:17:00.887Z",
              "content": "<p>Thank you for your response.</p>\n<p>I was in the silver zone this time, but thanks to you I learned that EDA is also important to get into the gold zone.</p>\n<p>Thank you.</p>",
              "rawMarkdown": "Thank you for your response.\n\nI was in the silver zone this time, but thanks to you I learned that EDA is also important to get into the gold zone.\n\nThank you.",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2322759,
      "postDate": "2023-06-29T13:34:52.950Z",
      "content": "<p><a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> Congratulations on your impressive 7th place finish!. Your feature engineering technique based on time differences between consecutive actions is intriguing.</p>",
      "rawMarkdown": "@rsakata Congratulations on your impressive 7th place finish!. Your feature engineering technique based on time differences between consecutive actions is intriguing.",
      "votes": 1
    },
    {
      "id": 2322692,
      "postDate": "2023-06-29T12:45:53.093Z",
      "content": "<p>Congratulations!<br>\nThank you for sharing your solution so quickly.<br>\nI am amazed at this solution that takes about 3 minutes of inference to get a high score!</p>",
      "rawMarkdown": "Congratulations!\nThank you for sharing your solution so quickly.\nI am amazed at this solution that takes about 3 minutes of inference to get a high score!",
      "votes": 1
    },
    {
      "id": 2322560,
      "postDate": "2023-06-29T10:42:54.300Z",
      "content": "<p>Amazing work! And congratulations on your gold medal!</p>\n<p>I have some questions about your pipeline - hope you could answer</p>\n<ol>\n<li>How do you train one LGBM to predict multiple questions? Is LGBM not similar to other tree boosting methods (XGB, Catboos, etc.) which have only one output head? I have not had a lot of experience with LGBM.</li>\n<li>How did you augment the second level_group label?</li>\n<li>Did you use early stopping when doing CV? If so, since each fold will have a different number of trees, how did you determine the <code>num_iterations</code> for the model trained on the entire dataset?</li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Amazing work! And congratulations on your gold medal!\n\nI have some questions about your pipeline - hope you could answer\n1. How do you train one LGBM to predict multiple questions? Is LGBM not similar to other tree boosting methods (XGB, Catboos, etc.) which have only one output head? I have not had a lot of experience with LGBM.\n2. How did you augment the second level_group label?\n3. Did you use early stopping when doing CV? If so, since each fold will have a different number of trees, how did you determine the `num_iterations` for the model trained on the entire dataset?\n\nThanks!",
      "votes": 1,
      "replies": [
        {
          "id": 2322665,
          "postDate": "2023-06-29T12:28:01.447Z",
          "content": "<p>Thank you for your comment and questions. Here are the answers.</p>\n<ol>\n<li>After computing the features, duplicate them for the number of questions and add the question numbers to the features.</li>\n<li>Same as 1.</li>\n<li>Yes. I used the simple average of best iterations when retrain models.</li>\n</ol>",
          "rawMarkdown": "Thank you for your comment and questions. Here are the answers.\n\n1. After computing the features, duplicate them for the number of questions and add the question numbers to the features.\n2. Same as 1.\n3. Yes. I used the simple average of best iterations when retrain models.",
          "votes": 4,
          "replies": [
            {
              "id": 2322856,
              "postDate": "2023-06-29T14:29:27.060Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2323998,
      "postDate": "2023-06-30T09:57:26.530Z",
      "content": "<p>Very solid score between CV, public LB and private LB. Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> </p>",
      "rawMarkdown": "Very solid score between CV, public LB and private LB. Congrats on results and thanks for sharing solution @rsakata ",
      "votes": 2
    },
    {
      "id": 2323645,
      "postDate": "2023-06-30T05:17:09.300Z",
      "content": "<p>Congratulations on your 7th place solution in the competition! Your feature engineering technique using time differences between consecutive actions is interesting. Including the host's raw data and adjusting for users who left the game midway were smart choices. Your CV and leaderboard scores show a strong performance. Well done!</p>",
      "rawMarkdown": "Congratulations on your 7th place solution in the competition! Your feature engineering technique using time differences between consecutive actions is interesting. Including the host's raw data and adjusting for users who left the game midway were smart choices. Your CV and leaderboard scores show a strong performance. Well done!",
      "votes": 2
    },
    {
      "id": 2322837,
      "postDate": "2023-06-29T14:18:46.580Z",
      "content": "<p>Thanks for sharing, this is really impressive. I learned a lot!</p>",
      "rawMarkdown": "Thanks for sharing, this is really impressive. I learned a lot!",
      "votes": 2
    },
    {
      "id": 2322684,
      "postDate": "2023-06-29T12:40:06.670Z",
      "content": "<p>Though, i was outside without my machine with me and trying to beat traffic on my way home, having it in mind to ensemble my best two models (Catboost, GradientBoost) and with  just 6 hours to go in this great competition as at that time. It's quite unfortunate, i couldn't implement this ensemble method due to this time constraint, but still to me i still consider <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> solution as the best ever, talk of its neatness, easy to understand, implemented ideas, creativity, and simplicity of this solution, i must say. Awesome !!!<br>\nA big congratulations 🎉🎊🍾㊗️ to <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> for this lovely solution. <br>\nThanks for sharing. </p>",
      "rawMarkdown": "Though, i was outside without my machine with me and trying to beat traffic on my way home, having it in mind to ensemble my best two models (Catboost, GradientBoost) and with  just 6 hours to go in this great competition as at that time. It's quite unfortunate, i couldn't implement this ensemble method due to this time constraint, but still to me i still consider @rsakata solution as the best ever, talk of its neatness, easy to understand, implemented ideas, creativity, and simplicity of this solution, i must say. Awesome !!!\nA big congratulations 🎉🎊🍾㊗️ to @rsakata for this lovely solution. \nThanks for sharing. ",
      "votes": 2
    },
    {
      "id": 2322645,
      "postDate": "2023-06-29T12:02:41.303Z",
      "content": "<p>Well done! Thanks for sharing your solution. We too found that 0.625 threshold for all questions was best.</p>",
      "rawMarkdown": "Well done! Thanks for sharing your solution. We too found that 0.625 threshold for all questions was best.",
      "votes": 2,
      "replies": [
        {
          "id": 2324718,
          "postDate": "2023-06-30T20:24:55.843Z",
          "content": "<p>Yes he is the best</p>",
          "rawMarkdown": "Yes he is the best"
        }
      ]
    },
    {
      "id": 2322553,
      "postDate": "2023-06-29T10:40:03.097Z",
      "content": "<p>congrats! that's impressive</p>",
      "rawMarkdown": "congrats! that's impressive",
      "votes": 2
    },
    {
      "id": 2322485,
      "postDate": "2023-06-29T10:05:41.313Z",
      "content": "<p><a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> congrats with 7 place and gold medal! It is very cool!</p>",
      "rawMarkdown": "@rsakata congrats with 7 place and gold medal! It is very cool!",
      "votes": 2
    },
    {
      "id": 2322450,
      "postDate": "2023-06-29T09:38:56.367Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> for securing 7th place 🎉</p>",
      "rawMarkdown": "Congratulations @rsakata for securing 7th place 🎉",
      "votes": 2
    },
    {
      "id": 2323635,
      "postDate": "2023-06-30T05:02:06.613Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 2323632,
      "postDate": "2023-06-30T05:01:03.543Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2324137,
      "postDate": "2023-06-30T12:08:48.250Z",
      "content": "<p>thank you and congrats</p>",
      "rawMarkdown": "thank you and congrats",
      "votes": 1
    },
    {
      "id": 2323739,
      "postDate": "2023-06-30T06:34:02.577Z",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2322540,
      "author_name": "Ya Xu",
      "author_url": "",
      "post_date": "2023-06-29T10:36:16.857000",
      "content": "<blockquote>\n  <p>Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.</p>\n</blockquote>\n<p>Wow, it's shocking to know you can build a model with such level of efficiency and competitiveness at the same time.</p>\n<p>I observed during my daily work that the list structure is far more efficient than dataframe when possible to use, but the grammar is very tricky and very difficult to master for large programs. It's very generous of you to make those notebooks public, I will learn a lot from them for sure.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2337344,
      "author_name": "song xiaozhe",
      "author_url": "",
      "post_date": "2023-07-10T04:24:47.510000",
      "content": "<p>Congratulations on your 7th place solution in the competition! <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> I notice in the second notebook when processing raw datasets, different data corresponding to different file_date is processed differently. For example: file_date == \"20220701\"; file_date &lt;= \"20200901\";and else are processed in a different way. May I know how you realize the split point and the corresponding appropriate processing method for datasets in different stages? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2337901,
          "author_name": "song xiaozhe",
          "author_url": "",
          "post_date": "2023-07-10T12:16:34.930000",
          "content": "<p>Another question is that, in your 3rd notebook, in the Feature Engineering cell, I found several codes as follows:</p>\n<pre><code>     i  ((df_session)):\n         keys[i]  self.valid_keys:   \n            \n            feature_idx = self.map_key[keys[i]]\n             level_group &lt;= :\n                features[feature_idx] += \n            \n            feature_idx += (self.map_key)\n             level_group &lt;=   i &gt; :\n                features[feature_idx] += values_time[i] - values_time[i-]\n            \n            feature_idx += (self.map_key)\n             i &lt; (df_session) - :\n                features[feature_idx] += values_time[i+] - values_time[i]\n</code></pre>\n<p>Does these code mean that the third level_group(level_group=2) is excluded during calculating the #cound,#bdiff,and #fdiff,  and why is that?   <br>\nAgain thank you for your great job and share, hope you could answer.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2338131,
              "author_name": "Jack (Japan)",
              "author_url": "",
              "post_date": "2023-07-10T14:52:30.093000",
              "content": "<p>Regarding the first question, it is simply the result of a series of modifications to the program to eliminate errors that were output due to changes in the data format.</p>\n<p>Regarding the second question, note that for the third level_group, bdiff and count are skipped, but fdiff is calculated. The reason for this is that the validation results showed that the contribution was small for bdiff and count, and the validation score for the third level_group was almost the same without including it.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2338190,
              "author_name": "song xiaozhe",
              "author_url": "",
              "post_date": "2023-07-10T15:47:15.593000",
              "content": "<p>Thanks  a lot! That really helps me!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2333948,
      "author_name": "Kota",
      "author_url": "",
      "post_date": "2023-07-07T10:03:32.580000",
      "content": "<p>Congratulations!!<br>\ni want to be like you,just study hard!<br>\nThank you for sharing tips.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2331492,
      "author_name": "Muhammad Awn",
      "author_url": "",
      "post_date": "2023-07-05T15:29:00.553000",
      "content": "<p>It's amazing just to be in the top 10! You'll most probably ace the next one 👍</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2329664,
      "author_name": "Hoping",
      "author_url": "",
      "post_date": "2023-07-04T12:24:06.603000",
      "content": "<p>Congratulations great job solo cash gold! Great work hope to learn from your excellent work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324717,
      "author_name": "Andriansyah1",
      "author_url": "",
      "post_date": "2023-06-30T20:24:20.727000",
      "content": "<p>Good experience</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324256,
      "author_name": "takeshi miura",
      "author_url": "",
      "post_date": "2023-06-30T13:51:29.937000",
      "content": "<p>Thank you for sharing your solution.<br>\nThis competition was my first experience participating in Kaggle.<br>\nSeeing your solution has further motivated me to continue participating in Kaggle.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323013,
      "author_name": "TAK",
      "author_url": "",
      "post_date": "2023-06-29T16:15:06.210000",
      "content": "<p>Congratulations, Jack! And I really thank you so much for sharing your great value solution. I am amazed at the model's score, efficiency, and stability all in one. I will make good use of it for my improvements.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2322804,
      "author_name": "konumaru",
      "author_url": "",
      "post_date": "2023-06-29T14:01:49.643000",
      "content": "<p>Thank you for publishing the solution.</p>\n<p>I have one question. <br>\nI am wondering what process you used to identify \"notification_click\" as important.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2323402,
          "author_name": "Jack (Japan)",
          "author_url": "",
          "post_date": "2023-06-29T23:28:26.723000",
          "content": "<p>Thanks for your comment.<br>\nIn looking at the data with my own eyes, \"notification_click\" seemed to be some sort of checkpoint event.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2323467,
              "author_name": "konumaru",
              "author_url": "",
              "post_date": "2023-06-30T01:17:00.887000",
              "content": "<p>Thank you for your response.</p>\n<p>I was in the silver zone this time, but thanks to you I learned that EDA is also important to get into the gold zone.</p>\n<p>Thank you.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2322759,
      "author_name": "Akshay Vyas",
      "author_url": "",
      "post_date": "2023-06-29T13:34:52.950000",
      "content": "<p><a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> Congratulations on your impressive 7th place finish!. Your feature engineering technique based on time differences between consecutive actions is intriguing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2322692,
      "author_name": "ino way",
      "author_url": "",
      "post_date": "2023-06-29T12:45:53.093000",
      "content": "<p>Congratulations!<br>\nThank you for sharing your solution so quickly.<br>\nI am amazed at this solution that takes about 3 minutes of inference to get a high score!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2322560,
      "author_name": "Hoang Nguyen",
      "author_url": "",
      "post_date": "2023-06-29T10:42:54.300000",
      "content": "<p>Amazing work! And congratulations on your gold medal!</p>\n<p>I have some questions about your pipeline - hope you could answer</p>\n<ol>\n<li>How do you train one LGBM to predict multiple questions? Is LGBM not similar to other tree boosting methods (XGB, Catboos, etc.) which have only one output head? I have not had a lot of experience with LGBM.</li>\n<li>How did you augment the second level_group label?</li>\n<li>Did you use early stopping when doing CV? If so, since each fold will have a different number of trees, how did you determine the <code>num_iterations</code> for the model trained on the entire dataset?</li>\n</ol>\n<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2322665,
          "author_name": "Jack (Japan)",
          "author_url": "",
          "post_date": "2023-06-29T12:28:01.447000",
          "content": "<p>Thank you for your comment and questions. Here are the answers.</p>\n<ol>\n<li>After computing the features, duplicate them for the number of questions and add the question numbers to the features.</li>\n<li>Same as 1.</li>\n<li>Yes. I used the simple average of best iterations when retrain models.</li>\n</ol>",
          "votes": 4,
          "replies": [
            {
              "id": 2322856,
              "author_name": "Hoang Nguyen",
              "author_url": "",
              "post_date": "2023-06-29T14:29:27.060000",
              "content": "<p>Thank you!</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2323998,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2023-06-30T09:57:26.530000",
      "content": "<p>Very solid score between CV, public LB and private LB. Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2323645,
      "author_name": "Pooja Chauhan",
      "author_url": "",
      "post_date": "2023-06-30T05:17:09.300000",
      "content": "<p>Congratulations on your 7th place solution in the competition! Your feature engineering technique using time differences between consecutive actions is interesting. Including the host's raw data and adjusting for users who left the game midway were smart choices. Your CV and leaderboard scores show a strong performance. Well done!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2322837,
      "author_name": "Xiao-Su (Frank) Hu",
      "author_url": "",
      "post_date": "2023-06-29T14:18:46.580000",
      "content": "<p>Thanks for sharing, this is really impressive. I learned a lot!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2322684,
      "author_name": "MICADEE",
      "author_url": "",
      "post_date": "2023-06-29T12:40:06.670000",
      "content": "<p>Though, i was outside without my machine with me and trying to beat traffic on my way home, having it in mind to ensemble my best two models (Catboost, GradientBoost) and with  just 6 hours to go in this great competition as at that time. It's quite unfortunate, i couldn't implement this ensemble method due to this time constraint, but still to me i still consider <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> solution as the best ever, talk of its neatness, easy to understand, implemented ideas, creativity, and simplicity of this solution, i must say. Awesome !!!<br>\nA big congratulations 🎉🎊🍾㊗️ to <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> for this lovely solution. <br>\nThanks for sharing. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2322645,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2023-06-29T12:02:41.303000",
      "content": "<p>Well done! Thanks for sharing your solution. We too found that 0.625 threshold for all questions was best.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2324718,
          "author_name": "Andriansyah1",
          "author_url": "",
          "post_date": "2023-06-30T20:24:55.843000",
          "content": "<p>Yes he is the best</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2322553,
      "author_name": "empty",
      "author_url": "",
      "post_date": "2023-06-29T10:40:03.097000",
      "content": "<p>congrats! that's impressive</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2322485,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "2023-06-29T10:05:41.313000",
      "content": "<p><a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> congrats with 7 place and gold medal! It is very cool!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2322450,
      "author_name": "Swapnil Chowdhury",
      "author_url": "",
      "post_date": "2023-06-29T09:38:56.367000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> for securing 7th place 🎉</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2323635,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T05:02:06.613000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323632,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-30T05:01:03.543000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2324137,
      "author_name": "ls",
      "author_url": "",
      "post_date": "2023-06-30T12:08:48.250000",
      "content": "<p>thank you and congrats</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323739,
      "author_name": "Johny Aldean",
      "author_url": "",
      "post_date": "2023-06-30T06:34:02.577000",
      "content": "<p>Congratulations! Thank you for sharing!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2322444": "I am pleased to have fought the long and hard competition with all of you here.\nHere I would like to outline my solution.\n\n## Overview\n\n- To make predictions for 18 questions, I trained 3 LightGBM models, one for each level_group. The reason I did not build a model for each question was primarily to reduce inference time.\n- Most of the features I have created are features based on the time difference between two consecutive actions. (More on this later.)\n- The CV score was improved by about 0.002 by adding raw data published by the competition host.\n- Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.\n\nThe notebooks reproducing my submission are as follows:\n- https://www.kaggle.com/code/rsakata/psp-1-save-data\n- https://www.kaggle.com/code/rsakata/psp-2-process-raw-data\n- https://www.kaggle.com/code/rsakata/psp-3-fe-and-train-lgb\n- https://www.kaggle.com/code/rsakata/psp-4-test-inference\n\n## Feature Engineering\n\nThe six variables (level, name, event_name, room_fqid, fqid, and text) were concatenated as aggregation keys, and the time difference from the previous or following record was summed for each key and used as the feature. If written in pandas-like code, \n`df.groupby(['level', 'name', 'event_name', 'room_fqid', 'fqid', 'text'])['elapsed_time_diff'].sum()`\n\nIn addition to the time difference from the previous or following records, the number of occurrences of each key is also added as a feature. Since these features can be calculated by sequentially reading the user's session, they can be calculated very efficiently by treating the data as the Python list instead of using Pandas.\n\nFurthermore, the record whose event_name is 'notification_click' is considered as a important event, and the time difference between the two events is added to the feature.\n\nThe procedures for calculating these features can be found by reading the third published notebook.\n\n## Modeling\n\nSince the variety of keys (combinations of six variables) is very large, I reduced features before training by excluding in advance rare combinations that appear only in a small number of sessions. However, since the number of features still amounted to several thousand, I first trained LightGBM with a large learning rate (0.1) and performed feature selection based on gain feature importance. The training was then performed again with a smaller learning rate (0.02) using 500 to 700 features.\n\nIn the second modeling, raw data published by the host (https://fielddaylab.wisc.edu/opengamedata/) was included in the training. Although I was unable to reproduce the host's train.csv file completely, but I was able to reproduce it approximately using the second published notebook.\n\nMany of the sessions included in this data were different in nature from the competition data because they did not complete the game until the end. In fact, users who left the game midway through tended to have lower percentages of correct responses. To reflect this difference, the maximum level of each session was added as a feature.\n\nWhen training the model for the last level_group, I augmented the label of the second level_group, which contributed to the improvement in accuracy. I believe that the reason for this is that overfitting was suppressed by using more data to determine the split point when splitting nodes of decision trees. However, for the first and second level_groups, this data augmentation method did not contribute to improve accuracy in local validation.\n\nThe CV/LB scores of my best submission is:\n- CV: 0.7034\n- Public LB: 0.703\n- Private LB: 0.703\n\n## Other Remarks\n\n- For stability of evaluation, 4-fold CV was repeated three times with different seeds.\n- Based on the validation results, the threshold was set at 0.625. No adjustment was made for each question.\n- To reduce inference time, models trained in CV were not used, but retrained models using all data were used for inference.\n",
    "2322540": ">Unexpectedly, the submission for the Efficiency Prize had the best score in Private Leaderboard amoung the selected sumissions. The inference time of that is approximately 3 minutes.\n\nWow, it's shocking to know you can build a model with such level of efficiency and competitiveness at the same time.\n\nI observed during my daily work that the list structure is far more efficient than dataframe when possible to use, but the grammar is very tricky and very difficult to master for large programs. It's very generous of you to make those notebooks public, I will learn a lot from them for sure.",
    "2337344": "Congratulations on your 7th place solution in the competition! @rsakata I notice in the second notebook when processing raw datasets, different data corresponding to different file_date is processed differently. For example: file_date == \"20220701\"; file_date <= \"20200901\";and else are processed in a different way. May I know how you realize the split point and the corresponding appropriate processing method for datasets in different stages? ",
    "2333948": "Congratulations!!\ni want to be like you,just study hard!\nThank you for sharing tips.",
    "2331492": "It's amazing just to be in the top 10! You'll most probably ace the next one 👍",
    "2329664": "Congratulations great job solo cash gold! Great work hope to learn from your excellent work.",
    "2324717": "Good experience",
    "2324256": "Thank you for sharing your solution.\nThis competition was my first experience participating in Kaggle.\nSeeing your solution has further motivated me to continue participating in Kaggle.",
    "2323013": "Congratulations, Jack! And I really thank you so much for sharing your great value solution. I am amazed at the model's score, efficiency, and stability all in one. I will make good use of it for my improvements.",
    "2322804": "Thank you for publishing the solution.\n\nI have one question. \nI am wondering what process you used to identify \"notification_click\" as important.",
    "2322759": "@rsakata Congratulations on your impressive 7th place finish!. Your feature engineering technique based on time differences between consecutive actions is intriguing.",
    "2322692": "Congratulations!\nThank you for sharing your solution so quickly.\nI am amazed at this solution that takes about 3 minutes of inference to get a high score!",
    "2322560": "Amazing work! And congratulations on your gold medal!\n\nI have some questions about your pipeline - hope you could answer\n1. How do you train one LGBM to predict multiple questions? Is LGBM not similar to other tree boosting methods (XGB, Catboos, etc.) which have only one output head? I have not had a lot of experience with LGBM.\n2. How did you augment the second level_group label?\n3. Did you use early stopping when doing CV? If so, since each fold will have a different number of trees, how did you determine the `num_iterations` for the model trained on the entire dataset?\n\nThanks!",
    "2323998": "Very solid score between CV, public LB and private LB. Congrats on results and thanks for sharing solution @rsakata ",
    "2323645": "Congratulations on your 7th place solution in the competition! Your feature engineering technique using time differences between consecutive actions is interesting. Including the host's raw data and adjusting for users who left the game midway were smart choices. Your CV and leaderboard scores show a strong performance. Well done!",
    "2322837": "Thanks for sharing, this is really impressive. I learned a lot!",
    "2322684": "Though, i was outside without my machine with me and trying to beat traffic on my way home, having it in mind to ensemble my best two models (Catboost, GradientBoost) and with  just 6 hours to go in this great competition as at that time. It's quite unfortunate, i couldn't implement this ensemble method due to this time constraint, but still to me i still consider @rsakata solution as the best ever, talk of its neatness, easy to understand, implemented ideas, creativity, and simplicity of this solution, i must say. Awesome !!!\nA big congratulations 🎉🎊🍾㊗️ to @rsakata for this lovely solution. \nThanks for sharing. ",
    "2322645": "Well done! Thanks for sharing your solution. We too found that 0.625 threshold for all questions was best.",
    "2322553": "congrats! that's impressive",
    "2322485": "@rsakata congrats with 7 place and gold medal! It is very cool!",
    "2322450": "Congratulations @rsakata for securing 7th place 🎉",
    "2323635": "",
    "2323632": "",
    "2324137": "thank you and congrats",
    "2323739": "Congratulations! Thank you for sharing!"
  }
}