{
  "id": 209629,
  "title": "[62nd solution, LB: 0.801] Single light GBM with 113 features",
  "url": "/competitions/riiid-test-answer-prediction/writeups/tomoo-inubushi-62nd-solution-lb-0-801-single-light",
  "author_name": "",
  "post_date": "2021-01-08T04:39:21.250Z",
  "votes": 33,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Many many thanks to the organizers and all kagglers.<br>\nThroughout this competition, I realized good kagglers are not just smart data scientists, but also proficient engineers. They have deep and broad knowledge about how to speed up processing, handle data memory efficiently, and debug submission error. I learned a lot from their notebook and comments. <br>\nI achieved 62nd (0.801 CV /0.799 public / 0.801 private) and got my first medal with a single LGB model. I tried blending with catboost and SAKT models, but it did not work well for me. I think it is mainly due to the time constraint to refine them.<br>\nThe code and model are opened <a href=\"https://www.kaggle.com/tomooinubushi/62nd-solution-lightgbm-single-model-lb-0-801\" target=\"_blank\">here</a>.</p>\n<p>I mainly focused on feature engineering and used 113 features. The feature importance of them was as follows</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fc09711551a883f6995032d91b6408bd9%2F__results___42_0.png?generation=1610076363046071&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F35c076ded96a2ef0d40ec307d0b8da51%2F__results___42_1.png?generation=1610076387230945&amp;alt=media\" alt=\"\"></p>\n<p>The meanings of abbreviations were as follows </p>\n<ul>\n<li>acc: answered_correctly</li>\n<li>exp: prior_question_had_explanation</li>\n<li>time: prior_question_elapsed_time</li>\n<li>count: number of trial</li>\n<li>count_lecture: number of lecture</li>\n<li>tfl: timestamp from last nth trial to that from n+1th trial</li>\n<li>dif: difficulty (1 - content-wise accuracy)</li>\n<li>dpt: difficulty point (add difficulty if users answer was correct)</li>\n<li>avg: average</li>\n<li>std: standard deviation</li>\n<li>sum: sum</li>\n<li>hist: historical average with temporal decay of gamma 0.75 </li>\n<li>hist2: historical average with temporal decay of gamma 0.25</li>\n<li>u: user-wise</li>\n<li>t: tag-wise</li>\n<li>tg: taggroup-wise</li>\n<li>p: part-wise</li>\n<li>c: content-wise</li>\n<li>b: bundle-wise</li>\n<li>uc: user-content-wise</li>\n<li>ub: user-bundle-wise</li>\n<li>up: user-part-wise</li>\n<li>time_per_count_u: timestamp/count_u</li>\n<li>count_u_nondiagnostic: user-wise count of non-diagnostic questions</li>\n<li>exp_avg_u_corrected: user-wise average of prior_question_had_explanation corrected for diagnostic questions. Th user cannot read explanation for diagnostic questions.</li>\n<li>pp_ratio: ratio of paid part (part1, 3, 4, 6, 7) trials per all trials, only paid users of SANTA app can solve them if it is not diagnostic questions</li>\n<li>listening_ratio: ratio of listening part trials per all trials</li>\n<li>relative_dif: relative difficulty of questions calculated with acc_avg_c / dif_avg_u</li>\n<li>notacc_sum: sum of failed trials</li>\n<li>wp_ts: trueskill feature of win probability</li>\n<li>umu_ts, usigma_ts, qmu_ts, qsigma_ts: trueskill features of mu, and sigma of user and content</li>\n</ul>\n<p>I learned a lot from many notebooks and comments of other kagglers. </p>\n<ul>\n<li>CV strategy, loop feature engineering, time series emulater by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a></li>\n<li>use of sqlite to prevent memory error from <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a>  and <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> </li>\n<li>use true skill features from <a href=\"https://www.kaggle.com/zyy2016\" target=\"_blank\">@zyy2016</a> and <a href=\"https://www.kaggle.com/neinun\" target=\"_blank\">@neinun</a> </li>\n<li>convert dataframe to numpy array memory-eficiently from <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> </li>\n<li>use of taggroup from <a href=\"https://www.kaggle.com/radema\" target=\"_blank\">@radema</a></li>\n</ul>\n<p>Please notify me if I missed anyone.</p>\n<p>Random thought about this competition.</p>\n<ul>\n<li>The quantity and quality of the dataset were very high. We saw good cv-lb relationship.</li>\n<li>The concept of submission API is very interesting. It is realistic and highly adequate for time-series prediction.</li>\n<li>I think the kaggle system needs some improvements. The computational power is so small to handle huge dataset of this competition. The error message of submission should be more user-friendly.</li>\n</ul>\n<p>Any comments and questions are welcome!</p>",
  "messages": [
    {
      "id": "1143788",
      "postDate": "01/08/2021 04:17:41",
      "content": "<p>Many many thanks to the organizers and all kagglers.<br>\nThroughout this competition, I realized good kagglers are not just smart data scientists, but also proficient engineers. They have deep and broad knowledge about how to speed up processing, handle data memory efficiently, and debug submission error. I learned a lot from their notebook and comments. <br>\nI achieved 62nd (0.801 CV /0.799 public / 0.801 private) and got my first medal with a single LGB model. I tried blending with catboost and SAKT models, but it did not work well for me. I think it is mainly due to the time constraint to refine them.<br>\nThe code and model are opened <a href=\"https://www.kaggle.com/tomooinubushi/62nd-solution-lightgbm-single-model-lb-0-801\" target=\"_blank\">here</a>.</p>\n<p>I mainly focused on feature engineering and used 113 features. The feature importance of them was as follows</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fc09711551a883f6995032d91b6408bd9%2F__results___42_0.png?generation=1610076363046071&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F35c076ded96a2ef0d40ec307d0b8da51%2F__results___42_1.png?generation=1610076387230945&amp;alt=media\" alt=\"\"></p>\n<p>The meanings of abbreviations were as follows </p>\n<ul>\n<li>acc: answered_correctly</li>\n<li>exp: prior_question_had_explanation</li>\n<li>time: prior_question_elapsed_time</li>\n<li>count: number of trial</li>\n<li>count_lecture: number of lecture</li>\n<li>tfl: timestamp from last nth trial to that from n+1th trial</li>\n<li>dif: difficulty (1 - content-wise accuracy)</li>\n<li>dpt: difficulty point (add difficulty if users answer was correct)</li>\n<li>avg: average</li>\n<li>std: standard deviation</li>\n<li>sum: sum</li>\n<li>hist: historical average with temporal decay of gamma 0.75 </li>\n<li>hist2: historical average with temporal decay of gamma 0.25</li>\n<li>u: user-wise</li>\n<li>t: tag-wise</li>\n<li>tg: taggroup-wise</li>\n<li>p: part-wise</li>\n<li>c: content-wise</li>\n<li>b: bundle-wise</li>\n<li>uc: user-content-wise</li>\n<li>ub: user-bundle-wise</li>\n<li>up: user-part-wise</li>\n<li>time_per_count_u: timestamp/count_u</li>\n<li>count_u_nondiagnostic: user-wise count of non-diagnostic questions</li>\n<li>exp_avg_u_corrected: user-wise average of prior_question_had_explanation corrected for diagnostic questions. Th user cannot read explanation for diagnostic questions.</li>\n<li>pp_ratio: ratio of paid part (part1, 3, 4, 6, 7) trials per all trials, only paid users of SANTA app can solve them if it is not diagnostic questions</li>\n<li>listening_ratio: ratio of listening part trials per all trials</li>\n<li>relative_dif: relative difficulty of questions calculated with acc_avg_c / dif_avg_u</li>\n<li>notacc_sum: sum of failed trials</li>\n<li>wp_ts: trueskill feature of win probability</li>\n<li>umu_ts, usigma_ts, qmu_ts, qsigma_ts: trueskill features of mu, and sigma of user and content</li>\n</ul>\n<p>I learned a lot from many notebooks and comments of other kagglers. </p>\n<ul>\n<li>CV strategy, loop feature engineering, time series emulater by <a href=\"https://www.kaggle.com/its7171\" target=\"_blank\">@its7171</a></li>\n<li>use of sqlite to prevent memory error from <a href=\"https://www.kaggle.com/higepon\" target=\"_blank\">@higepon</a>  and <a href=\"https://www.kaggle.com/calebeverett\" target=\"_blank\">@calebeverett</a> </li>\n<li>use true skill features from <a href=\"https://www.kaggle.com/zyy2016\" target=\"_blank\">@zyy2016</a> and <a href=\"https://www.kaggle.com/neinun\" target=\"_blank\">@neinun</a> </li>\n<li>convert dataframe to numpy array memory-eficiently from <a href=\"https://www.kaggle.com/markwijkhuizen\" target=\"_blank\">@markwijkhuizen</a> </li>\n<li>use of taggroup from <a href=\"https://www.kaggle.com/radema\" target=\"_blank\">@radema</a></li>\n</ul>\n<p>Please notify me if I missed anyone.</p>\n<p>Random thought about this competition.</p>\n<ul>\n<li>The quantity and quality of the dataset were very high. We saw good cv-lb relationship.</li>\n<li>The concept of submission API is very interesting. It is realistic and highly adequate for time-series prediction.</li>\n<li>I think the kaggle system needs some improvements. The computational power is so small to handle huge dataset of this competition. The error message of submission should be more user-friendly.</li>\n</ul>\n<p>Any comments and questions are welcome!</p>",
      "rawMarkdown": "Many many thanks to the organizers and all kagglers.\nThroughout this competition, I realized good kagglers are not just smart data scientists, but also proficient engineers. They have deep and broad knowledge about how to speed up processing, handle data memory efficiently, and debug submission error. I learned a lot from their notebook and comments. \nI achieved 62nd (0.801 CV /0.799 public / 0.801 private) and got my first medal with a single LGB model. I tried blending with catboost and SAKT models, but it did not work well for me. I think it is mainly due to the time constraint to refine them.\nThe code and model are opened [here](https://www.kaggle.com/tomooinubushi/62nd-solution-lightgbm-single-model-lb-0-801).\n\nI mainly focused on feature engineering and used 113 features. The feature importance of them was as follows\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fc09711551a883f6995032d91b6408bd9%2F__results___42_0.png?generation=1610076363046071&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F35c076ded96a2ef0d40ec307d0b8da51%2F__results___42_1.png?generation=1610076387230945&alt=media)\n\nThe meanings of abbreviations were as follows \n- acc: answered_correctly\n- exp: prior_question_had_explanation\n- time: prior_question_elapsed_time\n- count: number of trial\n- count_lecture: number of lecture\n- tfl: timestamp from last nth trial to that from n+1th trial\n- dif: difficulty (1 - content-wise accuracy)\n- dpt: difficulty point (add difficulty if users answer was correct)\n- avg: average\n- std: standard deviation\n- sum: sum\n- hist: historical average with temporal decay of gamma 0.75 \n- hist2: historical average with temporal decay of gamma 0.25\n- u: user-wise\n- t: tag-wise\n- tg: taggroup-wise\n- p: part-wise\n- c: content-wise\n- b: bundle-wise\n- uc: user-content-wise\n- ub: user-bundle-wise\n- up: user-part-wise\n- time_per_count_u: timestamp/count_u\n- count_u_nondiagnostic: user-wise count of non-diagnostic questions\n- exp_avg_u_corrected: user-wise average of prior_question_had_explanation corrected for diagnostic questions. Th user cannot read explanation for diagnostic questions.\n- pp_ratio: ratio of paid part (part1, 3, 4, 6, 7) trials per all trials, only paid users of SANTA app can solve them if it is not diagnostic questions\n- listening_ratio: ratio of listening part trials per all trials\n- relative_dif: relative difficulty of questions calculated with acc_avg_c / dif_avg_u\n- notacc_sum: sum of failed trials\n- wp_ts: trueskill feature of win probability\n- umu_ts, usigma_ts, qmu_ts, qsigma_ts: trueskill features of mu, and sigma of user and content\n\nI learned a lot from many notebooks and comments of other kagglers. \n- CV strategy, loop feature engineering, time series emulater by @its7171\n- use of sqlite to prevent memory error from @higepon  and @calebeverett \n- use true skill features from @zyy2016 and @neinun \n- convert dataframe to numpy array memory-eficiently from @markwijkhuizen \n- use of taggroup from @radema\n\nPlease notify me if I missed anyone.\n\nRandom thought about this competition.\n- The quantity and quality of the dataset were very high. We saw good cv-lb relationship.\n- The concept of submission API is very interesting. It is realistic and highly adequate for time-series prediction.\n- I think the kaggle system needs some improvements. The computational power is so small to handle huge dataset of this competition. The error message of submission should be more user-friendly.\n\nAny comments and questions are welcome!",
      "votes": null
    },
    {
      "id": "1143844",
      "postDate": "01/08/2021 05:10:02",
      "content": "<p>Congratulations!<br>\nI'm glad that I could contribute to the sqlite part a bit.</p>",
      "rawMarkdown": "Congratulations!\nI'm glad that I could contribute to the sqlite part a bit.",
      "votes": null
    },
    {
      "id": "1143880",
      "postDate": "01/08/2021 05:44:55",
      "content": "<p>congratulations. Glad to have contributed a bit.</p>",
      "rawMarkdown": "congratulations. Glad to have contributed a bit.",
      "votes": null
    },
    {
      "id": "1143956",
      "postDate": "01/08/2021 06:36:22",
      "content": "<p>Thank you very much.<br>\nI could not submit my results without your thoughtful comments.</p>",
      "rawMarkdown": "Thank you very much.\nI could not submit my results without your thoughtful comments.",
      "votes": null
    },
    {
      "id": "1143958",
      "postDate": "01/08/2021 06:37:33",
      "content": "<p>Thank you.<br>\nI utilized your code to add true skill features. It was very helpful for me.</p>",
      "rawMarkdown": "Thank you.\nI utilized your code to add true skill features. It was very helpful for me.",
      "votes": null
    },
    {
      "id": "1144025",
      "postDate": "01/08/2021 07:26:40",
      "content": "<p>Thanks for sharing and congratulations on the great result!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations on the great result!",
      "votes": null
    },
    {
      "id": "1144708",
      "postDate": "01/08/2021 16:15:17",
      "content": "<p>congratulations.A great job!</p>",
      "rawMarkdown": "congratulations.A great job!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1143844,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "01/08/2021 05:10:02",
      "content": "<p>Congratulations!<br>\nI'm glad that I could contribute to the sqlite part a bit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143956,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "01/08/2021 06:36:22",
          "content": "<p>Thank you very much.<br>\nI could not submit my results without your thoughtful comments.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1143880,
      "author_name": "neinun",
      "author_url": "",
      "post_date": "01/08/2021 05:44:55",
      "content": "<p>congratulations. Glad to have contributed a bit.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1143958,
          "author_name": "tomooinubushi",
          "author_url": "",
          "post_date": "01/08/2021 06:37:33",
          "content": "<p>Thank you.<br>\nI utilized your code to add true skill features. It was very helpful for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1144025,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "01/08/2021 07:26:40",
      "content": "<p>Thanks for sharing and congratulations on the great result!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1144708,
      "author_name": "zjjszj2",
      "author_url": "",
      "post_date": "01/08/2021 16:15:17",
      "content": "<p>congratulations.A great job!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1143788": "Many many thanks to the organizers and all kagglers.\nThroughout this competition, I realized good kagglers are not just smart data scientists, but also proficient engineers. They have deep and broad knowledge about how to speed up processing, handle data memory efficiently, and debug submission error. I learned a lot from their notebook and comments. \nI achieved 62nd (0.801 CV /0.799 public / 0.801 private) and got my first medal with a single LGB model. I tried blending with catboost and SAKT models, but it did not work well for me. I think it is mainly due to the time constraint to refine them.\nThe code and model are opened [here](https://www.kaggle.com/tomooinubushi/62nd-solution-lightgbm-single-model-lb-0-801).\n\nI mainly focused on feature engineering and used 113 features. The feature importance of them was as follows\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2Fc09711551a883f6995032d91b6408bd9%2F__results___42_0.png?generation=1610076363046071&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3754725%2F35c076ded96a2ef0d40ec307d0b8da51%2F__results___42_1.png?generation=1610076387230945&alt=media)\n\nThe meanings of abbreviations were as follows \n- acc: answered_correctly\n- exp: prior_question_had_explanation\n- time: prior_question_elapsed_time\n- count: number of trial\n- count_lecture: number of lecture\n- tfl: timestamp from last nth trial to that from n+1th trial\n- dif: difficulty (1 - content-wise accuracy)\n- dpt: difficulty point (add difficulty if users answer was correct)\n- avg: average\n- std: standard deviation\n- sum: sum\n- hist: historical average with temporal decay of gamma 0.75 \n- hist2: historical average with temporal decay of gamma 0.25\n- u: user-wise\n- t: tag-wise\n- tg: taggroup-wise\n- p: part-wise\n- c: content-wise\n- b: bundle-wise\n- uc: user-content-wise\n- ub: user-bundle-wise\n- up: user-part-wise\n- time_per_count_u: timestamp/count_u\n- count_u_nondiagnostic: user-wise count of non-diagnostic questions\n- exp_avg_u_corrected: user-wise average of prior_question_had_explanation corrected for diagnostic questions. Th user cannot read explanation for diagnostic questions.\n- pp_ratio: ratio of paid part (part1, 3, 4, 6, 7) trials per all trials, only paid users of SANTA app can solve them if it is not diagnostic questions\n- listening_ratio: ratio of listening part trials per all trials\n- relative_dif: relative difficulty of questions calculated with acc_avg_c / dif_avg_u\n- notacc_sum: sum of failed trials\n- wp_ts: trueskill feature of win probability\n- umu_ts, usigma_ts, qmu_ts, qsigma_ts: trueskill features of mu, and sigma of user and content\n\nI learned a lot from many notebooks and comments of other kagglers. \n- CV strategy, loop feature engineering, time series emulater by @its7171\n- use of sqlite to prevent memory error from @higepon  and @calebeverett \n- use true skill features from @zyy2016 and @neinun \n- convert dataframe to numpy array memory-eficiently from @markwijkhuizen \n- use of taggroup from @radema\n\nPlease notify me if I missed anyone.\n\nRandom thought about this competition.\n- The quantity and quality of the dataset were very high. We saw good cv-lb relationship.\n- The concept of submission API is very interesting. It is realistic and highly adequate for time-series prediction.\n- I think the kaggle system needs some improvements. The computational power is so small to handle huge dataset of this competition. The error message of submission should be more user-friendly.\n\nAny comments and questions are welcome!",
    "1143844": "Congratulations!\nI'm glad that I could contribute to the sqlite part a bit.",
    "1143880": "congratulations. Glad to have contributed a bit.",
    "1143956": "Thank you very much.\nI could not submit my results without your thoughtful comments.",
    "1143958": "Thank you.\nI utilized your code to add true skill features. It was very helpful for me.",
    "1144025": "Thanks for sharing and congratulations on the great result!",
    "1144708": "congratulations.A great job!"
  },
  "source": "meta"
}