{
  "id": 424329,
  "title": "2nd Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/mark4h-2nd-place-solution",
  "author_name": "",
  "post_date": "2023-07-13T13:26:46.448013500Z",
  "votes": 20,
  "comment_count": 8,
  "views": 0,
  "content": "<h1>2nd Place Solution</h1>\n<p>First, I would like to take the opportunity to thank The Learning Agency Lab for hosting the competition and the Kaggle team for making it happen.</p>\n<p>Here are the details of the 2nd place solution.</p>\n<p><strong>Summary</strong></p>\n<ul>\n<li>A single ‭LightGBM ‬model was used to predict all the questions (i.e. not separate models per question or level group)</li>\n<li>5 fold cross validation was used during development but for the final submission a single model was trained on all of the data</li>\n<li>The code was optimised to minimise the efficiency score<ul>\n<li>For the final submission the vast majority of the execution time was spent on the LightGBM prediction stage</li>\n<li>There was extensive use of numba and C for the feature generation code</li></ul></li>\n<li>The model contained ‬1296 features</li>\n<li>A Threshold value of 0.63 was used</li>\n</ul>\n<p>‭<strong>Features</strong></p>\n<p>A lot of the most important features were based on the time taken to complete a task or react the an event. One of the most important features (after some of the basic features such as the question number and the total count of events for a level group) was the amount of time the user spent looking at the report in level 1 (feature name: L‬G0_L1_first_report_open_duration‭).</p>\n<p>A plot of the feature importance (LightGBM gain) of the top features can be seen below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1199911%2F8bb9d994b0bb3f59dcff38f99eb95ebe%2Ffeature_importance.png?generation=1688479725194354&amp;alt=media\" alt=\"\"></p>\n<p><strong>Code</strong></p>\n<p>‭The code for each stage of the solution can be found here:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-0-preprocess-data\" target=\"_blank\">preprocess data</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-features-code\" target=\"_blank\">features code</a> (<a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-c-feature-code\" target=\"_blank\">features code utility script</a>)</li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-2-generate-features\" target=\"_blank\">generate features</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-3-train-model\" target=\"_blank\">train model</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-4-submission\" target=\"_blank\">submission</a></li>\n</ol>",
  "messages": [
    {
      "id": "2343194",
      "postDate": "07/13/2023 13:26:46",
      "content": "<h1>2nd Place Solution</h1>\n<p>First, I would like to take the opportunity to thank The Learning Agency Lab for hosting the competition and the Kaggle team for making it happen.</p>\n<p>Here are the details of the 2nd place solution.</p>\n<p><strong>Summary</strong></p>\n<ul>\n<li>A single ‭LightGBM ‬model was used to predict all the questions (i.e. not separate models per question or level group)</li>\n<li>5 fold cross validation was used during development but for the final submission a single model was trained on all of the data</li>\n<li>The code was optimised to minimise the efficiency score<ul>\n<li>For the final submission the vast majority of the execution time was spent on the LightGBM prediction stage</li>\n<li>There was extensive use of numba and C for the feature generation code</li></ul></li>\n<li>The model contained ‬1296 features</li>\n<li>A Threshold value of 0.63 was used</li>\n</ul>\n<p>‭<strong>Features</strong></p>\n<p>A lot of the most important features were based on the time taken to complete a task or react the an event. One of the most important features (after some of the basic features such as the question number and the total count of events for a level group) was the amount of time the user spent looking at the report in level 1 (feature name: L‬G0_L1_first_report_open_duration‭).</p>\n<p>A plot of the feature importance (LightGBM gain) of the top features can be seen below:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1199911%2F8bb9d994b0bb3f59dcff38f99eb95ebe%2Ffeature_importance.png?generation=1688479725194354&amp;alt=media\" alt=\"\"></p>\n<p><strong>Code</strong></p>\n<p>‭The code for each stage of the solution can be found here:</p>\n<ol>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-0-preprocess-data\" target=\"_blank\">preprocess data</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-features-code\" target=\"_blank\">features code</a> (<a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-c-feature-code\" target=\"_blank\">features code utility script</a>)</li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-2-generate-features\" target=\"_blank\">generate features</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-3-train-model\" target=\"_blank\">train model</a></li>\n<li><a href=\"https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-4-submission\" target=\"_blank\">submission</a></li>\n</ol>",
      "rawMarkdown": "# 2nd Place Solution\n\nFirst, I would like to take the opportunity to thank The Learning Agency Lab for hosting the competition and the Kaggle team for making it happen.\n\nHere are the details of the 2nd place solution.\n\n**Summary**\n\n- A single ‭LightGBM ‬model was used to predict all the questions (i.e. not separate models per question or level group)\n- 5 fold cross validation was used during development but for the final submission a single model was trained on all of the data\n- The code was optimised to minimise the efficiency score\n    - For the final submission the vast majority of the execution time was spent on the LightGBM prediction stage\n    - There was extensive use of numba and C for the feature generation code\n- The model contained ‬1296 features\n- A Threshold value of 0.63 was used\n\n‭**Features**\n\nA lot of the most important features were based on the time taken to complete a task or react the an event. One of the most important features (after some of the basic features such as the question number and the total count of events for a level group) was the amount of time the user spent looking at the report in level 1 (feature name: L‬G0_L1_first_report_open_duration‭).\n\nA plot of the feature importance (LightGBM gain) of the top features can be seen below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1199911%2F8bb9d994b0bb3f59dcff38f99eb95ebe%2Ffeature_importance.png?generation=1688479725194354&alt=media)\n\n**Code**\n\n‭The code for each stage of the solution can be found here:\n\n0. [preprocess data](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-0-preprocess-data)\n1. [features code](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-features-code) ([features code utility script](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-c-feature-code))\n2. [generate features](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-2-generate-features)\n3. [train model](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-3-train-model)\n4. [submission](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-4-submission)",
      "votes": null
    },
    {
      "id": "2343898",
      "postDate": "07/14/2023 05:02:31",
      "content": "<p>Congratulations on winning the second place. Thanks for sharing the detailed solution with useful notes.</p>",
      "rawMarkdown": "Congratulations on winning the second place. Thanks for sharing the detailed solution with useful notes.",
      "votes": null
    },
    {
      "id": "2344563",
      "postDate": "07/14/2023 15:07:31",
      "content": "<p>Awesome solution. I am glad you shared, I was real curious. Congrats for being 2nd on both LBs! In a way your solution looks similar to Jack (Japan)'s solution. No wonder the two of you trusted top of efficiency prize.</p>\n<p>Did you have a good CV LB correlation? We found that we had a great CV private LB correlation, but that public LB was more shaky, i.e. that some high CV models had a lower public LB but a great private LB. Fortunately filtering both on CV and public LB scores retained only good private LB scores for us.</p>",
      "rawMarkdown": "Awesome solution. I am glad you shared, I was real curious. Congrats for being 2nd on both LBs! In a way your solution looks similar to Jack (Japan)'s solution. No wonder the two of you trusted top of efficiency prize.\n\nDid you have a good CV LB correlation? We found that we had a great CV private LB correlation, but that public LB was more shaky, i.e. that some high CV models had a lower public LB but a great private LB. Fortunately filtering both on CV and public LB scores retained only good private LB scores for us.",
      "votes": null
    },
    {
      "id": "2345236",
      "postDate": "07/15/2023 08:10:50",
      "content": "<p>It's very impressive to use a single model, which would make much lower the training time.<br>\nDid using a single model with the question features improve your performance rather than using seperate models without them??</p>\n<p>Congrats and thank you for sharing your solution.</p>",
      "rawMarkdown": "It's very impressive to use a single model, which would make much lower the training time.\nDid using a single model with the question features improve your performance rather than using seperate models without them??\n\nCongrats and thank you for sharing your solution.",
      "votes": null
    },
    {
      "id": "2345319",
      "postDate": "07/15/2023 09:29:43",
      "content": "<p>Thanks. Well done on 1st place.</p>\n<p>I can’t really add much to the topic of CV-LB correlation. I only made a handful of submissions and they were all of essentially the same model (tuned to improve execution time), so I don’t have much of a spread to correlate.</p>",
      "rawMarkdown": "Thanks. Well done on 1st place.\n\nI can’t really add much to the topic of CV-LB correlation. I only made a handful of submissions and they were all of essentially the same model (tuned to improve execution time), so I don’t have much of a spread to correlate.",
      "votes": null
    },
    {
      "id": "2345322",
      "postDate": "07/15/2023 09:30:49",
      "content": "<p>On the question of single model vs separate models, my take was to use as simple a solution as possible (a single model), unless there was evidence that a more complex solution was needed (separate models). I never found any evidence that separate models were needed.</p>",
      "rawMarkdown": "On the question of single model vs separate models, my take was to use as simple a solution as possible (a single model), unless there was evidence that a more complex solution was needed (separate models). I never found any evidence that separate models were needed.",
      "votes": null
    },
    {
      "id": "2345357",
      "postDate": "07/15/2023 10:10:32",
      "content": "<p>Thanks.<br>\nWhat do you think applying this technique to a NN?</p>",
      "rawMarkdown": "Thanks.\nWhat do you think applying this technique to a NN?",
      "votes": null
    },
    {
      "id": "2370733",
      "postDate": "08/02/2023 15:50:41",
      "content": "<p>Thank you for sharing! Very elegant solution. I also read your 1st place solution post to the VSB Power Line Fault Detection competition and am very impressed.</p>\n<p>I have a question about your choice to use a single LightGBM vs ensemble, including for example some NN. Do you find there is no benefit in adding more models to ensemble with LightGBM, or is it simply more that time is better spent understanding the data and improving the features than creating additional models? Would you have been more likely to include other models if the dataset was smaller?</p>",
      "rawMarkdown": "Thank you for sharing! Very elegant solution. I also read your 1st place solution post to the VSB Power Line Fault Detection competition and am very impressed.\n\nI have a question about your choice to use a single LightGBM vs ensemble, including for example some NN. Do you find there is no benefit in adding more models to ensemble with LightGBM, or is it simply more that time is better spent understanding the data and improving the features than creating additional models? Would you have been more likely to include other models if the dataset was smaller?",
      "votes": null
    },
    {
      "id": "2383368",
      "postDate": "08/10/2023 10:18:45",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> and well done on your 4th place.</p>\n<p>I tend to find that it is better to focus on improving a single model than spread my efforts over multiple. That said, it can definitely be beneficial to add more models, particularly if they are different architectures. This is especially true if your are part of a team, when each member can focus on a different model. For small datasets, I think it still depends, I guess one advantage is that it typically requires less time/resources to train multiple models for small datasets which makes it more practical.</p>",
      "rawMarkdown": "Thanks @erijoel and well done on your 4th place.\n\nI tend to find that it is better to focus on improving a single model than spread my efforts over multiple. That said, it can definitely be beneficial to add more models, particularly if they are different architectures. This is especially true if your are part of a team, when each member can focus on a different model. For small datasets, I think it still depends, I guess one advantage is that it typically requires less time/resources to train multiple models for small datasets which makes it more practical.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2343898,
      "author_name": "crsuthikshnkumar",
      "author_url": "",
      "post_date": "07/14/2023 05:02:31",
      "content": "<p>Congratulations on winning the second place. Thanks for sharing the detailed solution with useful notes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2344563,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "07/14/2023 15:07:31",
      "content": "<p>Awesome solution. I am glad you shared, I was real curious. Congrats for being 2nd on both LBs! In a way your solution looks similar to Jack (Japan)'s solution. No wonder the two of you trusted top of efficiency prize.</p>\n<p>Did you have a good CV LB correlation? We found that we had a great CV private LB correlation, but that public LB was more shaky, i.e. that some high CV models had a lower public LB but a great private LB. Fortunately filtering both on CV and public LB scores retained only good private LB scores for us.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2345319,
          "author_name": "mark4h",
          "author_url": "",
          "post_date": "07/15/2023 09:29:43",
          "content": "<p>Thanks. Well done on 1st place.</p>\n<p>I can’t really add much to the topic of CV-LB correlation. I only made a handful of submissions and they were all of essentially the same model (tuned to improve execution time), so I don’t have much of a spread to correlate.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2345236,
      "author_name": "dongyk",
      "author_url": "",
      "post_date": "07/15/2023 08:10:50",
      "content": "<p>It's very impressive to use a single model, which would make much lower the training time.<br>\nDid using a single model with the question features improve your performance rather than using seperate models without them??</p>\n<p>Congrats and thank you for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2345322,
          "author_name": "mark4h",
          "author_url": "",
          "post_date": "07/15/2023 09:30:49",
          "content": "<p>On the question of single model vs separate models, my take was to use as simple a solution as possible (a single model), unless there was evidence that a more complex solution was needed (separate models). I never found any evidence that separate models were needed.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2345357,
              "author_name": "dongyk",
              "author_url": "",
              "post_date": "07/15/2023 10:10:32",
              "content": "<p>Thanks.<br>\nWhat do you think applying this technique to a NN?</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2370733,
      "author_name": "erijoel",
      "author_url": "",
      "post_date": "08/02/2023 15:50:41",
      "content": "<p>Thank you for sharing! Very elegant solution. I also read your 1st place solution post to the VSB Power Line Fault Detection competition and am very impressed.</p>\n<p>I have a question about your choice to use a single LightGBM vs ensemble, including for example some NN. Do you find there is no benefit in adding more models to ensemble with LightGBM, or is it simply more that time is better spent understanding the data and improving the features than creating additional models? Would you have been more likely to include other models if the dataset was smaller?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2383368,
          "author_name": "mark4h",
          "author_url": "",
          "post_date": "08/10/2023 10:18:45",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> and well done on your 4th place.</p>\n<p>I tend to find that it is better to focus on improving a single model than spread my efforts over multiple. That said, it can definitely be beneficial to add more models, particularly if they are different architectures. This is especially true if your are part of a team, when each member can focus on a different model. For small datasets, I think it still depends, I guess one advantage is that it typically requires less time/resources to train multiple models for small datasets which makes it more practical.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2343194": "# 2nd Place Solution\n\nFirst, I would like to take the opportunity to thank The Learning Agency Lab for hosting the competition and the Kaggle team for making it happen.\n\nHere are the details of the 2nd place solution.\n\n**Summary**\n\n- A single ‭LightGBM ‬model was used to predict all the questions (i.e. not separate models per question or level group)\n- 5 fold cross validation was used during development but for the final submission a single model was trained on all of the data\n- The code was optimised to minimise the efficiency score\n    - For the final submission the vast majority of the execution time was spent on the LightGBM prediction stage\n    - There was extensive use of numba and C for the feature generation code\n- The model contained ‬1296 features\n- A Threshold value of 0.63 was used\n\n‭**Features**\n\nA lot of the most important features were based on the time taken to complete a task or react the an event. One of the most important features (after some of the basic features such as the question number and the total count of events for a level group) was the amount of time the user spent looking at the report in level 1 (feature name: L‬G0_L1_first_report_open_duration‭).\n\nA plot of the feature importance (LightGBM gain) of the top features can be seen below:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1199911%2F8bb9d994b0bb3f59dcff38f99eb95ebe%2Ffeature_importance.png?generation=1688479725194354&alt=media)\n\n**Code**\n\n‭The code for each stage of the solution can be found here:\n\n0. [preprocess data](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-0-preprocess-data)\n1. [features code](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-features-code) ([features code utility script](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-1-c-feature-code))\n2. [generate features](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-2-generate-features)\n3. [train model](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-3-train-model)\n4. [submission](https://www.kaggle.com/mark4h/jowilder-2nd-place-solution-4-submission)",
    "2343898": "Congratulations on winning the second place. Thanks for sharing the detailed solution with useful notes.",
    "2344563": "Awesome solution. I am glad you shared, I was real curious. Congrats for being 2nd on both LBs! In a way your solution looks similar to Jack (Japan)'s solution. No wonder the two of you trusted top of efficiency prize.\n\nDid you have a good CV LB correlation? We found that we had a great CV private LB correlation, but that public LB was more shaky, i.e. that some high CV models had a lower public LB but a great private LB. Fortunately filtering both on CV and public LB scores retained only good private LB scores for us.",
    "2345236": "It's very impressive to use a single model, which would make much lower the training time.\nDid using a single model with the question features improve your performance rather than using seperate models without them??\n\nCongrats and thank you for sharing your solution.",
    "2345319": "Thanks. Well done on 1st place.\n\nI can’t really add much to the topic of CV-LB correlation. I only made a handful of submissions and they were all of essentially the same model (tuned to improve execution time), so I don’t have much of a spread to correlate.",
    "2345322": "On the question of single model vs separate models, my take was to use as simple a solution as possible (a single model), unless there was evidence that a more complex solution was needed (separate models). I never found any evidence that separate models were needed.",
    "2345357": "Thanks.\nWhat do you think applying this technique to a NN?",
    "2370733": "Thank you for sharing! Very elegant solution. I also read your 1st place solution post to the VSB Power Line Fault Detection competition and am very impressed.\n\nI have a question about your choice to use a single LightGBM vs ensemble, including for example some NN. Do you find there is no benefit in adding more models to ensemble with LightGBM, or is it simply more that time is better spent understanding the data and improving the features than creating additional models? Would you have been more likely to include other models if the dataset was smaller?",
    "2383368": "Thanks @erijoel and well done on your 4th place.\n\nI tend to find that it is better to focus on improving a single model than spread my efforts over multiple. That said, it can definitely be beneficial to add more models, particularly if they are different architectures. This is especially true if your are part of a team, when each member can focus on a different model. For small datasets, I think it still depends, I guess one advantage is that it typically requires less time/resources to train multiple models for small datasets which makes it more practical."
  },
  "source": "meta"
}