{
  "id": 398565,
  "title": "Achieve [CV 0.6914 | LB 0.694] with Only Four Features",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/398565",
  "author_name": "",
  "post_date": "2023-03-30T17:07:44.127472200Z",
  "votes": 68,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>Before diving into the manual feature engineering, I'm interested in seeing how far a DL-based model can go with a small set of features. After running experiments day and night, I temporarily can achieve <strong>CV 0.6914</strong> and LB 0.694 with only four features, including:</p>\n<ol>\n<li>Difference of <code>elapsed_time</code></li>\n<li><code>event_name</code></li>\n<li><code>name</code></li>\n<li><code>room_fqid</code></li>\n</ol>\n<p>First, <code>event_name</code> is combined with <code>name</code> to form a new feature <code>event_comb</code> (i.e., event combination). Then, I treat the difference of <code>elapsed_time</code> as the only numeric feature and the others categorical. The motivation behind the scene is that I hope the model can capture <strong>event-aware temporal patterns</strong>; that is, each time difference value is <strong>enriched by the event information and where the event take places</strong>. Finally, the model achieves the performance summarized as follows:</p>\n<table>\n<thead>\n<tr>\n<th>CV (GroupKFold with k=5)</th>\n<th>Holdout (Released Old Test Set)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6914</td>\n<td>0.6911</td>\n<td>0.694</td>\n</tr>\n</tbody>\n</table>\n<p>To be honest, I think I'm not good at model building and DL tuning, but I still want to share my experience and results I get so far. What's more, we should be cautious about the CV / LB gap to avoid overfitting on LB. I hope this can inspire you to build a more robust model!</p>\n<p>The inference part is <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">here</a>.<br>\nThe training part remains a work in progress and will be published within 2 ~ 3 days.</p>\n<p>Thanks a lot!</p>",
  "messages": [
    {
      "id": "2203292",
      "postDate": "03/30/2023 17:07:44",
      "content": "<p>Hi everyone,</p>\n<p>Before diving into the manual feature engineering, I'm interested in seeing how far a DL-based model can go with a small set of features. After running experiments day and night, I temporarily can achieve <strong>CV 0.6914</strong> and LB 0.694 with only four features, including:</p>\n<ol>\n<li>Difference of <code>elapsed_time</code></li>\n<li><code>event_name</code></li>\n<li><code>name</code></li>\n<li><code>room_fqid</code></li>\n</ol>\n<p>First, <code>event_name</code> is combined with <code>name</code> to form a new feature <code>event_comb</code> (i.e., event combination). Then, I treat the difference of <code>elapsed_time</code> as the only numeric feature and the others categorical. The motivation behind the scene is that I hope the model can capture <strong>event-aware temporal patterns</strong>; that is, each time difference value is <strong>enriched by the event information and where the event take places</strong>. Finally, the model achieves the performance summarized as follows:</p>\n<table>\n<thead>\n<tr>\n<th>CV (GroupKFold with k=5)</th>\n<th>Holdout (Released Old Test Set)</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.6914</td>\n<td>0.6911</td>\n<td>0.694</td>\n</tr>\n</tbody>\n</table>\n<p>To be honest, I think I'm not good at model building and DL tuning, but I still want to share my experience and results I get so far. What's more, we should be cautious about the CV / LB gap to avoid overfitting on LB. I hope this can inspire you to build a more robust model!</p>\n<p>The inference part is <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">here</a>.<br>\nThe training part remains a work in progress and will be published within 2 ~ 3 days.</p>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Hi everyone,\n\nBefore diving into the manual feature engineering, I'm interested in seeing how far a DL-based model can go with a small set of features. After running experiments day and night, I temporarily can achieve **CV 0.6914** and LB 0.694 with only four features, including:\n\n1. Difference of `elapsed_time`\n2. `event_name`\n3. `name`\n4. `room_fqid`\n\nFirst, `event_name` is combined with `name` to form a new feature `event_comb` (i.e., event combination). Then, I treat the difference of `elapsed_time` as the only numeric feature and the others categorical. The motivation behind the scene is that I hope the model can capture **event-aware temporal patterns**; that is, each time difference value is **enriched by the event information and where the event take places**. Finally, the model achieves the performance summarized as follows:\n\n| CV (GroupKFold with k=5) | Holdout (Released Old Test Set) | LB    |\n| ------------------------ | ------------------------------- | ----- |\n| 0.6914                   | 0.6911                          | 0.694 |\n\nTo be honest, I think I'm not good at model building and DL tuning, but I still want to share my experience and results I get so far. What's more, we should be cautious about the CV / LB gap to avoid overfitting on LB. I hope this can inspire you to build a more robust model!\n\nThe inference part is [here](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features).\nThe training part remains a work in progress and will be published within 2 ~ 3 days.\n\nThanks a lot!",
      "votes": null
    },
    {
      "id": "2203305",
      "postDate": "03/30/2023 17:27:48",
      "content": "<p>Very nice. How long does inference take when you submit (roughly)?</p>",
      "rawMarkdown": "Very nice. How long does inference take when you submit (roughly)?",
      "votes": null
    },
    {
      "id": "2203452",
      "postDate": "03/30/2023 20:04:05",
      "content": "<p>Great DL model. Well done</p>",
      "rawMarkdown": "Great DL model. Well done",
      "votes": null
    },
    {
      "id": "2203492",
      "postDate": "03/30/2023 21:21:10",
      "content": "<p>Nice .. thanks will be good to add this if works into an ensemble with a large feature rich GBM  😸 .</p>",
      "rawMarkdown": "Nice .. thanks will be good to add this if works into an ensemble with a large feature rich GBM  😸 .",
      "votes": null
    },
    {
      "id": "2203630",
      "postDate": "03/31/2023 02:00:43",
      "content": "<p>Nice! Combining <code>event_name</code> and <code>name</code> might be a good approach. <code>name</code> does not provide many information alone. </p>",
      "rawMarkdown": "Nice! Combining `event_name` and `name` might be a good approach. `name` does not provide many information alone.",
      "votes": null
    },
    {
      "id": "2203667",
      "postDate": "03/31/2023 03:09:06",
      "content": "<p>Thanks for sharing! What NN type did you use (suppose I can't wait for the training notebook haha)?</p>",
      "rawMarkdown": "Thanks for sharing! What NN type did you use (suppose I can't wait for the training notebook haha)?",
      "votes": null
    },
    {
      "id": "2203693",
      "postDate": "03/31/2023 03:56:14",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>It takes roughly 65 minutes to finish scoring (including data processing, simple feature engineering and model inference). And, there are 15 models in total (<em>i.e.</em>, 5-fold models for each <code>level_group</code>). </p>\n<p>Hope this helps, thanks a lot!</p>",
      "rawMarkdown": "Hi @cpmpml,\n\nIt takes roughly 65 minutes to finish scoring (including data processing, simple feature engineering and model inference). And, there are 15 models in total (*i.e.*, 5-fold models for each `level_group`). \n\nHope this helps, thanks a lot!",
      "votes": null
    },
    {
      "id": "2203697",
      "postDate": "03/31/2023 03:58:19",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,</p>\n<p>Thanks for your appreciation, I'll add more features to see whether DL-based models can catch up with tree-based models or not!</p>",
      "rawMarkdown": "Hi @cdeotte,\n\nThanks for your appreciation, I'll add more features to see whether DL-based models can catch up with tree-based models or not!",
      "votes": null
    },
    {
      "id": "2203703",
      "postDate": "03/31/2023 04:00:31",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>,</p>\n<p>Thanks for the comment. I'm also curious about how an ensemble of DL-based model and GBM can perform. Let's keep going!</p>",
      "rawMarkdown": "Hi @gauravbrills,\n\nThanks for the comment. I'm also curious about how an ensemble of DL-based model and GBM can perform. Let's keep going!",
      "votes": null
    },
    {
      "id": "2203722",
      "postDate": "03/31/2023 04:25:29",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a>,</p>\n<p>I think <code>name</code> can provide information about the <strong>status</strong> of the corresponding event. Also, there exist only 19 combinations of <code>event_name</code> (11 unique values) and <code>name</code> (6 unique values), not 66. Hence, training only one embedding table with the combination might be a little bit better than training two separate ones.</p>\n<p>I learned a lot from your notebook! Thanks for your sharing.</p>",
      "rawMarkdown": "Hi @chrisqiu,\n\nI think `name` can provide information about the **status** of the corresponding event. Also, there exist only 19 combinations of `event_name` (11 unique values) and `name` (6 unique values), not 66. Hence, training only one embedding table with the combination might be a little bit better than training two separate ones.\n\nI learned a lot from your notebook! Thanks for your sharing.",
      "votes": null
    },
    {
      "id": "2203746",
      "postDate": "03/31/2023 05:03:35",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>,</p>\n<p>I plot an overview of the model architecture shown as follows, </p>\n<p><a href=\"https://postimg.cc/R634wmSt\" target=\"_blank\"><img src=\"https://i.postimg.cc/MT5H2Zg9/event-aware-tconvb.png\" alt=\"event-aware-tconvb.png\"></a></p>\n<p>Hope this can help you understand. I'll publish the training part asap😂. Thanks a lot!</p>",
      "rawMarkdown": "Hi @hoangnguyen719,\n\nI plot an overview of the model architecture shown as follows, \n\n[![event-aware-tconvb.png](https://i.postimg.cc/MT5H2Zg9/event-aware-tconvb.png)](https://postimg.cc/R634wmSt)\n\nHope this can help you understand. I'll publish the training part asap😂. Thanks a lot!",
      "votes": null
    },
    {
      "id": "2205073",
      "postDate": "04/01/2023 08:15:18",
      "content": "<p>hi, do the difference of the elapsed time here means between 2 sucessive events based on the index?</p>",
      "rawMarkdown": "hi, do the difference of the elapsed time here means between 2 sucessive events based on the index?",
      "votes": null
    },
    {
      "id": "2205449",
      "postDate": "04/01/2023 15:10:48",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a>,</p>\n<p>That's true, just take the difference of <code>elapsed_time</code> between two successive events. However, please take care that <code>index</code> itself has some data errors confirmed by the competition host <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384342#2134312\" target=\"_blank\">here</a>. Thanks!</p>",
      "rawMarkdown": "Hi @narendra,\n\nThat's true, just take the difference of `elapsed_time` between two successive events. However, please take care that `index` itself has some data errors confirmed by the competition host [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384342#2134312). Thanks!",
      "votes": null
    },
    {
      "id": "2282098",
      "postDate": "05/31/2023 11:15:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>, </p>\n<p>Thanks for sharing this notebook. </p>\n<p>I am facing an issue. I am not able to get the pickle file. There is nothing at the path <code>/kaggle/input/cat2code-v2/cat2code</code>. Are others also facing the same issue? Could you please help me in finding out how to get the pickle file? </p>",
      "rawMarkdown": "Hi @abaojiang, \n\nThanks for sharing this notebook. \n\nI am facing an issue. I am not able to get the pickle file. There is nothing at the path `/kaggle/input/cat2code-v2/cat2code`. Are others also facing the same issue? Could you please help me in finding out how to get the pickle file?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2203305,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "03/30/2023 17:27:48",
      "content": "<p>Very nice. How long does inference take when you submit (roughly)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2203693,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/31/2023 03:56:14",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>,</p>\n<p>It takes roughly 65 minutes to finish scoring (including data processing, simple feature engineering and model inference). And, there are 15 models in total (<em>i.e.</em>, 5-fold models for each <code>level_group</code>). </p>\n<p>Hope this helps, thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2203452,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/30/2023 20:04:05",
      "content": "<p>Great DL model. Well done</p>",
      "votes": null,
      "replies": [
        {
          "id": 2203697,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/31/2023 03:58:19",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>,</p>\n<p>Thanks for your appreciation, I'll add more features to see whether DL-based models can catch up with tree-based models or not!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2203492,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "03/30/2023 21:21:10",
      "content": "<p>Nice .. thanks will be good to add this if works into an ensemble with a large feature rich GBM  😸 .</p>",
      "votes": null,
      "replies": [
        {
          "id": 2203703,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/31/2023 04:00:31",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a>,</p>\n<p>Thanks for the comment. I'm also curious about how an ensemble of DL-based model and GBM can perform. Let's keep going!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2203630,
      "author_name": "chrisqiu",
      "author_url": "",
      "post_date": "03/31/2023 02:00:43",
      "content": "<p>Nice! Combining <code>event_name</code> and <code>name</code> might be a good approach. <code>name</code> does not provide many information alone. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2203722,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/31/2023 04:25:29",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/chrisqiu\" target=\"_blank\">@chrisqiu</a>,</p>\n<p>I think <code>name</code> can provide information about the <strong>status</strong> of the corresponding event. Also, there exist only 19 combinations of <code>event_name</code> (11 unique values) and <code>name</code> (6 unique values), not 66. Hence, training only one embedding table with the combination might be a little bit better than training two separate ones.</p>\n<p>I learned a lot from your notebook! Thanks for your sharing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2203667,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "03/31/2023 03:09:06",
      "content": "<p>Thanks for sharing! What NN type did you use (suppose I can't wait for the training notebook haha)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2203746,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "03/31/2023 05:03:35",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>,</p>\n<p>I plot an overview of the model architecture shown as follows, </p>\n<p><a href=\"https://postimg.cc/R634wmSt\" target=\"_blank\"><img src=\"https://i.postimg.cc/MT5H2Zg9/event-aware-tconvb.png\" alt=\"event-aware-tconvb.png\"></a></p>\n<p>Hope this can help you understand. I'll publish the training part asap😂. Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2205073,
      "author_name": "narendra",
      "author_url": "",
      "post_date": "04/01/2023 08:15:18",
      "content": "<p>hi, do the difference of the elapsed time here means between 2 sucessive events based on the index?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2205449,
          "author_name": "abaojiang",
          "author_url": "",
          "post_date": "04/01/2023 15:10:48",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/narendra\" target=\"_blank\">@narendra</a>,</p>\n<p>That's true, just take the difference of <code>elapsed_time</code> between two successive events. However, please take care that <code>index</code> itself has some data errors confirmed by the competition host <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384342#2134312\" target=\"_blank\">here</a>. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2282098,
      "author_name": "tilliumshakespeare",
      "author_url": "",
      "post_date": "05/31/2023 11:15:59",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>, </p>\n<p>Thanks for sharing this notebook. </p>\n<p>I am facing an issue. I am not able to get the pickle file. There is nothing at the path <code>/kaggle/input/cat2code-v2/cat2code</code>. Are others also facing the same issue? Could you please help me in finding out how to get the pickle file? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2203292": "Hi everyone,\n\nBefore diving into the manual feature engineering, I'm interested in seeing how far a DL-based model can go with a small set of features. After running experiments day and night, I temporarily can achieve **CV 0.6914** and LB 0.694 with only four features, including:\n\n1. Difference of `elapsed_time`\n2. `event_name`\n3. `name`\n4. `room_fqid`\n\nFirst, `event_name` is combined with `name` to form a new feature `event_comb` (i.e., event combination). Then, I treat the difference of `elapsed_time` as the only numeric feature and the others categorical. The motivation behind the scene is that I hope the model can capture **event-aware temporal patterns**; that is, each time difference value is **enriched by the event information and where the event take places**. Finally, the model achieves the performance summarized as follows:\n\n| CV (GroupKFold with k=5) | Holdout (Released Old Test Set) | LB    |\n| ------------------------ | ------------------------------- | ----- |\n| 0.6914                   | 0.6911                          | 0.694 |\n\nTo be honest, I think I'm not good at model building and DL tuning, but I still want to share my experience and results I get so far. What's more, we should be cautious about the CV / LB gap to avoid overfitting on LB. I hope this can inspire you to build a more robust model!\n\nThe inference part is [here](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features).\nThe training part remains a work in progress and will be published within 2 ~ 3 days.\n\nThanks a lot!",
    "2203305": "Very nice. How long does inference take when you submit (roughly)?",
    "2203452": "Great DL model. Well done",
    "2203492": "Nice .. thanks will be good to add this if works into an ensemble with a large feature rich GBM  😸 .",
    "2203630": "Nice! Combining `event_name` and `name` might be a good approach. `name` does not provide many information alone.",
    "2203667": "Thanks for sharing! What NN type did you use (suppose I can't wait for the training notebook haha)?",
    "2203693": "Hi @cpmpml,\n\nIt takes roughly 65 minutes to finish scoring (including data processing, simple feature engineering and model inference). And, there are 15 models in total (*i.e.*, 5-fold models for each `level_group`). \n\nHope this helps, thanks a lot!",
    "2203697": "Hi @cdeotte,\n\nThanks for your appreciation, I'll add more features to see whether DL-based models can catch up with tree-based models or not!",
    "2203703": "Hi @gauravbrills,\n\nThanks for the comment. I'm also curious about how an ensemble of DL-based model and GBM can perform. Let's keep going!",
    "2203722": "Hi @chrisqiu,\n\nI think `name` can provide information about the **status** of the corresponding event. Also, there exist only 19 combinations of `event_name` (11 unique values) and `name` (6 unique values), not 66. Hence, training only one embedding table with the combination might be a little bit better than training two separate ones.\n\nI learned a lot from your notebook! Thanks for your sharing.",
    "2203746": "Hi @hoangnguyen719,\n\nI plot an overview of the model architecture shown as follows, \n\n[![event-aware-tconvb.png](https://i.postimg.cc/MT5H2Zg9/event-aware-tconvb.png)](https://postimg.cc/R634wmSt)\n\nHope this can help you understand. I'll publish the training part asap😂. Thanks a lot!",
    "2205073": "hi, do the difference of the elapsed time here means between 2 sucessive events based on the index?",
    "2205449": "Hi @narendra,\n\nThat's true, just take the difference of `elapsed_time` between two successive events. However, please take care that `index` itself has some data errors confirmed by the competition host [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/384342#2134312). Thanks!",
    "2282098": "Hi @abaojiang, \n\nThanks for sharing this notebook. \n\nI am facing an issue. I am not able to get the pickle file. There is nothing at the path `/kaggle/input/cat2code-v2/cat2code`. Are others also facing the same issue? Could you please help me in finding out how to get the pickle file?"
  },
  "source": "meta"
}