{
  "id": 420181,
  "title": "30 place | but our best 0.704/0.703 solution & Code what could have been - Gaurav's Part with help from my friends ",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420181",
  "author_name": "",
  "post_date": "2023-06-29T14:37:27.028832200Z",
  "votes": 20,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Thanks to Kaggle for hosting the comp and it was a long and bumpy ride . We were at 1st position most of the time but the adage of not selecting best CV striked (may have been case with most as most of the time lb and private had a difference )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2Fd270769e5b2cd735a81a7e2a01ba4f5d%2FScreenshot%202023-06-29%20at%208.09.34%20PM.png?generation=1688049640841410&amp;alt=media\" alt=\"\"></p>\n<p>Also had a blast working with friends <a href=\"https://www.kaggle.com/onelux\" target=\"_blank\">@onelux</a> <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a> <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> and shuzhe-deng</p>\n<h1><strong>Overview</strong></h1>\n<ul>\n<li>To make predictions for 18 questions, we first trained XGB across 5 folds for each level group and question and then trained the best iteration on full data for each of the 18 questions .</li>\n<li>Features were mostly based on our team member <a href=\"https://www.kaggle.com/leehomtabularhuang\" target=\"_blank\">@leehomtabularhuang</a> and <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a> and based elapsed time as is the case with most people . We had some meta features and some prior features as well that helped a lot will describe it in the next section .</li>\n<li>Out train time was mostly XGB on GPU was around 2-3 hours and inference took around 7 hours .</li>\n</ul>\n<h1><strong>Notebooks</strong></h1>\n<ul>\n<li>Feature engineering - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook</a></li>\n<li>Best Inference CV(0.70081) LB (0.702) Private (0.704) - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook</a> </li>\n<li>My best CV (0.7010775749344409), LB (0.705), Priva (0.703) - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-inference/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-inference/notebook</a></li>\n</ul>\n<p>hyperparms that worked best with bit low LR</p>\n<pre><code>xgb_params = {\n    :,\n    : ,\n    : ,\n    : ,\n    :,\n    : ,\n    : , \n    : , \n    :,\n    : , \n    : CFG.seed, \n    : ,\n    :,\n    :\n    } \n</code></pre>\n<h1><strong>Feature Engineering</strong></h1>\n<ul>\n<li>Most initial features were based on initially <a href=\"https://www.kaggle.com/onelux\" target=\"_blank\">@onelux</a> who was also part of our team </li>\n<li>Meta features to use probs from previous level groups were exactly similalr to <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041</a> solution and helped us greatly thanks to <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a></li>\n<li>Also using the full set of prior features (for second and third level group)suggested by.@wang5630900 helped though it did create a lot of features for our last level group around 10 k 😄 .</li>\n</ul>\n<pre><code>df2 = df2.join(df1,=, =)\ndf3 = df3.join(df2,=, =)\n</code></pre>\n<ul>\n<li><ul>\n<li>Recap features we also did add which worked fine with our best Private model . </li></ul></li>\n<li>I crafted some click agg functions with logic as below </li>\n</ul>\n<pre><code>def (group,field,alias,feature_suffix):  \n        aggs = []\n        for level in group:\n            clickagggs=[\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in ev_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in ev_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in name_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in name_feature],\n             ]\n            aggs = aggs +clickagggs  \n        return aggs\n</code></pre>\n<ul>\n<li>We also trimmed some of our features which were like redundant and applied only to specific level groups as you can see from our public shared notebook <a href=\"https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook</a></li>\n</ul>\n<p><strong>What did not work</strong> </p>\n<ul>\n<li>Mixing nn and GBT ensemble somehow did not give us benefits so we dropped it .</li>\n<li>Distance and angle based features we thought will work but didnt show promise :( .</li>\n<li>W2vec for text features and clusters on them were not taht effective . We also tried some embedding based clustters on bert but didnt help much .</li>\n<li>Clusters did visualize well based on clicks and co ordinates but somehow did not work that well .</li>\n</ul>\n<p><strong>Inference</strong></p>\n<ul>\n<li>Best threshold that worked for us was 0.605 mostly tuned on LB </li>\n<li>Full data trained models ~ 18 per question were used .</li>\n<li>For my part sorting by elapsed time worked post the change in data .</li>\n</ul>\n<h1><strong>Deciders</strong></h1>\n<p>So lastly our submission selection was too reliant on LB . our best SINGLE cv models below turned out good (with cv and correlation with private and public) but as we got 0.707 wit had ensemble we got tempted to select it  . Our best spread is as below. </p>\n<h2><em>What we should have selected</em></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F6ad7084d8af681bec864270a45e5d203%2FScreenshot%202023-06-29%20at%207.24.49%20PM.png?generation=1688048410286089&amp;alt=media\" alt=\"\"></p>\n<h2><em>What we did select</em></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F883463ce68b06916838a5390f611587b%2FScreenshot%202023-07-01%20at%206.11.43%20PM.png?generation=1688215338652572&amp;alt=media\" alt=\"\"></p>\n<p>Anyways hope we get a silver and some of us become Comp masters finally :)</p>",
  "messages": [
    {
      "id": "2322869",
      "postDate": "06/29/2023 14:37:27",
      "content": "<p>Thanks to Kaggle for hosting the comp and it was a long and bumpy ride . We were at 1st position most of the time but the adage of not selecting best CV striked (may have been case with most as most of the time lb and private had a difference )</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2Fd270769e5b2cd735a81a7e2a01ba4f5d%2FScreenshot%202023-06-29%20at%208.09.34%20PM.png?generation=1688049640841410&amp;alt=media\" alt=\"\"></p>\n<p>Also had a blast working with friends <a href=\"https://www.kaggle.com/onelux\" target=\"_blank\">@onelux</a> <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a> <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> and shuzhe-deng</p>\n<h1><strong>Overview</strong></h1>\n<ul>\n<li>To make predictions for 18 questions, we first trained XGB across 5 folds for each level group and question and then trained the best iteration on full data for each of the 18 questions .</li>\n<li>Features were mostly based on our team member <a href=\"https://www.kaggle.com/leehomtabularhuang\" target=\"_blank\">@leehomtabularhuang</a> and <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a> and based elapsed time as is the case with most people . We had some meta features and some prior features as well that helped a lot will describe it in the next section .</li>\n<li>Out train time was mostly XGB on GPU was around 2-3 hours and inference took around 7 hours .</li>\n</ul>\n<h1><strong>Notebooks</strong></h1>\n<ul>\n<li>Feature engineering - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook</a></li>\n<li>Best Inference CV(0.70081) LB (0.702) Private (0.704) - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook</a> </li>\n<li>My best CV (0.7010775749344409), LB (0.705), Priva (0.703) - <a href=\"https://www.kaggle.com/code/gauravbrills/psp-inference/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-inference/notebook</a></li>\n</ul>\n<p>hyperparms that worked best with bit low LR</p>\n<pre><code>xgb_params = {\n    :,\n    : ,\n    : ,\n    : ,\n    :,\n    : ,\n    : , \n    : , \n    :,\n    : , \n    : CFG.seed, \n    : ,\n    :,\n    :\n    } \n</code></pre>\n<h1><strong>Feature Engineering</strong></h1>\n<ul>\n<li>Most initial features were based on initially <a href=\"https://www.kaggle.com/onelux\" target=\"_blank\">@onelux</a> who was also part of our team </li>\n<li>Meta features to use probs from previous level groups were exactly similalr to <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041</a> solution and helped us greatly thanks to <a href=\"https://www.kaggle.com/wang5630900\" target=\"_blank\">@wang5630900</a></li>\n<li>Also using the full set of prior features (for second and third level group)suggested by.@wang5630900 helped though it did create a lot of features for our last level group around 10 k 😄 .</li>\n</ul>\n<pre><code>df2 = df2.join(df1,=, =)\ndf3 = df3.join(df2,=, =)\n</code></pre>\n<ul>\n<li><ul>\n<li>Recap features we also did add which worked fine with our best Private model . </li></ul></li>\n<li>I crafted some click agg functions with logic as below </li>\n</ul>\n<pre><code>def (group,field,alias,feature_suffix):  \n        aggs = []\n        for level in group:\n            clickagggs=[\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in ev_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in ev_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in name_feature],\n                *[pl.().((pl.() == c) &amp; (pl.(field)==level)).().(f) for c in name_feature],\n             ]\n            aggs = aggs +clickagggs  \n        return aggs\n</code></pre>\n<ul>\n<li>We also trimmed some of our features which were like redundant and applied only to specific level groups as you can see from our public shared notebook <a href=\"https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\" target=\"_blank\">https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook</a></li>\n</ul>\n<p><strong>What did not work</strong> </p>\n<ul>\n<li>Mixing nn and GBT ensemble somehow did not give us benefits so we dropped it .</li>\n<li>Distance and angle based features we thought will work but didnt show promise :( .</li>\n<li>W2vec for text features and clusters on them were not taht effective . We also tried some embedding based clustters on bert but didnt help much .</li>\n<li>Clusters did visualize well based on clicks and co ordinates but somehow did not work that well .</li>\n</ul>\n<p><strong>Inference</strong></p>\n<ul>\n<li>Best threshold that worked for us was 0.605 mostly tuned on LB </li>\n<li>Full data trained models ~ 18 per question were used .</li>\n<li>For my part sorting by elapsed time worked post the change in data .</li>\n</ul>\n<h1><strong>Deciders</strong></h1>\n<p>So lastly our submission selection was too reliant on LB . our best SINGLE cv models below turned out good (with cv and correlation with private and public) but as we got 0.707 wit had ensemble we got tempted to select it  . Our best spread is as below. </p>\n<h2><em>What we should have selected</em></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F6ad7084d8af681bec864270a45e5d203%2FScreenshot%202023-06-29%20at%207.24.49%20PM.png?generation=1688048410286089&amp;alt=media\" alt=\"\"></p>\n<h2><em>What we did select</em></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F883463ce68b06916838a5390f611587b%2FScreenshot%202023-07-01%20at%206.11.43%20PM.png?generation=1688215338652572&amp;alt=media\" alt=\"\"></p>\n<p>Anyways hope we get a silver and some of us become Comp masters finally :)</p>",
      "rawMarkdown": "Thanks to Kaggle for hosting the comp and it was a long and bumpy ride . We were at 1st position most of the time but the adage of not selecting best CV striked (may have been case with most as most of the time lb and private had a difference )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2Fd270769e5b2cd735a81a7e2a01ba4f5d%2FScreenshot%202023-06-29%20at%208.09.34%20PM.png?generation=1688049640841410&alt=media)\n\nAlso had a blast working with friends @onelux @wang5630900 @tonymarkchris and shuzhe-deng\n\n# **Overview**\n\n- To make predictions for 18 questions, we first trained XGB across 5 folds for each level group and question and then trained the best iteration on full data for each of the 18 questions .\n- Features were mostly based on our team member @leehomtabularhuang and @wang5630900 and based elapsed time as is the case with most people . We had some meta features and some prior features as well that helped a lot will describe it in the next section .\n- Out train time was mostly XGB on GPU was around 2-3 hours and inference took around 7 hours .\n\n# **Notebooks** \n- Feature engineering - https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\n- Best Inference CV(0.70081) LB (0.702) Private (0.704) - https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook \n- My best CV (0.7010775749344409), LB (0.705), Priva (0.703) - https://www.kaggle.com/code/gauravbrills/psp-inference/notebook\n\nhyperparms that worked best with bit low LR\n\n```python\nxgb_params = {\n    'n_estimators':1000,\n    'booster': 'gbtree',\n    'tree_method': 'hist',\n    'objective': 'binary:logistic',\n    'eval_metric':'logloss',\n    'learning_rate': 0.015,\n    'alpha': 8, \n    'max_depth': 4, \n    'subsample':0.8,\n    'colsample_bytree': 0.5, \n    'seed': CFG.seed, \n    'early_stopping_rounds': 90,\n    'tree_method':'gpu_hist',\n    'gpu_id':0\n    } \n```\n\n# **Feature Engineering**\n- Most initial features were based on initially @onelux who was also part of our team \n- Meta features to use probs from previous level groups were exactly similalr to https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041 solution and helped us greatly thanks to @wang5630900\n- Also using the full set of prior features (for second and third level group)suggested by.@wang5630900 helped though it did create a lot of features for our last level group around 10 k 😄 .\n```\ndf2 = df2.join(df1,on=\"session_id\", how='left')\ndf3 = df3.join(df2,on=\"session_id\", how='left')\n```\n- \n- Recap features we also did add which worked fine with our best Private model . \n- I crafted some click agg functions with logic as below \n\n```\ndef applyNewClickAggs(group,field,alias,feature_suffix):  \n        aggs = []\n        for level in group:\n            clickagggs=[\n                *[pl.col(\"event_name\").filter((pl.col(\"event_name\") == c) & (pl.col(field)==level)).count().alias(f\"{c}_EN_by_{alias}{level}_counts_{feature_suffix}\") for c in ev_feature],\n                *[pl.col(\"elapsed_time_diff\").filter((pl.col(\"event_name\") == c) & (pl.col(field)==level)).sum().alias(f\"{c}_ET_EN_by_{alias}{level}_sum_{feature_suffix}\") for c in ev_feature],\n                *[pl.col(\"event_name\").filter((pl.col(\"name\") == c) & (pl.col(field)==level)).count().alias(f\"{c}_EN_by_{alias}{level}_counts_{feature_suffix}\") for c in name_feature],\n                *[pl.col(\"elapsed_time_diff\").filter((pl.col(\"name\") == c) & (pl.col(field)==level)).sum().alias(f\"{c}_ET_EN_by_{alias}{level}_sum_{feature_suffix}\") for c in name_feature],\n             ]\n            aggs = aggs +clickagggs  \n        return aggs\n```\n- We also trimmed some of our features which were like redundant and applied only to specific level groups as you can see from our public shared notebook https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\n\n**What did not work** \n- Mixing nn and GBT ensemble somehow did not give us benefits so we dropped it .\n- Distance and angle based features we thought will work but didnt show promise :( .\n- W2vec for text features and clusters on them were not taht effective . We also tried some embedding based clustters on bert but didnt help much .\n- Clusters did visualize well based on clicks and co ordinates but somehow did not work that well .\n\n**Inference**\n\n- Best threshold that worked for us was 0.605 mostly tuned on LB \n- Full data trained models ~ 18 per question were used .\n- For my part sorting by elapsed time worked post the change in data .\n\n# **Deciders**\n\nSo lastly our submission selection was too reliant on LB . our best SINGLE cv models below turned out good (with cv and correlation with private and public) but as we got 0.707 wit had ensemble we got tempted to select it  . Our best spread is as below. \n\n## *What we should have selected*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F6ad7084d8af681bec864270a45e5d203%2FScreenshot%202023-06-29%20at%207.24.49%20PM.png?generation=1688048410286089&alt=media)\n\n## *What we did select*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F883463ce68b06916838a5390f611587b%2FScreenshot%202023-07-01%20at%206.11.43%20PM.png?generation=1688215338652572&alt=media)\n\nAnyways hope we get a silver and some of us become Comp masters finally :)",
      "votes": null
    },
    {
      "id": "2323005",
      "postDate": "06/29/2023 16:09:41",
      "content": "<p>Again, bummer to see huge shake-down, but big thumbs up for your hard work (I mean look at the number of your submissions!) Good luck in future competitions!</p>",
      "rawMarkdown": "Again, bummer to see huge shake-down, but big thumbs up for your hard work (I mean look at the number of your submissions!) Good luck in future competitions!",
      "votes": null
    },
    {
      "id": "2331458",
      "postDate": "07/05/2023 14:59:37",
      "content": "<p>Sure thanks <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>  and congrats to you </p>",
      "rawMarkdown": "Sure thanks @hoangnguyen719  and congrats to you",
      "votes": null
    },
    {
      "id": "2332064",
      "postDate": "07/06/2023 00:09:10",
      "content": "<p>Thanks for sharing your solution!</p>\n<p>I've built my model based on <a href=\"https://www.kaggle.com/code/leehomhuang/lb0-691-catboostbaseline-train/notebook\" target=\"_blank\">your teammate - Onelux's Catboost baseline model</a>, in which he uses Catboost. May I ask why your notebook uses XGBoost instead of Catboost? Does It perform better?</p>",
      "rawMarkdown": "Thanks for sharing your solution!\n\nI've built my model based on [your teammate - Onelux's Catboost baseline model](https://www.kaggle.com/code/leehomhuang/lb0-691-catboostbaseline-train/notebook), in which he uses Catboost. May I ask why your notebook uses XGBoost instead of Catboost? Does It perform better?",
      "votes": null
    },
    {
      "id": "2332429",
      "postDate": "07/06/2023 07:33:42",
      "content": "<p>Np <a href=\"https://www.kaggle.com/ncchen\" target=\"_blank\">@ncchen</a> .. On my side I observed that xgboost with GPU trained faster and gave better cv .. we did add more features though to Onelux public version which also helped .. somehow catboost and xg ensemble did not work so we and on my part skipped it .</p>",
      "rawMarkdown": "Np @ncchen .. On my side I observed that xgboost with GPU trained faster and gave better cv .. we did add more features though to Onelux public version which also helped .. somehow catboost and xg ensemble did not work so we and on my part skipped it .",
      "votes": null
    },
    {
      "id": "2341234",
      "postDate": "07/12/2023 01:07:59",
      "content": "<p><a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> what happened with your team? 👀</p>",
      "rawMarkdown": "gauravbrills what happened with your team? 👀",
      "votes": null
    },
    {
      "id": "2341472",
      "postDate": "07/12/2023 06:15:17",
      "content": "<p>just noticed this too 😨</p>",
      "rawMarkdown": "just noticed this too 😨",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2323005,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 16:09:41",
      "content": "<p>Again, bummer to see huge shake-down, but big thumbs up for your hard work (I mean look at the number of your submissions!) Good luck in future competitions!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2331458,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "07/05/2023 14:59:37",
          "content": "<p>Sure thanks <a href=\"https://www.kaggle.com/hoangnguyen719\" target=\"_blank\">@hoangnguyen719</a>  and congrats to you </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2332064,
      "author_name": "ncchen",
      "author_url": "",
      "post_date": "07/06/2023 00:09:10",
      "content": "<p>Thanks for sharing your solution!</p>\n<p>I've built my model based on <a href=\"https://www.kaggle.com/code/leehomhuang/lb0-691-catboostbaseline-train/notebook\" target=\"_blank\">your teammate - Onelux's Catboost baseline model</a>, in which he uses Catboost. May I ask why your notebook uses XGBoost instead of Catboost? Does It perform better?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2332429,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "07/06/2023 07:33:42",
          "content": "<p>Np <a href=\"https://www.kaggle.com/ncchen\" target=\"_blank\">@ncchen</a> .. On my side I observed that xgboost with GPU trained faster and gave better cv .. we did add more features though to Onelux public version which also helped .. somehow catboost and xg ensemble did not work so we and on my part skipped it .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2341234,
      "author_name": "kvlmll",
      "author_url": "",
      "post_date": "07/12/2023 01:07:59",
      "content": "<p><a href=\"https://www.kaggle.com/gauravbrills\" target=\"_blank\">@gauravbrills</a> what happened with your team? 👀</p>",
      "votes": null,
      "replies": [
        {
          "id": 2341472,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "07/12/2023 06:15:17",
          "content": "<p>just noticed this too 😨</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2322869": "Thanks to Kaggle for hosting the comp and it was a long and bumpy ride . We were at 1st position most of the time but the adage of not selecting best CV striked (may have been case with most as most of the time lb and private had a difference )\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2Fd270769e5b2cd735a81a7e2a01ba4f5d%2FScreenshot%202023-06-29%20at%208.09.34%20PM.png?generation=1688049640841410&alt=media)\n\nAlso had a blast working with friends @onelux @wang5630900 @tonymarkchris and shuzhe-deng\n\n# **Overview**\n\n- To make predictions for 18 questions, we first trained XGB across 5 folds for each level group and question and then trained the best iteration on full data for each of the 18 questions .\n- Features were mostly based on our team member @leehomtabularhuang and @wang5630900 and based elapsed time as is the case with most people . We had some meta features and some prior features as well that helped a lot will describe it in the next section .\n- Out train time was mostly XGB on GPU was around 2-3 hours and inference took around 7 hours .\n\n# **Notebooks** \n- Feature engineering - https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\n- Best Inference CV(0.70081) LB (0.702) Private (0.704) - https://www.kaggle.com/code/gauravbrills/psp-infer-single-0-70081/notebook \n- My best CV (0.7010775749344409), LB (0.705), Priva (0.703) - https://www.kaggle.com/code/gauravbrills/psp-inference/notebook\n\nhyperparms that worked best with bit low LR\n\n```python\nxgb_params = {\n    'n_estimators':1000,\n    'booster': 'gbtree',\n    'tree_method': 'hist',\n    'objective': 'binary:logistic',\n    'eval_metric':'logloss',\n    'learning_rate': 0.015,\n    'alpha': 8, \n    'max_depth': 4, \n    'subsample':0.8,\n    'colsample_bytree': 0.5, \n    'seed': CFG.seed, \n    'early_stopping_rounds': 90,\n    'tree_method':'gpu_hist',\n    'gpu_id':0\n    } \n```\n\n# **Feature Engineering**\n- Most initial features were based on initially @onelux who was also part of our team \n- Meta features to use probs from previous level groups were exactly similalr to https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/420041 solution and helped us greatly thanks to @wang5630900\n- Also using the full set of prior features (for second and third level group)suggested by.@wang5630900 helped though it did create a lot of features for our last level group around 10 k 😄 .\n```\ndf2 = df2.join(df1,on=\"session_id\", how='left')\ndf3 = df3.join(df2,on=\"session_id\", how='left')\n```\n- \n- Recap features we also did add which worked fine with our best Private model . \n- I crafted some click agg functions with logic as below \n\n```\ndef applyNewClickAggs(group,field,alias,feature_suffix):  \n        aggs = []\n        for level in group:\n            clickagggs=[\n                *[pl.col(\"event_name\").filter((pl.col(\"event_name\") == c) & (pl.col(field)==level)).count().alias(f\"{c}_EN_by_{alias}{level}_counts_{feature_suffix}\") for c in ev_feature],\n                *[pl.col(\"elapsed_time_diff\").filter((pl.col(\"event_name\") == c) & (pl.col(field)==level)).sum().alias(f\"{c}_ET_EN_by_{alias}{level}_sum_{feature_suffix}\") for c in ev_feature],\n                *[pl.col(\"event_name\").filter((pl.col(\"name\") == c) & (pl.col(field)==level)).count().alias(f\"{c}_EN_by_{alias}{level}_counts_{feature_suffix}\") for c in name_feature],\n                *[pl.col(\"elapsed_time_diff\").filter((pl.col(\"name\") == c) & (pl.col(field)==level)).sum().alias(f\"{c}_ET_EN_by_{alias}{level}_sum_{feature_suffix}\") for c in name_feature],\n             ]\n            aggs = aggs +clickagggs  \n        return aggs\n```\n- We also trimmed some of our features which were like redundant and applied only to specific level groups as you can see from our public shared notebook https://www.kaggle.com/code/gauravbrills/psp-xgb-2023-05-08/notebook\n\n**What did not work** \n- Mixing nn and GBT ensemble somehow did not give us benefits so we dropped it .\n- Distance and angle based features we thought will work but didnt show promise :( .\n- W2vec for text features and clusters on them were not taht effective . We also tried some embedding based clustters on bert but didnt help much .\n- Clusters did visualize well based on clicks and co ordinates but somehow did not work that well .\n\n**Inference**\n\n- Best threshold that worked for us was 0.605 mostly tuned on LB \n- Full data trained models ~ 18 per question were used .\n- For my part sorting by elapsed time worked post the change in data .\n\n# **Deciders**\n\nSo lastly our submission selection was too reliant on LB . our best SINGLE cv models below turned out good (with cv and correlation with private and public) but as we got 0.707 wit had ensemble we got tempted to select it  . Our best spread is as below. \n\n## *What we should have selected*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F6ad7084d8af681bec864270a45e5d203%2FScreenshot%202023-06-29%20at%207.24.49%20PM.png?generation=1688048410286089&alt=media)\n\n## *What we did select*\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F768166%2F883463ce68b06916838a5390f611587b%2FScreenshot%202023-07-01%20at%206.11.43%20PM.png?generation=1688215338652572&alt=media)\n\nAnyways hope we get a silver and some of us become Comp masters finally :)",
    "2323005": "Again, bummer to see huge shake-down, but big thumbs up for your hard work (I mean look at the number of your submissions!) Good luck in future competitions!",
    "2331458": "Sure thanks @hoangnguyen719  and congrats to you",
    "2332064": "Thanks for sharing your solution!\n\nI've built my model based on [your teammate - Onelux's Catboost baseline model](https://www.kaggle.com/code/leehomhuang/lb0-691-catboostbaseline-train/notebook), in which he uses Catboost. May I ask why your notebook uses XGBoost instead of Catboost? Does It perform better?",
    "2332429": "Np @ncchen .. On my side I observed that xgboost with GPU trained faster and gave better cv .. we did add more features though to Onelux public version which also helped .. somehow catboost and xg ensemble did not work so we and on my part skipped it .",
    "2341234": "gauravbrills what happened with your team? 👀",
    "2341472": "just noticed this too 😨"
  },
  "source": "meta"
}