{
  "id": 420349,
  "title": "4th Place Solution",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420349",
  "author_name": "Joel Erikanders",
  "post_date": "2023-06-30T11:04:21.412000",
  "votes": 48,
  "comment_count": 15,
  "views": 0,
  "content": "<h1>Acknowledgement</h1>\n<p>I'd like to thank the hosts for providing a very interesting and difficult project to work on for the past months. I am also grateful for all the public sharing on Kaggle, this has been an insane learning experience for me. Without all the public notebooks, discussion posts and old competition solutions available i would have had no chance in this competition.</p>\n<h1>Overview</h1>\n<ul>\n<li>Used most of the raw data for training, while validating only on the kaggle data.</li>\n<li>Ensemble of Transformer, XGBoost and Catboost, with 3 seeds and 5 folds each. </li>\n<li>Used a generic set of features based on time, index and screen_coor differences.</li>\n<li>Linear regression as a meta model.</li>\n<li>Thresholds have a big impact on LB score</li>\n</ul>\n<h1>Data</h1>\n<p>I used most of the raw data for training, including sessions that only completed level group 0-4 and 5-12. About ~38000 whole sessions and ~58000 sessions in total. Using the raw data increased CV by over 0.001. I validated only on the kaggle data. </p>\n<p>My initial data preprocessing is simply sorting by level group and index, same as what happens during inference. Also, my experiments indicated no benefit from using the hover durations, so after sorting i dropped the hover rows and re-indexed each session from 0 to len(session).</p>\n<h1>Transformer</h1>\n<p>I spent much of my time experimenting with transformers, which resulted in a light weight model that achieved 0.698 on the public and private LB, and 0.702 CV. </p>\n<pre><code>class (nn.Module):\n    def (self, num_cont_cols, embed_dim, num_layers, num_heads, max_seq_len):\n        (NN, self).()\n        self.emb_cont = nn.(\n            nn.(num_cont_cols, embed_dim//),\n            nn.(embed_dim//)\n        )\n        self.emb_cats = nn.(\n            nn.(max_seq_len + , embed_dim//),\n            nn.(embed_dim//)\n        )\n        encoder_layer = nn.(\n            d_model=embed_dim,\n            nhead=num_heads,\n            dim_feedforward=embed_dim,\n            dropout=,\n            batch_first=True,\n            activation=,\n        )\n        self.encoder = nn.(encoder_layer, num_layers=num_layers)\n        self.clf_heads = nn.([\n            nn.(embed_dim, out_dim) for out_dim in [, , ]\n        ])\n\n    def (self, x, grp):\n        emb_conts = self.(x[:, :, :-])\n        emb_cats = self.(x[:, :, -].(torch.int32))\n        x = torch.([emb_conts, emb_cats], dim=)\n        x = self.(x)\n        x = x.(dim=)\n        x = self.clf_heads[[, , ].(grp)](x)\n        return x.()\n</code></pre>\n<ul>\n<li>embed_dim: 64</li>\n<li>num_layers: 1</li>\n<li>num_heads: 8</li>\n<li>max_seq_len: 452 (explained below), though the sequences are cropped to 256</li>\n<li>I used the same single model for all questions</li>\n</ul>\n<p>I found the data easy to overfit with transformers, so in an attempt to improve the signal to noise ratio i did the following:</p>\n<ol>\n<li>Identify different points in the game by string concatenating event_name, level, name, page, fqid, room_fqid, text_fqid, in the dataframe.</li>\n<li>Some of these occur more than once in a session. Treat these as different points by enumerating them and adding the enumeration to their names.</li>\n<li>Filter out the rows with points that is present in over 0.999 of the sessions. This makes each session maximum 452 steps long. </li>\n<li>Create 6 feature columns:<br>\ntime difference, index difference, distance (cumulative distance moved, calculated from screen_coor's) difference, room_coor_x, room_coor_y and the categorical point column embedded.</li>\n</ol>\n<h1>XGBoost</h1>\n<p>This is my strongest single model with public LB 0.701, private 0.702, and 0.7029 CV. What stands out is that i flattened 5 of the transformer input columns (excluding the categorical column), and used all those values as individual features.</p>\n<p>The other features are mainly stats that can be found in public kernels, like mean and max time diff over the categoricals. The stats were calculated before applying the transformer input filtering.</p>\n<p>From early on I trained one model for each level group, inputing the question number as a feature. I found that CV increased by around 0.0002 compared to using a model for each question. This could be randomness, but i went with it since i thought 3 models instead of 18 would make my life easier during experimentation. Similar reasoning behind using one model for the transformer. </p>\n<h1>Catboost</h1>\n<p>Essentially looks the same as XGBoost. CV 0.7022.</p>\n<h1>Ensemble</h1>\n<p>I trained a linear regression meta model for each question, using the above models output probabilites as input, to produce the final predictions. I included probabilities of past questions and some future ones! For example the regression model trained to predict question 2 took probabilities on question 1-3 as input, to predict question 7 I used probabilities on question 1-13, and for question 16 i used probabilities on question 1-18. I took 3 seeds average before linear regression input to make it more robust. </p>\n<p>This finally results in public LB 0.702, private LB 0.703, CV 0.7044. </p>\n<h1>On threshold and submission selection</h1>\n<p>I tried to trust CV as much as possible, but the consistent gap between my CV and LB was suspicius until the last few days. Then I realized one reason could be my selected threshold was suboptimal on the test data.. I made some submissions with my highest CV solution, only changing the threshold, and noticed it was indeed suboptimal and caused more variation in LB score than most of my latest experiments. So in the end I selected 3 of the same solution, with different thresholds: 0.60 (best on LB), 0.62 (best during CV) and 0.64. Turned out 0.61 would have resulted in 0.704 private, but no regrets ;)</p>\n<p>Thank you for reading!</p>\n<h1>Code</h1>\n<p><a href=\"https://github.com/joelerikanders/pspgp/tree/main\" target=\"_blank\">Training code</a><br>\n<a href=\"https://www.kaggle.com/erijoel/4th-place-submission\" target=\"_blank\">Submission notebook</a></p>",
  "messages": [
    {
      "id": 2324073,
      "postDate": "2023-06-30T11:04:21.413Z",
      "content": "<h1>Acknowledgement</h1>\n<p>I'd like to thank the hosts for providing a very interesting and difficult project to work on for the past months. I am also grateful for all the public sharing on Kaggle, this has been an insane learning experience for me. Without all the public notebooks, discussion posts and old competition solutions available i would have had no chance in this competition.</p>\n<h1>Overview</h1>\n<ul>\n<li>Used most of the raw data for training, while validating only on the kaggle data.</li>\n<li>Ensemble of Transformer, XGBoost and Catboost, with 3 seeds and 5 folds each. </li>\n<li>Used a generic set of features based on time, index and screen_coor differences.</li>\n<li>Linear regression as a meta model.</li>\n<li>Thresholds have a big impact on LB score</li>\n</ul>\n<h1>Data</h1>\n<p>I used most of the raw data for training, including sessions that only completed level group 0-4 and 5-12. About ~38000 whole sessions and ~58000 sessions in total. Using the raw data increased CV by over 0.001. I validated only on the kaggle data. </p>\n<p>My initial data preprocessing is simply sorting by level group and index, same as what happens during inference. Also, my experiments indicated no benefit from using the hover durations, so after sorting i dropped the hover rows and re-indexed each session from 0 to len(session).</p>\n<h1>Transformer</h1>\n<p>I spent much of my time experimenting with transformers, which resulted in a light weight model that achieved 0.698 on the public and private LB, and 0.702 CV. </p>\n<pre><code>class (nn.Module):\n    def (self, num_cont_cols, embed_dim, num_layers, num_heads, max_seq_len):\n        (NN, self).()\n        self.emb_cont = nn.(\n            nn.(num_cont_cols, embed_dim//),\n            nn.(embed_dim//)\n        )\n        self.emb_cats = nn.(\n            nn.(max_seq_len + , embed_dim//),\n            nn.(embed_dim//)\n        )\n        encoder_layer = nn.(\n            d_model=embed_dim,\n            nhead=num_heads,\n            dim_feedforward=embed_dim,\n            dropout=,\n            batch_first=True,\n            activation=,\n        )\n        self.encoder = nn.(encoder_layer, num_layers=num_layers)\n        self.clf_heads = nn.([\n            nn.(embed_dim, out_dim) for out_dim in [, , ]\n        ])\n\n    def (self, x, grp):\n        emb_conts = self.(x[:, :, :-])\n        emb_cats = self.(x[:, :, -].(torch.int32))\n        x = torch.([emb_conts, emb_cats], dim=)\n        x = self.(x)\n        x = x.(dim=)\n        x = self.clf_heads[[, , ].(grp)](x)\n        return x.()\n</code></pre>\n<ul>\n<li>embed_dim: 64</li>\n<li>num_layers: 1</li>\n<li>num_heads: 8</li>\n<li>max_seq_len: 452 (explained below), though the sequences are cropped to 256</li>\n<li>I used the same single model for all questions</li>\n</ul>\n<p>I found the data easy to overfit with transformers, so in an attempt to improve the signal to noise ratio i did the following:</p>\n<ol>\n<li>Identify different points in the game by string concatenating event_name, level, name, page, fqid, room_fqid, text_fqid, in the dataframe.</li>\n<li>Some of these occur more than once in a session. Treat these as different points by enumerating them and adding the enumeration to their names.</li>\n<li>Filter out the rows with points that is present in over 0.999 of the sessions. This makes each session maximum 452 steps long. </li>\n<li>Create 6 feature columns:<br>\ntime difference, index difference, distance (cumulative distance moved, calculated from screen_coor's) difference, room_coor_x, room_coor_y and the categorical point column embedded.</li>\n</ol>\n<h1>XGBoost</h1>\n<p>This is my strongest single model with public LB 0.701, private 0.702, and 0.7029 CV. What stands out is that i flattened 5 of the transformer input columns (excluding the categorical column), and used all those values as individual features.</p>\n<p>The other features are mainly stats that can be found in public kernels, like mean and max time diff over the categoricals. The stats were calculated before applying the transformer input filtering.</p>\n<p>From early on I trained one model for each level group, inputing the question number as a feature. I found that CV increased by around 0.0002 compared to using a model for each question. This could be randomness, but i went with it since i thought 3 models instead of 18 would make my life easier during experimentation. Similar reasoning behind using one model for the transformer. </p>\n<h1>Catboost</h1>\n<p>Essentially looks the same as XGBoost. CV 0.7022.</p>\n<h1>Ensemble</h1>\n<p>I trained a linear regression meta model for each question, using the above models output probabilites as input, to produce the final predictions. I included probabilities of past questions and some future ones! For example the regression model trained to predict question 2 took probabilities on question 1-3 as input, to predict question 7 I used probabilities on question 1-13, and for question 16 i used probabilities on question 1-18. I took 3 seeds average before linear regression input to make it more robust. </p>\n<p>This finally results in public LB 0.702, private LB 0.703, CV 0.7044. </p>\n<h1>On threshold and submission selection</h1>\n<p>I tried to trust CV as much as possible, but the consistent gap between my CV and LB was suspicius until the last few days. Then I realized one reason could be my selected threshold was suboptimal on the test data.. I made some submissions with my highest CV solution, only changing the threshold, and noticed it was indeed suboptimal and caused more variation in LB score than most of my latest experiments. So in the end I selected 3 of the same solution, with different thresholds: 0.60 (best on LB), 0.62 (best during CV) and 0.64. Turned out 0.61 would have resulted in 0.704 private, but no regrets ;)</p>\n<p>Thank you for reading!</p>\n<h1>Code</h1>\n<p><a href=\"https://github.com/joelerikanders/pspgp/tree/main\" target=\"_blank\">Training code</a><br>\n<a href=\"https://www.kaggle.com/erijoel/4th-place-submission\" target=\"_blank\">Submission notebook</a></p>",
      "rawMarkdown": "# Acknowledgement\nI'd like to thank the hosts for providing a very interesting and difficult project to work on for the past months. I am also grateful for all the public sharing on Kaggle, this has been an insane learning experience for me. Without all the public notebooks, discussion posts and old competition solutions available i would have had no chance in this competition.\n\n# Overview\n- Used most of the raw data for training, while validating only on the kaggle data.\n- Ensemble of Transformer, XGBoost and Catboost, with 3 seeds and 5 folds each. \n- Used a generic set of features based on time, index and screen_coor differences.\n- Linear regression as a meta model.\n- Thresholds have a big impact on LB score\n\n# Data\nI used most of the raw data for training, including sessions that only completed level group 0-4 and 5-12. About ~38000 whole sessions and ~58000 sessions in total. Using the raw data increased CV by over 0.001. I validated only on the kaggle data. \n\nMy initial data preprocessing is simply sorting by level group and index, same as what happens during inference. Also, my experiments indicated no benefit from using the hover durations, so after sorting i dropped the hover rows and re-indexed each session from 0 to len(session).\n\n# Transformer\nI spent much of my time experimenting with transformers, which resulted in a light weight model that achieved 0.698 on the public and private LB, and 0.702 CV. \n\n\n```\nclass NN(nn.Module):\n    def __init__(self, num_cont_cols, embed_dim, num_layers, num_heads, max_seq_len):\n        super(NN, self).__init__()\n        self.emb_cont = nn.Sequential(\n            nn.Linear(num_cont_cols, embed_dim//2),\n            nn.LayerNorm(embed_dim//2)\n        )\n        self.emb_cats = nn.Sequential(\n            nn.Embedding(max_seq_len + 1, embed_dim//2),\n            nn.LayerNorm(embed_dim//2)\n        )\n        encoder_layer = nn.TransformerEncoderLayer(\n            d_model=embed_dim,\n            nhead=num_heads,\n            dim_feedforward=embed_dim,\n            dropout=0.1,\n            batch_first=True,\n            activation=\"relu\",\n        )\n        self.encoder = nn.TransformerEncoder(encoder_layer, num_layers=num_layers)\n        self.clf_heads = nn.ModuleList([\n            nn.Linear(embed_dim, out_dim) for out_dim in [3, 10, 5]\n        ])\n\n    def forward(self, x, grp):\n        emb_conts = self.emb_cont(x[:, :, :-1])\n        emb_cats = self.emb_cats(x[:, :, -1].type(torch.int32))\n        x = torch.cat([emb_conts, emb_cats], dim=2)\n        x = self.encoder(x)\n        x = x.mean(dim=1)\n        x = self.clf_heads[[\"0-4\", \"5-12\", \"13-22\"].index(grp)](x)\n        return x.unsqueeze(2)\n```\n\n- embed_dim: 64\n- num_layers: 1\n- num_heads: 8\n- max_seq_len: 452 (explained below), though the sequences are cropped to 256\n- I used the same single model for all questions\n\nI found the data easy to overfit with transformers, so in an attempt to improve the signal to noise ratio i did the following:\n1. Identify different points in the game by string concatenating event_name, level, name, page, fqid, room_fqid, text_fqid, in the dataframe.\n2. Some of these occur more than once in a session. Treat these as different points by enumerating them and adding the enumeration to their names.\n3. Filter out the rows with points that is present in over 0.999 of the sessions. This makes each session maximum 452 steps long. \n4. Create 6 feature columns:\ntime difference, index difference, distance (cumulative distance moved, calculated from screen_coor's) difference, room_coor_x, room_coor_y and the categorical point column embedded.\n\n# XGBoost\nThis is my strongest single model with public LB 0.701, private 0.702, and 0.7029 CV. What stands out is that i flattened 5 of the transformer input columns (excluding the categorical column), and used all those values as individual features.\n\nThe other features are mainly stats that can be found in public kernels, like mean and max time diff over the categoricals. The stats were calculated before applying the transformer input filtering.\n\nFrom early on I trained one model for each level group, inputing the question number as a feature. I found that CV increased by around 0.0002 compared to using a model for each question. This could be randomness, but i went with it since i thought 3 models instead of 18 would make my life easier during experimentation. Similar reasoning behind using one model for the transformer. \n\n# Catboost\nEssentially looks the same as XGBoost. CV 0.7022.\n\n# Ensemble\nI trained a linear regression meta model for each question, using the above models output probabilites as input, to produce the final predictions. I included probabilities of past questions and some future ones! For example the regression model trained to predict question 2 took probabilities on question 1-3 as input, to predict question 7 I used probabilities on question 1-13, and for question 16 i used probabilities on question 1-18. I took 3 seeds average before linear regression input to make it more robust. \n\nThis finally results in public LB 0.702, private LB 0.703, CV 0.7044. \n\n# On threshold and submission selection\nI tried to trust CV as much as possible, but the consistent gap between my CV and LB was suspicius until the last few days. Then I realized one reason could be my selected threshold was suboptimal on the test data.. I made some submissions with my highest CV solution, only changing the threshold, and noticed it was indeed suboptimal and caused more variation in LB score than most of my latest experiments. So in the end I selected 3 of the same solution, with different thresholds: 0.60 (best on LB), 0.62 (best during CV) and 0.64. Turned out 0.61 would have resulted in 0.704 private, but no regrets ;)\n\nThank you for reading!\n\n# Code\n[Training code](https://github.com/joelerikanders/pspgp/tree/main)\n[Submission notebook](https://www.kaggle.com/erijoel/4th-place-submission)",
      "votes": 48
    },
    {
      "id": 2325988,
      "postDate": "2023-07-01T18:39:30.567Z",
      "content": "<p>I considered that thresholds could be impactful (either due to randomness or some 'real' reason) but didn't officially test CV vs public LB thresholds. </p>\n<p>If wanting to really deep dive on CV higher than LB, though, nested CV might've been the right approach? I.e outer fold 1 does nested training for the sole purpose of determining threshold (and ensemble weights? And hyperparams?), then outputs 0 and 1 outer CV predictions based on fold 1 specific threshold value. So on for all folds and logically it would be more accurate or slightly pessimistic CV instead of literally optimistic (aka optimized) CV. </p>",
      "rawMarkdown": "I considered that thresholds could be impactful (either due to randomness or some 'real' reason) but didn't officially test CV vs public LB thresholds. \n\nIf wanting to really deep dive on CV higher than LB, though, nested CV might've been the right approach? I.e outer fold 1 does nested training for the sole purpose of determining threshold (and ensemble weights? And hyperparams?), then outputs 0 and 1 outer CV predictions based on fold 1 specific threshold value. So on for all folds and logically it would be more accurate or slightly pessimistic CV instead of literally optimistic (aka optimized) CV. ",
      "votes": 1,
      "replies": [
        {
          "id": 2326065,
          "postDate": "2023-07-01T21:10:39.327Z",
          "content": "<p>I think you're right that nested CV would have lead to more accurate or slightly pessimistic CV. Though I'm not sure whether or not it would have been a better approach, or lead to better thresholds in the end, maybe. I suppose it's a tradeoff between fast iteration and accuracy of the experiments</p>",
          "rawMarkdown": "I think you're right that nested CV would have lead to more accurate or slightly pessimistic CV. Though I'm not sure whether or not it would have been a better approach, or lead to better thresholds in the end, maybe. I suppose it's a tradeoff between fast iteration and accuracy of the experiments",
          "replies": [
            {
              "id": 2326093,
              "postDate": "2023-07-01T22:14:02.117Z",
              "content": "<p>Yeah, it wouldn't change the threshold picked, still probably pick based on regular CV. It just might have given more insight and peace of mind about CV vs LB difference. Which might not even be worth the time and effort, lol. ;)</p>",
              "rawMarkdown": "Yeah, it wouldn't change the threshold picked, still probably pick based on regular CV. It just might have given more insight and peace of mind about CV vs LB difference. Which might not even be worth the time and effort, lol. ;)"
            }
          ]
        }
      ]
    },
    {
      "id": 2325660,
      "postDate": "2023-07-01T14:02:40.183Z",
      "content": "<p><a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> Congratulations on your impressive achievement and thank you for sharing your journey! Your ensemble approach with Transformers, XGBoost, and Catboost, along with the meta-modeling technique, is quite insightful.</p>",
      "rawMarkdown": "@erijoel Congratulations on your impressive achievement and thank you for sharing your journey! Your ensemble approach with Transformers, XGBoost, and Catboost, along with the meta-modeling technique, is quite insightful.",
      "votes": 1
    },
    {
      "id": 2325489,
      "postDate": "2023-07-01T12:06:35.923Z",
      "content": "<p>The 4th place solution in the competition used an ensemble of Transformer, XGBoost, and Catboost models, along with linear regression as a meta-model. The participant applied feature engineering and experimented with different configurations to achieve high scores in both public and private leaderboards.</p>",
      "rawMarkdown": "The 4th place solution in the competition used an ensemble of Transformer, XGBoost, and Catboost models, along with linear regression as a meta-model. The participant applied feature engineering and experimented with different configurations to achieve high scores in both public and private leaderboards.",
      "votes": 1
    },
    {
      "id": 2325285,
      "postDate": "2023-07-01T08:22:05.337Z",
      "content": "<p>Congratulations on the solo gold!!</p>",
      "rawMarkdown": "Congratulations on the solo gold!!",
      "votes": 1
    },
    {
      "id": 2325113,
      "postDate": "2023-07-01T06:09:03.983Z",
      "content": "<p><a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> Congratulations. 4th place is quite a big achievement.</p>",
      "rawMarkdown": "@erijoel Congratulations. 4th place is quite a big achievement.",
      "votes": 1
    },
    {
      "id": 2324223,
      "postDate": "2023-06-30T13:31:13.083Z",
      "content": "<p>Congratulations on coming 4th <a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> . Keep achieving new milestones and medals 👍</p>",
      "rawMarkdown": "Congratulations on coming 4th @erijoel . Keep achieving new milestones and medals 👍",
      "votes": 1,
      "replies": [
        {
          "id": 2324247,
          "postDate": "2023-06-30T13:43:12.957Z",
          "content": "<p>Thank you 😄</p>",
          "rawMarkdown": "Thank you 😄",
          "votes": 1
        }
      ]
    },
    {
      "id": 2324136,
      "postDate": "2023-06-30T12:07:37.917Z",
      "content": "<p>Congrats your first gold medal !!</p>",
      "rawMarkdown": "Congrats your first gold medal !!",
      "votes": 1,
      "replies": [
        {
          "id": 2324174,
          "postDate": "2023-06-30T12:43:47.090Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!"
        }
      ]
    },
    {
      "id": 2343917,
      "postDate": "2023-07-14T05:14:56.113Z",
      "content": "<p>Congratulations and thank you for sharing the solution. It is interesting to review the solution and serves as a good learning resource. </p>",
      "rawMarkdown": "Congratulations and thank you for sharing the solution. It is interesting to review the solution and serves as a good learning resource. ",
      "votes": 2
    },
    {
      "id": 2326741,
      "postDate": "2023-07-02T11:14:39.087Z",
      "content": "<p>Thanks for sharing. I tried ensembling via linear regression per question, and while it improves Cv, LB and private is a bit lower (LB 0.704) than a single linear regression LB 0.705. I was surprised by that, and now I am surprised it worked for you.</p>",
      "rawMarkdown": "Thanks for sharing. I tried ensembling via linear regression per question, and while it improves Cv, LB and private is a bit lower (LB 0.704) than a single linear regression LB 0.705. I was surprised by that, and now I am surprised it worked for you.",
      "votes": 2,
      "replies": [
        {
          "id": 2327374,
          "postDate": "2023-07-02T22:56:34.347Z",
          "content": "<p>Thanks for the comment! Surprising indeed, do you think the difference between question wise and single model  score is due to randomness or something else? The only time I tried linear regression I did it question wise with previous probabilities the way I described</p>",
          "rawMarkdown": "Thanks for the comment! Surprising indeed, do you think the difference between question wise and single model  score is due to randomness or something else? The only time I tried linear regression I did it question wise with previous probabilities the way I described"
        }
      ]
    },
    {
      "id": 2325228,
      "postDate": "2023-07-01T07:30:17.983Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2325988,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-07-01T18:39:30.567000",
      "content": "<p>I considered that thresholds could be impactful (either due to randomness or some 'real' reason) but didn't officially test CV vs public LB thresholds. </p>\n<p>If wanting to really deep dive on CV higher than LB, though, nested CV might've been the right approach? I.e outer fold 1 does nested training for the sole purpose of determining threshold (and ensemble weights? And hyperparams?), then outputs 0 and 1 outer CV predictions based on fold 1 specific threshold value. So on for all folds and logically it would be more accurate or slightly pessimistic CV instead of literally optimistic (aka optimized) CV. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2326065,
          "author_name": "Joel Erikanders",
          "author_url": "",
          "post_date": "2023-07-01T21:10:39.327000",
          "content": "<p>I think you're right that nested CV would have lead to more accurate or slightly pessimistic CV. Though I'm not sure whether or not it would have been a better approach, or lead to better thresholds in the end, maybe. I suppose it's a tradeoff between fast iteration and accuracy of the experiments</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2326093,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-07-01T22:14:02.117000",
              "content": "<p>Yeah, it wouldn't change the threshold picked, still probably pick based on regular CV. It just might have given more insight and peace of mind about CV vs LB difference. Which might not even be worth the time and effort, lol. ;)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2325660,
      "author_name": "Akshay Vyas",
      "author_url": "",
      "post_date": "2023-07-01T14:02:40.183000",
      "content": "<p><a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> Congratulations on your impressive achievement and thank you for sharing your journey! Your ensemble approach with Transformers, XGBoost, and Catboost, along with the meta-modeling technique, is quite insightful.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325489,
      "author_name": "Pooja Chauhan",
      "author_url": "",
      "post_date": "2023-07-01T12:06:35.923000",
      "content": "<p>The 4th place solution in the competition used an ensemble of Transformer, XGBoost, and Catboost models, along with linear regression as a meta-model. The participant applied feature engineering and experimented with different configurations to achieve high scores in both public and private leaderboards.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325285,
      "author_name": "Priyanshu Chaudhary",
      "author_url": "",
      "post_date": "2023-07-01T08:22:05.337000",
      "content": "<p>Congratulations on the solo gold!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2325113,
      "author_name": "Pranav Jadhav",
      "author_url": "",
      "post_date": "2023-07-01T06:09:03.983000",
      "content": "<p><a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> Congratulations. 4th place is quite a big achievement.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324223,
      "author_name": "Swapnil Chowdhury",
      "author_url": "",
      "post_date": "2023-06-30T13:31:13.083000",
      "content": "<p>Congratulations on coming 4th <a href=\"https://www.kaggle.com/erijoel\" target=\"_blank\">@erijoel</a> . Keep achieving new milestones and medals 👍</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2324247,
          "author_name": "Joel Erikanders",
          "author_url": "",
          "post_date": "2023-06-30T13:43:12.957000",
          "content": "<p>Thank you 😄</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2324136,
      "author_name": "DongYK",
      "author_url": "",
      "post_date": "2023-06-30T12:07:37.917000",
      "content": "<p>Congrats your first gold medal !!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2324174,
          "author_name": "Joel Erikanders",
          "author_url": "",
          "post_date": "2023-06-30T12:43:47.090000",
          "content": "<p>Thanks a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2343917,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-07-14T05:14:56.113000",
      "content": "<p>Congratulations and thank you for sharing the solution. It is interesting to review the solution and serves as a good learning resource. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2326741,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2023-07-02T11:14:39.087000",
      "content": "<p>Thanks for sharing. I tried ensembling via linear regression per question, and while it improves Cv, LB and private is a bit lower (LB 0.704) than a single linear regression LB 0.705. I was surprised by that, and now I am surprised it worked for you.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2327374,
          "author_name": "Joel Erikanders",
          "author_url": "",
          "post_date": "2023-07-02T22:56:34.347000",
          "content": "<p>Thanks for the comment! Surprising indeed, do you think the difference between question wise and single model  score is due to randomness or something else? The only time I tried linear regression I did it question wise with previous probabilities the way I described</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2325228,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-07-01T07:30:17.983000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2324073": "# Acknowledgement\nI'd like to thank the hosts for providing a very interesting and difficult project to work on for the past months. I am also grateful for all the public sharing on Kaggle, this has been an insane learning experience for me. Without all the public notebooks, discussion posts and old competition solutions available i would have had no chance in this competition.\n\n# Overview\n- Used most of the raw data for training, while validating only on the kaggle data.\n- Ensemble of Transformer, XGBoost and Catboost, with 3 seeds and 5 folds each. \n- Used a generic set of features based on time, index and screen_coor differences.\n- Linear regression as a meta model.\n- Thresholds have a big impact on LB score\n\n# Data\nI used most of the raw data for training, including sessions that only completed level group 0-4 and 5-12. About ~38000 whole sessions and ~58000 sessions in total. Using the raw data increased CV by over 0.001. I validated only on the kaggle data. \n\nMy initial data preprocessing is simply sorting by level group and index, same as what happens during inference. Also, my experiments indicated no benefit from using the hover durations, so after sorting i dropped the hover rows and re-indexed each session from 0 to len(session).\n\n# Transformer\nI spent much of my time experimenting with transformers, which resulted in a light weight model that achieved 0.698 on the public and private LB, and 0.702 CV. \n\n\n```\nclass NN(nn.Module):\n    def __init__(self, num_cont_cols, embed_dim, num_layers, num_heads, max_seq_len):\n        super(NN, self).__init__()\n        self.emb_cont = nn.Sequential(\n            nn.Linear(num_cont_cols, embed_dim//2),\n            nn.LayerNorm(embed_dim//2)\n        )\n        self.emb_cats = nn.Sequential(\n            nn.Embedding(max_seq_len + 1, embed_dim//2),\n            nn.LayerNorm(embed_dim//2)\n        )\n        encoder_layer = nn.TransformerEncoderLayer(\n            d_model=embed_dim,\n            nhead=num_heads,\n            dim_feedforward=embed_dim,\n            dropout=0.1,\n            batch_first=True,\n            activation=\"relu\",\n        )\n        self.encoder = nn.TransformerEncoder(encoder_layer, num_layers=num_layers)\n        self.clf_heads = nn.ModuleList([\n            nn.Linear(embed_dim, out_dim) for out_dim in [3, 10, 5]\n        ])\n\n    def forward(self, x, grp):\n        emb_conts = self.emb_cont(x[:, :, :-1])\n        emb_cats = self.emb_cats(x[:, :, -1].type(torch.int32))\n        x = torch.cat([emb_conts, emb_cats], dim=2)\n        x = self.encoder(x)\n        x = x.mean(dim=1)\n        x = self.clf_heads[[\"0-4\", \"5-12\", \"13-22\"].index(grp)](x)\n        return x.unsqueeze(2)\n```\n\n- embed_dim: 64\n- num_layers: 1\n- num_heads: 8\n- max_seq_len: 452 (explained below), though the sequences are cropped to 256\n- I used the same single model for all questions\n\nI found the data easy to overfit with transformers, so in an attempt to improve the signal to noise ratio i did the following:\n1. Identify different points in the game by string concatenating event_name, level, name, page, fqid, room_fqid, text_fqid, in the dataframe.\n2. Some of these occur more than once in a session. Treat these as different points by enumerating them and adding the enumeration to their names.\n3. Filter out the rows with points that is present in over 0.999 of the sessions. This makes each session maximum 452 steps long. \n4. Create 6 feature columns:\ntime difference, index difference, distance (cumulative distance moved, calculated from screen_coor's) difference, room_coor_x, room_coor_y and the categorical point column embedded.\n\n# XGBoost\nThis is my strongest single model with public LB 0.701, private 0.702, and 0.7029 CV. What stands out is that i flattened 5 of the transformer input columns (excluding the categorical column), and used all those values as individual features.\n\nThe other features are mainly stats that can be found in public kernels, like mean and max time diff over the categoricals. The stats were calculated before applying the transformer input filtering.\n\nFrom early on I trained one model for each level group, inputing the question number as a feature. I found that CV increased by around 0.0002 compared to using a model for each question. This could be randomness, but i went with it since i thought 3 models instead of 18 would make my life easier during experimentation. Similar reasoning behind using one model for the transformer. \n\n# Catboost\nEssentially looks the same as XGBoost. CV 0.7022.\n\n# Ensemble\nI trained a linear regression meta model for each question, using the above models output probabilites as input, to produce the final predictions. I included probabilities of past questions and some future ones! For example the regression model trained to predict question 2 took probabilities on question 1-3 as input, to predict question 7 I used probabilities on question 1-13, and for question 16 i used probabilities on question 1-18. I took 3 seeds average before linear regression input to make it more robust. \n\nThis finally results in public LB 0.702, private LB 0.703, CV 0.7044. \n\n# On threshold and submission selection\nI tried to trust CV as much as possible, but the consistent gap between my CV and LB was suspicius until the last few days. Then I realized one reason could be my selected threshold was suboptimal on the test data.. I made some submissions with my highest CV solution, only changing the threshold, and noticed it was indeed suboptimal and caused more variation in LB score than most of my latest experiments. So in the end I selected 3 of the same solution, with different thresholds: 0.60 (best on LB), 0.62 (best during CV) and 0.64. Turned out 0.61 would have resulted in 0.704 private, but no regrets ;)\n\nThank you for reading!\n\n# Code\n[Training code](https://github.com/joelerikanders/pspgp/tree/main)\n[Submission notebook](https://www.kaggle.com/erijoel/4th-place-submission)",
    "2325988": "I considered that thresholds could be impactful (either due to randomness or some 'real' reason) but didn't officially test CV vs public LB thresholds. \n\nIf wanting to really deep dive on CV higher than LB, though, nested CV might've been the right approach? I.e outer fold 1 does nested training for the sole purpose of determining threshold (and ensemble weights? And hyperparams?), then outputs 0 and 1 outer CV predictions based on fold 1 specific threshold value. So on for all folds and logically it would be more accurate or slightly pessimistic CV instead of literally optimistic (aka optimized) CV. ",
    "2325660": "@erijoel Congratulations on your impressive achievement and thank you for sharing your journey! Your ensemble approach with Transformers, XGBoost, and Catboost, along with the meta-modeling technique, is quite insightful.",
    "2325489": "The 4th place solution in the competition used an ensemble of Transformer, XGBoost, and Catboost models, along with linear regression as a meta-model. The participant applied feature engineering and experimented with different configurations to achieve high scores in both public and private leaderboards.",
    "2325285": "Congratulations on the solo gold!!",
    "2325113": "@erijoel Congratulations. 4th place is quite a big achievement.",
    "2324223": "Congratulations on coming 4th @erijoel . Keep achieving new milestones and medals 👍",
    "2324136": "Congrats your first gold medal !!",
    "2343917": "Congratulations and thank you for sharing the solution. It is interesting to review the solution and serves as a good learning resource. ",
    "2326741": "Thanks for sharing. I tried ensembling via linear regression per question, and while it improves Cv, LB and private is a bit lower (LB 0.704) than a single linear regression LB 0.705. I was surprised by that, and now I am surprised it worked for you.",
    "2325228": ""
  }
}