{
  "id": 420077,
  "title": "13th place solution",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/takoi-13th-place-solution",
  "author_name": "",
  "post_date": "2023-06-29T10:37:44.580Z",
  "votes": 52,
  "comment_count": 18,
  "views": 0,
  "content": "<h1>13th place solution</h1>\n<p>First of all, I would like to thank the Kaggle community for sharing great ideas and engaging discussions. I would also like to thank the hosts for organizing this interesting task competition.</p>\n<h2>Summary</h2>\n<ul>\n<li>Ensemble of LightGBM and NN</li>\n<li>Cross validation: Nested cross validation<ul>\n<li>Training data: Data for which the first four digits of the session_id are 2200 or less.</li>\n<li>Validation data: Data for which the first four digits of the session_id are 2201 or more.</li>\n<li>I trained the model on the training data using a 5-fold cross validation strategy, and evaluated it on the validation data using predictions from all 5 trained models.</li>\n<li>For the final submission, I trained the model on the entire dataset using a 5-fold cross validation approach.</li></ul></li>\n</ul>\n<h2>LightGBM</h2>\n<ul>\n<li>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.</li>\n<li>For level group 13-22, I trained a distinct model for each target.</li>\n<li>Main features:<ul>\n<li>The count of categorical data for each session_id</li>\n<li>The statistical measures of numerical data for each session_id</li>\n<li>An aggregate of the next action taken </li></ul></li>\n<li>Scores<ul>\n<li>CV : 0.7032</li>\n<li>Public Score : 0.704</li>\n<li>Private Score : 0.701</li></ul></li>\n</ul>\n<h2>NN</h2>\n<ul>\n<li>Model: Transformer + GRU<ul>\n<li>The standalone Transformer didn't perform very well.</li>\n<li>The addition of GRU improved the score.</li></ul></li>\n<li>Trained with fewer features</li>\n<li>I trained a separate model for each level group.</li>\n<li>Scores<ul>\n<li>CV : 0.7010</li>\n<li>Public Score: 0.700</li>\n<li>Private Score: 0.700</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>LightGBM * 0.66 + NN * 0.34</li>\n<li>CV : 0.7053</li>\n<li>Public Score : 0.706</li>\n<li>Private Score : 0.702</li>\n</ul>",
  "messages": [
    {
      "id": "2322141",
      "postDate": "06/29/2023 05:25:46",
      "content": "<h1>13th place solution</h1>\n<p>First of all, I would like to thank the Kaggle community for sharing great ideas and engaging discussions. I would also like to thank the hosts for organizing this interesting task competition.</p>\n<h2>Summary</h2>\n<ul>\n<li>Ensemble of LightGBM and NN</li>\n<li>Cross validation: Nested cross validation<ul>\n<li>Training data: Data for which the first four digits of the session_id are 2200 or less.</li>\n<li>Validation data: Data for which the first four digits of the session_id are 2201 or more.</li>\n<li>I trained the model on the training data using a 5-fold cross validation strategy, and evaluated it on the validation data using predictions from all 5 trained models.</li>\n<li>For the final submission, I trained the model on the entire dataset using a 5-fold cross validation approach.</li></ul></li>\n</ul>\n<h2>LightGBM</h2>\n<ul>\n<li>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.</li>\n<li>For level group 13-22, I trained a distinct model for each target.</li>\n<li>Main features:<ul>\n<li>The count of categorical data for each session_id</li>\n<li>The statistical measures of numerical data for each session_id</li>\n<li>An aggregate of the next action taken </li></ul></li>\n<li>Scores<ul>\n<li>CV : 0.7032</li>\n<li>Public Score : 0.704</li>\n<li>Private Score : 0.701</li></ul></li>\n</ul>\n<h2>NN</h2>\n<ul>\n<li>Model: Transformer + GRU<ul>\n<li>The standalone Transformer didn't perform very well.</li>\n<li>The addition of GRU improved the score.</li></ul></li>\n<li>Trained with fewer features</li>\n<li>I trained a separate model for each level group.</li>\n<li>Scores<ul>\n<li>CV : 0.7010</li>\n<li>Public Score: 0.700</li>\n<li>Private Score: 0.700</li></ul></li>\n</ul>\n<h2>Ensemble</h2>\n<ul>\n<li>LightGBM * 0.66 + NN * 0.34</li>\n<li>CV : 0.7053</li>\n<li>Public Score : 0.706</li>\n<li>Private Score : 0.702</li>\n</ul>",
      "rawMarkdown": "# 13th place solution\nFirst of all, I would like to thank the Kaggle community for sharing great ideas and engaging discussions. I would also like to thank the hosts for organizing this interesting task competition.\n\n## Summary\n- Ensemble of LightGBM and NN\n- Cross validation: Nested cross validation\n    - Training data: Data for which the first four digits of the session_id are 2200 or less.\n    - Validation data: Data for which the first four digits of the session_id are 2201 or more.\n    - I trained the model on the training data using a 5-fold cross validation strategy, and evaluated it on the validation data using predictions from all 5 trained models.\n    - For the final submission, I trained the model on the entire dataset using a 5-fold cross validation approach.\n\n## LightGBM\n- I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.\n- For level group 13-22, I trained a distinct model for each target.\n- Main features:\n    - The count of categorical data for each session_id\n    - The statistical measures of numerical data for each session_id\n    - An aggregate of the next action taken \n- Scores\n    - CV : 0.7032\n    - Public Score : 0.704\n    - Private Score : 0.701\n\n\n## NN\n- Model: Transformer + GRU\n    - The standalone Transformer didn't perform very well.\n    - The addition of GRU improved the score.\n- Trained with fewer features\n- I trained a separate model for each level group.\n- Scores\n    - CV : 0.7010\n    - Public Score: 0.700\n    - Private Score: 0.700\n\n## Ensemble\n    - LightGBM * 0.66 + NN * 0.34\n    - CV : 0.7053\n    - Public Score : 0.706\n    - Private Score : 0.702",
      "votes": null
    },
    {
      "id": "2322148",
      "postDate": "06/29/2023 05:38:59",
      "content": "<p>Congratulations on your gold medal!</p>\n<p>May I ask why you split the training and validation data this way?</p>\n<blockquote>\n  <p>Cross validation: Nested cross validation<br>\n  Training data: Data for which the first four digits of the session_id are 2200 or less.<br>\n  Validation data: Data for which the first four digits of the session_id are 2201 or more.</p>\n</blockquote>",
      "rawMarkdown": "Congratulations on your gold medal!\n\nMay I ask why you split the training and validation data this way?\n\n>Cross validation: Nested cross validation\n>Training data: Data for which the first four digits of the session_id are 2200 or less.\n>Validation data: Data for which the first four digits of the session_id are 2201 or more.",
      "votes": null
    },
    {
      "id": "2322154",
      "postDate": "06/29/2023 05:46:22",
      "content": "<p>Congratulations on your solo gold! Can you share your NN architecture? For me, NN alone doesn't work that well, but extracting embedding from NN and add them into Tabular features works pretty help. So I am curious how NN looks like in your case. And it would be greater that you could share the training details, some remarks, etc.</p>",
      "rawMarkdown": "Congratulations on your solo gold! Can you share your NN architecture? For me, NN alone doesn't work that well, but extracting embedding from NN and add them into Tabular features works pretty help. So I am curious how NN looks like in your case. And it would be greater that you could share the training details, some remarks, etc.",
      "votes": null
    },
    {
      "id": "2322176",
      "postDate": "06/29/2023 06:05:14",
      "content": "<p>Congratulations on gold medal.</p>\n<blockquote>\n  <p>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.For level group 13-22, I trained a distinct model for each target.</p>\n</blockquote>\n<p>May I ask the reason for this strategy? I find training model for each level_group/level has a very different cv/lb correlation.</p>",
      "rawMarkdown": "Congratulations on gold medal.\n\n>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.For level group 13-22, I trained a distinct model for each target.\n\nMay I ask the reason for this strategy? I find training model for each level_group/level has a very different cv/lb correlation.",
      "votes": null
    },
    {
      "id": "2322194",
      "postDate": "06/29/2023 06:22:21",
      "content": "<p>Thank you for your comment! I used nested cross-validation to evaluate for the following two reasons:</p>\n<ul>\n<li><p>The evaluation metric for this time fluctuates significantly based on the threshold, and just changing the threshold also significantly changed the Public Score. I wanted to create the same situation locally as the Public Score (including the ensemble of 5-fold) and check the change in CV due to changing the threshold.</p></li>\n<li><p>For instance, in situations like the one described in the link below, I believed that nested cross-validation would be more reliable than standard cross-validation.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682</a></p></li>\n</ul>",
      "rawMarkdown": "Thank you for your comment! I used nested cross-validation to evaluate for the following two reasons:\n\n- The evaluation metric for this time fluctuates significantly based on the threshold, and just changing the threshold also significantly changed the Public Score. I wanted to create the same situation locally as the Public Score (including the ensemble of 5-fold) and check the change in CV due to changing the threshold.\n\n- For instance, in situations like the one described in the link below, I believed that nested cross-validation would be more reliable than standard cross-validation.\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682",
      "votes": null
    },
    {
      "id": "2322229",
      "postDate": "06/29/2023 06:49:04",
      "content": "<p>Thank you for your comment! Also, congratulations on your gold medal! <br>\nFor instance, I created the following model for level group 5-12.</p>\n<pre><code>\nnum_cols_level2 = [,,\n                  ,]\n\n\ncat_cols = [,,,,]\n\n\nclass MeanPooling(nn.Module):\n    def __init__(self):\n        super(MeanPooling, self).__init__()\n\n    def forward(self, last_hidden_state, attention_mask):\n        input_mask_expanded = attention_mask.unsqueeze(-1).expand(last_hidden_state.size()).float()\n        sum_embeddings = torch.sum(last_hidden_state * input_mask_expanded, 1)\n        sum_mask = input_mask_expanded.sum(1)\n        sum_mask = torch.clamp(sum_mask, =1e-9)\n        mean_embeddings = sum_embeddings / sum_mask\n        return mean_embeddings\n\nclass PspTransformerLevel2Model(nn.Module):\n    def __init__(\n        self, =0.2,\n        =4,\n        =7, name_embedding_size = 12,\n        =12, event_embedding_size = 12,\n        =130, fqid_embedding_size = 24,\n        =20, room_fqid_embedding_size = 12,\n        =24,level_embedding_size=12,\n        categorical_linear_size = 120,\n        numeraical_linear_size = 24,\n        model_size = 160,\n        nhead = 16,\n        =480,\n        =10):\n        super(PspTransformerLevel2Model, self).__init__()\n        self.name_embedding = nn.Embedding(=input_name_nunique, \n                                           =name_embedding_size)\n        self.event_embedding = nn.Embedding(=input_event_nunique, \n                                           =event_embedding_size)\n        self.fqid_embedding = nn.Embedding(=input_fqid_nunique, \n                                           =fqid_embedding_size)\n        self.room_fqid_embedding = nn.Embedding(=input_room_fqid_nunique, \n                                           =room_fqid_embedding_size)\n        self.level_embedding = nn.Embedding(=input_level_nunique, \n                                           =level_embedding_size)\n        self.categorical_linear = nn.Sequential(\n                nn.Linear(name_embedding_size + event_embedding_size + \n                          fqid_embedding_size + room_fqid_embedding_size + \n                          level_embedding_size, categorical_linear_size),\n                nn.LayerNorm(categorical_linear_size)\n            )\n        self.numerical_linear  = nn.Sequential(\n                nn.Linear(input_numerical_size, numeraical_linear_size),\n                nn.LayerNorm(numeraical_linear_size)\n            )\n\n        self.linear1  = nn.Sequential(\n                nn.Linear(categorical_linear_size + numeraical_linear_size, \n                          model_size),\n                nn.LayerNorm(model_size),\n            )\n        self.transformer_encoder = TransformerEncoder(\n            encoder_layer = nn.TransformerEncoderLayer(=model_size, \n                                                       =nhead,\n                                                       =dim_feedforward ,\n                                                       =dropout),\n                                                     =1)\n        self.gru = nn.GRU(model_size, model_size,\n                            num_layers = 1, \n                            =,\n                            =)\n        self.linear_out  = nn.Sequential(\n                nn.Linear(model_size, \n                          out_size)\n            )\n        self.pool = MeanPooling()\n\n    def forward(self, numerical_array, name_array,event_array,\n                fqid_array, room_fqid_array, level_array, \n                mask,mask_for_pooling):\n\n        name_embedding = self.name_embedding(name_array)\n        event_embedding = self.event_embedding(event_array)\n        fqid_embedding = self.fqid_embedding(fqid_array)\n        room_fqid_embedding = self.room_fqid_embedding(room_fqid_array)\n        level_embedding = self.level_embedding(level_array)\n        categorical_emedding = torch.cat([name_embedding,\n                                          event_embedding,\n                                          fqid_embedding, \n                                          room_fqid_embedding,\n                                          level_embedding\n                                          ], =2)\n        categorical_emedding = self.categorical_linear(categorical_emedding)\n        numerical_embedding = self.numerical_linear(numerical_array)\n        concat_embedding = torch.cat([categorical_emedding,\n                                      numerical_embedding],=2)\n        concat_embedding = self.linear1(concat_embedding)\n        concat_embedding  = concat_embedding.permute(1,0,2).contiguous()\n        output = self.transformer_encoder(concat_embedding, \n                                          =mask)\n        output = output.permute(1,0,2).contiguous()\n        output,_ = self.gru(output)\n        output = self.pool(output,mask_for_pooling)\n        output = self.linear_out(output)\n        return output\n</code></pre>\n<p>The main improvements I made are listed below. Through a series of experiments, I found that the following practices were advantageous:</p>\n<ul>\n<li>Keeping the model size small:<ul>\n<li>For instance, I kept the number of layers for both the Transformer and GRU at 1.</li>\n<li>I didn't enlarge the encoder layer or hidden size.</li>\n<li>I also maintained a relatively small embedding size for categorical features.</li></ul></li>\n<li>Using fewer features:<ul>\n<li>For instance, I didn't use features such as screen_coor_x, screen_coor_y, hover_duration, and text.</li></ul></li>\n</ul>",
      "rawMarkdown": "Thank you for your comment! Also, congratulations on your gold medal! \nFor instance, I created the following model for level group 5-12.\n\n```\n# Numerical features used for NN\nnum_cols_level2 = [\"elapsed_time_log1p\",\"elapsed_time_diff_log1p\",\n                  \"room_coor_x\",\"room_coor_y\"]\n\n# Categorical features used for NN\ncat_cols = [\"event_name\",\"name\",\"fqid\",\"room_fqid\",\"level\"]\n\n\nclass MeanPooling(nn.Module):\n    def __init__(self):\n        super(MeanPooling, self).__init__()\n        \n    def forward(self, last_hidden_state, attention_mask):\n        input_mask_expanded = attention_mask.unsqueeze(-1).expand(last_hidden_state.size()).float()\n        sum_embeddings = torch.sum(last_hidden_state * input_mask_expanded, 1)\n        sum_mask = input_mask_expanded.sum(1)\n        sum_mask = torch.clamp(sum_mask, min=1e-9)\n        mean_embeddings = sum_embeddings / sum_mask\n        return mean_embeddings\n\nclass PspTransformerLevel2Model(nn.Module):\n    def __init__(\n        self, dropout=0.2,\n        input_numerical_size=4,\n        input_name_nunique=7, name_embedding_size = 12,\n        input_event_nunique=12, event_embedding_size = 12,\n        input_fqid_nunique=130, fqid_embedding_size = 24,\n        input_room_fqid_nunique=20, room_fqid_embedding_size = 12,\n        input_level_nunique=24,level_embedding_size=12,\n        categorical_linear_size = 120,\n        numeraical_linear_size = 24,\n        model_size = 160,\n        nhead = 16,\n        dim_feedforward=480,\n        out_size=10):\n        super(PspTransformerLevel2Model, self).__init__()\n        self.name_embedding = nn.Embedding(num_embeddings=input_name_nunique, \n                                           embedding_dim=name_embedding_size)\n        self.event_embedding = nn.Embedding(num_embeddings=input_event_nunique, \n                                           embedding_dim=event_embedding_size)\n        self.fqid_embedding = nn.Embedding(num_embeddings=input_fqid_nunique, \n                                           embedding_dim=fqid_embedding_size)\n        self.room_fqid_embedding = nn.Embedding(num_embeddings=input_room_fqid_nunique, \n                                           embedding_dim=room_fqid_embedding_size)\n        self.level_embedding = nn.Embedding(num_embeddings=input_level_nunique, \n                                           embedding_dim=level_embedding_size)\n        self.categorical_linear = nn.Sequential(\n                nn.Linear(name_embedding_size + event_embedding_size + \n                          fqid_embedding_size + room_fqid_embedding_size + \n                          level_embedding_size, categorical_linear_size),\n                nn.LayerNorm(categorical_linear_size)\n            )\n        self.numerical_linear  = nn.Sequential(\n                nn.Linear(input_numerical_size, numeraical_linear_size),\n                nn.LayerNorm(numeraical_linear_size)\n            )\n        \n        self.linear1  = nn.Sequential(\n                nn.Linear(categorical_linear_size + numeraical_linear_size, \n                          model_size),\n                nn.LayerNorm(model_size),\n            )\n        self.transformer_encoder = TransformerEncoder(\n            encoder_layer = nn.TransformerEncoderLayer(d_model=model_size, \n                                                       nhead=nhead,\n                                                       dim_feedforward=dim_feedforward ,\n                                                       dropout=dropout),\n                                                     num_layers=1)\n        self.gru = nn.GRU(model_size, model_size,\n                            num_layers = 1, \n                            batch_first=True,\n                            bidirectional=True)\n        self.linear_out  = nn.Sequential(\n                nn.Linear(model_size*2, \n                          out_size)\n            )\n        self.pool = MeanPooling()\n    \n    def forward(self, numerical_array, name_array,event_array,\n                fqid_array, room_fqid_array, level_array, \n                mask,mask_for_pooling):\n    \n        name_embedding = self.name_embedding(name_array)\n        event_embedding = self.event_embedding(event_array)\n        fqid_embedding = self.fqid_embedding(fqid_array)\n        room_fqid_embedding = self.room_fqid_embedding(room_fqid_array)\n        level_embedding = self.level_embedding(level_array)\n        categorical_emedding = torch.cat([name_embedding,\n                                          event_embedding,\n                                          fqid_embedding, \n                                          room_fqid_embedding,\n                                          level_embedding\n                                          ], axis=2)\n        categorical_emedding = self.categorical_linear(categorical_emedding)\n        numerical_embedding = self.numerical_linear(numerical_array)\n        concat_embedding = torch.cat([categorical_emedding,\n                                      numerical_embedding],axis=2)\n        concat_embedding = self.linear1(concat_embedding)\n        concat_embedding  = concat_embedding.permute(1,0,2).contiguous()\n        output = self.transformer_encoder(concat_embedding, \n                                          src_key_padding_mask=mask)\n        output = output.permute(1,0,2).contiguous()\n        output,_ = self.gru(output)\n        output = self.pool(output,mask_for_pooling)\n        output = self.linear_out(output)\n        return output\n\n```\n\n\nThe main improvements I made are listed below. Through a series of experiments, I found that the following practices were advantageous:\n\n- Keeping the model size small:\n    - For instance, I kept the number of layers for both the Transformer and GRU at 1.\n    - I didn't enlarge the encoder layer or hidden size.\n    - I also maintained a relatively small embedding size for categorical features.\n- Using fewer features:\n    - For instance, I didn't use features such as screen_coor_x, screen_coor_y, hover_duration, and text.",
      "votes": null
    },
    {
      "id": "2322243",
      "postDate": "06/29/2023 06:59:53",
      "content": "<p>Thank you for your comment! <br>\nRegarding level groups 0-4 and 5-12, I adopted the above method because creating a single model for them all, rather than creating a model for each target, resulted in better CV and public score.</p>",
      "rawMarkdown": "Thank you for your comment! \nRegarding level groups 0-4 and 5-12, I adopted the above method because creating a single model for them all, rather than creating a model for each target, resulted in better CV and public score.",
      "votes": null
    },
    {
      "id": "2322245",
      "postDate": "06/29/2023 07:05:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a>,</p>\n<p>Congratulations on your solo gold medal! I have some questions about your solutions described as follows:</p>\n<ol>\n<li>For LightGBM, how many features are considered for each <code>level_group</code> model?</li>\n<li>For NN, how did you choose the sequence length of the event records?</li>\n</ol>\n<p>Thanks a lot for the sharing!</p>",
      "rawMarkdown": "Hi @takoihiraokazu,\n\nCongratulations on your solo gold medal! I have some questions about your solutions described as follows:\n1. For LightGBM, how many features are considered for each `level_group` model?\n2. For NN, how did you choose the sequence length of the event records?\n\nThanks a lot for the sharing!",
      "votes": null
    },
    {
      "id": "2322499",
      "postDate": "06/29/2023 10:11:40",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> congrats on solo gold medal!</p>",
      "rawMarkdown": "takoihiraokazu congrats on solo gold medal!",
      "votes": null
    },
    {
      "id": "2322729",
      "postDate": "06/29/2023 13:08:32",
      "content": "<p>Thanks!</p>\n<ol>\n<li><p>For the LightGBM model, the number of features considered for each level_group model is as follows: For level_group 0-4, I consider 1570 features. For level_group 5-12, I consider 3992 features. And for level_group 13-22, I consider 5290 features.</p></li>\n<li><p>As for the Neural Network, the sequence length of the event records was determined based on cross-validation results. Specifically, for level_group 0-4, I chose a sequence length of 250. For level_group 5-12, the sequence length is 500. And for level_group 13-22, the sequence length is 800.</p></li>\n</ol>",
      "rawMarkdown": "Thanks!\n\n1. For the LightGBM model, the number of features considered for each level_group model is as follows: For level_group 0-4, I consider 1570 features. For level_group 5-12, I consider 3992 features. And for level_group 13-22, I consider 5290 features.\n\n2. As for the Neural Network, the sequence length of the event records was determined based on cross-validation results. Specifically, for level_group 0-4, I chose a sequence length of 250. For level_group 5-12, the sequence length is 500. And for level_group 13-22, the sequence length is 800.",
      "votes": null
    },
    {
      "id": "2322764",
      "postDate": "06/29/2023 13:37:04",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> Congratulations on your 13th place finish! Your strategy of training separate models for different level groups and incorporating features like categorical and numerical data counts demonstrates a thoughtful approach. </p>",
      "rawMarkdown": "takoihiraokazu Congratulations on your 13th place finish! Your strategy of training separate models for different level groups and incorporating features like categorical and numerical data counts demonstrates a thoughtful approach.",
      "votes": null
    },
    {
      "id": "2322853",
      "postDate": "06/29/2023 14:28:12",
      "content": "<p>Thanks for the quick reply, what a robust DL-based model! One more question if you don't mind sharing, did you use early stopping to choose the best checkpoints or just let the training process converge to some satisfactory position?</p>",
      "rawMarkdown": "Thanks for the quick reply, what a robust DL-based model! One more question if you don't mind sharing, did you use early stopping to choose the best checkpoints or just let the training process converge to some satisfactory position?",
      "votes": null
    },
    {
      "id": "2323866",
      "postDate": "06/30/2023 07:54:35",
      "content": "<p>Congratulations on your 13th place solution! It's impressive that you combined LightGBM and neural networks (NN) in an ensemble approach. Your use of nested cross-validation and careful splitting of training and validation data demonstrates a thoughtful approach. Well done!</p>",
      "rawMarkdown": "Congratulations on your 13th place solution! It's impressive that you combined LightGBM and neural networks (NN) in an ensemble approach. Your use of nested cross-validation and careful splitting of training and validation data demonstrates a thoughtful approach. Well done!",
      "votes": null
    },
    {
      "id": "2324140",
      "postDate": "06/30/2023 12:10:16",
      "content": "<p>+1 congrats on gold</p>",
      "rawMarkdown": "1 congrats on gold",
      "votes": null
    },
    {
      "id": "2324257",
      "postDate": "06/30/2023 13:51:48",
      "content": "<p>I didn't exactly use early stopping. Instead, I set the number of epochs to 20 and saved the checkpoint that produced the best validation score. The choice to use 20 epochs was based on several experiments, where this setting resulted in the best cross-validation performance.\"</p>",
      "rawMarkdown": "I didn't exactly use early stopping. Instead, I set the number of epochs to 20 and saved the checkpoint that produced the best validation score. The choice to use 20 epochs was based on several experiments, where this setting resulted in the best cross-validation performance.\"",
      "votes": null
    },
    {
      "id": "2324389",
      "postDate": "06/30/2023 15:10:00",
      "content": "<p>As mentioned in the NN section, you train separate model for each <code>level_group</code>. If I'm not mistaken, the model checkpoints used to obtain the best CV performance could be different for different <code>level_group</code>s. For example, epoch 15 for  <code>level_group</code> \"0-4\", 17 for \"5-12\" and 11 for \"13-22\" can lead to the best overall Macro-F1 on local CV. Then, what's the better way to find out this combination. During the competition, I found it hard to figure out an efficient way to determine which checkpoints to choose. Thanks a lot for your clarification and patience 😓</p>",
      "rawMarkdown": "As mentioned in the NN section, you train separate model for each `level_group`. If I'm not mistaken, the model checkpoints used to obtain the best CV performance could be different for different `level_group`s. For example, epoch 15 for  `level_group` \"0-4\", 17 for \"5-12\" and 11 for \"13-22\" can lead to the best overall Macro-F1 on local CV. Then, what's the better way to find out this combination. During the competition, I found it hard to figure out an efficient way to determine which checkpoints to choose. Thanks a lot for your clarification and patience 😓",
      "votes": null
    },
    {
      "id": "2324447",
      "postDate": "06/30/2023 15:54:09",
      "content": "<p>Thank you! Your practices are interesting. My NN was not super good, maybe it's because I kept the architecture and the input features complicated.</p>",
      "rawMarkdown": "Thank you! Your practices are interesting. My NN was not super good, maybe it's because I kept the architecture and the input features complicated.",
      "votes": null
    },
    {
      "id": "2324819",
      "postDate": "06/30/2023 22:40:26",
      "content": "<p>Yes, as you said, the optimal epoch does differ for each level group. As for the checkpoints, I decided on the one where the AUC was highest. Even then, if the AUC improved for each level group, the final macro-f1 score also improved.</p>",
      "rawMarkdown": "Yes, as you said, the optimal epoch does differ for each level group. As for the checkpoints, I decided on the one where the AUC was highest. Even then, if the AUC improved for each level group, the final macro-f1 score also improved.",
      "votes": null
    },
    {
      "id": "2325133",
      "postDate": "07/01/2023 06:21:08",
      "content": "<p>I see! What I did was to choose those with the highest Macro-F1 @ 0.63 (but, I knew optimizing locally on <code>level_group</code> models might not trasfer to global optimization…). I forgot to try to select checkpoints with other metrics. I'll go back to give them a try and do some late submissions. </p>\n<p>Thanks for the reply and kindness. Good luck with your next competition!</p>",
      "rawMarkdown": "I see! What I did was to choose those with the highest Macro-F1 @ 0.63 (but, I knew optimizing locally on `level_group` models might not trasfer to global optimization...). I forgot to try to select checkpoints with other metrics. I'll go back to give them a try and do some late submissions. \n\nThanks for the reply and kindness. Good luck with your next competition!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2322148,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 05:38:59",
      "content": "<p>Congratulations on your gold medal!</p>\n<p>May I ask why you split the training and validation data this way?</p>\n<blockquote>\n  <p>Cross validation: Nested cross validation<br>\n  Training data: Data for which the first four digits of the session_id are 2200 or less.<br>\n  Validation data: Data for which the first four digits of the session_id are 2201 or more.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 2322194,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/29/2023 06:22:21",
          "content": "<p>Thank you for your comment! I used nested cross-validation to evaluate for the following two reasons:</p>\n<ul>\n<li><p>The evaluation metric for this time fluctuates significantly based on the threshold, and just changing the threshold also significantly changed the Public Score. I wanted to create the same situation locally as the Public Score (including the ensemble of 5-fold) and check the change in CV due to changing the threshold.</p></li>\n<li><p>For instance, in situations like the one described in the link below, I believed that nested cross-validation would be more reliable than standard cross-validation.<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682</a></p></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2322154,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "06/29/2023 05:46:22",
      "content": "<p>Congratulations on your solo gold! Can you share your NN architecture? For me, NN alone doesn't work that well, but extracting embedding from NN and add them into Tabular features works pretty help. So I am curious how NN looks like in your case. And it would be greater that you could share the training details, some remarks, etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322229,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/29/2023 06:49:04",
          "content": "<p>Thank you for your comment! Also, congratulations on your gold medal! <br>\nFor instance, I created the following model for level group 5-12.</p>\n<pre><code>\nnum_cols_level2 = [,,\n                  ,]\n\n\ncat_cols = [,,,,]\n\n\nclass MeanPooling(nn.Module):\n    def __init__(self):\n        super(MeanPooling, self).__init__()\n\n    def forward(self, last_hidden_state, attention_mask):\n        input_mask_expanded = attention_mask.unsqueeze(-1).expand(last_hidden_state.size()).float()\n        sum_embeddings = torch.sum(last_hidden_state * input_mask_expanded, 1)\n        sum_mask = input_mask_expanded.sum(1)\n        sum_mask = torch.clamp(sum_mask, =1e-9)\n        mean_embeddings = sum_embeddings / sum_mask\n        return mean_embeddings\n\nclass PspTransformerLevel2Model(nn.Module):\n    def __init__(\n        self, =0.2,\n        =4,\n        =7, name_embedding_size = 12,\n        =12, event_embedding_size = 12,\n        =130, fqid_embedding_size = 24,\n        =20, room_fqid_embedding_size = 12,\n        =24,level_embedding_size=12,\n        categorical_linear_size = 120,\n        numeraical_linear_size = 24,\n        model_size = 160,\n        nhead = 16,\n        =480,\n        =10):\n        super(PspTransformerLevel2Model, self).__init__()\n        self.name_embedding = nn.Embedding(=input_name_nunique, \n                                           =name_embedding_size)\n        self.event_embedding = nn.Embedding(=input_event_nunique, \n                                           =event_embedding_size)\n        self.fqid_embedding = nn.Embedding(=input_fqid_nunique, \n                                           =fqid_embedding_size)\n        self.room_fqid_embedding = nn.Embedding(=input_room_fqid_nunique, \n                                           =room_fqid_embedding_size)\n        self.level_embedding = nn.Embedding(=input_level_nunique, \n                                           =level_embedding_size)\n        self.categorical_linear = nn.Sequential(\n                nn.Linear(name_embedding_size + event_embedding_size + \n                          fqid_embedding_size + room_fqid_embedding_size + \n                          level_embedding_size, categorical_linear_size),\n                nn.LayerNorm(categorical_linear_size)\n            )\n        self.numerical_linear  = nn.Sequential(\n                nn.Linear(input_numerical_size, numeraical_linear_size),\n                nn.LayerNorm(numeraical_linear_size)\n            )\n\n        self.linear1  = nn.Sequential(\n                nn.Linear(categorical_linear_size + numeraical_linear_size, \n                          model_size),\n                nn.LayerNorm(model_size),\n            )\n        self.transformer_encoder = TransformerEncoder(\n            encoder_layer = nn.TransformerEncoderLayer(=model_size, \n                                                       =nhead,\n                                                       =dim_feedforward ,\n                                                       =dropout),\n                                                     =1)\n        self.gru = nn.GRU(model_size, model_size,\n                            num_layers = 1, \n                            =,\n                            =)\n        self.linear_out  = nn.Sequential(\n                nn.Linear(model_size, \n                          out_size)\n            )\n        self.pool = MeanPooling()\n\n    def forward(self, numerical_array, name_array,event_array,\n                fqid_array, room_fqid_array, level_array, \n                mask,mask_for_pooling):\n\n        name_embedding = self.name_embedding(name_array)\n        event_embedding = self.event_embedding(event_array)\n        fqid_embedding = self.fqid_embedding(fqid_array)\n        room_fqid_embedding = self.room_fqid_embedding(room_fqid_array)\n        level_embedding = self.level_embedding(level_array)\n        categorical_emedding = torch.cat([name_embedding,\n                                          event_embedding,\n                                          fqid_embedding, \n                                          room_fqid_embedding,\n                                          level_embedding\n                                          ], =2)\n        categorical_emedding = self.categorical_linear(categorical_emedding)\n        numerical_embedding = self.numerical_linear(numerical_array)\n        concat_embedding = torch.cat([categorical_emedding,\n                                      numerical_embedding],=2)\n        concat_embedding = self.linear1(concat_embedding)\n        concat_embedding  = concat_embedding.permute(1,0,2).contiguous()\n        output = self.transformer_encoder(concat_embedding, \n                                          =mask)\n        output = output.permute(1,0,2).contiguous()\n        output,_ = self.gru(output)\n        output = self.pool(output,mask_for_pooling)\n        output = self.linear_out(output)\n        return output\n</code></pre>\n<p>The main improvements I made are listed below. Through a series of experiments, I found that the following practices were advantageous:</p>\n<ul>\n<li>Keeping the model size small:<ul>\n<li>For instance, I kept the number of layers for both the Transformer and GRU at 1.</li>\n<li>I didn't enlarge the encoder layer or hidden size.</li>\n<li>I also maintained a relatively small embedding size for categorical features.</li></ul></li>\n<li>Using fewer features:<ul>\n<li>For instance, I didn't use features such as screen_coor_x, screen_coor_y, hover_duration, and text.</li></ul></li>\n</ul>",
          "votes": null,
          "replies": [
            {
              "id": 2324447,
              "author_name": "shinomoriaoshi",
              "author_url": "",
              "post_date": "06/30/2023 15:54:09",
              "content": "<p>Thank you! Your practices are interesting. My NN was not super good, maybe it's because I kept the architecture and the input features complicated.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2322176,
      "author_name": "ataraxian",
      "author_url": "",
      "post_date": "06/29/2023 06:05:14",
      "content": "<p>Congratulations on gold medal.</p>\n<blockquote>\n  <p>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.For level group 13-22, I trained a distinct model for each target.</p>\n</blockquote>\n<p>May I ask the reason for this strategy? I find training model for each level_group/level has a very different cv/lb correlation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322243,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/29/2023 06:59:53",
          "content": "<p>Thank you for your comment! <br>\nRegarding level groups 0-4 and 5-12, I adopted the above method because creating a single model for them all, rather than creating a model for each target, resulted in better CV and public score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2322245,
      "author_name": "abaojiang",
      "author_url": "",
      "post_date": "06/29/2023 07:05:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a>,</p>\n<p>Congratulations on your solo gold medal! I have some questions about your solutions described as follows:</p>\n<ol>\n<li>For LightGBM, how many features are considered for each <code>level_group</code> model?</li>\n<li>For NN, how did you choose the sequence length of the event records?</li>\n</ol>\n<p>Thanks a lot for the sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322729,
          "author_name": "takoihiraokazu",
          "author_url": "",
          "post_date": "06/29/2023 13:08:32",
          "content": "<p>Thanks!</p>\n<ol>\n<li><p>For the LightGBM model, the number of features considered for each level_group model is as follows: For level_group 0-4, I consider 1570 features. For level_group 5-12, I consider 3992 features. And for level_group 13-22, I consider 5290 features.</p></li>\n<li><p>As for the Neural Network, the sequence length of the event records was determined based on cross-validation results. Specifically, for level_group 0-4, I chose a sequence length of 250. For level_group 5-12, the sequence length is 500. And for level_group 13-22, the sequence length is 800.</p></li>\n</ol>",
          "votes": null,
          "replies": [
            {
              "id": 2322853,
              "author_name": "abaojiang",
              "author_url": "",
              "post_date": "06/29/2023 14:28:12",
              "content": "<p>Thanks for the quick reply, what a robust DL-based model! One more question if you don't mind sharing, did you use early stopping to choose the best checkpoints or just let the training process converge to some satisfactory position?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2324257,
                  "author_name": "takoihiraokazu",
                  "author_url": "",
                  "post_date": "06/30/2023 13:51:48",
                  "content": "<p>I didn't exactly use early stopping. Instead, I set the number of epochs to 20 and saved the checkpoint that produced the best validation score. The choice to use 20 epochs was based on several experiments, where this setting resulted in the best cross-validation performance.\"</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2324389,
                      "author_name": "abaojiang",
                      "author_url": "",
                      "post_date": "06/30/2023 15:10:00",
                      "content": "<p>As mentioned in the NN section, you train separate model for each <code>level_group</code>. If I'm not mistaken, the model checkpoints used to obtain the best CV performance could be different for different <code>level_group</code>s. For example, epoch 15 for  <code>level_group</code> \"0-4\", 17 for \"5-12\" and 11 for \"13-22\" can lead to the best overall Macro-F1 on local CV. Then, what's the better way to find out this combination. During the competition, I found it hard to figure out an efficient way to determine which checkpoints to choose. Thanks a lot for your clarification and patience 😓</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2324819,
                          "author_name": "takoihiraokazu",
                          "author_url": "",
                          "post_date": "06/30/2023 22:40:26",
                          "content": "<p>Yes, as you said, the optimal epoch does differ for each level group. As for the checkpoints, I decided on the one where the AUC was highest. Even then, if the AUC improved for each level group, the final macro-f1 score also improved.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2325133,
                              "author_name": "abaojiang",
                              "author_url": "",
                              "post_date": "07/01/2023 06:21:08",
                              "content": "<p>I see! What I did was to choose those with the highest Macro-F1 @ 0.63 (but, I knew optimizing locally on <code>level_group</code> models might not trasfer to global optimization…). I forgot to try to select checkpoints with other metrics. I'll go back to give them a try and do some late submissions. </p>\n<p>Thanks for the reply and kindness. Good luck with your next competition!</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2322499,
      "author_name": "serangu",
      "author_url": "",
      "post_date": "06/29/2023 10:11:40",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> congrats on solo gold medal!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2324140,
          "author_name": "joek47",
          "author_url": "",
          "post_date": "06/30/2023 12:10:16",
          "content": "<p>+1 congrats on gold</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2322764,
      "author_name": "akshayvyas02",
      "author_url": "",
      "post_date": "06/29/2023 13:37:04",
      "content": "<p><a href=\"https://www.kaggle.com/takoihiraokazu\" target=\"_blank\">@takoihiraokazu</a> Congratulations on your 13th place finish! Your strategy of training separate models for different level groups and incorporating features like categorical and numerical data counts demonstrates a thoughtful approach. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2323866,
      "author_name": "poojach7611",
      "author_url": "",
      "post_date": "06/30/2023 07:54:35",
      "content": "<p>Congratulations on your 13th place solution! It's impressive that you combined LightGBM and neural networks (NN) in an ensemble approach. Your use of nested cross-validation and careful splitting of training and validation data demonstrates a thoughtful approach. Well done!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2322141": "# 13th place solution\nFirst of all, I would like to thank the Kaggle community for sharing great ideas and engaging discussions. I would also like to thank the hosts for organizing this interesting task competition.\n\n## Summary\n- Ensemble of LightGBM and NN\n- Cross validation: Nested cross validation\n    - Training data: Data for which the first four digits of the session_id are 2200 or less.\n    - Validation data: Data for which the first four digits of the session_id are 2201 or more.\n    - I trained the model on the training data using a 5-fold cross validation strategy, and evaluated it on the validation data using predictions from all 5 trained models.\n    - For the final submission, I trained the model on the entire dataset using a 5-fold cross validation approach.\n\n## LightGBM\n- I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.\n- For level group 13-22, I trained a distinct model for each target.\n- Main features:\n    - The count of categorical data for each session_id\n    - The statistical measures of numerical data for each session_id\n    - An aggregate of the next action taken \n- Scores\n    - CV : 0.7032\n    - Public Score : 0.704\n    - Private Score : 0.701\n\n\n## NN\n- Model: Transformer + GRU\n    - The standalone Transformer didn't perform very well.\n    - The addition of GRU improved the score.\n- Trained with fewer features\n- I trained a separate model for each level group.\n- Scores\n    - CV : 0.7010\n    - Public Score: 0.700\n    - Private Score: 0.700\n\n## Ensemble\n    - LightGBM * 0.66 + NN * 0.34\n    - CV : 0.7053\n    - Public Score : 0.706\n    - Private Score : 0.702",
    "2322148": "Congratulations on your gold medal!\n\nMay I ask why you split the training and validation data this way?\n\n>Cross validation: Nested cross validation\n>Training data: Data for which the first four digits of the session_id are 2200 or less.\n>Validation data: Data for which the first four digits of the session_id are 2201 or more.",
    "2322154": "Congratulations on your solo gold! Can you share your NN architecture? For me, NN alone doesn't work that well, but extracting embedding from NN and add them into Tabular features works pretty help. So I am curious how NN looks like in your case. And it would be greater that you could share the training details, some remarks, etc.",
    "2322176": "Congratulations on gold medal.\n\n>I trained one model for level group 0-4 and another for level group 5-12. For these, I included features representing the target and trained a single model.For level group 13-22, I trained a distinct model for each target.\n\nMay I ask the reason for this strategy? I find training model for each level_group/level has a very different cv/lb correlation.",
    "2322194": "Thank you for your comment! I used nested cross-validation to evaluate for the following two reasons:\n\n- The evaluation metric for this time fluctuates significantly based on the threshold, and just changing the threshold also significantly changed the Public Score. I wanted to create the same situation locally as the Public Score (including the ensemble of 5-fold) and check the change in CV due to changing the threshold.\n\n- For instance, in situations like the one described in the link below, I believed that nested cross-validation would be more reliable than standard cross-validation.\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682",
    "2322229": "Thank you for your comment! Also, congratulations on your gold medal! \nFor instance, I created the following model for level group 5-12.\n\n```\n# Numerical features used for NN\nnum_cols_level2 = [\"elapsed_time_log1p\",\"elapsed_time_diff_log1p\",\n                  \"room_coor_x\",\"room_coor_y\"]\n\n# Categorical features used for NN\ncat_cols = [\"event_name\",\"name\",\"fqid\",\"room_fqid\",\"level\"]\n\n\nclass MeanPooling(nn.Module):\n    def __init__(self):\n        super(MeanPooling, self).__init__()\n        \n    def forward(self, last_hidden_state, attention_mask):\n        input_mask_expanded = attention_mask.unsqueeze(-1).expand(last_hidden_state.size()).float()\n        sum_embeddings = torch.sum(last_hidden_state * input_mask_expanded, 1)\n        sum_mask = input_mask_expanded.sum(1)\n        sum_mask = torch.clamp(sum_mask, min=1e-9)\n        mean_embeddings = sum_embeddings / sum_mask\n        return mean_embeddings\n\nclass PspTransformerLevel2Model(nn.Module):\n    def __init__(\n        self, dropout=0.2,\n        input_numerical_size=4,\n        input_name_nunique=7, name_embedding_size = 12,\n        input_event_nunique=12, event_embedding_size = 12,\n        input_fqid_nunique=130, fqid_embedding_size = 24,\n        input_room_fqid_nunique=20, room_fqid_embedding_size = 12,\n        input_level_nunique=24,level_embedding_size=12,\n        categorical_linear_size = 120,\n        numeraical_linear_size = 24,\n        model_size = 160,\n        nhead = 16,\n        dim_feedforward=480,\n        out_size=10):\n        super(PspTransformerLevel2Model, self).__init__()\n        self.name_embedding = nn.Embedding(num_embeddings=input_name_nunique, \n                                           embedding_dim=name_embedding_size)\n        self.event_embedding = nn.Embedding(num_embeddings=input_event_nunique, \n                                           embedding_dim=event_embedding_size)\n        self.fqid_embedding = nn.Embedding(num_embeddings=input_fqid_nunique, \n                                           embedding_dim=fqid_embedding_size)\n        self.room_fqid_embedding = nn.Embedding(num_embeddings=input_room_fqid_nunique, \n                                           embedding_dim=room_fqid_embedding_size)\n        self.level_embedding = nn.Embedding(num_embeddings=input_level_nunique, \n                                           embedding_dim=level_embedding_size)\n        self.categorical_linear = nn.Sequential(\n                nn.Linear(name_embedding_size + event_embedding_size + \n                          fqid_embedding_size + room_fqid_embedding_size + \n                          level_embedding_size, categorical_linear_size),\n                nn.LayerNorm(categorical_linear_size)\n            )\n        self.numerical_linear  = nn.Sequential(\n                nn.Linear(input_numerical_size, numeraical_linear_size),\n                nn.LayerNorm(numeraical_linear_size)\n            )\n        \n        self.linear1  = nn.Sequential(\n                nn.Linear(categorical_linear_size + numeraical_linear_size, \n                          model_size),\n                nn.LayerNorm(model_size),\n            )\n        self.transformer_encoder = TransformerEncoder(\n            encoder_layer = nn.TransformerEncoderLayer(d_model=model_size, \n                                                       nhead=nhead,\n                                                       dim_feedforward=dim_feedforward ,\n                                                       dropout=dropout),\n                                                     num_layers=1)\n        self.gru = nn.GRU(model_size, model_size,\n                            num_layers = 1, \n                            batch_first=True,\n                            bidirectional=True)\n        self.linear_out  = nn.Sequential(\n                nn.Linear(model_size*2, \n                          out_size)\n            )\n        self.pool = MeanPooling()\n    \n    def forward(self, numerical_array, name_array,event_array,\n                fqid_array, room_fqid_array, level_array, \n                mask,mask_for_pooling):\n    \n        name_embedding = self.name_embedding(name_array)\n        event_embedding = self.event_embedding(event_array)\n        fqid_embedding = self.fqid_embedding(fqid_array)\n        room_fqid_embedding = self.room_fqid_embedding(room_fqid_array)\n        level_embedding = self.level_embedding(level_array)\n        categorical_emedding = torch.cat([name_embedding,\n                                          event_embedding,\n                                          fqid_embedding, \n                                          room_fqid_embedding,\n                                          level_embedding\n                                          ], axis=2)\n        categorical_emedding = self.categorical_linear(categorical_emedding)\n        numerical_embedding = self.numerical_linear(numerical_array)\n        concat_embedding = torch.cat([categorical_emedding,\n                                      numerical_embedding],axis=2)\n        concat_embedding = self.linear1(concat_embedding)\n        concat_embedding  = concat_embedding.permute(1,0,2).contiguous()\n        output = self.transformer_encoder(concat_embedding, \n                                          src_key_padding_mask=mask)\n        output = output.permute(1,0,2).contiguous()\n        output,_ = self.gru(output)\n        output = self.pool(output,mask_for_pooling)\n        output = self.linear_out(output)\n        return output\n\n```\n\n\nThe main improvements I made are listed below. Through a series of experiments, I found that the following practices were advantageous:\n\n- Keeping the model size small:\n    - For instance, I kept the number of layers for both the Transformer and GRU at 1.\n    - I didn't enlarge the encoder layer or hidden size.\n    - I also maintained a relatively small embedding size for categorical features.\n- Using fewer features:\n    - For instance, I didn't use features such as screen_coor_x, screen_coor_y, hover_duration, and text.",
    "2322243": "Thank you for your comment! \nRegarding level groups 0-4 and 5-12, I adopted the above method because creating a single model for them all, rather than creating a model for each target, resulted in better CV and public score.",
    "2322245": "Hi @takoihiraokazu,\n\nCongratulations on your solo gold medal! I have some questions about your solutions described as follows:\n1. For LightGBM, how many features are considered for each `level_group` model?\n2. For NN, how did you choose the sequence length of the event records?\n\nThanks a lot for the sharing!",
    "2322499": "takoihiraokazu congrats on solo gold medal!",
    "2322729": "Thanks!\n\n1. For the LightGBM model, the number of features considered for each level_group model is as follows: For level_group 0-4, I consider 1570 features. For level_group 5-12, I consider 3992 features. And for level_group 13-22, I consider 5290 features.\n\n2. As for the Neural Network, the sequence length of the event records was determined based on cross-validation results. Specifically, for level_group 0-4, I chose a sequence length of 250. For level_group 5-12, the sequence length is 500. And for level_group 13-22, the sequence length is 800.",
    "2322764": "takoihiraokazu Congratulations on your 13th place finish! Your strategy of training separate models for different level groups and incorporating features like categorical and numerical data counts demonstrates a thoughtful approach.",
    "2322853": "Thanks for the quick reply, what a robust DL-based model! One more question if you don't mind sharing, did you use early stopping to choose the best checkpoints or just let the training process converge to some satisfactory position?",
    "2323866": "Congratulations on your 13th place solution! It's impressive that you combined LightGBM and neural networks (NN) in an ensemble approach. Your use of nested cross-validation and careful splitting of training and validation data demonstrates a thoughtful approach. Well done!",
    "2324140": "1 congrats on gold",
    "2324257": "I didn't exactly use early stopping. Instead, I set the number of epochs to 20 and saved the checkpoint that produced the best validation score. The choice to use 20 epochs was based on several experiments, where this setting resulted in the best cross-validation performance.\"",
    "2324389": "As mentioned in the NN section, you train separate model for each `level_group`. If I'm not mistaken, the model checkpoints used to obtain the best CV performance could be different for different `level_group`s. For example, epoch 15 for  `level_group` \"0-4\", 17 for \"5-12\" and 11 for \"13-22\" can lead to the best overall Macro-F1 on local CV. Then, what's the better way to find out this combination. During the competition, I found it hard to figure out an efficient way to determine which checkpoints to choose. Thanks a lot for your clarification and patience 😓",
    "2324447": "Thank you! Your practices are interesting. My NN was not super good, maybe it's because I kept the architecture and the input features complicated.",
    "2324819": "Yes, as you said, the optimal epoch does differ for each level group. As for the checkpoints, I decided on the one where the AUC was highest. Even then, if the AUC improved for each level group, the final macro-f1 score also improved.",
    "2325133": "I see! What I did was to choose those with the highest Macro-F1 @ 0.63 (but, I knew optimizing locally on `level_group` models might not trasfer to global optimization...). I forgot to try to select checkpoints with other metrics. I'll go back to give them a try and do some late submissions. \n\nThanks for the reply and kindness. Good luck with your next competition!"
  },
  "source": "meta"
}