{
  "id": 420105,
  "title": "[0.67] 742nd Place Solution - Single Transformers Model Comparison Between Data Aggregation and None",
  "url": "/competitions/predict-student-performance-from-game-play/writeups/phu-gia-hoang-0-67-742nd-place-solution-single-tra",
  "author_name": "",
  "post_date": "2023-06-29T08:04:44.263Z",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<h1><strong>742nd Place Solution</strong></h1>\n<p>First of all, congratulations to the first 4 teams who emerged victorious in this long and challenging competition. This was my very first challenge since I created this account 4 years ago, during my first year at university for a random course at that time.</p>\n<p>In this competition, I made every effort to implement transformers to solve the problem, resulting in two solutions with different approaches:</p>\n<h1>1. Sequence-N Aprroach</h1>\n<ul>\n<li>Codes: <a href=\"https://www.kaggle.com/phuhoang26/psp-llama\" target=\"_blank\">Llama</a></li>\n<li>Model's weights: <a href=\"https://www.kaggle.com/datasets/phuhoang26/psp-seqtf-llama\" target=\"_blank\">Llama-W</a></li>\n<li>Public score: 0.616</li>\n<li>Private score: 0.621</li>\n</ul>\n<p>This approach was inspired by the <a href=\"https://www.kaggle.com/competitions/data-science-bowl-2019/discussion/127891\" target=\"_blank\">3rd solution of limerobot in the Data Science Bowl 2019</a> by <a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a>.<br>\nIn this approach, I utilized N information (rows) of session_id as input to my model. The pipeline for this approach is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F8c31538c1619c8e330bdde51b078247f%2Fdrawio.png?generation=1688024191708037&amp;alt=media\" alt=\"Sequence-N data pipeline\"></p>\n<p>The rationale behind this approach was to leverage the sequential characteristics of the data. After preprocessing and grouping the data by sessions, it was fed into a language model. The model was trained using StratifiedKFold (K=5).</p>\n<p>Instead of using BERT, as in the original codes of limerobot, I employed <a href=\"https://arxiv.org/abs/2302.13971\" target=\"_blank\">Llama</a>, a state-of-the-art language model developed by Facebook. Llama claims to perform well or even better than previous models (with billions of parameters) while having a significantly smaller number of parameters (4M) for the same tasks.</p>\n<p>I trained this approach with different sequence lengths. For level_group = '0-4', I trained with sequence length equals to 256; level_group = '5-12' is 528; and level_group = '13-22' is 808. These number was the output of EDA (counting number of rows at quantile 0.9)</p>\n<p>I also experimented with different classifiers to learn from the output of Llama, but the scores did not show significant improvement. These are the classifiers I tried:</p>\n<ul>\n<li>Llama + <a href=\"https://arxiv.org/abs/1408.5882\" target=\"_blank\">1D CNN</a></li>\n<li>Llama + XGBOOST</li>\n<li>Llama + LSTM</li>\n</ul>\n<p><strong>Reasons for the failure of this approach:</strong></p>\n<ul>\n<li>For processing CATS features (categorical features), I employed a mapping technique (e.g., cutsence_click as 1, personal_click as 2, navigate_click as 3, etc.). This technique works well when the tokens are carefully customized, as done in Glove or Fasttext, where similar tokens are embedded close to each other in the feature space (e.g., run and walk, hit and slap, brown and yellow).</li>\n<li>For NUMS features (numerical features), I only filled the NaN values with 0, without performing feature selection or generation. This may have contributed to the failure of the model.</li>\n</ul>\n<h1>2. Gated Tab Transformer</h1>\n<ul>\n<li>Codes: <a href=\"https://www.kaggle.com/phuhoang26/gated-tab-transformer\" target=\"_blank\">GATED-TAB-TF</a></li>\n<li>Model's weights: <a href=\"https://www.kaggle.com/datasets/phuhoang26/gated-tf-psp\" target=\"_blank\">GATED-TAB-TF-W</a></li>\n<li>Public score: 0.674</li>\n<li>Private score: 0.67</li>\n</ul>\n<p>Paper of this approach: <a href=\"https://arxiv.org/abs/2201.00199\" target=\"_blank\">The GatedTabTransformer. An enhanced deep learning architecture for tabular modeling</a></p>\n<p>Data aggregation ideas and codes derive from <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> notebook, you can see <a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">here</a></p>\n<p>this approach, I used aggregated information of each session_id to feed into my model. Here is this approach's pipeline:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F4d9e348215f7856a575c1fb8db55c739%2F2.drawio.png?generation=1688024800924132&amp;alt=media\" alt=\"\"></p>\n<h1>Conclusion</h1>\n<p>Initially, I believed that the Sequence-N approach would outperform the Gated Tab Transformers. However, the scores of the latter model turned out to be significantly better than the former, and I have yet to determine the reason for this.</p>\n<p>Furthermore, as I state earlier, I train the Sequence-N approach with quite long sequence. However, in the submit versions, I saw that the smaller sequence length the better! I cannot explain this.</p>\n<p>If you have any questions, comments, or contributions, please feel free to contact me or leave a comment here.</p>\n<p>Thank you for taking the time to read this!</p>",
  "messages": [
    {
      "id": "2322329",
      "postDate": "06/29/2023 08:01:29",
      "content": "<h1><strong>742nd Place Solution</strong></h1>\n<p>First of all, congratulations to the first 4 teams who emerged victorious in this long and challenging competition. This was my very first challenge since I created this account 4 years ago, during my first year at university for a random course at that time.</p>\n<p>In this competition, I made every effort to implement transformers to solve the problem, resulting in two solutions with different approaches:</p>\n<h1>1. Sequence-N Aprroach</h1>\n<ul>\n<li>Codes: <a href=\"https://www.kaggle.com/phuhoang26/psp-llama\" target=\"_blank\">Llama</a></li>\n<li>Model's weights: <a href=\"https://www.kaggle.com/datasets/phuhoang26/psp-seqtf-llama\" target=\"_blank\">Llama-W</a></li>\n<li>Public score: 0.616</li>\n<li>Private score: 0.621</li>\n</ul>\n<p>This approach was inspired by the <a href=\"https://www.kaggle.com/competitions/data-science-bowl-2019/discussion/127891\" target=\"_blank\">3rd solution of limerobot in the Data Science Bowl 2019</a> by <a href=\"https://www.kaggle.com/limerobot\" target=\"_blank\">@limerobot</a>.<br>\nIn this approach, I utilized N information (rows) of session_id as input to my model. The pipeline for this approach is as follows:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F8c31538c1619c8e330bdde51b078247f%2Fdrawio.png?generation=1688024191708037&amp;alt=media\" alt=\"Sequence-N data pipeline\"></p>\n<p>The rationale behind this approach was to leverage the sequential characteristics of the data. After preprocessing and grouping the data by sessions, it was fed into a language model. The model was trained using StratifiedKFold (K=5).</p>\n<p>Instead of using BERT, as in the original codes of limerobot, I employed <a href=\"https://arxiv.org/abs/2302.13971\" target=\"_blank\">Llama</a>, a state-of-the-art language model developed by Facebook. Llama claims to perform well or even better than previous models (with billions of parameters) while having a significantly smaller number of parameters (4M) for the same tasks.</p>\n<p>I trained this approach with different sequence lengths. For level_group = '0-4', I trained with sequence length equals to 256; level_group = '5-12' is 528; and level_group = '13-22' is 808. These number was the output of EDA (counting number of rows at quantile 0.9)</p>\n<p>I also experimented with different classifiers to learn from the output of Llama, but the scores did not show significant improvement. These are the classifiers I tried:</p>\n<ul>\n<li>Llama + <a href=\"https://arxiv.org/abs/1408.5882\" target=\"_blank\">1D CNN</a></li>\n<li>Llama + XGBOOST</li>\n<li>Llama + LSTM</li>\n</ul>\n<p><strong>Reasons for the failure of this approach:</strong></p>\n<ul>\n<li>For processing CATS features (categorical features), I employed a mapping technique (e.g., cutsence_click as 1, personal_click as 2, navigate_click as 3, etc.). This technique works well when the tokens are carefully customized, as done in Glove or Fasttext, where similar tokens are embedded close to each other in the feature space (e.g., run and walk, hit and slap, brown and yellow).</li>\n<li>For NUMS features (numerical features), I only filled the NaN values with 0, without performing feature selection or generation. This may have contributed to the failure of the model.</li>\n</ul>\n<h1>2. Gated Tab Transformer</h1>\n<ul>\n<li>Codes: <a href=\"https://www.kaggle.com/phuhoang26/gated-tab-transformer\" target=\"_blank\">GATED-TAB-TF</a></li>\n<li>Model's weights: <a href=\"https://www.kaggle.com/datasets/phuhoang26/gated-tf-psp\" target=\"_blank\">GATED-TAB-TF-W</a></li>\n<li>Public score: 0.674</li>\n<li>Private score: 0.67</li>\n</ul>\n<p>Paper of this approach: <a href=\"https://arxiv.org/abs/2201.00199\" target=\"_blank\">The GatedTabTransformer. An enhanced deep learning architecture for tabular modeling</a></p>\n<p>Data aggregation ideas and codes derive from <a href=\"https://www.kaggle.com/pourchot\" target=\"_blank\">@pourchot</a> notebook, you can see <a href=\"https://www.kaggle.com/code/pourchot/simple-xgb\" target=\"_blank\">here</a></p>\n<p>this approach, I used aggregated information of each session_id to feed into my model. Here is this approach's pipeline:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F4d9e348215f7856a575c1fb8db55c739%2F2.drawio.png?generation=1688024800924132&amp;alt=media\" alt=\"\"></p>\n<h1>Conclusion</h1>\n<p>Initially, I believed that the Sequence-N approach would outperform the Gated Tab Transformers. However, the scores of the latter model turned out to be significantly better than the former, and I have yet to determine the reason for this.</p>\n<p>Furthermore, as I state earlier, I train the Sequence-N approach with quite long sequence. However, in the submit versions, I saw that the smaller sequence length the better! I cannot explain this.</p>\n<p>If you have any questions, comments, or contributions, please feel free to contact me or leave a comment here.</p>\n<p>Thank you for taking the time to read this!</p>",
      "rawMarkdown": "# **742nd Place Solution**\n\nFirst of all, congratulations to the first 4 teams who emerged victorious in this long and challenging competition. This was my very first challenge since I created this account 4 years ago, during my first year at university for a random course at that time.\n\nIn this competition, I made every effort to implement transformers to solve the problem, resulting in two solutions with different approaches:\n\n# 1. Sequence-N Aprroach\n- Codes: [Llama](https://www.kaggle.com/phuhoang26/psp-llama)\n- Model's weights: [Llama-W](https://www.kaggle.com/datasets/phuhoang26/psp-seqtf-llama)\n- Public score: 0.616\n- Private score: 0.621\n\nThis approach was inspired by the [3rd solution of limerobot in the Data Science Bowl 2019](https://www.kaggle.com/competitions/data-science-bowl-2019/discussion/127891) by @limerobot.\nIn this approach, I utilized N information (rows) of session_id as input to my model. The pipeline for this approach is as follows:\n\n![Sequence-N data pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F8c31538c1619c8e330bdde51b078247f%2Fdrawio.png?generation=1688024191708037&alt=media)\n\nThe rationale behind this approach was to leverage the sequential characteristics of the data. After preprocessing and grouping the data by sessions, it was fed into a language model. The model was trained using StratifiedKFold (K=5).\n\nInstead of using BERT, as in the original codes of limerobot, I employed [Llama](https://arxiv.org/abs/2302.13971), a state-of-the-art language model developed by Facebook. Llama claims to perform well or even better than previous models (with billions of parameters) while having a significantly smaller number of parameters (4M) for the same tasks.\n\nI trained this approach with different sequence lengths. For level_group = '0-4', I trained with sequence length equals to 256; level_group = '5-12' is 528; and level_group = '13-22' is 808. These number was the output of EDA (counting number of rows at quantile 0.9)\n\nI also experimented with different classifiers to learn from the output of Llama, but the scores did not show significant improvement. These are the classifiers I tried:\n- Llama + [1D CNN](https://arxiv.org/abs/1408.5882)\n- Llama + XGBOOST\n- Llama + LSTM\n\n**Reasons for the failure of this approach:**\n- For processing CATS features (categorical features), I employed a mapping technique (e.g., cutsence_click as 1, personal_click as 2, navigate_click as 3, etc.). This technique works well when the tokens are carefully customized, as done in Glove or Fasttext, where similar tokens are embedded close to each other in the feature space (e.g., run and walk, hit and slap, brown and yellow).\n- For NUMS features (numerical features), I only filled the NaN values with 0, without performing feature selection or generation. This may have contributed to the failure of the model.\n\n# 2. Gated Tab Transformer\n- Codes: [GATED-TAB-TF](https://www.kaggle.com/phuhoang26/gated-tab-transformer)\n- Model's weights: [GATED-TAB-TF-W](https://www.kaggle.com/datasets/phuhoang26/gated-tf-psp)\n- Public score: 0.674\n- Private score: 0.67\n\nPaper of this approach: [The GatedTabTransformer. An enhanced deep learning architecture for tabular modeling](https://arxiv.org/abs/2201.00199)\n\nData aggregation ideas and codes derive from @pourchot notebook, you can see [here](https://www.kaggle.com/code/pourchot/simple-xgb)\n\nthis approach, I used aggregated information of each session_id to feed into my model. Here is this approach's pipeline:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F4d9e348215f7856a575c1fb8db55c739%2F2.drawio.png?generation=1688024800924132&alt=media)\n\n# Conclusion\nInitially, I believed that the Sequence-N approach would outperform the Gated Tab Transformers. However, the scores of the latter model turned out to be significantly better than the former, and I have yet to determine the reason for this.\n\nFurthermore, as I state earlier, I train the Sequence-N approach with quite long sequence. However, in the submit versions, I saw that the smaller sequence length the better! I cannot explain this.\n\nIf you have any questions, comments, or contributions, please feel free to contact me or leave a comment here.\n\nThank you for taking the time to read this!",
      "votes": null
    },
    {
      "id": "2322333",
      "postDate": "06/29/2023 08:02:28",
      "content": "<p>I am cleaning my codes, stay tuned! :3 </p>",
      "rawMarkdown": "I am cleaning my codes, stay tuned! :3",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2322333,
      "author_name": "phuhoang26",
      "author_url": "",
      "post_date": "06/29/2023 08:02:28",
      "content": "<p>I am cleaning my codes, stay tuned! :3 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2322329": "# **742nd Place Solution**\n\nFirst of all, congratulations to the first 4 teams who emerged victorious in this long and challenging competition. This was my very first challenge since I created this account 4 years ago, during my first year at university for a random course at that time.\n\nIn this competition, I made every effort to implement transformers to solve the problem, resulting in two solutions with different approaches:\n\n# 1. Sequence-N Aprroach\n- Codes: [Llama](https://www.kaggle.com/phuhoang26/psp-llama)\n- Model's weights: [Llama-W](https://www.kaggle.com/datasets/phuhoang26/psp-seqtf-llama)\n- Public score: 0.616\n- Private score: 0.621\n\nThis approach was inspired by the [3rd solution of limerobot in the Data Science Bowl 2019](https://www.kaggle.com/competitions/data-science-bowl-2019/discussion/127891) by @limerobot.\nIn this approach, I utilized N information (rows) of session_id as input to my model. The pipeline for this approach is as follows:\n\n![Sequence-N data pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F8c31538c1619c8e330bdde51b078247f%2Fdrawio.png?generation=1688024191708037&alt=media)\n\nThe rationale behind this approach was to leverage the sequential characteristics of the data. After preprocessing and grouping the data by sessions, it was fed into a language model. The model was trained using StratifiedKFold (K=5).\n\nInstead of using BERT, as in the original codes of limerobot, I employed [Llama](https://arxiv.org/abs/2302.13971), a state-of-the-art language model developed by Facebook. Llama claims to perform well or even better than previous models (with billions of parameters) while having a significantly smaller number of parameters (4M) for the same tasks.\n\nI trained this approach with different sequence lengths. For level_group = '0-4', I trained with sequence length equals to 256; level_group = '5-12' is 528; and level_group = '13-22' is 808. These number was the output of EDA (counting number of rows at quantile 0.9)\n\nI also experimented with different classifiers to learn from the output of Llama, but the scores did not show significant improvement. These are the classifiers I tried:\n- Llama + [1D CNN](https://arxiv.org/abs/1408.5882)\n- Llama + XGBOOST\n- Llama + LSTM\n\n**Reasons for the failure of this approach:**\n- For processing CATS features (categorical features), I employed a mapping technique (e.g., cutsence_click as 1, personal_click as 2, navigate_click as 3, etc.). This technique works well when the tokens are carefully customized, as done in Glove or Fasttext, where similar tokens are embedded close to each other in the feature space (e.g., run and walk, hit and slap, brown and yellow).\n- For NUMS features (numerical features), I only filled the NaN values with 0, without performing feature selection or generation. This may have contributed to the failure of the model.\n\n# 2. Gated Tab Transformer\n- Codes: [GATED-TAB-TF](https://www.kaggle.com/phuhoang26/gated-tab-transformer)\n- Model's weights: [GATED-TAB-TF-W](https://www.kaggle.com/datasets/phuhoang26/gated-tf-psp)\n- Public score: 0.674\n- Private score: 0.67\n\nPaper of this approach: [The GatedTabTransformer. An enhanced deep learning architecture for tabular modeling](https://arxiv.org/abs/2201.00199)\n\nData aggregation ideas and codes derive from @pourchot notebook, you can see [here](https://www.kaggle.com/code/pourchot/simple-xgb)\n\nthis approach, I used aggregated information of each session_id to feed into my model. Here is this approach's pipeline:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3563184%2F4d9e348215f7856a575c1fb8db55c739%2F2.drawio.png?generation=1688024800924132&alt=media)\n\n# Conclusion\nInitially, I believed that the Sequence-N approach would outperform the Gated Tab Transformers. However, the scores of the latter model turned out to be significantly better than the former, and I have yet to determine the reason for this.\n\nFurthermore, as I state earlier, I train the Sequence-N approach with quite long sequence. However, in the submit versions, I saw that the smaller sequence length the better! I cannot explain this.\n\nIf you have any questions, comments, or contributions, please feel free to contact me or leave a comment here.\n\nThank you for taking the time to read this!",
    "2322333": "I am cleaning my codes, stay tuned! :3"
  },
  "source": "meta"
}