{
  "id": 419652,
  "title": "Catboost vs Transformer - Close!!",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/419652",
  "author_name": "",
  "post_date": "2023-06-27T01:42:01.656569Z",
  "votes": -1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>This is third Kaggle competition in the last one month (all joined in last 10 days before completion!!). In the previous two competitions (AMP Parkinson, Freezing of Gait Prediction), there was a possibility of using Transformers - the first one though did not have enough data, with regard to second one, didn't not much time to create a new model when I realized so..</p>\n<p>So, I decided to use Transformers for this competition - the advantage is no feature engineering is required - just use the data as they are - except for few I have taken from public notebooks (like dealt elapsed time). This means the model is very flexible and can even extend to future versions of the game…</p>\n<p>Here are the score comparisons with catboost model (model with around 0.7 LB score - A Big Thanks to VADIM KAMAEV - else I would not have even participated in this competition) which had lots of feature engineering done.  (No hyper tuning done for the transformer model - with just basic model from Keras Time Sequence Tutorial)</p>\n<p>Catboost vs Transformer (F1 Macro Scores)</p>\n<ol>\n<li><p>0.6195 - 0.5864</p></li>\n<li><p>0.4803 - 0.4595</p></li>\n<li><p>0.5477 - 0.4876</p></li>\n<li><p>0.6284 - 0.5942</p></li>\n<li><p>0.6270 - 0.6170</p></li>\n<li><p>0.5903 - 0.6010</p></li>\n<li><p>0.5679 - 0.5822 </p></li>\n<li><p>0.533  - 0.5496 </p></li>\n<li><p>0.5728 - 0.5846 </p></li>\n<li><p>0.616 - 0.6031</p></li>\n<li><p>0.5461 -0.5771 </p></li>\n<li><p>0.5023 -0.504</p></li>\n<li><p>0.5707 -0.5891 </p></li>\n<li><p>0.5745 - 0.5645 </p></li>\n<li><p>0.6347 - 0.6047 </p></li>\n<li><p>0.5134 - 0.4881</p></li>\n<li><p>0.5170 - 0.5185</p></li>\n<li><p>0.5162 - 0.4646 </p></li>\n</ol>\n<p>But alas, I could not make even single submission successfully with transformer models due to submission timing out after 9 hours </p>\n<p>Have anyone found a method to do faster inference with Transformer models - while advantage is no feature engineering is done, it takes a lot of time to do inference due to batching operations involved in transformer models..</p>",
  "messages": [
    {
      "id": "2319212",
      "postDate": "06/27/2023 01:42:01",
      "content": "<p>This is third Kaggle competition in the last one month (all joined in last 10 days before completion!!). In the previous two competitions (AMP Parkinson, Freezing of Gait Prediction), there was a possibility of using Transformers - the first one though did not have enough data, with regard to second one, didn't not much time to create a new model when I realized so..</p>\n<p>So, I decided to use Transformers for this competition - the advantage is no feature engineering is required - just use the data as they are - except for few I have taken from public notebooks (like dealt elapsed time). This means the model is very flexible and can even extend to future versions of the game…</p>\n<p>Here are the score comparisons with catboost model (model with around 0.7 LB score - A Big Thanks to VADIM KAMAEV - else I would not have even participated in this competition) which had lots of feature engineering done.  (No hyper tuning done for the transformer model - with just basic model from Keras Time Sequence Tutorial)</p>\n<p>Catboost vs Transformer (F1 Macro Scores)</p>\n<ol>\n<li><p>0.6195 - 0.5864</p></li>\n<li><p>0.4803 - 0.4595</p></li>\n<li><p>0.5477 - 0.4876</p></li>\n<li><p>0.6284 - 0.5942</p></li>\n<li><p>0.6270 - 0.6170</p></li>\n<li><p>0.5903 - 0.6010</p></li>\n<li><p>0.5679 - 0.5822 </p></li>\n<li><p>0.533  - 0.5496 </p></li>\n<li><p>0.5728 - 0.5846 </p></li>\n<li><p>0.616 - 0.6031</p></li>\n<li><p>0.5461 -0.5771 </p></li>\n<li><p>0.5023 -0.504</p></li>\n<li><p>0.5707 -0.5891 </p></li>\n<li><p>0.5745 - 0.5645 </p></li>\n<li><p>0.6347 - 0.6047 </p></li>\n<li><p>0.5134 - 0.4881</p></li>\n<li><p>0.5170 - 0.5185</p></li>\n<li><p>0.5162 - 0.4646 </p></li>\n</ol>\n<p>But alas, I could not make even single submission successfully with transformer models due to submission timing out after 9 hours </p>\n<p>Have anyone found a method to do faster inference with Transformer models - while advantage is no feature engineering is done, it takes a lot of time to do inference due to batching operations involved in transformer models..</p>",
      "rawMarkdown": "This is third Kaggle competition in the last one month (all joined in last 10 days before completion!!). In the previous two competitions (AMP Parkinson, Freezing of Gait Prediction), there was a possibility of using Transformers - the first one though did not have enough data, with regard to second one, didn't not much time to create a new model when I realized so..\n\nSo, I decided to use Transformers for this competition - the advantage is no feature engineering is required - just use the data as they are - except for few I have taken from public notebooks (like dealt elapsed time). This means the model is very flexible and can even extend to future versions of the game...\n\nHere are the score comparisons with catboost model (model with around 0.7 LB score - A Big Thanks to VADIM KAMAEV - else I would not have even participated in this competition) which had lots of feature engineering done.  (No hyper tuning done for the transformer model - with just basic model from Keras Time Sequence Tutorial)\n\nCatboost vs Transformer (F1 Macro Scores)\n\n1. 0.6195 - 0.5864\n2. 0.4803 - 0.4595\n3. 0.5477 - 0.4876\n\n4. 0.6284 - 0.5942\n5. 0.6270 - 0.6170\n6. 0.5903 - 0.6010\n7. 0.5679 - 0.5822 \n8. 0.533  - 0.5496 \n9. 0.5728 - 0.5846 \n10. 0.616 - 0.6031\n11. 0.5461 -0.5771 \n12. 0.5023 -0.504\n13. 0.5707 -0.5891 \n\n14. 0.5745 - 0.5645 \n15. 0.6347 - 0.6047 \n16. 0.5134 - 0.4881\n17. 0.5170 - 0.5185\n18. 0.5162 - 0.4646 \n\nBut alas, I could not make even single submission successfully with transformer models due to submission timing out after 9 hours \n\nHave anyone found a method to do faster inference with Transformer models - while advantage is no feature engineering is done, it takes a lot of time to do inference due to batching operations involved in transformer models..",
      "votes": null
    },
    {
      "id": "2319249",
      "postDate": "06/27/2023 02:32:38",
      "content": "<p>You may want to compare the overall F1 score, and not per question, as this is what matters at the end. </p>",
      "rawMarkdown": "You may want to compare the overall F1 score, and not per question, as this is what matters at the end.",
      "votes": null
    },
    {
      "id": "2319307",
      "postDate": "06/27/2023 04:04:34",
      "content": "<p>Yes, for a long time, I focussed on individual f1 scores while training transformer models, and did not notice that the overall scores could be more than individual scores - the above catboost individual scores give overall LB score of 0.7 and so, I hope the overall score of transformer model should also be in that range (though we will never know as even three questions from transformer model times out in scoring….)</p>",
      "rawMarkdown": "Yes, for a long time, I focussed on individual f1 scores while training transformer models, and did not notice that the overall scores could be more than individual scores - the above catboost individual scores give overall LB score of 0.7 and so, I hope the overall score of transformer model should also be in that range (though we will never know as even three questions from transformer model times out in scoring....)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2319249,
      "author_name": "kononenko",
      "author_url": "",
      "post_date": "06/27/2023 02:32:38",
      "content": "<p>You may want to compare the overall F1 score, and not per question, as this is what matters at the end. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2319307,
          "author_name": "murugesann",
          "author_url": "",
          "post_date": "06/27/2023 04:04:34",
          "content": "<p>Yes, for a long time, I focussed on individual f1 scores while training transformer models, and did not notice that the overall scores could be more than individual scores - the above catboost individual scores give overall LB score of 0.7 and so, I hope the overall score of transformer model should also be in that range (though we will never know as even three questions from transformer model times out in scoring….)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2319212": "This is third Kaggle competition in the last one month (all joined in last 10 days before completion!!). In the previous two competitions (AMP Parkinson, Freezing of Gait Prediction), there was a possibility of using Transformers - the first one though did not have enough data, with regard to second one, didn't not much time to create a new model when I realized so..\n\nSo, I decided to use Transformers for this competition - the advantage is no feature engineering is required - just use the data as they are - except for few I have taken from public notebooks (like dealt elapsed time). This means the model is very flexible and can even extend to future versions of the game...\n\nHere are the score comparisons with catboost model (model with around 0.7 LB score - A Big Thanks to VADIM KAMAEV - else I would not have even participated in this competition) which had lots of feature engineering done.  (No hyper tuning done for the transformer model - with just basic model from Keras Time Sequence Tutorial)\n\nCatboost vs Transformer (F1 Macro Scores)\n\n1. 0.6195 - 0.5864\n2. 0.4803 - 0.4595\n3. 0.5477 - 0.4876\n\n4. 0.6284 - 0.5942\n5. 0.6270 - 0.6170\n6. 0.5903 - 0.6010\n7. 0.5679 - 0.5822 \n8. 0.533  - 0.5496 \n9. 0.5728 - 0.5846 \n10. 0.616 - 0.6031\n11. 0.5461 -0.5771 \n12. 0.5023 -0.504\n13. 0.5707 -0.5891 \n\n14. 0.5745 - 0.5645 \n15. 0.6347 - 0.6047 \n16. 0.5134 - 0.4881\n17. 0.5170 - 0.5185\n18. 0.5162 - 0.4646 \n\nBut alas, I could not make even single submission successfully with transformer models due to submission timing out after 9 hours \n\nHave anyone found a method to do faster inference with Transformer models - while advantage is no feature engineering is done, it takes a lot of time to do inference due to batching operations involved in transformer models..",
    "2319249": "You may want to compare the overall F1 score, and not per question, as this is what matters at the end.",
    "2319307": "Yes, for a long time, I focussed on individual f1 scores while training transformer models, and did not notice that the overall scores could be more than individual scores - the above catboost individual scores give overall LB score of 0.7 and so, I hope the overall score of transformer model should also be in that range (though we will never know as even three questions from transformer model times out in scoring....)"
  },
  "source": "meta"
}