{
  "id": 328859,
  "title": "It seems as if deep learning models are not performing well...",
  "url": "/competitions/amex-default-prediction/discussion/328859",
  "author_name": "",
  "post_date": "2022-06-03T09:59:57.328165500Z",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In another topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846\" target=\"_blank\">Model Performance Ranking</a>, the best models are xx-boost models rather than deep learning models implemented with PyTorch or TensorFlow, that's why?</p>",
  "messages": [
    {
      "id": "1810149",
      "postDate": "06/03/2022 09:59:57",
      "content": "<p>In another topic <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846\" target=\"_blank\">Model Performance Ranking</a>, the best models are xx-boost models rather than deep learning models implemented with PyTorch or TensorFlow, that's why?</p>",
      "rawMarkdown": "In another topic [Model Performance Ranking](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846), the best models are xx-boost models rather than deep learning models implemented with PyTorch or TensorFlow, that's why?",
      "votes": null
    },
    {
      "id": "1810422",
      "postDate": "06/03/2022 14:45:17",
      "content": "<p>My public time series deep learning models GRU and Transformer, <strong>do not</strong> use any feature engineering. They use the competition data as is and achieve LB 0.789. (Also their architecture is very simple and do not use NN tricks). Remember that the first GBT without feature engineering only achieved LB 0.783. If we add feature engineering to the GRU and Transformer (and/or improve architecture and/or use NN tricks) we can boost their performance. </p>",
      "rawMarkdown": "My public time series deep learning models GRU and Transformer, **do not** use any feature engineering. They use the competition data as is and achieve LB 0.789. (Also their architecture is very simple and do not use NN tricks). Remember that the first GBT without feature engineering only achieved LB 0.783. If we add feature engineering to the GRU and Transformer (and/or improve architecture and/or use NN tricks) we can boost their performance.",
      "votes": null
    },
    {
      "id": "1810428",
      "postDate": "06/03/2022 14:53:26",
      "content": "<p>Note that NN benefit from data augmentation, learning schedules, unsupervised pretraining tasks, knowledge distillation, different loss functions, multitask learning, and many other techniques that we cannot do with GBT.</p>\n<p>My best Transformer with GRU head offline beats my best GBT 😄 And both have much room for improvement !!</p>",
      "rawMarkdown": "Note that NN benefit from data augmentation, learning schedules, unsupervised pretraining tasks, knowledge distillation, different loss functions, multitask learning, and many other techniques that we cannot do with GBT.\n\nMy best Transformer with GRU head offline beats my best GBT 😄 And both have much room for improvement !!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1810422,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/03/2022 14:45:17",
      "content": "<p>My public time series deep learning models GRU and Transformer, <strong>do not</strong> use any feature engineering. They use the competition data as is and achieve LB 0.789. (Also their architecture is very simple and do not use NN tricks). Remember that the first GBT without feature engineering only achieved LB 0.783. If we add feature engineering to the GRU and Transformer (and/or improve architecture and/or use NN tricks) we can boost their performance. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1810428,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "06/03/2022 14:53:26",
          "content": "<p>Note that NN benefit from data augmentation, learning schedules, unsupervised pretraining tasks, knowledge distillation, different loss functions, multitask learning, and many other techniques that we cannot do with GBT.</p>\n<p>My best Transformer with GRU head offline beats my best GBT 😄 And both have much room for improvement !!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1810149": "In another topic [Model Performance Ranking](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328846), the best models are xx-boost models rather than deep learning models implemented with PyTorch or TensorFlow, that's why?",
    "1810422": "My public time series deep learning models GRU and Transformer, **do not** use any feature engineering. They use the competition data as is and achieve LB 0.789. (Also their architecture is very simple and do not use NN tricks). Remember that the first GBT without feature engineering only achieved LB 0.783. If we add feature engineering to the GRU and Transformer (and/or improve architecture and/or use NN tricks) we can boost their performance.",
    "1810428": "Note that NN benefit from data augmentation, learning schedules, unsupervised pretraining tasks, knowledge distillation, different loss functions, multitask learning, and many other techniques that we cannot do with GBT.\n\nMy best Transformer with GRU head offline beats my best GBT 😄 And both have much room for improvement !!"
  },
  "source": "meta"
}