{
  "id": 343966,
  "title": "[Paper Summary] Why do tree-based models still outperform deep learning on tabular data?",
  "url": "/competitions/amex-default-prediction/discussion/343966",
  "author_name": "",
  "post_date": "2022-08-13T09:52:26.874149800Z",
  "votes": 31,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Saw this paper discussing the performance of NNs and tree-based models on tabular data, thought it might be of interest: <a href=\"https://arxiv.org/abs/2207.08815\" target=\"_blank\">https://arxiv.org/abs/2207.08815</a></p>\n<p>Summary:</p>\n<ul>\n<li><p>Created a benchmark meta-dataset (45 tabular datasets across domains + some pre-processing of the data to remove NAs etc)</p></li>\n<li><p>Compared performance of NNs (MLP, Resnet) and Transformers (FT_Transformer and SAINT) against tree-based models (sklearn RandomForest, sklearn GradientBoostingTrees and XGBoost) with a wide random hyper-parameter search over a range of search iterations<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2Ffba63f33840dba63d29ccf276be92625%2F1.png?generation=1660382415384929&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2F64592b2db12897d95482687932d3c201%2F2.png?generation=1660382426259832&amp;alt=media\" alt=\"\"><br>\n[Top: Datasets containing only numerical features, Bottom: Datasets containing both numerical and categorical. Dotted lines represent performance using default parameters]</p></li>\n<li><p>Tuning hyper-parameters does not make the NNs or transformers state of the art, tree ensembles still perform best for every random search budget</p></li>\n<li><p>Categorical features aren't the only hurdle for NN performance. A large performance gap relative to tree based methods is still observed on the datasets containing only numerical features</p></li>\n<li><p>Investigated NN performance by altering the datasets to probe the reason for performance gap:</p></li>\n<li><p>Smoothing the target function in the datasets revealed that NN perform worse as they are biased toward smooth target functions, whereas trees can fit to more irregular functions which the tabular datasets contain</p></li>\n<li><p>MLP architectures are less robust to uninformative features, manually dropping them prior reduced the performance gap between tree methods. Adding uninformative features widened this gap.</p></li>\n<li><p>Resnet is rotationally invariant as a learner but the structure of tabular data is not. When the data was randomly rotated, performance of NNs were higher than tree models, suggesting that rotational invariance in the learner is not desirable.</p></li>\n</ul>",
  "messages": [
    {
      "id": "1896948",
      "postDate": "08/13/2022 09:52:26",
      "content": "<p>Saw this paper discussing the performance of NNs and tree-based models on tabular data, thought it might be of interest: <a href=\"https://arxiv.org/abs/2207.08815\" target=\"_blank\">https://arxiv.org/abs/2207.08815</a></p>\n<p>Summary:</p>\n<ul>\n<li><p>Created a benchmark meta-dataset (45 tabular datasets across domains + some pre-processing of the data to remove NAs etc)</p></li>\n<li><p>Compared performance of NNs (MLP, Resnet) and Transformers (FT_Transformer and SAINT) against tree-based models (sklearn RandomForest, sklearn GradientBoostingTrees and XGBoost) with a wide random hyper-parameter search over a range of search iterations<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2Ffba63f33840dba63d29ccf276be92625%2F1.png?generation=1660382415384929&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2F64592b2db12897d95482687932d3c201%2F2.png?generation=1660382426259832&amp;alt=media\" alt=\"\"><br>\n[Top: Datasets containing only numerical features, Bottom: Datasets containing both numerical and categorical. Dotted lines represent performance using default parameters]</p></li>\n<li><p>Tuning hyper-parameters does not make the NNs or transformers state of the art, tree ensembles still perform best for every random search budget</p></li>\n<li><p>Categorical features aren't the only hurdle for NN performance. A large performance gap relative to tree based methods is still observed on the datasets containing only numerical features</p></li>\n<li><p>Investigated NN performance by altering the datasets to probe the reason for performance gap:</p></li>\n<li><p>Smoothing the target function in the datasets revealed that NN perform worse as they are biased toward smooth target functions, whereas trees can fit to more irregular functions which the tabular datasets contain</p></li>\n<li><p>MLP architectures are less robust to uninformative features, manually dropping them prior reduced the performance gap between tree methods. Adding uninformative features widened this gap.</p></li>\n<li><p>Resnet is rotationally invariant as a learner but the structure of tabular data is not. When the data was randomly rotated, performance of NNs were higher than tree models, suggesting that rotational invariance in the learner is not desirable.</p></li>\n</ul>",
      "rawMarkdown": "Saw this paper discussing the performance of NNs and tree-based models on tabular data, thought it might be of interest: [https://arxiv.org/abs/2207.08815](https://arxiv.org/abs/2207.08815)\n\nSummary:\n* Created a benchmark meta-dataset (45 tabular datasets across domains + some pre-processing of the data to remove NAs etc)\n* Compared performance of NNs (MLP, Resnet) and Transformers (FT_Transformer and SAINT) against tree-based models (sklearn RandomForest, sklearn GradientBoostingTrees and XGBoost) with a wide random hyper-parameter search over a range of search iterations\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2Ffba63f33840dba63d29ccf276be92625%2F1.png?generation=1660382415384929&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2F64592b2db12897d95482687932d3c201%2F2.png?generation=1660382426259832&alt=media)\n[Top: Datasets containing only numerical features, Bottom: Datasets containing both numerical and categorical. Dotted lines represent performance using default parameters]\n* Tuning hyper-parameters does not make the NNs or transformers state of the art, tree ensembles still perform best for every random search budget\n* Categorical features aren't the only hurdle for NN performance. A large performance gap relative to tree based methods is still observed on the datasets containing only numerical features\n\n* Investigated NN performance by altering the datasets to probe the reason for performance gap:\n* Smoothing the target function in the datasets revealed that NN perform worse as they are biased toward smooth target functions, whereas trees can fit to more irregular functions which the tabular datasets contain\n* MLP architectures are less robust to uninformative features, manually dropping them prior reduced the performance gap between tree methods. Adding uninformative features widened this gap.\n* Resnet is rotationally invariant as a learner but the structure of tabular data is not. When the data was randomly rotated, performance of NNs were higher than tree models, suggesting that rotational invariance in the learner is not desirable.",
      "votes": null
    },
    {
      "id": "1897170",
      "postDate": "08/13/2022 13:53:54",
      "content": "<p>I've shared some thoughts about this before <a href=\"https://www.kaggle.com/competitions/avito-demand-prediction/discussion/57085#331588\" target=\"_blank\">here</a>. My favorite way to think about this is to focus on the difference between hierarchical representation learning on homogenous, low signal features (NN) and interactive partitioning on heterogenous, high signal features (GBT).  </p>\n<blockquote>\n  <p>It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned.</p>\n</blockquote>\n<p>But just like in that competition, this data has a non-tabular aspect. In the end, I expect to see well-constructed neural networks that are within ~.001-.002 of the best GBT models if not on par with their performance.</p>",
      "rawMarkdown": "I've shared some thoughts about this before [here](https://www.kaggle.com/competitions/avito-demand-prediction/discussion/57085#331588). My favorite way to think about this is to focus on the difference between hierarchical representation learning on homogenous, low signal features (NN) and interactive partitioning on heterogenous, high signal features (GBT).  \n\n> It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned.\n\nBut just like in that competition, this data has a non-tabular aspect. In the end, I expect to see well-constructed neural networks that are within ~.001-.002 of the best GBT models if not on par with their performance.",
      "votes": null
    },
    {
      "id": "1897548",
      "postDate": "08/13/2022 21:34:32",
      "content": "<p>Thank you for sharing this thought! The comparison between price and pixel is fascinating. <br>\nJust curious, why did you say this data has a non-tabular aspect? </p>",
      "rawMarkdown": "Thank you for sharing this thought! The comparison between price and pixel is fascinating. \nJust curious, why did you say this data has a non-tabular aspect?",
      "votes": null
    },
    {
      "id": "1897585",
      "postDate": "08/13/2022 22:21:25",
      "content": "<p>I'd usually define tabular as traditional rows and columns format, i.e. each data point has a static row representation with no spatial structure. Here each data point is really a time series of customer statements, so non-tabular. But of course the literal storage format is still tabular so fair to describe it that way. I'm basically using non-tabular to mean -- it would make sense to train a model with data points represented as 2D+ tensors instead of 1D arrays.</p>",
      "rawMarkdown": "I'd usually define tabular as traditional rows and columns format, i.e. each data point has a static row representation with no spatial structure. Here each data point is really a time series of customer statements, so non-tabular. But of course the literal storage format is still tabular so fair to describe it that way. I'm basically using non-tabular to mean -- it would make sense to train a model with data points represented as 2D+ tensors instead of 1D arrays.",
      "votes": null
    },
    {
      "id": "1898162",
      "postDate": "08/14/2022 11:06:37",
      "content": "<p>AFAIR I stopped to read the paper when they said that they truncate number of rows to 10k.</p>",
      "rawMarkdown": "AFAIR I stopped to read the paper when they said that they truncate number of rows to 10k.",
      "votes": null
    },
    {
      "id": "1898345",
      "postDate": "08/14/2022 14:26:58",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": null
    },
    {
      "id": "1898611",
      "postDate": "08/14/2022 17:09:37",
      "content": "<p>What about outliers in the training data? A tree based method can isolate and model an outlier without impacting predictions for other points, by placing splits on either side of the outlier. Is this a natural thing for NN to do?</p>",
      "rawMarkdown": "What about outliers in the training data? A tree based method can isolate and model an outlier without impacting predictions for other points, by placing splits on either side of the outlier. Is this a natural thing for NN to do?",
      "votes": null
    },
    {
      "id": "1898884",
      "postDate": "08/14/2022 23:28:01",
      "content": "<p>Thanks this is helpful in at-least telling what is the best NN we can maybe ensemble with GBDT's .. SAINT was trying but do they have a code repo also with the experiments or will lookup at Paperswithcode .</p>",
      "rawMarkdown": "Thanks this is helpful in at-least telling what is the best NN we can maybe ensemble with GBDT's .. SAINT was trying but do they have a code repo also with the experiments or will lookup at Paperswithcode .",
      "votes": null
    },
    {
      "id": "1898991",
      "postDate": "08/15/2022 01:37:20",
      "content": "<p>You could have a check with the Riiid competition, in which SAINT-like models are heavily used, like this open kernel: <a href=\"https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run\" target=\"_blank\">https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run</a></p>",
      "rawMarkdown": "You could have a check with the Riiid competition, in which SAINT-like models are heavily used, like this open kernel: https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run",
      "votes": null
    },
    {
      "id": "1899537",
      "postDate": "08/15/2022 10:32:20",
      "content": "<p>There is also a new architecture called Hopular ; its based on  Hopfield Networks. This architecture is basically a Transformer that, in the case of categorical features, find the best embedding for them. It's Worth taking a look at the paper, however methods like LGBM still performs same as Popular or even better.<br>\nLink to the paper : <a href=\"https://arxiv.org/abs/2206.00664\" target=\"_blank\">https://arxiv.org/abs/2206.00664</a></p>",
      "rawMarkdown": "There is also a new architecture called Hopular ; its based on  Hopfield Networks. This architecture is basically a Transformer that, in the case of categorical features, find the best embedding for them. It's Worth taking a look at the paper, however methods like LGBM still performs same as Popular or even better.\nLink to the paper : https://arxiv.org/abs/2206.00664",
      "votes": null
    },
    {
      "id": "1900272",
      "postDate": "08/15/2022 21:16:14",
      "content": "<p>Thanks was able to check it out but so far didnt give good results compared to Tabnet </p>",
      "rawMarkdown": "Thanks was able to check it out but so far didnt give good results compared to Tabnet",
      "votes": null
    },
    {
      "id": "1900277",
      "postDate": "08/15/2022 21:23:03",
      "content": "<p>Thanks for posting. Just what I was looking for!</p>",
      "rawMarkdown": "Thanks for posting. Just what I was looking for!",
      "votes": null
    },
    {
      "id": "1900801",
      "postDate": "08/16/2022 09:05:45",
      "content": "<p>Thanks for posting.</p>",
      "rawMarkdown": "Thanks for posting.",
      "votes": null
    },
    {
      "id": "1903225",
      "postDate": "08/17/2022 07:53:05",
      "content": "<p>Thanks, this is helpful</p>",
      "rawMarkdown": "Thanks, this is helpful",
      "votes": null
    },
    {
      "id": "1906460",
      "postDate": "08/19/2022 22:45:29",
      "content": "<p>I really like the approach of this research in general. I hope that one day we would find a better NN architecture that will be competitive to GBMs on tabular data. <br>\n(It would make one hell of an ensemble ;) )</p>\n<p>The Devastator.</p>",
      "rawMarkdown": "I really like the approach of this research in general. I hope that one day we would find a better NN architecture that will be competitive to GBMs on tabular data. \n(It would make one hell of an ensemble ;) )\n\nThe Devastator.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1897170,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "08/13/2022 13:53:54",
      "content": "<p>I've shared some thoughts about this before <a href=\"https://www.kaggle.com/competitions/avito-demand-prediction/discussion/57085#331588\" target=\"_blank\">here</a>. My favorite way to think about this is to focus on the difference between hierarchical representation learning on homogenous, low signal features (NN) and interactive partitioning on heterogenous, high signal features (GBT).  </p>\n<blockquote>\n  <p>It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned.</p>\n</blockquote>\n<p>But just like in that competition, this data has a non-tabular aspect. In the end, I expect to see well-constructed neural networks that are within ~.001-.002 of the best GBT models if not on par with their performance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1897548,
          "author_name": "raphael1123",
          "author_url": "",
          "post_date": "08/13/2022 21:34:32",
          "content": "<p>Thank you for sharing this thought! The comparison between price and pixel is fascinating. <br>\nJust curious, why did you say this data has a non-tabular aspect? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1897585,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "08/13/2022 22:21:25",
          "content": "<p>I'd usually define tabular as traditional rows and columns format, i.e. each data point has a static row representation with no spatial structure. Here each data point is really a time series of customer statements, so non-tabular. But of course the literal storage format is still tabular so fair to describe it that way. I'm basically using non-tabular to mean -- it would make sense to train a model with data points represented as 2D+ tensors instead of 1D arrays.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1898162,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "08/14/2022 11:06:37",
      "content": "<p>AFAIR I stopped to read the paper when they said that they truncate number of rows to 10k.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1898345,
      "author_name": "",
      "author_url": "",
      "post_date": "08/14/2022 14:26:58",
      "content": "<p>Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1898611,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "08/14/2022 17:09:37",
      "content": "<p>What about outliers in the training data? A tree based method can isolate and model an outlier without impacting predictions for other points, by placing splits on either side of the outlier. Is this a natural thing for NN to do?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1898884,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "08/14/2022 23:28:01",
      "content": "<p>Thanks this is helpful in at-least telling what is the best NN we can maybe ensemble with GBDT's .. SAINT was trying but do they have a code repo also with the experiments or will lookup at Paperswithcode .</p>",
      "votes": null,
      "replies": [
        {
          "id": 1898991,
          "author_name": "lihaorocky",
          "author_url": "",
          "post_date": "08/15/2022 01:37:20",
          "content": "<p>You could have a check with the Riiid competition, in which SAINT-like models are heavily used, like this open kernel: <a href=\"https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run\" target=\"_blank\">https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1900272,
          "author_name": "gauravbrills",
          "author_url": "",
          "post_date": "08/15/2022 21:16:14",
          "content": "<p>Thanks was able to check it out but so far didnt give good results compared to Tabnet </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1899537,
      "author_name": "rayanaay",
      "author_url": "",
      "post_date": "08/15/2022 10:32:20",
      "content": "<p>There is also a new architecture called Hopular ; its based on  Hopfield Networks. This architecture is basically a Transformer that, in the case of categorical features, find the best embedding for them. It's Worth taking a look at the paper, however methods like LGBM still performs same as Popular or even better.<br>\nLink to the paper : <a href=\"https://arxiv.org/abs/2206.00664\" target=\"_blank\">https://arxiv.org/abs/2206.00664</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1900277,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "08/15/2022 21:23:03",
      "content": "<p>Thanks for posting. Just what I was looking for!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1900801,
      "author_name": "",
      "author_url": "",
      "post_date": "08/16/2022 09:05:45",
      "content": "<p>Thanks for posting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1903225,
      "author_name": "",
      "author_url": "",
      "post_date": "08/17/2022 07:53:05",
      "content": "<p>Thanks, this is helpful</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1906460,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "08/19/2022 22:45:29",
      "content": "<p>I really like the approach of this research in general. I hope that one day we would find a better NN architecture that will be competitive to GBMs on tabular data. <br>\n(It would make one hell of an ensemble ;) )</p>\n<p>The Devastator.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1896948": "Saw this paper discussing the performance of NNs and tree-based models on tabular data, thought it might be of interest: [https://arxiv.org/abs/2207.08815](https://arxiv.org/abs/2207.08815)\n\nSummary:\n* Created a benchmark meta-dataset (45 tabular datasets across domains + some pre-processing of the data to remove NAs etc)\n* Compared performance of NNs (MLP, Resnet) and Transformers (FT_Transformer and SAINT) against tree-based models (sklearn RandomForest, sklearn GradientBoostingTrees and XGBoost) with a wide random hyper-parameter search over a range of search iterations\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2Ffba63f33840dba63d29ccf276be92625%2F1.png?generation=1660382415384929&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7653735%2F64592b2db12897d95482687932d3c201%2F2.png?generation=1660382426259832&alt=media)\n[Top: Datasets containing only numerical features, Bottom: Datasets containing both numerical and categorical. Dotted lines represent performance using default parameters]\n* Tuning hyper-parameters does not make the NNs or transformers state of the art, tree ensembles still perform best for every random search budget\n* Categorical features aren't the only hurdle for NN performance. A large performance gap relative to tree based methods is still observed on the datasets containing only numerical features\n\n* Investigated NN performance by altering the datasets to probe the reason for performance gap:\n* Smoothing the target function in the datasets revealed that NN perform worse as they are biased toward smooth target functions, whereas trees can fit to more irregular functions which the tabular datasets contain\n* MLP architectures are less robust to uninformative features, manually dropping them prior reduced the performance gap between tree methods. Adding uninformative features widened this gap.\n* Resnet is rotationally invariant as a learner but the structure of tabular data is not. When the data was randomly rotated, performance of NNs were higher than tree models, suggesting that rotational invariance in the learner is not desirable.",
    "1897170": "I've shared some thoughts about this before [here](https://www.kaggle.com/competitions/avito-demand-prediction/discussion/57085#331588). My favorite way to think about this is to focus on the difference between hierarchical representation learning on homogenous, low signal features (NN) and interactive partitioning on heterogenous, high signal features (GBT).  \n\n> It's harder to see how hierarchical layering would be as naturally useful for most tabular tasks. For example, something like \"price\" is already a very well-described feature that doesn't need to be intricately manipulated to have a lot of predictive value, whereas individual pixels in an image need a lot of work to be leveraged as signals. This is why you'll often see that shallow neural networks outperform deeper ones on tabular data. And to capture explicit, direct feature interactions, it's more natural to use tree based methods (especially gradient-boosted trees) that allow for repeated interactive partitioning that derives from meaningful features as is instead of from more complex representations that need to be learned.\n\nBut just like in that competition, this data has a non-tabular aspect. In the end, I expect to see well-constructed neural networks that are within ~.001-.002 of the best GBT models if not on par with their performance.",
    "1897548": "Thank you for sharing this thought! The comparison between price and pixel is fascinating. \nJust curious, why did you say this data has a non-tabular aspect?",
    "1897585": "I'd usually define tabular as traditional rows and columns format, i.e. each data point has a static row representation with no spatial structure. Here each data point is really a time series of customer statements, so non-tabular. But of course the literal storage format is still tabular so fair to describe it that way. I'm basically using non-tabular to mean -- it would make sense to train a model with data points represented as 2D+ tensors instead of 1D arrays.",
    "1898162": "AFAIR I stopped to read the paper when they said that they truncate number of rows to 10k.",
    "1898345": "Thank you for sharing",
    "1898611": "What about outliers in the training data? A tree based method can isolate and model an outlier without impacting predictions for other points, by placing splits on either side of the outlier. Is this a natural thing for NN to do?",
    "1898884": "Thanks this is helpful in at-least telling what is the best NN we can maybe ensemble with GBDT's .. SAINT was trying but do they have a code repo also with the experiments or will lookup at Paperswithcode .",
    "1898991": "You could have a check with the Riiid competition, in which SAINT-like models are heavily used, like this open kernel: https://www.kaggle.com/code/shivanandmn/saint-training-using-pytorch-success-run",
    "1899537": "There is also a new architecture called Hopular ; its based on  Hopfield Networks. This architecture is basically a Transformer that, in the case of categorical features, find the best embedding for them. It's Worth taking a look at the paper, however methods like LGBM still performs same as Popular or even better.\nLink to the paper : https://arxiv.org/abs/2206.00664",
    "1900272": "Thanks was able to check it out but so far didnt give good results compared to Tabnet",
    "1900277": "Thanks for posting. Just what I was looking for!",
    "1900801": "Thanks for posting.",
    "1903225": "Thanks, this is helpful",
    "1906460": "I really like the approach of this research in general. I hope that one day we would find a better NN architecture that will be competitive to GBMs on tabular data. \n(It would make one hell of an ensemble ;) )\n\nThe Devastator."
  },
  "source": "meta"
}