{
  "id": 200938,
  "title": "Is TabNet better than GBT?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/200938",
  "author_name": "",
  "post_date": "2020-12-02T12:55:13.164731600Z",
  "votes": 28,
  "comment_count": 17,
  "views": 0,
  "content": "<p>TabNet looks promising with from this paper (it claims it can be better than Gradient Boosted Trees)<br>\n<a href=\"https://arxiv.org/pdf/1908.07442.pdf\" target=\"_blank\">https://arxiv.org/pdf/1908.07442.pdf</a></p>\n<p>I've tried the following implementation:<br>\n<a href=\"https://dreamquark-ai.github.io/tabnet/\" target=\"_blank\">https://dreamquark-ai.github.io/tabnet/</a><br>\n<a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a></p>\n<p>But it does not overperform any LightGBM/XGBoost/CatBoost. On 10M data I've got:</p>\n<ul>\n<li>LGB, Local CV 0.774</li>\n<li>CatBoost, Local CV 0.773</li>\n<li>XGBoost, Local CV 0.771</li>\n<li>TabNet, Local CV 0.761</li>\n</ul>\n<p>Anyone with similar or better results with TabNet and willing to share about?<br>\n<strong>Update</strong>: Inference is fast even with CPU.</p>",
  "messages": [
    {
      "id": "1099567",
      "postDate": "12/02/2020 12:55:13",
      "content": "<p>TabNet looks promising with from this paper (it claims it can be better than Gradient Boosted Trees)<br>\n<a href=\"https://arxiv.org/pdf/1908.07442.pdf\" target=\"_blank\">https://arxiv.org/pdf/1908.07442.pdf</a></p>\n<p>I've tried the following implementation:<br>\n<a href=\"https://dreamquark-ai.github.io/tabnet/\" target=\"_blank\">https://dreamquark-ai.github.io/tabnet/</a><br>\n<a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a></p>\n<p>But it does not overperform any LightGBM/XGBoost/CatBoost. On 10M data I've got:</p>\n<ul>\n<li>LGB, Local CV 0.774</li>\n<li>CatBoost, Local CV 0.773</li>\n<li>XGBoost, Local CV 0.771</li>\n<li>TabNet, Local CV 0.761</li>\n</ul>\n<p>Anyone with similar or better results with TabNet and willing to share about?<br>\n<strong>Update</strong>: Inference is fast even with CPU.</p>",
      "rawMarkdown": "TabNet looks promising with from this paper (it claims it can be better than Gradient Boosted Trees)\nhttps://arxiv.org/pdf/1908.07442.pdf\n\nI've tried the following implementation:\nhttps://dreamquark-ai.github.io/tabnet/\nhttps://github.com/dreamquark-ai/tabnet\n\nBut it does not overperform any LightGBM/XGBoost/CatBoost. On 10M data I've got:\n* LGB, Local CV 0.774\n* CatBoost, Local CV 0.773\n* XGBoost, Local CV 0.771\n* TabNet, Local CV 0.761\n\nAnyone with similar or better results with TabNet and willing to share about?\n**Update**: Inference is fast even with CPU.",
      "votes": null
    },
    {
      "id": "1099655",
      "postDate": "12/02/2020 14:09:49",
      "content": "<p>Training results for TabNet:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe198fe772854b3f61ac79867194f3673%2Ftabnet.png?generation=1606918180046909&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Training results for TabNet:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe198fe772854b3f61ac79867194f3673%2Ftabnet.png?generation=1606918180046909&alt=media)",
      "votes": null
    },
    {
      "id": "1100336",
      "postDate": "12/03/2020 02:24:45",
      "content": "<p>The same here.</p>\n<ul>\n<li>LGB Local CV AUC 0.767</li>\n<li>Tabnet Local CV AUC 0.762</li>\n</ul>\n<p>I used this public notebook (<a href=\"https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\" target=\"_blank\">https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel</a>) with some modifications.<br>\nI only used train[-2**24:] for training tabnet due to the limitation of pytorch sampler.<br>\nI still think that tabnet could be useful for model ensemble but I'm not sure.</p>",
      "rawMarkdown": "The same here.\n- LGB Local CV AUC 0.767\n- Tabnet Local CV AUC 0.762\n\nI used this public notebook (https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel) with some modifications.\nI only used train[-2**24:] for training tabnet due to the limitation of pytorch sampler.\nI still think that tabnet could be useful for model ensemble but I'm not sure.",
      "votes": null
    },
    {
      "id": "1100633",
      "postDate": "12/03/2020 07:55:22",
      "content": "<p>Thanks for the report. Interesting to see similar results.</p>",
      "rawMarkdown": "Thanks for the report. Interesting to see similar results.",
      "votes": null
    },
    {
      "id": "1100892",
      "postDate": "12/03/2020 12:44:44",
      "content": "<p>Hello,did u train on the whole train set?</p>",
      "rawMarkdown": "Hello,did u train on the whole train set?",
      "votes": null
    },
    {
      "id": "1100903",
      "postDate": "12/03/2020 12:58:31",
      "content": "<p>No just 10M.</p>",
      "rawMarkdown": "No just 10M.",
      "votes": null
    },
    {
      "id": "1100923",
      "postDate": "12/03/2020 13:21:28",
      "content": "<p>Thanks,I think your feature engineering is great.</p>",
      "rawMarkdown": "Thanks,I think your feature engineering is great.",
      "votes": null
    },
    {
      "id": "1102391",
      "postDate": "12/04/2020 21:55:15",
      "content": "<p>Thank you for sharing !  I noticed your LB is much higher than your CVs.  <br>\nIs your LB result an ensemble of models, a (happy) result of training on more data, or is it due to your CV strategy? </p>",
      "rawMarkdown": "Thank you for sharing !  I noticed your LB is much higher than your CVs.  \nIs your LB result an ensemble of models, a (happy) result of training on more data, or is it due to your CV strategy?",
      "votes": null
    },
    {
      "id": "1102407",
      "postDate": "12/04/2020 22:13:56",
      "content": "<p>It's an ensemble.</p>",
      "rawMarkdown": "It's an ensemble.",
      "votes": null
    },
    {
      "id": "1102943",
      "postDate": "12/05/2020 13:54:14",
      "content": "<p>On structured data  like this Lightgbm will beat tabnet anyday on the same feature subset.</p>",
      "rawMarkdown": "On structured data  like this Lightgbm will beat tabnet anyday on the same feature subset.",
      "votes": null
    },
    {
      "id": "1103675",
      "postDate": "12/06/2020 05:46:24",
      "content": "<p>Maybe the hyper parameter shake the model, So did you try search for magic parameter?</p>",
      "rawMarkdown": "Maybe the hyper parameter shake the model, So did you try search for magic parameter?",
      "votes": null
    },
    {
      "id": "1103685",
      "postDate": "12/06/2020 06:00:13",
      "content": "<p>The advantage of tabnet is you can do self-supervised pre-training before you move on to classification. And even though we don't have access to test, we do have a 100M dataset to work with. Can't do that with tree based models.</p>",
      "rawMarkdown": "The advantage of tabnet is you can do self-supervised pre-training before you move on to classification. And even though we don't have access to test, we do have a 100M dataset to work with. Can't do that with tree based models.",
      "votes": null
    },
    {
      "id": "1103693",
      "postDate": "12/06/2020 06:10:12",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, in my limited experiment, LGBM seems to score better than TabNet. What are your LB scores for those CV's listed? if you don't mind me asking, how are you ensembling to get your high LB score from those  models?</p>",
      "rawMarkdown": "mpware, in my limited experiment, LGBM seems to score better than TabNet. What are your LB scores for those CV's listed? if you don't mind me asking, how are you ensembling to get your high LB score from those  models?",
      "votes": null
    },
    {
      "id": "1103694",
      "postDate": "12/06/2020 06:21:41",
      "content": "<p>Using SAINT and TabNet, i believe is the ensemble coming from.</p>",
      "rawMarkdown": "Using SAINT and TabNet, i believe is the ensemble coming from.",
      "votes": null
    },
    {
      "id": "1103811",
      "postDate": "12/06/2020 10:16:48",
      "content": "<p>TabNet LB is 0.772 with CV=0.761 (trained with 10M data). I'm not sure TabNet could beat other XGB for this competition. I will try to continue to play with it but I'm also facing severe OOM with TabNet.  </p>",
      "rawMarkdown": "TabNet LB is 0.772 with CV=0.761 (trained with 10M data). I'm not sure TabNet could beat other XGB for this competition. I will try to continue to play with it but I'm also facing severe OOM with TabNet.",
      "votes": null
    },
    {
      "id": "1103855",
      "postDate": "12/06/2020 11:18:48",
      "content": "<blockquote>\n  <p>Update: Inference is fast even with CPU.</p>\n</blockquote>\n<p>Are you saying that you made inference for lgbm and SAINT combined on CPU?</p>",
      "rawMarkdown": ">Update: Inference is fast even with CPU.\n\nAre you saying that you made inference for lgbm and SAINT combined on CPU?",
      "votes": null
    },
    {
      "id": "1103861",
      "postDate": "12/06/2020 11:24:03",
      "content": "<p>No, I meant TabNet inference is fast on CPU.</p>",
      "rawMarkdown": "No, I meant TabNet inference is fast on CPU.",
      "votes": null
    },
    {
      "id": "1195486",
      "postDate": "02/10/2021 20:50:04",
      "content": "<p>cool thanks for sharing</p>",
      "rawMarkdown": "cool thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1099655,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "12/02/2020 14:09:49",
      "content": "<p>Training results for TabNet:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe198fe772854b3f61ac79867194f3673%2Ftabnet.png?generation=1606918180046909&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1103685,
          "author_name": "authman",
          "author_url": "",
          "post_date": "12/06/2020 06:00:13",
          "content": "<p>The advantage of tabnet is you can do self-supervised pre-training before you move on to classification. And even though we don't have access to test, we do have a 100M dataset to work with. Can't do that with tree based models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100336,
      "author_name": "tomooinubushi",
      "author_url": "",
      "post_date": "12/03/2020 02:24:45",
      "content": "<p>The same here.</p>\n<ul>\n<li>LGB Local CV AUC 0.767</li>\n<li>Tabnet Local CV AUC 0.762</li>\n</ul>\n<p>I used this public notebook (<a href=\"https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\" target=\"_blank\">https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel</a>) with some modifications.<br>\nI only used train[-2**24:] for training tabnet due to the limitation of pytorch sampler.<br>\nI still think that tabnet could be useful for model ensemble but I'm not sure.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100633,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/03/2020 07:55:22",
          "content": "<p>Thanks for the report. Interesting to see similar results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1195486,
          "author_name": "maciejgronczynski",
          "author_url": "",
          "post_date": "02/10/2021 20:50:04",
          "content": "<p>cool thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100892,
      "author_name": "zekun98",
      "author_url": "",
      "post_date": "12/03/2020 12:44:44",
      "content": "<p>Hello,did u train on the whole train set?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1100903,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/03/2020 12:58:31",
          "content": "<p>No just 10M.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1100923,
          "author_name": "zekun98",
          "author_url": "",
          "post_date": "12/03/2020 13:21:28",
          "content": "<p>Thanks,I think your feature engineering is great.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1102391,
      "author_name": "raphael1123",
      "author_url": "",
      "post_date": "12/04/2020 21:55:15",
      "content": "<p>Thank you for sharing !  I noticed your LB is much higher than your CVs.  <br>\nIs your LB result an ensemble of models, a (happy) result of training on more data, or is it due to your CV strategy? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1102407,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/04/2020 22:13:56",
          "content": "<p>It's an ensemble.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1102943,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "12/05/2020 13:54:14",
      "content": "<p>On structured data  like this Lightgbm will beat tabnet anyday on the same feature subset.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103675,
      "author_name": "biubiug",
      "author_url": "",
      "post_date": "12/06/2020 05:46:24",
      "content": "<p>Maybe the hyper parameter shake the model, So did you try search for magic parameter?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1103693,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "12/06/2020 06:10:12",
      "content": "<p><a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a>, in my limited experiment, LGBM seems to score better than TabNet. What are your LB scores for those CV's listed? if you don't mind me asking, how are you ensembling to get your high LB score from those  models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1103694,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/06/2020 06:21:41",
          "content": "<p>Using SAINT and TabNet, i believe is the ensemble coming from.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103811,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/06/2020 10:16:48",
          "content": "<p>TabNet LB is 0.772 with CV=0.761 (trained with 10M data). I'm not sure TabNet could beat other XGB for this competition. I will try to continue to play with it but I'm also facing severe OOM with TabNet.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103855,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "12/06/2020 11:18:48",
          "content": "<blockquote>\n  <p>Update: Inference is fast even with CPU.</p>\n</blockquote>\n<p>Are you saying that you made inference for lgbm and SAINT combined on CPU?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1103861,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "12/06/2020 11:24:03",
          "content": "<p>No, I meant TabNet inference is fast on CPU.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1099567": "TabNet looks promising with from this paper (it claims it can be better than Gradient Boosted Trees)\nhttps://arxiv.org/pdf/1908.07442.pdf\n\nI've tried the following implementation:\nhttps://dreamquark-ai.github.io/tabnet/\nhttps://github.com/dreamquark-ai/tabnet\n\nBut it does not overperform any LightGBM/XGBoost/CatBoost. On 10M data I've got:\n* LGB, Local CV 0.774\n* CatBoost, Local CV 0.773\n* XGBoost, Local CV 0.771\n* TabNet, Local CV 0.761\n\nAnyone with similar or better results with TabNet and willing to share about?\n**Update**: Inference is fast even with CPU.",
    "1099655": "Training results for TabNet:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F698363%2Fe198fe772854b3f61ac79867194f3673%2Ftabnet.png?generation=1606918180046909&alt=media)",
    "1100336": "The same here.\n- LGB Local CV AUC 0.767\n- Tabnet Local CV AUC 0.762\n\nI used this public notebook (https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel) with some modifications.\nI only used train[-2**24:] for training tabnet due to the limitation of pytorch sampler.\nI still think that tabnet could be useful for model ensemble but I'm not sure.",
    "1100633": "Thanks for the report. Interesting to see similar results.",
    "1100892": "Hello,did u train on the whole train set?",
    "1100903": "No just 10M.",
    "1100923": "Thanks,I think your feature engineering is great.",
    "1102391": "Thank you for sharing !  I noticed your LB is much higher than your CVs.  \nIs your LB result an ensemble of models, a (happy) result of training on more data, or is it due to your CV strategy?",
    "1102407": "It's an ensemble.",
    "1102943": "On structured data  like this Lightgbm will beat tabnet anyday on the same feature subset.",
    "1103675": "Maybe the hyper parameter shake the model, So did you try search for magic parameter?",
    "1103685": "The advantage of tabnet is you can do self-supervised pre-training before you move on to classification. And even though we don't have access to test, we do have a 100M dataset to work with. Can't do that with tree based models.",
    "1103693": "mpware, in my limited experiment, LGBM seems to score better than TabNet. What are your LB scores for those CV's listed? if you don't mind me asking, how are you ensembling to get your high LB score from those  models?",
    "1103694": "Using SAINT and TabNet, i believe is the ensemble coming from.",
    "1103811": "TabNet LB is 0.772 with CV=0.761 (trained with 10M data). I'm not sure TabNet could beat other XGB for this competition. I will try to continue to play with it but I'm also facing severe OOM with TabNet.",
    "1103855": ">Update: Inference is fast even with CPU.\n\nAre you saying that you made inference for lgbm and SAINT combined on CPU?",
    "1103861": "No, I meant TabNet inference is fast on CPU.",
    "1195486": "cool thanks for sharing"
  },
  "source": "meta"
}