{
  "id": 192466,
  "title": "Are you using a TabNet?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/192466",
  "author_name": "",
  "post_date": "2020-10-21T16:49:40.620099100Z",
  "votes": null,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've just red about TabNet. Here is it's implementation on PyTorch <a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a>. Will it be useful for this competition?</p>",
  "messages": [
    {
      "id": "1056392",
      "postDate": "10/21/2020 16:49:40",
      "content": "<p>I've just red about TabNet. Here is it's implementation on PyTorch <a href=\"https://github.com/dreamquark-ai/tabnet\" target=\"_blank\">https://github.com/dreamquark-ai/tabnet</a>. Will it be useful for this competition?</p>",
      "rawMarkdown": "I've just red about TabNet. Here is it's implementation on PyTorch https://github.com/dreamquark-ai/tabnet. Will it be useful for this competition?",
      "votes": null
    },
    {
      "id": "1056921",
      "postDate": "10/22/2020 07:45:58",
      "content": "<p>I've made a starter kernel : <a href=\"https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\" target=\"_blank\">https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel</a><br>\nIt is much slower than trees and provides lower results. It seems not well suited for very large data</p>",
      "rawMarkdown": "I've made a starter kernel : https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\nIt is much slower than trees and provides lower results. It seems not well suited for very large data",
      "votes": null
    },
    {
      "id": "1057662",
      "postDate": "10/22/2020 20:46:44",
      "content": "<p>I thought it would be worth a shot as well, and when comparing it to other methods I had the same experience as <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> : slower and performing worse. Perhaps later on when I have had the time to redesign the architecture and features I will try again.</p>",
      "rawMarkdown": "I thought it would be worth a shot as well, and when comparing it to other methods I had the same experience as @alexj21 : slower and performing worse. Perhaps later on when I have had the time to redesign the architecture and features I will try again.",
      "votes": null
    },
    {
      "id": "1057695",
      "postDate": "10/22/2020 22:03:55",
      "content": "<p>This was my experience as well, although I used a TabNetRegressor instead of a TabNetClassifier. Didn't submit to the LB because validation scores were only around ~0.72-0.73</p>",
      "rawMarkdown": "This was my experience as well, although I used a TabNetRegressor instead of a TabNetClassifier. Didn't submit to the LB because validation scores were only around ~0.72-0.73",
      "votes": null
    },
    {
      "id": "1057958",
      "postDate": "10/23/2020 07:26:31",
      "content": "<p>Depending on your validation scheme, local 0.73 could lead to a nice LB score.<br>\nAre you using trees right now ?</p>",
      "rawMarkdown": "Depending on your validation scheme, local 0.73 could lead to a nice LB score.\nAre you using trees right now ?",
      "votes": null
    },
    {
      "id": "1057962",
      "postDate": "10/23/2020 07:32:08",
      "content": "<p>Yeah if I remember well I've seen a notebook in which the user redesigned the network<br>\nIt would be an elegant solution. I would have reached top 7% in OSIC challenge with a single tabnet if I had chosen it as final solution</p>",
      "rawMarkdown": "Yeah if I remember well I've seen a notebook in which the user redesigned the network\nIt would be an elegant solution. I would have reached top 7% in OSIC challenge with a single tabnet if I had chosen it as final solution",
      "votes": null
    },
    {
      "id": "1058366",
      "postDate": "10/23/2020 16:02:02",
      "content": "<p>Yep, just a single LightGBM model with a lot of feature engineering. </p>",
      "rawMarkdown": "Yep, just a single LightGBM model with a lot of feature engineering.",
      "votes": null
    },
    {
      "id": "1058463",
      "postDate": "10/23/2020 18:33:57",
      "content": "<p>Great Going and Good luck! I hope that single lgbm is not trained on whole data. In case it's, What's the peak mem requirement and time it takes to do a pass? Ty!</p>",
      "rawMarkdown": "Great Going and Good luck! I hope that single lgbm is not trained on whole data. In case it's, What's the peak mem requirement and time it takes to do a pass? Ty!",
      "votes": null
    },
    {
      "id": "1058469",
      "postDate": "10/23/2020 18:43:04",
      "content": "<p>Thanks! Nah it's only trained on a small subset of the data--this is how I implement the train/validation split:</p>\n<pre><code>print(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n</code></pre>\n<p>Regarding memory, it is a bit difficult to keep within the limits, but there's a ton of tricks you can implement to conserve it, including gc, doing memory-intensive calculations in a separate notebook, and the like. It currently takes about two hours to train a model, and then about four hours to see results on the LB after using that saved model in a separate run to generate predictions.</p>",
      "rawMarkdown": "Thanks! Nah it's only trained on a small subset of the data--this is how I implement the train/validation split:\n\n```\nprint(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n```\nRegarding memory, it is a bit difficult to keep within the limits, but there's a ton of tricks you can implement to conserve it, including gc, doing memory-intensive calculations in a separate notebook, and the like. It currently takes about two hours to train a model, and then about four hours to see results on the LB after using that saved model in a separate run to generate predictions.",
      "votes": null
    },
    {
      "id": "1058475",
      "postDate": "10/23/2020 18:53:10",
      "content": "<p>Thanks for sharing Nick! I guess, I am over thinking about the problem statement 😅!</p>",
      "rawMarkdown": "Thanks for sharing Nick! I guess, I am over thinking about the problem statement 😅!",
      "votes": null
    },
    {
      "id": "1063663",
      "postDate": "10/29/2020 07:31:48",
      "content": "<p>Thak you for sharing! <br>\n When creating features for validation, do you use combined_df or combined_df + trainig? </p>",
      "rawMarkdown": "Thak you for sharing! \n When creating features for validation, do you use combined_df or combined_df + trainig?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1056921,
      "author_name": "alexj21",
      "author_url": "",
      "post_date": "10/22/2020 07:45:58",
      "content": "<p>I've made a starter kernel : <a href=\"https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\" target=\"_blank\">https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel</a><br>\nIt is much slower than trees and provides lower results. It seems not well suited for very large data</p>",
      "votes": null,
      "replies": [
        {
          "id": 1057695,
          "author_name": "misfyre",
          "author_url": "",
          "post_date": "10/22/2020 22:03:55",
          "content": "<p>This was my experience as well, although I used a TabNetRegressor instead of a TabNetClassifier. Didn't submit to the LB because validation scores were only around ~0.72-0.73</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1057958,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/23/2020 07:26:31",
          "content": "<p>Depending on your validation scheme, local 0.73 could lead to a nice LB score.<br>\nAre you using trees right now ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058366,
          "author_name": "misfyre",
          "author_url": "",
          "post_date": "10/23/2020 16:02:02",
          "content": "<p>Yep, just a single LightGBM model with a lot of feature engineering. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058463,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/23/2020 18:33:57",
          "content": "<p>Great Going and Good luck! I hope that single lgbm is not trained on whole data. In case it's, What's the peak mem requirement and time it takes to do a pass? Ty!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058469,
          "author_name": "misfyre",
          "author_url": "",
          "post_date": "10/23/2020 18:43:04",
          "content": "<p>Thanks! Nah it's only trained on a small subset of the data--this is how I implement the train/validation split:</p>\n<pre><code>print(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n</code></pre>\n<p>Regarding memory, it is a bit difficult to keep within the limits, but there's a ton of tricks you can implement to conserve it, including gc, doing memory-intensive calculations in a separate notebook, and the like. It currently takes about two hours to train a model, and then about four hours to see results on the LB after using that saved model in a separate run to generate predictions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1058475,
          "author_name": "adityaecdrid",
          "author_url": "",
          "post_date": "10/23/2020 18:53:10",
          "content": "<p>Thanks for sharing Nick! I guess, I am over thinking about the problem statement 😅!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1063663,
          "author_name": "d1348k",
          "author_url": "",
          "post_date": "10/29/2020 07:31:48",
          "content": "<p>Thak you for sharing! <br>\n When creating features for validation, do you use combined_df or combined_df + trainig? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1057662,
      "author_name": "isaacllorente",
      "author_url": "",
      "post_date": "10/22/2020 20:46:44",
      "content": "<p>I thought it would be worth a shot as well, and when comparing it to other methods I had the same experience as <a href=\"https://www.kaggle.com/alexj21\" target=\"_blank\">@alexj21</a> : slower and performing worse. Perhaps later on when I have had the time to redesign the architecture and features I will try again.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1057962,
          "author_name": "alexj21",
          "author_url": "",
          "post_date": "10/23/2020 07:32:08",
          "content": "<p>Yeah if I remember well I've seen a notebook in which the user redesigned the network<br>\nIt would be an elegant solution. I would have reached top 7% in OSIC challenge with a single tabnet if I had chosen it as final solution</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1056392": "I've just red about TabNet. Here is it's implementation on PyTorch https://github.com/dreamquark-ai/tabnet. Will it be useful for this competition?",
    "1056921": "I've made a starter kernel : https://www.kaggle.com/alexj21/riiid-tabnet-starter-kernel\nIt is much slower than trees and provides lower results. It seems not well suited for very large data",
    "1057662": "I thought it would be worth a shot as well, and when comparing it to other methods I had the same experience as @alexj21 : slower and performing worse. Perhaps later on when I have had the time to redesign the architecture and features I will try again.",
    "1057695": "This was my experience as well, although I used a TabNetRegressor instead of a TabNetClassifier. Didn't submit to the LB because validation scores were only around ~0.72-0.73",
    "1057958": "Depending on your validation scheme, local 0.73 could lead to a nice LB score.\nAre you using trees right now ?",
    "1057962": "Yeah if I remember well I've seen a notebook in which the user redesigned the network\nIt would be an elegant solution. I would have reached top 7% in OSIC challenge with a single tabnet if I had chosen it as final solution",
    "1058366": "Yep, just a single LightGBM model with a lot of feature engineering.",
    "1058463": "Great Going and Good luck! I hope that single lgbm is not trained on whole data. In case it's, What's the peak mem requirement and time it takes to do a pass? Ty!",
    "1058469": "Thanks! Nah it's only trained on a small subset of the data--this is how I implement the train/validation split:\n\n```\nprint(\"[1] Create Training Set\")\ntraining = combined_df.groupby(\"user_id\").tail(24)\ncombined_df = combined_df.drop(training.index)\n\nprint(\"[2] Split Training Set into Validation\")\nvalidation = training.groupby(\"user_id\").tail(6)\ntraining = training.drop(validation.index)\n```\nRegarding memory, it is a bit difficult to keep within the limits, but there's a ton of tricks you can implement to conserve it, including gc, doing memory-intensive calculations in a separate notebook, and the like. It currently takes about two hours to train a model, and then about four hours to see results on the LB after using that saved model in a separate run to generate predictions.",
    "1058475": "Thanks for sharing Nick! I guess, I am over thinking about the problem statement 😅!",
    "1063663": "Thak you for sharing! \n When creating features for validation, do you use combined_df or combined_df + trainig?"
  },
  "source": "meta"
}