{
  "id": 551431,
  "title": "Sharing the code using PreTrain TabNet",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/551431",
  "author_name": "MJeremy",
  "post_date": "2024-12-13T05:34:03.120000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Pretraining in Tabnet can help the model learn a rich representation of the features.</p>\n<pre><code>        unsupervised_model = TabNetPretrainer(\n            =torch.optim.Adam,\n            =dict(lr=2e-2),\n            = # \n        )\n\n        unsupervised_model.fit(\n            =X_train_np,\n            =0.8, # use 80% data  predict the rest\n            =60\n        )\n\n        params = {\n                : 64,              # Width of the decision prediction layer\n                : 64,              # Width of the attention embedding  each \n                : 5,           # Number of steps  the architecture\n                : 1.5,           # Coefficient  feature selection regularization\n                : 2,     # Number of independent GLU layer  each GLU block\n                : 2,          # Number of shared GLU layer  each GLU block\n                : 1e-4,  # Sparsity regularization\n                : torch.optim.Adam,\n                : dict(=2e-2, =1e-5),\n                : ,\n                : dict(=, =10, =1e-5, =0.5),\n                : torch.optim.lr_scheduler.ReduceLROnPlateau,\n                : 1,\n                :   torch.cuda.is_available()  \n        }\n        model = TabNetRegressor(**params)\n\n        model.fit(\n            X_train_np, \n            y_train_np.reshape(-1, 1),\n            eval_set=[(X_val_np, y_val_np.reshape(-1, 1))],\n            eval_metric=[],\n            =200,\n            =10,  \n            =256,\n            =unsupervised_model\n        )\n</code></pre>",
  "messages": [
    {
      "id": 3070853,
      "postDate": "2024-12-13T05:34:03.120Z",
      "content": "<p>Pretraining in Tabnet can help the model learn a rich representation of the features.</p>\n<pre><code>        unsupervised_model = TabNetPretrainer(\n            =torch.optim.Adam,\n            =dict(lr=2e-2),\n            = # \n        )\n\n        unsupervised_model.fit(\n            =X_train_np,\n            =0.8, # use 80% data  predict the rest\n            =60\n        )\n\n        params = {\n                : 64,              # Width of the decision prediction layer\n                : 64,              # Width of the attention embedding  each \n                : 5,           # Number of steps  the architecture\n                : 1.5,           # Coefficient  feature selection regularization\n                : 2,     # Number of independent GLU layer  each GLU block\n                : 2,          # Number of shared GLU layer  each GLU block\n                : 1e-4,  # Sparsity regularization\n                : torch.optim.Adam,\n                : dict(=2e-2, =1e-5),\n                : ,\n                : dict(=, =10, =1e-5, =0.5),\n                : torch.optim.lr_scheduler.ReduceLROnPlateau,\n                : 1,\n                :   torch.cuda.is_available()  \n        }\n        model = TabNetRegressor(**params)\n\n        model.fit(\n            X_train_np, \n            y_train_np.reshape(-1, 1),\n            eval_set=[(X_val_np, y_val_np.reshape(-1, 1))],\n            eval_metric=[],\n            =200,\n            =10,  \n            =256,\n            =unsupervised_model\n        )\n</code></pre>",
      "rawMarkdown": "Pretraining in Tabnet can help the model learn a rich representation of the features.\n\n```\n        unsupervised_model = TabNetPretrainer(\n            optimizer_fn=torch.optim.Adam,\n            optimizer_params=dict(lr=2e-2),\n            mask_type='entmax' # \"sparsemax\"\n        )\n        \n        unsupervised_model.fit(\n            X_train=X_train_np,\n            pretraining_ratio=0.8, # use 80% data to predict the rest\n            max_epochs=60\n        )\n\n        params = {\n                'n_d': 64,              # Width of the decision prediction layer\n                'n_a': 64,              # Width of the attention embedding for each step\n                'n_steps': 5,           # Number of steps in the architecture\n                'gamma': 1.5,           # Coefficient for feature selection regularization\n                'n_independent': 2,     # Number of independent GLU layer in each GLU block\n                'n_shared': 2,          # Number of shared GLU layer in each GLU block\n                'lambda_sparse': 1e-4,  # Sparsity regularization\n                'optimizer_fn': torch.optim.Adam,\n                'optimizer_params': dict(lr=2e-2, weight_decay=1e-5),\n                'mask_type': 'entmax',\n                'scheduler_params': dict(mode=\"min\", patience=10, min_lr=1e-5, factor=0.5),\n                'scheduler_fn': torch.optim.lr_scheduler.ReduceLROnPlateau,\n                'verbose': 1,\n                'device_name': 'cuda' if torch.cuda.is_available() else 'cpu'\n        }\n        model = TabNetRegressor(**params)\n        \n        model.fit(\n            X_train_np, \n            y_train_np.reshape(-1, 1),\n            eval_set=[(X_val_np, y_val_np.reshape(-1, 1))],\n            eval_metric=['mse'],\n            max_epochs=200,\n            patience=10,  \n            batch_size=256,\n            from_unsupervised=unsupervised_model\n        )\n```",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3070853": "Pretraining in Tabnet can help the model learn a rich representation of the features.\n\n```\n        unsupervised_model = TabNetPretrainer(\n            optimizer_fn=torch.optim.Adam,\n            optimizer_params=dict(lr=2e-2),\n            mask_type='entmax' # \"sparsemax\"\n        )\n        \n        unsupervised_model.fit(\n            X_train=X_train_np,\n            pretraining_ratio=0.8, # use 80% data to predict the rest\n            max_epochs=60\n        )\n\n        params = {\n                'n_d': 64,              # Width of the decision prediction layer\n                'n_a': 64,              # Width of the attention embedding for each step\n                'n_steps': 5,           # Number of steps in the architecture\n                'gamma': 1.5,           # Coefficient for feature selection regularization\n                'n_independent': 2,     # Number of independent GLU layer in each GLU block\n                'n_shared': 2,          # Number of shared GLU layer in each GLU block\n                'lambda_sparse': 1e-4,  # Sparsity regularization\n                'optimizer_fn': torch.optim.Adam,\n                'optimizer_params': dict(lr=2e-2, weight_decay=1e-5),\n                'mask_type': 'entmax',\n                'scheduler_params': dict(mode=\"min\", patience=10, min_lr=1e-5, factor=0.5),\n                'scheduler_fn': torch.optim.lr_scheduler.ReduceLROnPlateau,\n                'verbose': 1,\n                'device_name': 'cuda' if torch.cuda.is_available() else 'cpu'\n        }\n        model = TabNetRegressor(**params)\n        \n        model.fit(\n            X_train_np, \n            y_train_np.reshape(-1, 1),\n            eval_set=[(X_val_np, y_val_np.reshape(-1, 1))],\n            eval_metric=['mse'],\n            max_epochs=200,\n            patience=10,  \n            batch_size=256,\n            from_unsupervised=unsupervised_model\n        )\n```"
  }
}