{
  "id": 549186,
  "title": "Some feature tuning technics about TabNet",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/549186",
  "author_name": "MJeremy",
  "post_date": "2024-12-01T05:12:07.372000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In the original paper of TabNet, the authors shared the tunning rules, summarised below:</p>\n<p><strong>Number of Steps (Nsteps):</strong></p>\n<ul>\n<li>Optimal range: 3 to 10.</li>\n<li>Higher Nsteps suits datasets with more information-bearing features.</li>\n<li>Too high Nsteps can cause ill-conditioned matrices and overfitting, leading to poor generalization.</li>\n</ul>\n<p><strong>Dimensionality of Features (Nd) and Attention (Na):</strong></p>\n<ul>\n<li>Set Nd = Na for a balanced trade-off between performance and complexity.</li>\n<li>Too high values may lead to overfitting and poor generalization.</li>\n</ul>\n<p><strong>Sparsity Regularization (γ):</strong></p>\n<ul>\n<li>Higher Nsteps typically requires a larger γ.</li>\n<li>Optimal tuning of γ significantly affects performance.</li>\n</ul>\n<p><strong>Batch Size:</strong></p>\n<ul>\n<li>Use a large batch size if memory allows (1–10% of total dataset size).</li>\n<li>Virtual batch size should be much smaller for stability.</li>\n</ul>\n<p><strong>Learning Rate:</strong></p>\n<ul>\n<li>Start with a large initial learning rate.</li>\n<li>Gradually decay the learning rate during training for convergence.</li>\n</ul>\n<hr>\n<p>Beyond this, does TabNet actually help? I trained and validated TabNet alone, the result of QWK is actually quite sub-optimal.</p>",
  "messages": [
    {
      "id": 3059817,
      "postDate": "2024-12-01T05:12:07.373Z",
      "content": "<p>In the original paper of TabNet, the authors shared the tunning rules, summarised below:</p>\n<p><strong>Number of Steps (Nsteps):</strong></p>\n<ul>\n<li>Optimal range: 3 to 10.</li>\n<li>Higher Nsteps suits datasets with more information-bearing features.</li>\n<li>Too high Nsteps can cause ill-conditioned matrices and overfitting, leading to poor generalization.</li>\n</ul>\n<p><strong>Dimensionality of Features (Nd) and Attention (Na):</strong></p>\n<ul>\n<li>Set Nd = Na for a balanced trade-off between performance and complexity.</li>\n<li>Too high values may lead to overfitting and poor generalization.</li>\n</ul>\n<p><strong>Sparsity Regularization (γ):</strong></p>\n<ul>\n<li>Higher Nsteps typically requires a larger γ.</li>\n<li>Optimal tuning of γ significantly affects performance.</li>\n</ul>\n<p><strong>Batch Size:</strong></p>\n<ul>\n<li>Use a large batch size if memory allows (1–10% of total dataset size).</li>\n<li>Virtual batch size should be much smaller for stability.</li>\n</ul>\n<p><strong>Learning Rate:</strong></p>\n<ul>\n<li>Start with a large initial learning rate.</li>\n<li>Gradually decay the learning rate during training for convergence.</li>\n</ul>\n<hr>\n<p>Beyond this, does TabNet actually help? I trained and validated TabNet alone, the result of QWK is actually quite sub-optimal.</p>",
      "rawMarkdown": "In the original paper of TabNet, the authors shared the tunning rules, summarised below:\n\n__Number of Steps (Nsteps):__\n\n- Optimal range: 3 to 10.\n- Higher Nsteps suits datasets with more information-bearing features.\n- Too high Nsteps can cause ill-conditioned matrices and overfitting, leading to poor generalization.\n\n__Dimensionality of Features (Nd) and Attention (Na):__\n- Set Nd = Na for a balanced trade-off between performance and complexity.\n- Too high values may lead to overfitting and poor generalization.\n\n__Sparsity Regularization (γ):__\n- Higher Nsteps typically requires a larger γ.\n- Optimal tuning of γ significantly affects performance.\n\n__Batch Size:__\n- Use a large batch size if memory allows (1–10% of total dataset size).\n- Virtual batch size should be much smaller for stability.\n\n__Learning Rate:__\n- Start with a large initial learning rate.\n- Gradually decay the learning rate during training for convergence.\n\n\n---\n\nBeyond this, does TabNet actually help? I trained and validated TabNet alone, the result of QWK is actually quite sub-optimal.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3059817": "In the original paper of TabNet, the authors shared the tunning rules, summarised below:\n\n__Number of Steps (Nsteps):__\n\n- Optimal range: 3 to 10.\n- Higher Nsteps suits datasets with more information-bearing features.\n- Too high Nsteps can cause ill-conditioned matrices and overfitting, leading to poor generalization.\n\n__Dimensionality of Features (Nd) and Attention (Na):__\n- Set Nd = Na for a balanced trade-off between performance and complexity.\n- Too high values may lead to overfitting and poor generalization.\n\n__Sparsity Regularization (γ):__\n- Higher Nsteps typically requires a larger γ.\n- Optimal tuning of γ significantly affects performance.\n\n__Batch Size:__\n- Use a large batch size if memory allows (1–10% of total dataset size).\n- Virtual batch size should be much smaller for stability.\n\n__Learning Rate:__\n- Start with a large initial learning rate.\n- Gradually decay the learning rate during training for convergence.\n\n\n---\n\nBeyond this, does TabNet actually help? I trained and validated TabNet alone, the result of QWK is actually quite sub-optimal."
  }
}