{
  "id": 339195,
  "title": "What is your best XGBoost score?",
  "url": "/competitions/amex-default-prediction/discussion/339195",
  "author_name": "The Devastator",
  "post_date": "2022-07-23T18:15:37.640000",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi All,</p>\n<p>I am trying to optimize XGB &amp; Catboost hyperparameters and got stuck.<br>\n(The optimization had been running for +3 days at the moment)</p>\n<p>Some questions to those of you that use XGB:</p>\n<ul>\n<li>What is your LB score?</li>\n<li>Are you also willing to share your hyperparameters? </li>\n</ul>\n<p>Thanks! </p>",
  "messages": [
    {
      "id": 1868127,
      "postDate": "2022-07-23T18:15:37.640Z",
      "content": "<p>Hi All,</p>\n<p>I am trying to optimize XGB &amp; Catboost hyperparameters and got stuck.<br>\n(The optimization had been running for +3 days at the moment)</p>\n<p>Some questions to those of you that use XGB:</p>\n<ul>\n<li>What is your LB score?</li>\n<li>Are you also willing to share your hyperparameters? </li>\n</ul>\n<p>Thanks! </p>",
      "rawMarkdown": "Hi All,\n\nI am trying to optimize XGB & Catboost hyperparameters and got stuck.\n(The optimization had been running for +3 days at the moment)\n\nSome questions to those of you that use XGB:\n- What is your LB score?\n- Are you also willing to share your hyperparameters? \n\nThanks! \n",
      "votes": 11
    },
    {
      "id": 1868412,
      "postDate": "2022-07-24T00:42:32.170Z",
      "content": "<p>If you are optimizing 14 parameters simultaneously as in <a href=\"https://www.kaggle.com/code/thedevastator/the-fine-art-of-hyperparameter-tuning\" target=\"_blank\"><strong>here</strong></a>, I am not surprised it has taken a while. My suggestion is at the very least not to optimize the learning rate or the regularization parameters. I go with a fixed learning rate (usually 0.1, sometimes 0.05), and know that I will get a small boost after finding the parameters and running it with smaller <code>eta</code>.</p>",
      "rawMarkdown": "If you are optimizing 14 parameters simultaneously as in [**here**](https://www.kaggle.com/code/thedevastator/the-fine-art-of-hyperparameter-tuning), I am not surprised it has taken a while. My suggestion is at the very least not to optimize the learning rate or the regularization parameters. I go with a fixed learning rate (usually 0.1, sometimes 0.05), and know that I will get a small boost after finding the parameters and running it with smaller `eta`.",
      "votes": 5,
      "replies": [
        {
          "id": 1869227,
          "postDate": "2022-07-24T15:31:42.613Z",
          "content": "<p>Thank you, I will try this!</p>",
          "rawMarkdown": "Thank you, I will try this!"
        }
      ]
    },
    {
      "id": 1869276,
      "postDate": "2022-07-24T16:16:11.090Z",
      "content": "<p>Best CV model: CV - 0.79774 / LB - 0.796 (I freaked out when saw the LB result..)<br>\nBest LB model: CV - 0.79743 / LB - 0.797</p>\n<p>5 folds, seed 42.</p>",
      "rawMarkdown": "Best CV model: CV - 0.79774 / LB - 0.796 (I freaked out when saw the LB result..)\nBest LB model: CV - 0.79743 / LB - 0.797\n\n5 folds, seed 42.",
      "votes": 3
    },
    {
      "id": 1868320,
      "postDate": "2022-07-23T22:56:54.963Z",
      "content": "<p>Not sure how helpful this is, but my best XGB CVs at .7971 average (that's with early stopping so optimistic, but in line with the way I think most people report CV here). Haven't submitted it as a single model. </p>\n<p>Lower learning rate and column sampling per tree + higher early stopping rounds all seem to help. I'm very sold on the former, but somewhat skeptical of the latter --  are tiny improvements over a large number of rounds just noise? The way people are using DART seems to run the same risk to me, though things seem to stabilize out in aggregate.</p>",
      "rawMarkdown": "Not sure how helpful this is, but my best XGB CVs at .7971 average (that's with early stopping so optimistic, but in line with the way I think most people report CV here). Haven't submitted it as a single model. \n\nLower learning rate and column sampling per tree + higher early stopping rounds all seem to help. I'm very sold on the former, but somewhat skeptical of the latter --  are tiny improvements over a large number of rounds just noise? The way people are using DART seems to run the same risk to me, though things seem to stabilize out in aggregate.",
      "votes": 3
    },
    {
      "id": 1868324,
      "postDate": "2022-07-23T23:04:23.607Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1868412,
      "author_name": "Tilii",
      "author_url": "",
      "post_date": "2022-07-24T00:42:32.170000",
      "content": "<p>If you are optimizing 14 parameters simultaneously as in <a href=\"https://www.kaggle.com/code/thedevastator/the-fine-art-of-hyperparameter-tuning\" target=\"_blank\"><strong>here</strong></a>, I am not surprised it has taken a while. My suggestion is at the very least not to optimize the learning rate or the regularization parameters. I go with a fixed learning rate (usually 0.1, sometimes 0.05), and know that I will get a small boost after finding the parameters and running it with smaller <code>eta</code>.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1869227,
          "author_name": "The Devastator",
          "author_url": "",
          "post_date": "2022-07-24T15:31:42.613000",
          "content": "<p>Thank you, I will try this!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1869276,
      "author_name": "Dmitry Uarov",
      "author_url": "",
      "post_date": "2022-07-24T16:16:11.090000",
      "content": "<p>Best CV model: CV - 0.79774 / LB - 0.796 (I freaked out when saw the LB result..)<br>\nBest LB model: CV - 0.79743 / LB - 0.797</p>\n<p>5 folds, seed 42.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1868320,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2022-07-23T22:56:54.963000",
      "content": "<p>Not sure how helpful this is, but my best XGB CVs at .7971 average (that's with early stopping so optimistic, but in line with the way I think most people report CV here). Haven't submitted it as a single model. </p>\n<p>Lower learning rate and column sampling per tree + higher early stopping rounds all seem to help. I'm very sold on the former, but somewhat skeptical of the latter --  are tiny improvements over a large number of rounds just noise? The way people are using DART seems to run the same risk to me, though things seem to stabilize out in aggregate.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1868324,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-23T23:04:23.607000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1868127": "Hi All,\n\nI am trying to optimize XGB & Catboost hyperparameters and got stuck.\n(The optimization had been running for +3 days at the moment)\n\nSome questions to those of you that use XGB:\n- What is your LB score?\n- Are you also willing to share your hyperparameters? \n\nThanks! \n",
    "1868412": "If you are optimizing 14 parameters simultaneously as in [**here**](https://www.kaggle.com/code/thedevastator/the-fine-art-of-hyperparameter-tuning), I am not surprised it has taken a while. My suggestion is at the very least not to optimize the learning rate or the regularization parameters. I go with a fixed learning rate (usually 0.1, sometimes 0.05), and know that I will get a small boost after finding the parameters and running it with smaller `eta`.",
    "1869276": "Best CV model: CV - 0.79774 / LB - 0.796 (I freaked out when saw the LB result..)\nBest LB model: CV - 0.79743 / LB - 0.797\n\n5 folds, seed 42.",
    "1868320": "Not sure how helpful this is, but my best XGB CVs at .7971 average (that's with early stopping so optimistic, but in line with the way I think most people report CV here). Haven't submitted it as a single model. \n\nLower learning rate and column sampling per tree + higher early stopping rounds all seem to help. I'm very sold on the former, but somewhat skeptical of the latter --  are tiny improvements over a large number of rounds just noise? The way people are using DART seems to run the same risk to me, though things seem to stabilize out in aggregate.",
    "1868324": ""
  }
}