{
  "id": 55626,
  "title": "Does hyperparameter tuning in lightgbm really helpful in this competition?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55626",
  "author_name": "",
  "post_date": "2018-04-29T21:40:13.000244Z",
  "votes": 11,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi All,</p>\n\n<p>Till now, I only choose a set of parameters which I think 'making sense' to train my lightgbm model and I tried to do some hyperparameter tuning as well but the improvement is less than 0.0003. Does anyone try hyperparameter tuning seriously and get a significant in result, such as 0.001 improvement? Thank you.</p>",
  "messages": [
    {
      "id": "320793",
      "postDate": "04/29/2018 21:40:13",
      "content": "<p>Hi All,</p>\n\n<p>Till now, I only choose a set of parameters which I think 'making sense' to train my lightgbm model and I tried to do some hyperparameter tuning as well but the improvement is less than 0.0003. Does anyone try hyperparameter tuning seriously and get a significant in result, such as 0.001 improvement? Thank you.</p>",
      "rawMarkdown": "Hi All,\n\nTill now, I only choose a set of parameters which I think 'making sense' to train my lightgbm model and I tried to do some hyperparameter tuning as well but the improvement is less than 0.0003. Does anyone try hyperparameter tuning seriously and get a significant in result, such as 0.001 improvement? Thank you.",
      "votes": null
    },
    {
      "id": "320803",
      "postDate": "04/29/2018 22:34:28",
      "content": "<p>I tuned hyperparameters with cross validation on the first million rows (~1 minute/run), then used those parameters on the full 185 million rows (~1 day/run).</p>\n\n<p>In my model \"num_leaves\" and \"feature_fraction\" turned out to be important parameters.</p>",
      "rawMarkdown": "I tuned hyperparameters with cross validation on the first million rows (~1 minute/run), then used those parameters on the full 185 million rows (~1 day/run).\n\nIn my model \"num_leaves\" and \"feature_fraction\" turned out to be important parameters.",
      "votes": null
    },
    {
      "id": "320836",
      "postDate": "04/30/2018 01:33:17",
      "content": "<p>spent 3 day on tuning   .9804 to .9807.   I gave up on tuning already.</p>",
      "rawMarkdown": "spent 3 day on tuning   .9804 to .9807.   I gave up on tuning already.",
      "votes": null
    },
    {
      "id": "320851",
      "postDate": "04/30/2018 03:04:13",
      "content": "<p>As for me, 0.0003 improvement is not a slight improvement lol. </p>",
      "rawMarkdown": "As for me, 0.0003 improvement is not a slight improvement lol.",
      "votes": null
    },
    {
      "id": "321023",
      "postDate": "04/30/2018 12:27:24",
      "content": "<p>I achieved good improvements to my lgbm model by dropping the learning rate to a low value like 0.01 then setting the number of boost rounds very high (~10K) and using early stopping rounds around 50. This gave me an improvement of 0.0008 to the leaderboard AUC when compared to a learning rate of 0.2 with the same early stopping logic.</p>",
      "rawMarkdown": "I achieved good improvements to my lgbm model by dropping the learning rate to a low value like 0.01 then setting the number of boost rounds very high (~10K) and using early stopping rounds around 50. This gave me an improvement of 0.0008 to the leaderboard AUC when compared to a learning rate of 0.2 with the same early stopping logic.",
      "votes": null
    },
    {
      "id": "321037",
      "postDate": "04/30/2018 12:57:50",
      "content": "<p>Good point, but how long will learning rate = 0.001 run? several hours?</p>",
      "rawMarkdown": "Good point, but how long will learning rate = 0.001 run? several hours?",
      "votes": null
    },
    {
      "id": "321043",
      "postDate": "04/30/2018 13:05:15",
      "content": "<p>Yeah my last model took 2.4 hours to train, so it maybe worth waiting till the very end to lower the learning rate, once you are happy with all the other model parameters. </p>",
      "rawMarkdown": "Yeah my last model took 2.4 hours to train, so it maybe worth waiting till the very end to lower the learning rate, once you are happy with all the other model parameters.",
      "votes": null
    },
    {
      "id": "321313",
      "postDate": "05/01/2018 02:43:19",
      "content": "<p>have you read this <a href=\"https://www.kaggle.com/nanomathias/bayesian-optimization-of-xgboost-lb-0-9769/notebook\">notebook</a>? The author applied Bayesian search and reached a good score; however, it is computationally expensive.</p>",
      "rawMarkdown": "have you read this [notebook][1]? The author applied Bayesian search and reached a good score; however, it is computationally expensive.\n\n\n  [1]: https://www.kaggle.com/nanomathias/bayesian-optimization-of-xgboost-lb-0-9769/notebook",
      "votes": null
    },
    {
      "id": "321403",
      "postDate": "05/01/2018 07:10:38",
      "content": "<p>Thanks for your sharing! I try to drop the learning rate to 0.05, the local AUC improve ~0.0002, but the public AUC has no improvement.</p>",
      "rawMarkdown": "Thanks for your sharing! I try to drop the learning rate to 0.05, the local AUC improve ~0.0002, but the public AUC has no improvement.",
      "votes": null
    },
    {
      "id": "321597",
      "postDate": "05/01/2018 16:23:15",
      "content": "<p>based on my experience, bayesian optimization is no better than my manual tunning....  </p>",
      "rawMarkdown": "based on my experience, bayesian optimization is no better than my manual tunning....",
      "votes": null
    },
    {
      "id": "321611",
      "postDate": "05/01/2018 16:55:24",
      "content": "<p>Which is already a pretty good result, given that you dont have to do it manually :)</p>",
      "rawMarkdown": "Which is already a pretty good result, given that you dont have to do it manually :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 320803,
      "author_name": "danieljbrooks",
      "author_url": "",
      "post_date": "04/29/2018 22:34:28",
      "content": "<p>I tuned hyperparameters with cross validation on the first million rows (~1 minute/run), then used those parameters on the full 185 million rows (~1 day/run).</p>\n\n<p>In my model \"num_leaves\" and \"feature_fraction\" turned out to be important parameters.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320836,
      "author_name": "marcuslin",
      "author_url": "",
      "post_date": "04/30/2018 01:33:17",
      "content": "<p>spent 3 day on tuning   .9804 to .9807.   I gave up on tuning already.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320851,
      "author_name": "hhh920406",
      "author_url": "",
      "post_date": "04/30/2018 03:04:13",
      "content": "<p>As for me, 0.0003 improvement is not a slight improvement lol. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 321023,
      "author_name": "climbercarmich",
      "author_url": "",
      "post_date": "04/30/2018 12:27:24",
      "content": "<p>I achieved good improvements to my lgbm model by dropping the learning rate to a low value like 0.01 then setting the number of boost rounds very high (~10K) and using early stopping rounds around 50. This gave me an improvement of 0.0008 to the leaderboard AUC when compared to a learning rate of 0.2 with the same early stopping logic.</p>",
      "votes": null,
      "replies": [
        {
          "id": 321037,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/30/2018 12:57:50",
          "content": "<p>Good point, but how long will learning rate = 0.001 run? several hours?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321043,
          "author_name": "climbercarmich",
          "author_url": "",
          "post_date": "04/30/2018 13:05:15",
          "content": "<p>Yeah my last model took 2.4 hours to train, so it maybe worth waiting till the very end to lower the learning rate, once you are happy with all the other model parameters. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321403,
          "author_name": "hhh920406",
          "author_url": "",
          "post_date": "05/01/2018 07:10:38",
          "content": "<p>Thanks for your sharing! I try to drop the learning rate to 0.05, the local AUC improve ~0.0002, but the public AUC has no improvement.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 321313,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "05/01/2018 02:43:19",
      "content": "<p>have you read this <a href=\"https://www.kaggle.com/nanomathias/bayesian-optimization-of-xgboost-lb-0-9769/notebook\">notebook</a>? The author applied Bayesian search and reached a good score; however, it is computationally expensive.</p>",
      "votes": null,
      "replies": [
        {
          "id": 321597,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "05/01/2018 16:23:15",
          "content": "<p>based on my experience, bayesian optimization is no better than my manual tunning....  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321611,
          "author_name": "malten",
          "author_url": "",
          "post_date": "05/01/2018 16:55:24",
          "content": "<p>Which is already a pretty good result, given that you dont have to do it manually :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "320793": "Hi All,\n\nTill now, I only choose a set of parameters which I think 'making sense' to train my lightgbm model and I tried to do some hyperparameter tuning as well but the improvement is less than 0.0003. Does anyone try hyperparameter tuning seriously and get a significant in result, such as 0.001 improvement? Thank you.",
    "320803": "I tuned hyperparameters with cross validation on the first million rows (~1 minute/run), then used those parameters on the full 185 million rows (~1 day/run).\n\nIn my model \"num_leaves\" and \"feature_fraction\" turned out to be important parameters.",
    "320836": "spent 3 day on tuning   .9804 to .9807.   I gave up on tuning already.",
    "320851": "As for me, 0.0003 improvement is not a slight improvement lol.",
    "321023": "I achieved good improvements to my lgbm model by dropping the learning rate to a low value like 0.01 then setting the number of boost rounds very high (~10K) and using early stopping rounds around 50. This gave me an improvement of 0.0008 to the leaderboard AUC when compared to a learning rate of 0.2 with the same early stopping logic.",
    "321037": "Good point, but how long will learning rate = 0.001 run? several hours?",
    "321043": "Yeah my last model took 2.4 hours to train, so it maybe worth waiting till the very end to lower the learning rate, once you are happy with all the other model parameters.",
    "321313": "have you read this [notebook][1]? The author applied Bayesian search and reached a good score; however, it is computationally expensive.\n\n\n  [1]: https://www.kaggle.com/nanomathias/bayesian-optimization-of-xgboost-lb-0-9769/notebook",
    "321403": "Thanks for your sharing! I try to drop the learning rate to 0.05, the local AUC improve ~0.0002, but the public AUC has no improvement.",
    "321597": "based on my experience, bayesian optimization is no better than my manual tunning....",
    "321611": "Which is already a pretty good result, given that you dont have to do it manually :)"
  },
  "source": "meta"
}