{
  "id": 552195,
  "title": "LGB training does not seem to converge",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/552195",
  "author_name": "",
  "post_date": "2024-12-18T06:44:26.765549900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I did a time series train val split on &gt; 40milllion data point with LGB. I increased the n_estimators from 100 to over 1000, even though the validation score is improving (rmse), the training still does not seem to converge, and the PB actually got worse.</p>\n<p>Wonder is this situation met by others?</p>",
  "messages": [
    {
      "id": "3074878",
      "postDate": "12/18/2024 06:44:26",
      "content": "<p>I did a time series train val split on &gt; 40milllion data point with LGB. I increased the n_estimators from 100 to over 1000, even though the validation score is improving (rmse), the training still does not seem to converge, and the PB actually got worse.</p>\n<p>Wonder is this situation met by others?</p>",
      "rawMarkdown": "I did a time series train val split on > 40milllion data point with LGB. I increased the n_estimators from 100 to over 1000, even though the validation score is improving (rmse), the training still does not seem to converge, and the PB actually got worse.\n\nWonder is this situation met by others?",
      "votes": null
    },
    {
      "id": "3074894",
      "postDate": "12/18/2024 07:18:22",
      "content": "<p>The same as me.More n_estimators will turn to lower LB score becaues of overfitting.For me, the best n_estimators are between 200~300.And I use all dates (0~1698) to train gbdt model.</p>",
      "rawMarkdown": "The same as me.More n_estimators will turn to lower LB score becaues of overfitting.For me, the best n_estimators are between 200~300.And I use all dates (0~1698) to train gbdt model.",
      "votes": null
    },
    {
      "id": "3074974",
      "postDate": "12/18/2024 09:33:40",
      "content": "<p>I wonder is it actually overfitting, if it is, the validation score should reduce, but mine is actually improving. (I have early stopping for that as well)</p>",
      "rawMarkdown": "I wonder is it actually overfitting, if it is, the validation score should reduce, but mine is actually improving. (I have early stopping for that as well)",
      "votes": null
    },
    {
      "id": "3075019",
      "postDate": "12/18/2024 11:00:30",
      "content": "<p>One of difficulty for this competion is to find a good cv strategy which is a certain correlation with LB. Maybe a model perform well in cv but not in LB because the data is a non-stationary time series.</p>",
      "rawMarkdown": "One of difficulty for this competion is to find a good cv strategy which is a certain correlation with LB. Maybe a model perform well in cv but not in LB because the data is a non-stationary time series.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3074894,
      "author_name": "i2nfinit3y",
      "author_url": "",
      "post_date": "12/18/2024 07:18:22",
      "content": "<p>The same as me.More n_estimators will turn to lower LB score becaues of overfitting.For me, the best n_estimators are between 200~300.And I use all dates (0~1698) to train gbdt model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3074974,
          "author_name": "zhangyue199",
          "author_url": "",
          "post_date": "12/18/2024 09:33:40",
          "content": "<p>I wonder is it actually overfitting, if it is, the validation score should reduce, but mine is actually improving. (I have early stopping for that as well)</p>",
          "votes": null,
          "replies": [
            {
              "id": 3075019,
              "author_name": "i2nfinit3y",
              "author_url": "",
              "post_date": "12/18/2024 11:00:30",
              "content": "<p>One of difficulty for this competion is to find a good cv strategy which is a certain correlation with LB. Maybe a model perform well in cv but not in LB because the data is a non-stationary time series.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3074878": "I did a time series train val split on > 40milllion data point with LGB. I increased the n_estimators from 100 to over 1000, even though the validation score is improving (rmse), the training still does not seem to converge, and the PB actually got worse.\n\nWonder is this situation met by others?",
    "3074894": "The same as me.More n_estimators will turn to lower LB score becaues of overfitting.For me, the best n_estimators are between 200~300.And I use all dates (0~1698) to train gbdt model.",
    "3074974": "I wonder is it actually overfitting, if it is, the validation score should reduce, but mine is actually improving. (I have early stopping for that as well)",
    "3075019": "One of difficulty for this competion is to find a good cv strategy which is a certain correlation with LB. Maybe a model perform well in cv but not in LB because the data is a non-stationary time series."
  },
  "source": "meta"
}