{
  "id": 580564,
  "title": "Too early stop on lightgbm",
  "url": "/competitions/drw-crypto-market-prediction/discussion/580564",
  "author_name": "",
  "post_date": "2025-05-25T03:27:19.869113300Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>When i use lightgbm model and early stopping, it always stop after a few runs.<br>\nFor example</p>\n<pre><code> lightgbm  lgb\n\nX = df[X_cols]\ny = df[]\n\nsplit_index = ((df) * )   \nX_train, X_val = X.iloc[:split_index], X.iloc[split_index:]\ny_train, y_val = y.iloc[:split_index], y.iloc[split_index:]\n\ntrain_data = lgb.Dataset(X_train, label=y_train)\nval_data = lgb.Dataset(X_val, label=y_val)\n\nparams = {\n        : ,\n        : ,\n        : -,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : \n    }\n\nmodel = lgb.train(\n        params,\n        train_data,\n        valid_sets=[train_data, val_data],\n        valid_names=[, ],\n        num_boost_round=,\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=),\n            lgb.log_evaluation(period=)\n        ]\n    )\n</code></pre>\n<pre><code>Training  validation scores dons rmse: 1.34649\n[100]    valid_0s rmse: 1.19002\n</code></pre>",
  "messages": [
    {
      "id": "3208971",
      "postDate": "05/25/2025 03:27:19",
      "content": "<p>When i use lightgbm model and early stopping, it always stop after a few runs.<br>\nFor example</p>\n<pre><code> lightgbm  lgb\n\nX = df[X_cols]\ny = df[]\n\nsplit_index = ((df) * )   \nX_train, X_val = X.iloc[:split_index], X.iloc[split_index:]\ny_train, y_val = y.iloc[:split_index], y.iloc[split_index:]\n\ntrain_data = lgb.Dataset(X_train, label=y_train)\nval_data = lgb.Dataset(X_val, label=y_val)\n\nparams = {\n        : ,\n        : ,\n        : -,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : ,\n        : \n    }\n\nmodel = lgb.train(\n        params,\n        train_data,\n        valid_sets=[train_data, val_data],\n        valid_names=[, ],\n        num_boost_round=,\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=),\n            lgb.log_evaluation(period=)\n        ]\n    )\n</code></pre>\n<pre><code>Training  validation scores dons rmse: 1.34649\n[100]    valid_0s rmse: 1.19002\n</code></pre>",
      "rawMarkdown": "When i use lightgbm model and early stopping, it always stop after a few runs.\nFor example\n```python\nimport lightgbm as lgb\n\nX = df[X_cols]\ny = df['label']\n    \nsplit_index = int(len(df) * 0.8)   # here i tried 0.5 but still same problem\nX_train, X_val = X.iloc[:split_index], X.iloc[split_index:]\ny_train, y_val = y.iloc[:split_index], y.iloc[split_index:]\n    \ntrain_data = lgb.Dataset(X_train, label=y_train)\nval_data = lgb.Dataset(X_val, label=y_val)\n\nparams = {\n        \"objective\": \"regression\",\n        \"metric\": \"rmse\",\n        \"verbosity\": -1,\n        \"boosting_type\": \"gbdt\",\n        'num_leaves': 128,\n        'max_depth': 16,\n        \"learning_rate\": 0.01,\n        \"num_leaves\": 64,\n        \"feature_fraction\": 0.8,\n        \"bagging_fraction\": 0.8,\n        \"bagging_freq\": 5,\n        \"seed\": 42\n    }\n\nmodel = lgb.train(\n        params,\n        train_data,\n        valid_sets=[train_data, val_data],\n        valid_names=[\"train\", \"valid\"],\n        num_boost_round=1000,\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=100),\n            lgb.log_evaluation(period=50)\n        ]\n    )\n\n```\n```bash\nTraining until validation scores don't improve for 100 rounds\n[50]\tvalid_0's rmse: 1.34649\n[100]\tvalid_0's rmse: 1.35127\nEarly stopping, best iteration is:\n[1]\tvalid_0's rmse: 1.19002\n```",
      "votes": null
    },
    {
      "id": "3209013",
      "postDate": "05/25/2025 05:18:42",
      "content": "<p>Please look into your model parameters and the eval-metric <a href=\"https://www.kaggle.com/ivantang86\" target=\"_blank\">@ivantang86</a> <br>\nI think these need to be modified.</p>",
      "rawMarkdown": "Please look into your model parameters and the eval-metric @ivantang86 \nI think these need to be modified.",
      "votes": null
    },
    {
      "id": "3211364",
      "postDate": "05/28/2025 10:47:36",
      "content": "<p>Currently facing the same problem and seeking a solution too</p>",
      "rawMarkdown": "Currently facing the same problem and seeking a solution too",
      "votes": null
    },
    {
      "id": "3227668",
      "postDate": "06/19/2025 06:58:57",
      "content": "<p>could be an underfitting problem. try to set the \"min_split_gain\" param a small number like 1e-10</p>",
      "rawMarkdown": "could be an underfitting problem. try to set the \"min_split_gain\" param a small number like 1e-10",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3209013,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "05/25/2025 05:18:42",
      "content": "<p>Please look into your model parameters and the eval-metric <a href=\"https://www.kaggle.com/ivantang86\" target=\"_blank\">@ivantang86</a> <br>\nI think these need to be modified.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3211364,
      "author_name": "olowedaniel",
      "author_url": "",
      "post_date": "05/28/2025 10:47:36",
      "content": "<p>Currently facing the same problem and seeking a solution too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3227668,
      "author_name": "melodyd",
      "author_url": "",
      "post_date": "06/19/2025 06:58:57",
      "content": "<p>could be an underfitting problem. try to set the \"min_split_gain\" param a small number like 1e-10</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3208971": "When i use lightgbm model and early stopping, it always stop after a few runs.\nFor example\n```python\nimport lightgbm as lgb\n\nX = df[X_cols]\ny = df['label']\n    \nsplit_index = int(len(df) * 0.8)   # here i tried 0.5 but still same problem\nX_train, X_val = X.iloc[:split_index], X.iloc[split_index:]\ny_train, y_val = y.iloc[:split_index], y.iloc[split_index:]\n    \ntrain_data = lgb.Dataset(X_train, label=y_train)\nval_data = lgb.Dataset(X_val, label=y_val)\n\nparams = {\n        \"objective\": \"regression\",\n        \"metric\": \"rmse\",\n        \"verbosity\": -1,\n        \"boosting_type\": \"gbdt\",\n        'num_leaves': 128,\n        'max_depth': 16,\n        \"learning_rate\": 0.01,\n        \"num_leaves\": 64,\n        \"feature_fraction\": 0.8,\n        \"bagging_fraction\": 0.8,\n        \"bagging_freq\": 5,\n        \"seed\": 42\n    }\n\nmodel = lgb.train(\n        params,\n        train_data,\n        valid_sets=[train_data, val_data],\n        valid_names=[\"train\", \"valid\"],\n        num_boost_round=1000,\n        callbacks=[\n            lgb.early_stopping(stopping_rounds=100),\n            lgb.log_evaluation(period=50)\n        ]\n    )\n\n```\n```bash\nTraining until validation scores don't improve for 100 rounds\n[50]\tvalid_0's rmse: 1.34649\n[100]\tvalid_0's rmse: 1.35127\nEarly stopping, best iteration is:\n[1]\tvalid_0's rmse: 1.19002\n```",
    "3209013": "Please look into your model parameters and the eval-metric @ivantang86 \nI think these need to be modified.",
    "3211364": "Currently facing the same problem and seeking a solution too",
    "3227668": "could be an underfitting problem. try to set the \"min_split_gain\" param a small number like 1e-10"
  },
  "source": "meta"
}