{
  "id": 498363,
  "title": "How to Train a Model",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/498363",
  "author_name": "",
  "post_date": "2024-04-28T00:53:08.698742800Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone, I'm trying to train a model locally, but due to memory constraints, I have to split the data into chunks to train the LGBM model. However, what surprises me is that the results obtained from training with chunks and training with the data unified in the LGBM model are very different. The model trained with chunks performs very poorly. Does anyone know why?<br>\nHere's the initial code:<br>\nax = cv.split(df_train, y, groups=weeks)<br>\nfor idx_train, idx_valid in ax:<br>\n    X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  <br>\n X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]<br>\n X_train[cat_cols] = X_train[cat_cols].astype(\"category\")<br>\n X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\n model.fit(<br>\n         X_train, y_train,<br>\n         eval_set=[(X_valid, y_valid)],<br>\n         callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])</p>\n<p>fitted_models_lgb.append(model)<br>\nand the code for chunked training:<br>\nax = cv.split(df_train, y, groups=weeks)<br>\nN = int(100000)<br>\nNt = int(N / 5)<br>\nfor idx_train_all, idx_valid_all in ax:<br>\n L = len(idx_train_all)<br>\n Nu = int((L / N) + 1)</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\n for i in range(Nu):<br>\n     idx_train = idx_train_all[i * N:(i + 1) * N]<br>\n     idx_valid = idx_valid_all[i * Nt:(i + 1) * Nt]<br>\n      X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  #<br>\n     X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]<br>\n      X_train[cat_cols] = X_train[cat_cols].astype(\"category\")<br>\n     X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")<br>\n      model.fit(<br>\n             X_train, y_train,<br>\n             eval_set=[(X_valid, y_valid)],<br>\n             callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])<br>\n  fitted_models_lgb.append(model)<br>\nI made sure that the code that stores the model is placed outside of the chunked training loop, and that the other parameters such as params_gpu 、cv…. are exactly the same</p>",
  "messages": [
    {
      "id": "2779982",
      "postDate": "04/28/2024 00:53:08",
      "content": "<p>Hello everyone, I'm trying to train a model locally, but due to memory constraints, I have to split the data into chunks to train the LGBM model. However, what surprises me is that the results obtained from training with chunks and training with the data unified in the LGBM model are very different. The model trained with chunks performs very poorly. Does anyone know why?<br>\nHere's the initial code:<br>\nax = cv.split(df_train, y, groups=weeks)<br>\nfor idx_train, idx_valid in ax:<br>\n    X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  <br>\n X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]<br>\n X_train[cat_cols] = X_train[cat_cols].astype(\"category\")<br>\n X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\n model.fit(<br>\n         X_train, y_train,<br>\n         eval_set=[(X_valid, y_valid)],<br>\n         callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])</p>\n<p>fitted_models_lgb.append(model)<br>\nand the code for chunked training:<br>\nax = cv.split(df_train, y, groups=weeks)<br>\nN = int(100000)<br>\nNt = int(N / 5)<br>\nfor idx_train_all, idx_valid_all in ax:<br>\n L = len(idx_train_all)<br>\n Nu = int((L / N) + 1)</p>\n<p>model = lgb.LGBMClassifier(**params_gpu)<br>\n for i in range(Nu):<br>\n     idx_train = idx_train_all[i * N:(i + 1) * N]<br>\n     idx_valid = idx_valid_all[i * Nt:(i + 1) * Nt]<br>\n      X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  #<br>\n     X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]<br>\n      X_train[cat_cols] = X_train[cat_cols].astype(\"category\")<br>\n     X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")<br>\n      model.fit(<br>\n             X_train, y_train,<br>\n             eval_set=[(X_valid, y_valid)],<br>\n             callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])<br>\n  fitted_models_lgb.append(model)<br>\nI made sure that the code that stores the model is placed outside of the chunked training loop, and that the other parameters such as params_gpu 、cv…. are exactly the same</p>",
      "rawMarkdown": "Hello everyone, I'm trying to train a model locally, but due to memory constraints, I have to split the data into chunks to train the LGBM model. However, what surprises me is that the results obtained from training with chunks and training with the data unified in the LGBM model are very different. The model trained with chunks performs very poorly. Does anyone know why?\n\nHere's the initial code:\n \nax = cv.split(df_train, y, groups=weeks)\nfor idx_train, idx_valid in ax:\n\n\n\n        X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  \n        X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]\n        X_train[cat_cols] = X_train[cat_cols].astype(\"category\")\n        X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")\n        \n        model = lgb.LGBMClassifier(**params_gpu)\n        model.fit(\n                X_train, y_train,\n                eval_set=[(X_valid, y_valid)],\n                callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\n        \n        fitted_models_lgb.append(model)\n\n\n\nand the code for chunked training:\n\nax = cv.split(df_train, y, groups=weeks)\nN = int(100000)\nNt = int(N / 5)\nfor idx_train_all, idx_valid_all in ax:\n        L = len(idx_train_all)\n        Nu = int((L / N) + 1)\n\n\n        \n        model = lgb.LGBMClassifier(**params_gpu)\n        for i in range(Nu):\n            idx_train = idx_train_all[i * N:(i + 1) * N]\n            idx_valid = idx_valid_all[i * Nt:(i + 1) * Nt]\n\n            X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  #\n            X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]\n\n            X_train[cat_cols] = X_train[cat_cols].astype(\"category\")\n            X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")\n\n            model.fit(\n                    X_train, y_train,\n                    eval_set=[(X_valid, y_valid)],\n                    callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\n         fitted_models_lgb.append(model)\n\n\nI made sure that the code that stores the model is placed outside of the chunked training loop, and that the other parameters such as params_gpu 、cv.... are exactly the same",
      "votes": null
    },
    {
      "id": "2779988",
      "postDate": "04/28/2024 00:58:06",
      "content": "<p>Here's the code:</p>",
      "rawMarkdown": "Here's the code:",
      "votes": null
    },
    {
      "id": "2779992",
      "postDate": "04/28/2024 01:03:05",
      "content": "<p>and the chunk</p>",
      "rawMarkdown": "and the chunk",
      "votes": null
    },
    {
      "id": "2780476",
      "postDate": "04/28/2024 08:49:33",
      "content": "<p>I don't have solution to this but I will say that I have experienced exactly the same, training in chunks has huge -ve impact on the performance. <br>\nInstead of investigating, I ended up using public notebooks which are usually able to feed all the data to model in one go without hitting memory issues, we just have be efficient with memory utilization while preprocessing/aggregating. </p>",
      "rawMarkdown": "I don't have solution to this but I will say that I have experienced exactly the same, training in chunks has huge -ve impact on the performance. \nInstead of investigating, I ended up using public notebooks which are usually able to feed all the data to model in one go without hitting memory issues, we just have be efficient with memory utilization while preprocessing/aggregating.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2779988,
      "author_name": "shuchangyu",
      "author_url": "",
      "post_date": "04/28/2024 00:58:06",
      "content": "<p>Here's the code:</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2779992,
      "author_name": "shuchangyu",
      "author_url": "",
      "post_date": "04/28/2024 01:03:05",
      "content": "<p>and the chunk</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2780476,
      "author_name": "jabranzahid",
      "author_url": "",
      "post_date": "04/28/2024 08:49:33",
      "content": "<p>I don't have solution to this but I will say that I have experienced exactly the same, training in chunks has huge -ve impact on the performance. <br>\nInstead of investigating, I ended up using public notebooks which are usually able to feed all the data to model in one go without hitting memory issues, we just have be efficient with memory utilization while preprocessing/aggregating. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2779982": "Hello everyone, I'm trying to train a model locally, but due to memory constraints, I have to split the data into chunks to train the LGBM model. However, what surprises me is that the results obtained from training with chunks and training with the data unified in the LGBM model are very different. The model trained with chunks performs very poorly. Does anyone know why?\n\nHere's the initial code:\n \nax = cv.split(df_train, y, groups=weeks)\nfor idx_train, idx_valid in ax:\n\n\n\n        X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  \n        X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]\n        X_train[cat_cols] = X_train[cat_cols].astype(\"category\")\n        X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")\n        \n        model = lgb.LGBMClassifier(**params_gpu)\n        model.fit(\n                X_train, y_train,\n                eval_set=[(X_valid, y_valid)],\n                callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\n        \n        fitted_models_lgb.append(model)\n\n\n\nand the code for chunked training:\n\nax = cv.split(df_train, y, groups=weeks)\nN = int(100000)\nNt = int(N / 5)\nfor idx_train_all, idx_valid_all in ax:\n        L = len(idx_train_all)\n        Nu = int((L / N) + 1)\n\n\n        \n        model = lgb.LGBMClassifier(**params_gpu)\n        for i in range(Nu):\n            idx_train = idx_train_all[i * N:(i + 1) * N]\n            idx_valid = idx_valid_all[i * Nt:(i + 1) * Nt]\n\n            X_train, y_train = df_train.iloc[idx_train], y.iloc[idx_train]  #\n            X_valid, y_valid = df_train.iloc[idx_valid], y.iloc[idx_valid]\n\n            X_train[cat_cols] = X_train[cat_cols].astype(\"category\")\n            X_valid[cat_cols] = X_valid[cat_cols].astype(\"category\")\n\n            model.fit(\n                    X_train, y_train,\n                    eval_set=[(X_valid, y_valid)],\n                    callbacks=[lgb.log_evaluation(200), lgb.early_stopping(100)])\n         fitted_models_lgb.append(model)\n\n\nI made sure that the code that stores the model is placed outside of the chunked training loop, and that the other parameters such as params_gpu 、cv.... are exactly the same",
    "2779988": "Here's the code:",
    "2779992": "and the chunk",
    "2780476": "I don't have solution to this but I will say that I have experienced exactly the same, training in chunks has huge -ve impact on the performance. \nInstead of investigating, I ended up using public notebooks which are usually able to feed all the data to model in one go without hitting memory issues, we just have be efficient with memory utilization while preprocessing/aggregating."
  },
  "source": "meta"
}