{
  "id": 498418,
  "title": " Training Sklearn/XGB models without OOM",
  "url": "/competitions/leash-BELKA/discussion/498418",
  "author_name": "",
  "post_date": "2024-04-28T08:58:07.275968400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hi,<br>\nI want to train Sklearn/XGB models on the whole dataset, but after converting the molecules into ECFP embeddings, the table is to large to fit in memory (OOM).<br>\nDo you have any good ressources that show how to train a model incrementally using chunks of the dataframe ?<br>\nI have read some references to <code>partial_fit</code> (sklearn), but I am certain some of you have some prior experience to share.<br>\nThanks 😉</p>",
  "messages": [
    {
      "id": "2780482",
      "postDate": "04/28/2024 08:58:07",
      "content": "<p>Hi,<br>\nI want to train Sklearn/XGB models on the whole dataset, but after converting the molecules into ECFP embeddings, the table is to large to fit in memory (OOM).<br>\nDo you have any good ressources that show how to train a model incrementally using chunks of the dataframe ?<br>\nI have read some references to <code>partial_fit</code> (sklearn), but I am certain some of you have some prior experience to share.<br>\nThanks 😉</p>",
      "rawMarkdown": "Hi,\nI want to train Sklearn/XGB models on the whole dataset, but after converting the molecules into ECFP embeddings, the table is to large to fit in memory (OOM).\nDo you have any good ressources that show how to train a model incrementally using chunks of the dataframe ?\nI have read some references to `partial_fit` (sklearn), but I am certain some of you have some prior experience to share.\nThanks 😉",
      "votes": null
    },
    {
      "id": "2781073",
      "postDate": "04/28/2024 14:44:29",
      "content": "<p>You may want to try np.packbits to help with size. Check <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion posts on their experiments. They have some great tips and tricks in there that help a lot.</p>",
      "rawMarkdown": "You may want to try np.packbits to help with size. Check @hengck23 discussion posts on their experiments. They have some great tips and tricks in there that help a lot.",
      "votes": null
    },
    {
      "id": "2781236",
      "postDate": "04/28/2024 16:58:49",
      "content": "<p>XGB models can work with sparce matrix data. Convert ECFP  to sparse matrix and test how much rows it can manage.<br>\nfrom scipy import sparse<br>\nsparse.csr_matrix(ecfp_data)</p>",
      "rawMarkdown": "XGB models can work with sparce matrix data. Convert ECFP  to sparse matrix and test how much rows it can manage.\nfrom scipy import sparse\nsparse.csr_matrix(ecfp_data)",
      "votes": null
    },
    {
      "id": "2781344",
      "postDate": "04/28/2024 18:18:43",
      "content": "<p>On Kaggle, (and probably on local too), it's also helpful to gather and format the data (in chunks) first, then save to file, then end program. Then start new program that just loads the data and starts training. Start with like 1GB or 5GB data size in the format you're testing (sparse matrix or packbits), and see how much it explodes when doing the prework and first tree.</p>\n<p>Not sure for XGBoost, but for LGBM you can also load the dataset into a dataset object, save that and discard the raw data entirely. In the case of fingerprint data, not sure that would help any, but for larger numeric features, that should go from orig size to max_bin-based data size (1 byte if max_bin=255 as is default for CPU training)</p>",
      "rawMarkdown": "On Kaggle, (and probably on local too), it's also helpful to gather and format the data (in chunks) first, then save to file, then end program. Then start new program that just loads the data and starts training. Start with like 1GB or 5GB data size in the format you're testing (sparse matrix or packbits), and see how much it explodes when doing the prework and first tree.\n\nNot sure for XGBoost, but for LGBM you can also load the dataset into a dataset object, save that and discard the raw data entirely. In the case of fingerprint data, not sure that would help any, but for larger numeric features, that should go from orig size to max_bin-based data size (1 byte if max_bin=255 as is default for CPU training)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2781073,
      "author_name": "chemdatafarmer",
      "author_url": "",
      "post_date": "04/28/2024 14:44:29",
      "content": "<p>You may want to try np.packbits to help with size. Check <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> discussion posts on their experiments. They have some great tips and tricks in there that help a lot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2781236,
      "author_name": "ricopue",
      "author_url": "",
      "post_date": "04/28/2024 16:58:49",
      "content": "<p>XGB models can work with sparce matrix data. Convert ECFP  to sparse matrix and test how much rows it can manage.<br>\nfrom scipy import sparse<br>\nsparse.csr_matrix(ecfp_data)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2781344,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "04/28/2024 18:18:43",
      "content": "<p>On Kaggle, (and probably on local too), it's also helpful to gather and format the data (in chunks) first, then save to file, then end program. Then start new program that just loads the data and starts training. Start with like 1GB or 5GB data size in the format you're testing (sparse matrix or packbits), and see how much it explodes when doing the prework and first tree.</p>\n<p>Not sure for XGBoost, but for LGBM you can also load the dataset into a dataset object, save that and discard the raw data entirely. In the case of fingerprint data, not sure that would help any, but for larger numeric features, that should go from orig size to max_bin-based data size (1 byte if max_bin=255 as is default for CPU training)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2780482": "Hi,\nI want to train Sklearn/XGB models on the whole dataset, but after converting the molecules into ECFP embeddings, the table is to large to fit in memory (OOM).\nDo you have any good ressources that show how to train a model incrementally using chunks of the dataframe ?\nI have read some references to `partial_fit` (sklearn), but I am certain some of you have some prior experience to share.\nThanks 😉",
    "2781073": "You may want to try np.packbits to help with size. Check @hengck23 discussion posts on their experiments. They have some great tips and tricks in there that help a lot.",
    "2781236": "XGB models can work with sparce matrix data. Convert ECFP  to sparse matrix and test how much rows it can manage.\nfrom scipy import sparse\nsparse.csr_matrix(ecfp_data)",
    "2781344": "On Kaggle, (and probably on local too), it's also helpful to gather and format the data (in chunks) first, then save to file, then end program. Then start new program that just loads the data and starts training. Start with like 1GB or 5GB data size in the format you're testing (sparse matrix or packbits), and see how much it explodes when doing the prework and first tree.\n\nNot sure for XGBoost, but for LGBM you can also load the dataset into a dataset object, save that and discard the raw data entirely. In the case of fingerprint data, not sure that would help any, but for larger numeric features, that should go from orig size to max_bin-based data size (1 byte if max_bin=255 as is default for CPU training)"
  },
  "source": "meta"
}