{
  "id": 331887,
  "title": "Memory efficient ways to save predictions of various models",
  "url": "/competitions/amex-default-prediction/discussion/331887",
  "author_name": "",
  "post_date": "2022-06-19T09:42:22.530229500Z",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n<p>As we all know this competition has large dataset. My initial train file shape is (458913, 918). <br>\nSo far what I do is I concatenate the oof prediction and test prediction of each model to train and test file respectively. <br>\nNow the problem is if I have trained 100 models my train file size becomes (458913, 1018) and similarly test file also increases. <br>\nThis has made my code quite slow. I am thinking to change it,  so that each prediction is saved as separate file instead of being concatenated to train and test. </p>\n<p>I am new here, so I would like to ask experts here how you save your model predictions on disc. <br>\n(what file format is memory efficient to save 1D array?)<br>\nAlso best practices on managing these submissions. </p>\n<p>Thanks,</p>",
  "messages": [
    {
      "id": "1825392",
      "postDate": "06/19/2022 09:42:22",
      "content": "<p>Hi,</p>\n<p>As we all know this competition has large dataset. My initial train file shape is (458913, 918). <br>\nSo far what I do is I concatenate the oof prediction and test prediction of each model to train and test file respectively. <br>\nNow the problem is if I have trained 100 models my train file size becomes (458913, 1018) and similarly test file also increases. <br>\nThis has made my code quite slow. I am thinking to change it,  so that each prediction is saved as separate file instead of being concatenated to train and test. </p>\n<p>I am new here, so I would like to ask experts here how you save your model predictions on disc. <br>\n(what file format is memory efficient to save 1D array?)<br>\nAlso best practices on managing these submissions. </p>\n<p>Thanks,</p>",
      "rawMarkdown": "Hi,\n\nAs we all know this competition has large dataset. My initial train file shape is (458913, 918). \nSo far what I do is I concatenate the oof prediction and test prediction of each model to train and test file respectively. \nNow the problem is if I have trained 100 models my train file size becomes (458913, 1018) and similarly test file also increases. \nThis has made my code quite slow. I am thinking to change it,  so that each prediction is saved as separate file instead of being concatenated to train and test. \n\nI am new here, so I would like to ask experts here how you save your model predictions on disc. \n(what file format is memory efficient to save 1D array?)\nAlso best practices on managing these submissions. \n\nThanks,",
      "votes": null
    },
    {
      "id": "1825752",
      "postDate": "06/19/2022 17:11:45",
      "content": "<p>I'd definitely recommend saving oof and test predictions in standalone files. It avoids the unnecessary clutter/speed bottleneck you mentioned and also makes it easier to track data provenance. If you want to be rigorous, you can have all model training runs generate output in one place: prediction files alongside logs that document features used, hyperparameters, and evaluation results. When you need to combine different predictions later for ensembling etc. it should be easy to merge them back together by including a <code>user_id</code> column or guaranteeing that order is preserved. </p>\n<p>There are a bunch of good options for efficient storage, both in terms of read/write time and disk space cost. In python, <code>Pickle</code> is a great serialized storage option -- it lets you store python objects like <code>pandas</code> dataframes or <code>numpy</code> arrays in a format that's close to how they're represented in RAM, allowing a very fast read/write process (much faster than writing or reading csvs!). <code>Feather</code> is another great option that's also more portable across languages than pickle. </p>",
      "rawMarkdown": "I'd definitely recommend saving oof and test predictions in standalone files. It avoids the unnecessary clutter/speed bottleneck you mentioned and also makes it easier to track data provenance. If you want to be rigorous, you can have all model training runs generate output in one place: prediction files alongside logs that document features used, hyperparameters, and evaluation results. When you need to combine different predictions later for ensembling etc. it should be easy to merge them back together by including a `user_id` column or guaranteeing that order is preserved. \n\nThere are a bunch of good options for efficient storage, both in terms of read/write time and disk space cost. In python, `Pickle` is a great serialized storage option -- it lets you store python objects like `pandas` dataframes or `numpy` arrays in a format that's close to how they're represented in RAM, allowing a very fast read/write process (much faster than writing or reading csvs!). `Feather` is another great option that's also more portable across languages than pickle.",
      "votes": null
    },
    {
      "id": "1825779",
      "postDate": "06/19/2022 17:43:40",
      "content": "<p>I recommend using pickle files for the same. They are very useful and can be stored easily. I recommend storing different predictions in separate files and then combining them upon completion</p>",
      "rawMarkdown": "I recommend using pickle files for the same. They are very useful and can be stored easily. I recommend storing different predictions in separate files and then combining them upon completion",
      "votes": null
    },
    {
      "id": "1825806",
      "postDate": "06/19/2022 18:12:00",
      "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , thanks for your suggestion.</p>",
      "rawMarkdown": "ravi20076 , thanks for your suggestion.",
      "votes": null
    },
    {
      "id": "1825811",
      "postDate": "06/19/2022 18:21:18",
      "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> , thanks for your suggestion, there are file formats like pkl, feather, parquet. I think each has their own benefit in term of how we want to access it. Like as mentioned by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">this</a> post that csv is good for when we want to read row by row, parquet stores data column wise. Another option is uncompressed NumPy file (.npy) [which I have not tried yet].  That being said these options are good to store DataFrame or 2D array, but I don't know which will be good to store 1D array (1D array since we will be storing our each model prediction separately).  </p>",
      "rawMarkdown": "aquatic , thanks for your suggestion, there are file formats like pkl, feather, parquet. I think each has their own benefit in term of how we want to access it. Like as mentioned by @cdeotte in [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) post that csv is good for when we want to read row by row, parquet stores data column wise. Another option is uncompressed NumPy file (.npy) [which I have not tried yet].  That being said these options are good to store DataFrame or 2D array, but I don't know which will be good to store 1D array (1D array since we will be storing our each model prediction separately).",
      "votes": null
    },
    {
      "id": "1825849",
      "postDate": "06/19/2022 19:23:05",
      "content": "<p>All of these options should also work fine for storing a 1D array. And you may also find it convenient to keep using 2D arrays so that you can have a <code>row_id</code> or <code>user_id</code> column to help ensure consistent ordering.</p>",
      "rawMarkdown": "All of these options should also work fine for storing a 1D array. And you may also find it convenient to keep using 2D arrays so that you can have a `row_id` or `user_id` column to help ensure consistent ordering.",
      "votes": null
    },
    {
      "id": "1826122",
      "postDate": "06/20/2022 04:19:03",
      "content": "<p>thanks <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> , will try it.</p>",
      "rawMarkdown": "thanks @aquatic , will try it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1825752,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "06/19/2022 17:11:45",
      "content": "<p>I'd definitely recommend saving oof and test predictions in standalone files. It avoids the unnecessary clutter/speed bottleneck you mentioned and also makes it easier to track data provenance. If you want to be rigorous, you can have all model training runs generate output in one place: prediction files alongside logs that document features used, hyperparameters, and evaluation results. When you need to combine different predictions later for ensembling etc. it should be easy to merge them back together by including a <code>user_id</code> column or guaranteeing that order is preserved. </p>\n<p>There are a bunch of good options for efficient storage, both in terms of read/write time and disk space cost. In python, <code>Pickle</code> is a great serialized storage option -- it lets you store python objects like <code>pandas</code> dataframes or <code>numpy</code> arrays in a format that's close to how they're represented in RAM, allowing a very fast read/write process (much faster than writing or reading csvs!). <code>Feather</code> is another great option that's also more portable across languages than pickle. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1825811,
          "author_name": "raj401",
          "author_url": "",
          "post_date": "06/19/2022 18:21:18",
          "content": "<p><a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> , thanks for your suggestion, there are file formats like pkl, feather, parquet. I think each has their own benefit in term of how we want to access it. Like as mentioned by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054\" target=\"_blank\">this</a> post that csv is good for when we want to read row by row, parquet stores data column wise. Another option is uncompressed NumPy file (.npy) [which I have not tried yet].  That being said these options are good to store DataFrame or 2D array, but I don't know which will be good to store 1D array (1D array since we will be storing our each model prediction separately).  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1825849,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "06/19/2022 19:23:05",
          "content": "<p>All of these options should also work fine for storing a 1D array. And you may also find it convenient to keep using 2D arrays so that you can have a <code>row_id</code> or <code>user_id</code> column to help ensure consistent ordering.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1826122,
          "author_name": "raj401",
          "author_url": "",
          "post_date": "06/20/2022 04:19:03",
          "content": "<p>thanks <a href=\"https://www.kaggle.com/aquatic\" target=\"_blank\">@aquatic</a> , will try it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1825779,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "06/19/2022 17:43:40",
      "content": "<p>I recommend using pickle files for the same. They are very useful and can be stored easily. I recommend storing different predictions in separate files and then combining them upon completion</p>",
      "votes": null,
      "replies": [
        {
          "id": 1825806,
          "author_name": "raj401",
          "author_url": "",
          "post_date": "06/19/2022 18:12:00",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> , thanks for your suggestion.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1825392": "Hi,\n\nAs we all know this competition has large dataset. My initial train file shape is (458913, 918). \nSo far what I do is I concatenate the oof prediction and test prediction of each model to train and test file respectively. \nNow the problem is if I have trained 100 models my train file size becomes (458913, 1018) and similarly test file also increases. \nThis has made my code quite slow. I am thinking to change it,  so that each prediction is saved as separate file instead of being concatenated to train and test. \n\nI am new here, so I would like to ask experts here how you save your model predictions on disc. \n(what file format is memory efficient to save 1D array?)\nAlso best practices on managing these submissions. \n\nThanks,",
    "1825752": "I'd definitely recommend saving oof and test predictions in standalone files. It avoids the unnecessary clutter/speed bottleneck you mentioned and also makes it easier to track data provenance. If you want to be rigorous, you can have all model training runs generate output in one place: prediction files alongside logs that document features used, hyperparameters, and evaluation results. When you need to combine different predictions later for ensembling etc. it should be easy to merge them back together by including a `user_id` column or guaranteeing that order is preserved. \n\nThere are a bunch of good options for efficient storage, both in terms of read/write time and disk space cost. In python, `Pickle` is a great serialized storage option -- it lets you store python objects like `pandas` dataframes or `numpy` arrays in a format that's close to how they're represented in RAM, allowing a very fast read/write process (much faster than writing or reading csvs!). `Feather` is another great option that's also more portable across languages than pickle.",
    "1825779": "I recommend using pickle files for the same. They are very useful and can be stored easily. I recommend storing different predictions in separate files and then combining them upon completion",
    "1825806": "ravi20076 , thanks for your suggestion.",
    "1825811": "aquatic , thanks for your suggestion, there are file formats like pkl, feather, parquet. I think each has their own benefit in term of how we want to access it. Like as mentioned by @cdeotte in [this](https://www.kaggle.com/competitions/amex-default-prediction/discussion/328054) post that csv is good for when we want to read row by row, parquet stores data column wise. Another option is uncompressed NumPy file (.npy) [which I have not tried yet].  That being said these options are good to store DataFrame or 2D array, but I don't know which will be good to store 1D array (1D array since we will be storing our each model prediction separately).",
    "1825849": "All of these options should also work fine for storing a 1D array. And you may also find it convenient to keep using 2D arrays so that you can have a `row_id` or `user_id` column to help ensure consistent ordering.",
    "1826122": "thanks @aquatic , will try it."
  },
  "source": "meta"
}