{
  "id": 486259,
  "title": "Handing common memory errors in submission",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/486259",
  "author_name": "",
  "post_date": "2024-03-24T06:56:23.776708700Z",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello all,</p>\n<p>I think many of us used to face/ are facing memory errors in submission. This is common with such a large dataset and multiple models in prediction phase too. I may suggest the below to perhaps thwart these issues and submit effectively.</p>\n<ol>\n<li>I suggest one to peruse successful public kernels (successfully submitted code) and benchmark with one's own code. This will help the individual and team to remove any bugs in their code and create a successful submission process</li>\n<li>One may choose to edit a successful kernel with his/ her model structure and submit, another way to thwart the submission memory issues</li>\n<li>Delete objects no longer in use, perhaps regularly as well. This is an important step to conserve memory</li>\n<li>While using a large number of models and a large dataset, try and use Polars lazy frame instead of pandas. I think this is a better data handling mechanism and seldom creates memory issues </li>\n<li>Predict in batches, I may put up a pseudo code for this. I have executed a similar process in my inference kernels also-<br>\na. Consider n rows at a time. 'n' could correspond to a small number, say 5000/ 10000. Your test set ideally could be a polars lazy frame, so you <strong>collect only n rows and selected columns to conserve memory</strong></li>\n</ol>\n<ul>\n<li>Use the models to predict these n rows</li>\n<li>Store the predictions in the working directory</li>\n<li>Purge off interim tables </li>\n<li>Proceed to take up the next batch</li>\n<li>Repeat this till all batches are covered </li>\n<li>Load the files from the working directory and append them to create a submission file. </li>\n</ul>\n<p>I hope these tips may help one and all thwart memory issues. I suggest one could submit models in batches if one works with multiple scripts/ models to fathom reasons for potential errors and bugs in specific areas of the code.</p>\n<p>Wishing you the best for the challenge and otherwise too! All the best and happy learning!</p>",
  "messages": [
    {
      "id": "2713394",
      "postDate": "03/24/2024 06:56:23",
      "content": "<p>Hello all,</p>\n<p>I think many of us used to face/ are facing memory errors in submission. This is common with such a large dataset and multiple models in prediction phase too. I may suggest the below to perhaps thwart these issues and submit effectively.</p>\n<ol>\n<li>I suggest one to peruse successful public kernels (successfully submitted code) and benchmark with one's own code. This will help the individual and team to remove any bugs in their code and create a successful submission process</li>\n<li>One may choose to edit a successful kernel with his/ her model structure and submit, another way to thwart the submission memory issues</li>\n<li>Delete objects no longer in use, perhaps regularly as well. This is an important step to conserve memory</li>\n<li>While using a large number of models and a large dataset, try and use Polars lazy frame instead of pandas. I think this is a better data handling mechanism and seldom creates memory issues </li>\n<li>Predict in batches, I may put up a pseudo code for this. I have executed a similar process in my inference kernels also-<br>\na. Consider n rows at a time. 'n' could correspond to a small number, say 5000/ 10000. Your test set ideally could be a polars lazy frame, so you <strong>collect only n rows and selected columns to conserve memory</strong></li>\n</ol>\n<ul>\n<li>Use the models to predict these n rows</li>\n<li>Store the predictions in the working directory</li>\n<li>Purge off interim tables </li>\n<li>Proceed to take up the next batch</li>\n<li>Repeat this till all batches are covered </li>\n<li>Load the files from the working directory and append them to create a submission file. </li>\n</ul>\n<p>I hope these tips may help one and all thwart memory issues. I suggest one could submit models in batches if one works with multiple scripts/ models to fathom reasons for potential errors and bugs in specific areas of the code.</p>\n<p>Wishing you the best for the challenge and otherwise too! All the best and happy learning!</p>",
      "rawMarkdown": "Hello all,\n\nI think many of us used to face/ are facing memory errors in submission. This is common with such a large dataset and multiple models in prediction phase too. I may suggest the below to perhaps thwart these issues and submit effectively.\n\n1. I suggest one to peruse successful public kernels (successfully submitted code) and benchmark with one's own code. This will help the individual and team to remove any bugs in their code and create a successful submission process\n2. One may choose to edit a successful kernel with his/ her model structure and submit, another way to thwart the submission memory issues\n3. Delete objects no longer in use, perhaps regularly as well. This is an important step to conserve memory\n4. While using a large number of models and a large dataset, try and use Polars lazy frame instead of pandas. I think this is a better data handling mechanism and seldom creates memory issues \n5. Predict in batches, I may put up a pseudo code for this. I have executed a similar process in my inference kernels also-\na. Consider n rows at a time. 'n' could correspond to a small number, say 5000/ 10000. Your test set ideally could be a polars lazy frame, so you **collect only n rows and selected columns to conserve memory**\n- Use the models to predict these n rows\n- Store the predictions in the working directory\n- Purge off interim tables \n- Proceed to take up the next batch\n- Repeat this till all batches are covered \n- Load the files from the working directory and append them to create a submission file. \n\nI hope these tips may help one and all thwart memory issues. I suggest one could submit models in batches if one works with multiple scripts/ models to fathom reasons for potential errors and bugs in specific areas of the code.\n\nWishing you the best for the challenge and otherwise too! All the best and happy learning!",
      "votes": null
    },
    {
      "id": "2713662",
      "postDate": "03/24/2024 10:40:13",
      "content": "<p>Hi,</p>\n<p>Good post Ravi, a few more suggestions:</p>\n<p>If someone is using the cpu for inference, but the data doesn't fit into ram, you can enable the NVIDIA T4 and use cuDF or CuPy to move a part of the data to each gpu (2x15GB) and copy it back to ram when you need it, the I/O is going to be way faster than using the hdd for offloading. Edit: (Not long ago you would get 2 logical cores instead of 4 when an accelerator was enable, so if you are reading this in the future, better check for yourself how many logical cores you have with and without accelerator).</p>\n<p>You can also avoid the overhead of saving the predicted batches to disk and loading it back, if you use uint32 for case_id and fp32 for the score, that is a under 11MB of memory for the whole leaderboard dataset.</p>\n<p>When preprocessing, avoid doing operations to a lot of columns at same time.</p>\n<p>Be careful with python list, it can get out of control really fast when you add tons of objects, the references adds up quickly, try doing the same thing using NumPy.</p>",
      "rawMarkdown": "Hi,\n\nGood post Ravi, a few more suggestions:\n\nIf someone is using the cpu for inference, but the data doesn't fit into ram, you can enable the NVIDIA T4 and use cuDF or CuPy to move a part of the data to each gpu (2x15GB) and copy it back to ram when you need it, the I/O is going to be way faster than using the hdd for offloading. Edit: (Not long ago you would get 2 logical cores instead of 4 when an accelerator was enable, so if you are reading this in the future, better check for yourself how many logical cores you have with and without accelerator).\n\nYou can also avoid the overhead of saving the predicted batches to disk and loading it back, if you use uint32 for case_id and fp32 for the score, that is a under 11MB of memory for the whole leaderboard dataset.\n\nWhen preprocessing, avoid doing operations to a lot of columns at same time.\n\nBe careful with python list, it can get out of control really fast when you add tons of objects, the references adds up quickly, try doing the same thing using NumPy.",
      "votes": null
    },
    {
      "id": "2713666",
      "postDate": "03/24/2024 10:42:28",
      "content": "<p>Very correct, using cuDF is a good way to use pandas syntax and GPU. Also one may use shrink_dtypes() in polars to optimize the datatype of the selected columns <a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a> </p>",
      "rawMarkdown": "Very correct, using cuDF is a good way to use pandas syntax and GPU. Also one may use shrink_dtypes() in polars to optimize the datatype of the selected columns @enriquezaf",
      "votes": null
    },
    {
      "id": "2715577",
      "postDate": "03/25/2024 15:13:12",
      "content": "<p>Thank you Ravi and Antonio. <br>\nRavi, I have been following you to get me through this first Kaggle competition of mine and have learned a lot. I have a few questions for either of you. <br>\nWhile we predict in batches, is it also possible to train the model in batches? Will I need to use init_model to do so? Also, is it a good idea in the first place to do so?</p>",
      "rawMarkdown": "Thank you Ravi and Antonio. \nRavi, I have been following you to get me through this first Kaggle competition of mine and have learned a lot. I have a few questions for either of you. \nWhile we predict in batches, is it also possible to train the model in batches? Will I need to use init_model to do so? Also, is it a good idea in the first place to do so?",
      "votes": null
    },
    {
      "id": "2716066",
      "postDate": "03/25/2024 19:55:17",
      "content": "<p>This is a tricky question <a href=\"https://www.kaggle.com/varuniraothumsi\" target=\"_blank\">@varuniraothumsi</a> <br>\nTechnically it is possible, but will it be effective?? <br>\nFolding the data for cv is also training in a batch in a way…</p>",
      "rawMarkdown": "This is a tricky question @varuniraothumsi \nTechnically it is possible, but will it be effective?? \nFolding the data for cv is also training in a batch in a way...",
      "votes": null
    },
    {
      "id": "2716218",
      "postDate": "03/25/2024 22:23:49",
      "content": "<p>Use a pipeline for all the processing steps and then train the model. Save these 2 objects (pipeline and model) and load them into a clean notebook and use it just for submission. In conclusion you will have a notebook that reads the test set, process it and predict - all this with all the ram available. If you handled the train set with available ram, this method will assure that the full test set will be handled too, as they should be about the same sizes.</p>",
      "rawMarkdown": "Use a pipeline for all the processing steps and then train the model. Save these 2 objects (pipeline and model) and load them into a clean notebook and use it just for submission. In conclusion you will have a notebook that reads the test set, process it and predict - all this with all the ram available. If you handled the train set with available ram, this method will assure that the full test set will be handled too, as they should be about the same sizes.",
      "votes": null
    },
    {
      "id": "2716279",
      "postDate": "03/25/2024 23:38:38",
      "content": "<p>Good ideas. I haven't really thought about saving data to working directory. Another idea is to use reduce_mem_usage function for pandas Data frames memory optimization. It was shared in the discussion somewhere.</p>",
      "rawMarkdown": "Good ideas. I haven't really thought about saving data to working directory. Another idea is to use reduce_mem_usage function for pandas Data frames memory optimization. It was shared in the discussion somewhere.",
      "votes": null
    },
    {
      "id": "2716556",
      "postDate": "03/26/2024 05:09:15",
      "content": "<p>This is the best way to submit for code submissions <br>\nI usually never train models in my inference kernel <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> </p>",
      "rawMarkdown": "This is the best way to submit for code submissions \nI usually never train models in my inference kernel @eu1234",
      "votes": null
    },
    {
      "id": "2719095",
      "postDate": "03/27/2024 13:32:11",
      "content": "<p>Appreciate your response. I have been wondering the same, if it will be effective. But, unless I try I would not know. Knowing your extensive experience in this field, have you tried it? </p>\n<p>I am quite new to the field and hence my knowledge is limited. But, when we fold the data for cv, isn't a clone of the model trained, instead of the original model? So, from my understanding, the only way to retain the cv models is to save the cv model and later reload it. Kindly clarify my understanding if it is incorrect.</p>",
      "rawMarkdown": "Appreciate your response. I have been wondering the same, if it will be effective. But, unless I try I would not know. Knowing your extensive experience in this field, have you tried it? \n\nI am quite new to the field and hence my knowledge is limited. But, when we fold the data for cv, isn't a clone of the model trained, instead of the original model? So, from my understanding, the only way to retain the cv models is to save the cv model and later reload it. Kindly clarify my understanding if it is incorrect.",
      "votes": null
    },
    {
      "id": "2719199",
      "postDate": "03/27/2024 15:03:24",
      "content": "<p>You can always save the model using joblib, you can import it in a dataset and use it in an inference kernel <a href=\"https://www.kaggle.com/varuniraothumsi\" target=\"_blank\">@varuniraothumsi</a> </p>",
      "rawMarkdown": "You can always save the model using joblib, you can import it in a dataset and use it in an inference kernel @varuniraothumsi",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2713662,
      "author_name": "enriquezaf",
      "author_url": "",
      "post_date": "03/24/2024 10:40:13",
      "content": "<p>Hi,</p>\n<p>Good post Ravi, a few more suggestions:</p>\n<p>If someone is using the cpu for inference, but the data doesn't fit into ram, you can enable the NVIDIA T4 and use cuDF or CuPy to move a part of the data to each gpu (2x15GB) and copy it back to ram when you need it, the I/O is going to be way faster than using the hdd for offloading. Edit: (Not long ago you would get 2 logical cores instead of 4 when an accelerator was enable, so if you are reading this in the future, better check for yourself how many logical cores you have with and without accelerator).</p>\n<p>You can also avoid the overhead of saving the predicted batches to disk and loading it back, if you use uint32 for case_id and fp32 for the score, that is a under 11MB of memory for the whole leaderboard dataset.</p>\n<p>When preprocessing, avoid doing operations to a lot of columns at same time.</p>\n<p>Be careful with python list, it can get out of control really fast when you add tons of objects, the references adds up quickly, try doing the same thing using NumPy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2713666,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "03/24/2024 10:42:28",
          "content": "<p>Very correct, using cuDF is a good way to use pandas syntax and GPU. Also one may use shrink_dtypes() in polars to optimize the datatype of the selected columns <a href=\"https://www.kaggle.com/enriquezaf\" target=\"_blank\">@enriquezaf</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2715577,
      "author_name": "varuniraothumsi",
      "author_url": "",
      "post_date": "03/25/2024 15:13:12",
      "content": "<p>Thank you Ravi and Antonio. <br>\nRavi, I have been following you to get me through this first Kaggle competition of mine and have learned a lot. I have a few questions for either of you. <br>\nWhile we predict in batches, is it also possible to train the model in batches? Will I need to use init_model to do so? Also, is it a good idea in the first place to do so?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2716066,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "03/25/2024 19:55:17",
          "content": "<p>This is a tricky question <a href=\"https://www.kaggle.com/varuniraothumsi\" target=\"_blank\">@varuniraothumsi</a> <br>\nTechnically it is possible, but will it be effective?? <br>\nFolding the data for cv is also training in a batch in a way…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2719095,
              "author_name": "varuniraothumsi",
              "author_url": "",
              "post_date": "03/27/2024 13:32:11",
              "content": "<p>Appreciate your response. I have been wondering the same, if it will be effective. But, unless I try I would not know. Knowing your extensive experience in this field, have you tried it? </p>\n<p>I am quite new to the field and hence my knowledge is limited. But, when we fold the data for cv, isn't a clone of the model trained, instead of the original model? So, from my understanding, the only way to retain the cv models is to save the cv model and later reload it. Kindly clarify my understanding if it is incorrect.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2719199,
                  "author_name": "ravi20076",
                  "author_url": "",
                  "post_date": "03/27/2024 15:03:24",
                  "content": "<p>You can always save the model using joblib, you can import it in a dataset and use it in an inference kernel <a href=\"https://www.kaggle.com/varuniraothumsi\" target=\"_blank\">@varuniraothumsi</a> </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2716218,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "03/25/2024 22:23:49",
      "content": "<p>Use a pipeline for all the processing steps and then train the model. Save these 2 objects (pipeline and model) and load them into a clean notebook and use it just for submission. In conclusion you will have a notebook that reads the test set, process it and predict - all this with all the ram available. If you handled the train set with available ram, this method will assure that the full test set will be handled too, as they should be about the same sizes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2716556,
          "author_name": "ravi20076",
          "author_url": "",
          "post_date": "03/26/2024 05:09:15",
          "content": "<p>This is the best way to submit for code submissions <br>\nI usually never train models in my inference kernel <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2716279,
      "author_name": "matousfamera",
      "author_url": "",
      "post_date": "03/25/2024 23:38:38",
      "content": "<p>Good ideas. I haven't really thought about saving data to working directory. Another idea is to use reduce_mem_usage function for pandas Data frames memory optimization. It was shared in the discussion somewhere.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2713394": "Hello all,\n\nI think many of us used to face/ are facing memory errors in submission. This is common with such a large dataset and multiple models in prediction phase too. I may suggest the below to perhaps thwart these issues and submit effectively.\n\n1. I suggest one to peruse successful public kernels (successfully submitted code) and benchmark with one's own code. This will help the individual and team to remove any bugs in their code and create a successful submission process\n2. One may choose to edit a successful kernel with his/ her model structure and submit, another way to thwart the submission memory issues\n3. Delete objects no longer in use, perhaps regularly as well. This is an important step to conserve memory\n4. While using a large number of models and a large dataset, try and use Polars lazy frame instead of pandas. I think this is a better data handling mechanism and seldom creates memory issues \n5. Predict in batches, I may put up a pseudo code for this. I have executed a similar process in my inference kernels also-\na. Consider n rows at a time. 'n' could correspond to a small number, say 5000/ 10000. Your test set ideally could be a polars lazy frame, so you **collect only n rows and selected columns to conserve memory**\n- Use the models to predict these n rows\n- Store the predictions in the working directory\n- Purge off interim tables \n- Proceed to take up the next batch\n- Repeat this till all batches are covered \n- Load the files from the working directory and append them to create a submission file. \n\nI hope these tips may help one and all thwart memory issues. I suggest one could submit models in batches if one works with multiple scripts/ models to fathom reasons for potential errors and bugs in specific areas of the code.\n\nWishing you the best for the challenge and otherwise too! All the best and happy learning!",
    "2713662": "Hi,\n\nGood post Ravi, a few more suggestions:\n\nIf someone is using the cpu for inference, but the data doesn't fit into ram, you can enable the NVIDIA T4 and use cuDF or CuPy to move a part of the data to each gpu (2x15GB) and copy it back to ram when you need it, the I/O is going to be way faster than using the hdd for offloading. Edit: (Not long ago you would get 2 logical cores instead of 4 when an accelerator was enable, so if you are reading this in the future, better check for yourself how many logical cores you have with and without accelerator).\n\nYou can also avoid the overhead of saving the predicted batches to disk and loading it back, if you use uint32 for case_id and fp32 for the score, that is a under 11MB of memory for the whole leaderboard dataset.\n\nWhen preprocessing, avoid doing operations to a lot of columns at same time.\n\nBe careful with python list, it can get out of control really fast when you add tons of objects, the references adds up quickly, try doing the same thing using NumPy.",
    "2713666": "Very correct, using cuDF is a good way to use pandas syntax and GPU. Also one may use shrink_dtypes() in polars to optimize the datatype of the selected columns @enriquezaf",
    "2715577": "Thank you Ravi and Antonio. \nRavi, I have been following you to get me through this first Kaggle competition of mine and have learned a lot. I have a few questions for either of you. \nWhile we predict in batches, is it also possible to train the model in batches? Will I need to use init_model to do so? Also, is it a good idea in the first place to do so?",
    "2716066": "This is a tricky question @varuniraothumsi \nTechnically it is possible, but will it be effective?? \nFolding the data for cv is also training in a batch in a way...",
    "2716218": "Use a pipeline for all the processing steps and then train the model. Save these 2 objects (pipeline and model) and load them into a clean notebook and use it just for submission. In conclusion you will have a notebook that reads the test set, process it and predict - all this with all the ram available. If you handled the train set with available ram, this method will assure that the full test set will be handled too, as they should be about the same sizes.",
    "2716279": "Good ideas. I haven't really thought about saving data to working directory. Another idea is to use reduce_mem_usage function for pandas Data frames memory optimization. It was shared in the discussion somewhere.",
    "2716556": "This is the best way to submit for code submissions \nI usually never train models in my inference kernel @eu1234",
    "2719095": "Appreciate your response. I have been wondering the same, if it will be effective. But, unless I try I would not know. Knowing your extensive experience in this field, have you tried it? \n\nI am quite new to the field and hence my knowledge is limited. But, when we fold the data for cv, isn't a clone of the model trained, instead of the original model? So, from my understanding, the only way to retain the cv models is to save the cv model and later reload it. Kindly clarify my understanding if it is incorrect.",
    "2719199": "You can always save the model using joblib, you can import it in a dataset and use it in an inference kernel @varuniraothumsi"
  },
  "source": "meta"
}