{
  "id": 491529,
  "title": "Out of Memory error on scoring stage?",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/491529",
  "author_name": "",
  "post_date": "2024-04-06T07:30:49.162531500Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I understand that there are many topics related to memory issues. However, I still can't grasp what is going wrong. After submission notebook runs successfully. But as far as I understand this run is just saving notebook version, like \"Run all &amp; save\" option. Through the notebook I delete unused variables, call garbage collector and perform final prediction with batches. Top memory consuming variables looks like this:<br>\n699.2s    3575                        base_train: 51.2 MiB<br>\n699.2s    3576                              data: 44.0 MiB<br>\n699.2s    3577                                 y: 17.5 MiB<br>\n699.2s    3578                           y_train: 14.0 MiB<br>\n699.2s    3579                                 s: 14.0 MiB<br>\n699.2s    3580                          base_val: 10.2 MiB<br>\n699.2s    3581                             y_val:  2.8 MiB<br>\n699.2s    3582                              base:  2.6 MiB<br>\n699.2s    3583                         base_test:  2.6 MiB<br>\n699.2s    3584                            y_test: 715.6 KiB</p>\n<p>Whereas test data is loaded into variable 'test_data', so I infer that in the run test_data is still 10 rows. <br>\nAnd after notebook run, scoring fails immediately with OOM error, so under assumption that in scoring stage notebook runs again with full test data, then it would fail before even it reaches the data collection part. So the question is, what is wrong with notebook and is there a way to understand what is happening during scoring stage?</p>",
  "messages": [
    {
      "id": "2738215",
      "postDate": "04/06/2024 07:30:49",
      "content": "<p>I understand that there are many topics related to memory issues. However, I still can't grasp what is going wrong. After submission notebook runs successfully. But as far as I understand this run is just saving notebook version, like \"Run all &amp; save\" option. Through the notebook I delete unused variables, call garbage collector and perform final prediction with batches. Top memory consuming variables looks like this:<br>\n699.2s    3575                        base_train: 51.2 MiB<br>\n699.2s    3576                              data: 44.0 MiB<br>\n699.2s    3577                                 y: 17.5 MiB<br>\n699.2s    3578                           y_train: 14.0 MiB<br>\n699.2s    3579                                 s: 14.0 MiB<br>\n699.2s    3580                          base_val: 10.2 MiB<br>\n699.2s    3581                             y_val:  2.8 MiB<br>\n699.2s    3582                              base:  2.6 MiB<br>\n699.2s    3583                         base_test:  2.6 MiB<br>\n699.2s    3584                            y_test: 715.6 KiB</p>\n<p>Whereas test data is loaded into variable 'test_data', so I infer that in the run test_data is still 10 rows. <br>\nAnd after notebook run, scoring fails immediately with OOM error, so under assumption that in scoring stage notebook runs again with full test data, then it would fail before even it reaches the data collection part. So the question is, what is wrong with notebook and is there a way to understand what is happening during scoring stage?</p>",
      "rawMarkdown": "I understand that there are many topics related to memory issues. However, I still can't grasp what is going wrong. After submission notebook runs successfully. But as far as I understand this run is just saving notebook version, like \"Run all & save\" option. Through the notebook I delete unused variables, call garbage collector and perform final prediction with batches. Top memory consuming variables looks like this:\n699.2s\t3575\t                    base_train: 51.2 MiB\n699.2s\t3576\t                          data: 44.0 MiB\n699.2s\t3577\t                             y: 17.5 MiB\n699.2s\t3578\t                       y_train: 14.0 MiB\n699.2s\t3579\t                             s: 14.0 MiB\n699.2s\t3580\t                      base_val: 10.2 MiB\n699.2s\t3581\t                         y_val:  2.8 MiB\n699.2s\t3582\t                          base:  2.6 MiB\n699.2s\t3583\t                     base_test:  2.6 MiB\n699.2s\t3584\t                        y_test: 715.6 KiB\n\nWhereas test data is loaded into variable 'test_data', so I infer that in the run test_data is still 10 rows. \nAnd after notebook run, scoring fails immediately with OOM error, so under assumption that in scoring stage notebook runs again with full test data, then it would fail before even it reaches the data collection part. So the question is, what is wrong with notebook and is there a way to understand what is happening during scoring stage?",
      "votes": null
    },
    {
      "id": "2738417",
      "postDate": "04/06/2024 10:30:24",
      "content": "<p>It is difficult to point exactly when the notebook is failing. but in most of the cases it is fails due to test data size.  there are certain cases, for example, you may have train data file loaded and also the test data file loaded at the same time, than features extraction also loaded of those file at that same time. make cause this issue.</p>\n<p>But u can create the same situation artificially to see what and where it may be failing. in ur demo notebook, make test same as train. so in stead of 10 test data yo will be using the entire train data as a testing data. now this demo notebook will run on train data and test data (which is also the train data). Here u can check memory wise, if everything is going fine or not. this should give somewhat clear picture as I read somewhere that test data size is comparable to train data.</p>",
      "rawMarkdown": "It is difficult to point exactly when the notebook is failing. but in most of the cases it is fails due to test data size.  there are certain cases, for example, you may have train data file loaded and also the test data file loaded at the same time, than features extraction also loaded of those file at that same time. make cause this issue.\n\nBut u can create the same situation artificially to see what and where it may be failing. in ur demo notebook, make test same as train. so in stead of 10 test data yo will be using the entire train data as a testing data. now this demo notebook will run on train data and test data (which is also the train data). Here u can check memory wise, if everything is going fine or not. this should give somewhat clear picture as I read somewhere that test data size is comparable to train data.",
      "votes": null
    },
    {
      "id": "2739755",
      "postDate": "04/07/2024 08:41:33",
      "content": "<p>You can upload your trained model as a private dataset, make test datset and predict with the model in your notebook , and then submit. Hope this will help.</p>",
      "rawMarkdown": "You can upload your trained model as a private dataset, make test datset and predict with the model in your notebook , and then submit. Hope this will help.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2738417,
      "author_name": "shreyas9181",
      "author_url": "",
      "post_date": "04/06/2024 10:30:24",
      "content": "<p>It is difficult to point exactly when the notebook is failing. but in most of the cases it is fails due to test data size.  there are certain cases, for example, you may have train data file loaded and also the test data file loaded at the same time, than features extraction also loaded of those file at that same time. make cause this issue.</p>\n<p>But u can create the same situation artificially to see what and where it may be failing. in ur demo notebook, make test same as train. so in stead of 10 test data yo will be using the entire train data as a testing data. now this demo notebook will run on train data and test data (which is also the train data). Here u can check memory wise, if everything is going fine or not. this should give somewhat clear picture as I read somewhere that test data size is comparable to train data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2739755,
      "author_name": "sani84",
      "author_url": "",
      "post_date": "04/07/2024 08:41:33",
      "content": "<p>You can upload your trained model as a private dataset, make test datset and predict with the model in your notebook , and then submit. Hope this will help.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2738215": "I understand that there are many topics related to memory issues. However, I still can't grasp what is going wrong. After submission notebook runs successfully. But as far as I understand this run is just saving notebook version, like \"Run all & save\" option. Through the notebook I delete unused variables, call garbage collector and perform final prediction with batches. Top memory consuming variables looks like this:\n699.2s\t3575\t                    base_train: 51.2 MiB\n699.2s\t3576\t                          data: 44.0 MiB\n699.2s\t3577\t                             y: 17.5 MiB\n699.2s\t3578\t                       y_train: 14.0 MiB\n699.2s\t3579\t                             s: 14.0 MiB\n699.2s\t3580\t                      base_val: 10.2 MiB\n699.2s\t3581\t                         y_val:  2.8 MiB\n699.2s\t3582\t                          base:  2.6 MiB\n699.2s\t3583\t                     base_test:  2.6 MiB\n699.2s\t3584\t                        y_test: 715.6 KiB\n\nWhereas test data is loaded into variable 'test_data', so I infer that in the run test_data is still 10 rows. \nAnd after notebook run, scoring fails immediately with OOM error, so under assumption that in scoring stage notebook runs again with full test data, then it would fail before even it reaches the data collection part. So the question is, what is wrong with notebook and is there a way to understand what is happening during scoring stage?",
    "2738417": "It is difficult to point exactly when the notebook is failing. but in most of the cases it is fails due to test data size.  there are certain cases, for example, you may have train data file loaded and also the test data file loaded at the same time, than features extraction also loaded of those file at that same time. make cause this issue.\n\nBut u can create the same situation artificially to see what and where it may be failing. in ur demo notebook, make test same as train. so in stead of 10 test data yo will be using the entire train data as a testing data. now this demo notebook will run on train data and test data (which is also the train data). Here u can check memory wise, if everything is going fine or not. this should give somewhat clear picture as I read somewhere that test data size is comparable to train data.",
    "2739755": "You can upload your trained model as a private dataset, make test datset and predict with the model in your notebook , and then submit. Hope this will help."
  },
  "source": "meta"
}