{
  "id": 496734,
  "title": "Submission scoring time taken",
  "url": "/competitions/home-credit-credit-risk-model-stability/discussion/496734",
  "author_name": "",
  "post_date": "2024-04-22T09:18:54.006466400Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>How long is your submission scoring time? Mine ran on GPU - has been taking more than 4 hours… Normal?</p>",
  "messages": [
    {
      "id": "2767353",
      "postDate": "04/22/2024 09:18:54",
      "content": "<p>How long is your submission scoring time? Mine ran on GPU - has been taking more than 4 hours… Normal?</p>",
      "rawMarkdown": "How long is your submission scoring time? Mine ran on GPU - has been taking more than 4 hours... Normal?",
      "votes": null
    },
    {
      "id": "2767661",
      "postDate": "04/22/2024 13:19:08",
      "content": "<p>May I ask if you have trained inference separation? If not, is it copying open source code? When running that code offline, in order to save GPU, the amount of data and the number of iterators were reduced</p>\n<p>Let me also talk about my results only inference here: 700 features, CPU, 20 models, 3 hours.</p>",
      "rawMarkdown": "May I ask if you have trained inference separation? If not, is it copying open source code? When running that code offline, in order to save GPU, the amount of data and the number of iterators were reduced\n\nLet me also talk about my results only inference here: 700 features, CPU, 20 models, 3 hours.",
      "votes": null
    },
    {
      "id": "2767910",
      "postDate": "04/22/2024 15:11:40",
      "content": "<p>let me separate them and see.</p>",
      "rawMarkdown": "let me separate them and see.",
      "votes": null
    },
    {
      "id": "2767996",
      "postDate": "04/22/2024 16:00:03",
      "content": "<p>I am also observing the same, if I run my notebook, it just takes about 15 mins (training + prediction on dummy test data) but when I submit, the scoring takes several hours (400 features).<br>\nWhat I understand the difference between running the notebook and submission is that the submission has much larger test file compared to dummy test data. Is this difference responsible for longer scoring time ?</p>\n<p>I see you mentioned separation of training and inference, do you mean training in a separate notebook, saving as a dataset and loading into a different notebook ? should that make prediction on actual test data quicker ? but why ?</p>\n<p>thanks</p>\n<p>Edit : its one of the public notebook I have copied</p>",
      "rawMarkdown": "I am also observing the same, if I run my notebook, it just takes about 15 mins (training + prediction on dummy test data) but when I submit, the scoring takes several hours (400 features).\nWhat I understand the difference between running the notebook and submission is that the submission has much larger test file compared to dummy test data. Is this difference responsible for longer scoring time ?\n\nI see you mentioned separation of training and inference, do you mean training in a separate notebook, saving as a dataset and loading into a different notebook ? should that make prediction on actual test data quicker ? but why ?\n\nthanks\n\nEdit : its one of the public notebook I have copied",
      "votes": null
    },
    {
      "id": "2768561",
      "postDate": "04/22/2024 23:45:01",
      "content": "<p>If the training and inference of the model are separated, you can save time on retraining the model when you submit it for running</p>",
      "rawMarkdown": "If the training and inference of the model are separated, you can save time on retraining the model when you submit it for running",
      "votes": null
    },
    {
      "id": "2768812",
      "postDate": "04/23/2024 03:37:37",
      "content": "<p>sorry but my point was, the training isn't taking more than 15 mins, so why does the submission take that long. My suspicion is, its the amount of test data prediction that is taking that long during scoring process and if thats true then separating out prediction in a separate notebook shouldn't make any difference. I will try to separate prediction to see if that makes any difference as I do have a separate notebook just for submission purpose.</p>",
      "rawMarkdown": "sorry but my point was, the training isn't taking more than 15 mins, so why does the submission take that long. My suspicion is, its the amount of test data prediction that is taking that long during scoring process and if thats true then separating out prediction in a separate notebook shouldn't make any difference. I will try to separate prediction to see if that makes any difference as I do have a separate notebook just for submission purpose.",
      "votes": null
    },
    {
      "id": "2768963",
      "postDate": "04/23/2024 05:23:04",
      "content": "<p>Are you looking at open source code? When training the model, a portion of the data was selected, and the number of iterators was also significantly reduced.</p>",
      "rawMarkdown": "Are you looking at open source code? When training the model, a portion of the data was selected, and the number of iterators was also significantly reduced.",
      "votes": null
    },
    {
      "id": "2768972",
      "postDate": "04/23/2024 05:27:01",
      "content": "<p>example:   <a href=\"https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble\" target=\"_blank\">https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble</a></p>\n<p>In [6]</p>\n<p>\"\"\"<br>\nsample = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")<br>\ndevice='gpu'</p>\n<h1>n_samples=200000</h1>\n<p>n_est=6000<br>\nDRY_RUN = True if sample.shape[0] == 10 else False   <br>\nif DRY_RUN:<br>\n    device='cpu'<br>\n    df_train = df_train.iloc[:50000]<br>\n    #n_samples=10000<br>\n    n_est=600<br>\nprint(device)</p>\n<p>\"\"\"</p>",
      "rawMarkdown": "example:   https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble\n\nIn [6]\n\n\"\"\"\nsample = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\ndevice='gpu'\n#n_samples=200000\nn_est=6000\nDRY_RUN = True if sample.shape[0] == 10 else False   \nif DRY_RUN:\n    device='cpu'\n    df_train = df_train.iloc[:50000]\n    #n_samples=10000\n    n_est=600\nprint(device)\n\n\"\"\"",
      "votes": null
    },
    {
      "id": "2769670",
      "postDate": "04/23/2024 13:44:23",
      "content": "<p>ah I see :) I had only yet reviewed the code until the dry run bit, now I understand. Thanks for clarifying !</p>",
      "rawMarkdown": "ah I see :) I had only yet reviewed the code until the dry run bit, now I understand. Thanks for clarifying !",
      "votes": null
    },
    {
      "id": "2775064",
      "postDate": "04/25/2024 13:12:25",
      "content": "<p>It all depends what you are doing at each notebook run.<br>\nIf you are loading a pre-processed train set - you are not losing time for processing it, but during submission you read and process the entire test set from scratch, every time.</p>",
      "rawMarkdown": "It all depends what you are doing at each notebook run.\nIf you are loading a pre-processed train set - you are not losing time for processing it, but during submission you read and process the entire test set from scratch, every time.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2767661,
      "author_name": "yunsuxiaozi",
      "author_url": "",
      "post_date": "04/22/2024 13:19:08",
      "content": "<p>May I ask if you have trained inference separation? If not, is it copying open source code? When running that code offline, in order to save GPU, the amount of data and the number of iterators were reduced</p>\n<p>Let me also talk about my results only inference here: 700 features, CPU, 20 models, 3 hours.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2767910,
          "author_name": "skw1990",
          "author_url": "",
          "post_date": "04/22/2024 15:11:40",
          "content": "<p>let me separate them and see.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2767996,
          "author_name": "jabranzahid",
          "author_url": "",
          "post_date": "04/22/2024 16:00:03",
          "content": "<p>I am also observing the same, if I run my notebook, it just takes about 15 mins (training + prediction on dummy test data) but when I submit, the scoring takes several hours (400 features).<br>\nWhat I understand the difference between running the notebook and submission is that the submission has much larger test file compared to dummy test data. Is this difference responsible for longer scoring time ?</p>\n<p>I see you mentioned separation of training and inference, do you mean training in a separate notebook, saving as a dataset and loading into a different notebook ? should that make prediction on actual test data quicker ? but why ?</p>\n<p>thanks</p>\n<p>Edit : its one of the public notebook I have copied</p>",
          "votes": null,
          "replies": [
            {
              "id": 2768561,
              "author_name": "yunsuxiaozi",
              "author_url": "",
              "post_date": "04/22/2024 23:45:01",
              "content": "<p>If the training and inference of the model are separated, you can save time on retraining the model when you submit it for running</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2768812,
                  "author_name": "jabranzahid",
                  "author_url": "",
                  "post_date": "04/23/2024 03:37:37",
                  "content": "<p>sorry but my point was, the training isn't taking more than 15 mins, so why does the submission take that long. My suspicion is, its the amount of test data prediction that is taking that long during scoring process and if thats true then separating out prediction in a separate notebook shouldn't make any difference. I will try to separate prediction to see if that makes any difference as I do have a separate notebook just for submission purpose.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2768963,
                      "author_name": "yunsuxiaozi",
                      "author_url": "",
                      "post_date": "04/23/2024 05:23:04",
                      "content": "<p>Are you looking at open source code? When training the model, a portion of the data was selected, and the number of iterators was also significantly reduced.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2768972,
                          "author_name": "yunsuxiaozi",
                          "author_url": "",
                          "post_date": "04/23/2024 05:27:01",
                          "content": "<p>example:   <a href=\"https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble\" target=\"_blank\">https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble</a></p>\n<p>In [6]</p>\n<p>\"\"\"<br>\nsample = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")<br>\ndevice='gpu'</p>\n<h1>n_samples=200000</h1>\n<p>n_est=6000<br>\nDRY_RUN = True if sample.shape[0] == 10 else False   <br>\nif DRY_RUN:<br>\n    device='cpu'<br>\n    df_train = df_train.iloc[:50000]<br>\n    #n_samples=10000<br>\n    n_est=600<br>\nprint(device)</p>\n<p>\"\"\"</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2769670,
                              "author_name": "jabranzahid",
                              "author_url": "",
                              "post_date": "04/23/2024 13:44:23",
                              "content": "<p>ah I see :) I had only yet reviewed the code until the dry run bit, now I understand. Thanks for clarifying !</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2775064,
      "author_name": "eu1234",
      "author_url": "",
      "post_date": "04/25/2024 13:12:25",
      "content": "<p>It all depends what you are doing at each notebook run.<br>\nIf you are loading a pre-processed train set - you are not losing time for processing it, but during submission you read and process the entire test set from scratch, every time.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2767353": "How long is your submission scoring time? Mine ran on GPU - has been taking more than 4 hours... Normal?",
    "2767661": "May I ask if you have trained inference separation? If not, is it copying open source code? When running that code offline, in order to save GPU, the amount of data and the number of iterators were reduced\n\nLet me also talk about my results only inference here: 700 features, CPU, 20 models, 3 hours.",
    "2767910": "let me separate them and see.",
    "2767996": "I am also observing the same, if I run my notebook, it just takes about 15 mins (training + prediction on dummy test data) but when I submit, the scoring takes several hours (400 features).\nWhat I understand the difference between running the notebook and submission is that the submission has much larger test file compared to dummy test data. Is this difference responsible for longer scoring time ?\n\nI see you mentioned separation of training and inference, do you mean training in a separate notebook, saving as a dataset and loading into a different notebook ? should that make prediction on actual test data quicker ? but why ?\n\nthanks\n\nEdit : its one of the public notebook I have copied",
    "2768561": "If the training and inference of the model are separated, you can save time on retraining the model when you submit it for running",
    "2768812": "sorry but my point was, the training isn't taking more than 15 mins, so why does the submission take that long. My suspicion is, its the amount of test data prediction that is taking that long during scoring process and if thats true then separating out prediction in a separate notebook shouldn't make any difference. I will try to separate prediction to see if that makes any difference as I do have a separate notebook just for submission purpose.",
    "2768963": "Are you looking at open source code? When training the model, a portion of the data was selected, and the number of iterators was also significantly reduced.",
    "2768972": "example:   https://www.kaggle.com/code/jeffersusus/home-credit-lgb-cat-ensemble\n\nIn [6]\n\n\"\"\"\nsample = pd.read_csv(\"/kaggle/input/home-credit-credit-risk-model-stability/sample_submission.csv\")\ndevice='gpu'\n#n_samples=200000\nn_est=6000\nDRY_RUN = True if sample.shape[0] == 10 else False   \nif DRY_RUN:\n    device='cpu'\n    df_train = df_train.iloc[:50000]\n    #n_samples=10000\n    n_est=600\nprint(device)\n\n\"\"\"",
    "2769670": "ah I see :) I had only yet reviewed the code until the dry run bit, now I understand. Thanks for clarifying !",
    "2775064": "It all depends what you are doing at each notebook run.\nIf you are loading a pre-processed train set - you are not losing time for processing it, but during submission you read and process the entire test set from scratch, every time."
  },
  "source": "meta"
}