{
  "id": 78409,
  "title": "prediction Time in Stage 2",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78409",
  "author_name": "",
  "post_date": "2019-01-23T11:48:17.606961800Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Test file have ~56k rows in stage 1 and ~376k rows in stage 2. We will need more prediction time in stage2.  this mean the model that we used in stage 1 will can not be used? We need to simplify our model? </p>",
  "messages": [
    {
      "id": "460306",
      "postDate": "01/23/2019 11:48:17",
      "content": "<p>Test file have ~56k rows in stage 1 and ~376k rows in stage 2. We will need more prediction time in stage2.  this mean the model that we used in stage 1 will can not be used? We need to simplify our model? </p>",
      "rawMarkdown": "Test file have ~56k rows in stage 1 and ~376k rows in stage 2. We will need more prediction time in stage2.  this mean the model that we used in stage 1 will can not be used? We need to simplify our model?",
      "votes": null
    },
    {
      "id": "460307",
      "postDate": "01/23/2019 11:52:23",
      "content": "<p>You most probably have to optimize your code in such a way that you don't lose enough performance while finishing the code under 7200s.</p>",
      "rawMarkdown": "You most probably have to optimize your code in such a way that you don't lose enough performance while finishing the code under 7200s.",
      "votes": null
    },
    {
      "id": "460311",
      "postDate": "01/23/2019 12:24:31",
      "content": "<p>Probably need to test how much more time it will take to do the 320k extra predictions each fold, and the feature engineering time on the whole test set.</p>",
      "rawMarkdown": "Probably need to test how much more time it will take to do the 320k extra predictions each fold, and the feature engineering time on the whole test set.",
      "votes": null
    },
    {
      "id": "460409",
      "postDate": "01/23/2019 16:33:00",
      "content": "<p>Can we just save the model we have trained at stage 1 and load them back to do the prediction at stage 2? (rather than to train the model from zero at stage 2)</p>",
      "rawMarkdown": "Can we just save the model we have trained at stage 1 and load them back to do the prediction at stage 2? (rather than to train the model from zero at stage 2)",
      "votes": null
    },
    {
      "id": "460659",
      "postDate": "01/24/2019 07:13:21",
      "content": "<p>I calculated how long it takes to predict values. It takes about 20 sec with 5 fold in 1st stage test data.</p>",
      "rawMarkdown": "I calculated how long it takes to predict values. It takes about 20 sec with 5 fold in 1st stage test data.",
      "votes": null
    },
    {
      "id": "461050",
      "postDate": "01/25/2019 04:58:06",
      "content": "<p>For my 6 fold cross validation, it takes about 4.5 to 5 seconds per fold for inference. That makes it about 27 to 30 s for inference on 56 K samples. So for 376 K samples (approximately  7 times increase) , it takes 210 seconds or so.  So for inference alone, there is about a 3 min (~ 180s ) increase due to the larger test set in the second stage.  For data loading and  pre-processing of the larger test set, I am going to allow for another 2 mins or so. So roughly, 5 minutes of additional time is required to accommodate the larger test set in the second stage. So, currently your model must fit within 6900s to barely scrape through the 7200s time limit. But, there is also a huge variance in kernel run-times for identical models. My best model (identical except for random seed) takes anywhere between 6800s and 7200s to run. Given that the organizers have indicated that there will be no leniency in total time allowed (7200 s for everything : pre-processing/training/inference/submission. Model will be rejected even if it exceeds by a second), I will feel comfortable only if my current model completes in under 6500s. Anything more than that, you are taking a chance that it will be rejected.</p>",
      "rawMarkdown": "For my 6 fold cross validation, it takes about 4.5 to 5 seconds per fold for inference. That makes it about 27 to 30 s for inference on 56 K samples. So for 376 K samples (approximately  7 times increase) , it takes 210 seconds or so.  So for inference alone, there is about a 3 min (~ 180s ) increase due to the larger test set in the second stage.  For data loading and  pre-processing of the larger test set, I am going to allow for another 2 mins or so. So roughly, 5 minutes of additional time is required to accommodate the larger test set in the second stage. So, currently your model must fit within 6900s to barely scrape through the 7200s time limit. But, there is also a huge variance in kernel run-times for identical models. My best model (identical except for random seed) takes anywhere between 6800s and 7200s to run. Given that the organizers have indicated that there will be no leniency in total time allowed (7200 s for everything : pre-processing/training/inference/submission. Model will be rejected even if it exceeds by a second), I will feel comfortable only if my current model completes in under 6500s. Anything more than that, you are taking a chance that it will be rejected.",
      "votes": null
    },
    {
      "id": "461125",
      "postDate": "01/25/2019 09:50:40",
      "content": "<p>Thanks for your answer. I will make my model under 6900s.</p>",
      "rawMarkdown": "Thanks for your answer. I will make my model under 6900s.",
      "votes": null
    },
    {
      "id": "461703",
      "postDate": "01/26/2019 19:29:13",
      "content": "<p>a simple trick you can do to check time for your final kernels is to do:\n<code>\ntest_df = pd.concat([test_df]*7).reset_index()\n</code>\nThis will let you know the run time. make sure it is &lt;7200s</p>",
      "rawMarkdown": "a simple trick you can do to check time for your final kernels is to do:\n```\ntest_df = pd.concat([test_df]*7).reset_index()\n```\nThis will let you know the run time. make sure it is &lt;7200s",
      "votes": null
    },
    {
      "id": "462290",
      "postDate": "01/28/2019 03:57:23",
      "content": "<p>I tried and find it is not work. we can use <br>\n <code>pd.concat([test, test, test,test,test,test, test]).reset_index()</code></p>",
      "rawMarkdown": "I tried and find it is not work. we can use  \n `pd.concat([test, test, test,test,test,test, test]).reset_index()`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 460307,
      "author_name": "satian",
      "author_url": "",
      "post_date": "01/23/2019 11:52:23",
      "content": "<p>You most probably have to optimize your code in such a way that you don't lose enough performance while finishing the code under 7200s.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460311,
      "author_name": "m7catsue",
      "author_url": "",
      "post_date": "01/23/2019 12:24:31",
      "content": "<p>Probably need to test how much more time it will take to do the 320k extra predictions each fold, and the feature engineering time on the whole test set.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460409,
      "author_name": "songzy12",
      "author_url": "",
      "post_date": "01/23/2019 16:33:00",
      "content": "<p>Can we just save the model we have trained at stage 1 and load them back to do the prediction at stage 2? (rather than to train the model from zero at stage 2)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 460659,
      "author_name": "tamanyan",
      "author_url": "",
      "post_date": "01/24/2019 07:13:21",
      "content": "<p>I calculated how long it takes to predict values. It takes about 20 sec with 5 fold in 1st stage test data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 461050,
      "author_name": "kagsen",
      "author_url": "",
      "post_date": "01/25/2019 04:58:06",
      "content": "<p>For my 6 fold cross validation, it takes about 4.5 to 5 seconds per fold for inference. That makes it about 27 to 30 s for inference on 56 K samples. So for 376 K samples (approximately  7 times increase) , it takes 210 seconds or so.  So for inference alone, there is about a 3 min (~ 180s ) increase due to the larger test set in the second stage.  For data loading and  pre-processing of the larger test set, I am going to allow for another 2 mins or so. So roughly, 5 minutes of additional time is required to accommodate the larger test set in the second stage. So, currently your model must fit within 6900s to barely scrape through the 7200s time limit. But, there is also a huge variance in kernel run-times for identical models. My best model (identical except for random seed) takes anywhere between 6800s and 7200s to run. Given that the organizers have indicated that there will be no leniency in total time allowed (7200 s for everything : pre-processing/training/inference/submission. Model will be rejected even if it exceeds by a second), I will feel comfortable only if my current model completes in under 6500s. Anything more than that, you are taking a chance that it will be rejected.</p>",
      "votes": null,
      "replies": [
        {
          "id": 461125,
          "author_name": "linxid615",
          "author_url": "",
          "post_date": "01/25/2019 09:50:40",
          "content": "<p>Thanks for your answer. I will make my model under 6900s.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 461703,
      "author_name": "mlwhiz",
      "author_url": "",
      "post_date": "01/26/2019 19:29:13",
      "content": "<p>a simple trick you can do to check time for your final kernels is to do:\n<code>\ntest_df = pd.concat([test_df]*7).reset_index()\n</code>\nThis will let you know the run time. make sure it is &lt;7200s</p>",
      "votes": null,
      "replies": [
        {
          "id": 462290,
          "author_name": "linxid615",
          "author_url": "",
          "post_date": "01/28/2019 03:57:23",
          "content": "<p>I tried and find it is not work. we can use <br>\n <code>pd.concat([test, test, test,test,test,test, test]).reset_index()</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "460306": "Test file have ~56k rows in stage 1 and ~376k rows in stage 2. We will need more prediction time in stage2.  this mean the model that we used in stage 1 will can not be used? We need to simplify our model?",
    "460307": "You most probably have to optimize your code in such a way that you don't lose enough performance while finishing the code under 7200s.",
    "460311": "Probably need to test how much more time it will take to do the 320k extra predictions each fold, and the feature engineering time on the whole test set.",
    "460409": "Can we just save the model we have trained at stage 1 and load them back to do the prediction at stage 2? (rather than to train the model from zero at stage 2)",
    "460659": "I calculated how long it takes to predict values. It takes about 20 sec with 5 fold in 1st stage test data.",
    "461050": "For my 6 fold cross validation, it takes about 4.5 to 5 seconds per fold for inference. That makes it about 27 to 30 s for inference on 56 K samples. So for 376 K samples (approximately  7 times increase) , it takes 210 seconds or so.  So for inference alone, there is about a 3 min (~ 180s ) increase due to the larger test set in the second stage.  For data loading and  pre-processing of the larger test set, I am going to allow for another 2 mins or so. So roughly, 5 minutes of additional time is required to accommodate the larger test set in the second stage. So, currently your model must fit within 6900s to barely scrape through the 7200s time limit. But, there is also a huge variance in kernel run-times for identical models. My best model (identical except for random seed) takes anywhere between 6800s and 7200s to run. Given that the organizers have indicated that there will be no leniency in total time allowed (7200 s for everything : pre-processing/training/inference/submission. Model will be rejected even if it exceeds by a second), I will feel comfortable only if my current model completes in under 6500s. Anything more than that, you are taking a chance that it will be rejected.",
    "461125": "Thanks for your answer. I will make my model under 6900s.",
    "461703": "a simple trick you can do to check time for your final kernels is to do:\n```\ntest_df = pd.concat([test_df]*7).reset_index()\n```\nThis will let you know the run time. make sure it is &lt;7200s",
    "462290": "I tried and find it is not work. we can use  \n `pd.concat([test, test, test,test,test,test, test]).reset_index()`"
  },
  "source": "meta"
}