{
  "id": 195148,
  "title": "Guessing the private test set batch size",
  "url": "/competitions/riiid-test-answer-prediction/discussion/195148",
  "author_name": "",
  "post_date": "2020-11-03T18:15:21.089732500Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Running my submission notebook (includes training, features generation and inference ) takes approximately 15, 20 minutes at max. However when I submit it, it times-out. </p>\n<p>So in an attempt to understand why, I came across this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192779#1058493\" target=\"_blank\">thread that talks about how we should optimize code for small batches. </a></p>\n<p>This made me attempt to guess which batch size private test set is using during inference. To do this, I created a simulation: Preprocessing and predicting values of a dataframe of 2,500,000 rows using different batch sizes, and calculating the time it takes in total for every batch size.</p>\n<p>The result was this (x: Batch or chunk size, y: time in seconds):<br>\n<img src=\"https://i.imgur.com/9DqvPj9.png\" alt=\"\"><br>\nAs you see, the smaller the batch, the more time in seconds preprocessing and inference takes. For a batch size of 10,000 it would take 441 seconds (~7.5 Minutes) to process 2.5M rows. </p>\n<p>For batch size of 50, it would take 12,380 seconds (~206 Minutes ~ 3.43 Hours). <br>\nAdding 206 Minutes + 15 Minutes (loading data) + 20 Minutes (Training and preprocessing) = 241 Minutes, which is far less than 9 hours (540 Minutes).</p>\n<p>So why is my notebook timing-out? Is it ran on different hardware during submission or am I missing something ? If anyone experimented with such a thing, please provide us with your input!</p>",
  "messages": [
    {
      "id": "1068770",
      "postDate": "11/03/2020 18:15:21",
      "content": "<p>Running my submission notebook (includes training, features generation and inference ) takes approximately 15, 20 minutes at max. However when I submit it, it times-out. </p>\n<p>So in an attempt to understand why, I came across this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192779#1058493\" target=\"_blank\">thread that talks about how we should optimize code for small batches. </a></p>\n<p>This made me attempt to guess which batch size private test set is using during inference. To do this, I created a simulation: Preprocessing and predicting values of a dataframe of 2,500,000 rows using different batch sizes, and calculating the time it takes in total for every batch size.</p>\n<p>The result was this (x: Batch or chunk size, y: time in seconds):<br>\n<img src=\"https://i.imgur.com/9DqvPj9.png\" alt=\"\"><br>\nAs you see, the smaller the batch, the more time in seconds preprocessing and inference takes. For a batch size of 10,000 it would take 441 seconds (~7.5 Minutes) to process 2.5M rows. </p>\n<p>For batch size of 50, it would take 12,380 seconds (~206 Minutes ~ 3.43 Hours). <br>\nAdding 206 Minutes + 15 Minutes (loading data) + 20 Minutes (Training and preprocessing) = 241 Minutes, which is far less than 9 hours (540 Minutes).</p>\n<p>So why is my notebook timing-out? Is it ran on different hardware during submission or am I missing something ? If anyone experimented with such a thing, please provide us with your input!</p>",
      "rawMarkdown": "Running my submission notebook (includes training, features generation and inference ) takes approximately 15, 20 minutes at max. However when I submit it, it times-out. \n\nSo in an attempt to understand why, I came across this [thread that talks about how we should optimize code for small batches. ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192779#1058493)\n\nThis made me attempt to guess which batch size private test set is using during inference. To do this, I created a simulation: Preprocessing and predicting values of a dataframe of 2,500,000 rows using different batch sizes, and calculating the time it takes in total for every batch size.\n\nThe result was this (x: Batch or chunk size, y: time in seconds):\n![](https://i.imgur.com/9DqvPj9.png)\nAs you see, the smaller the batch, the more time in seconds preprocessing and inference takes. For a batch size of 10,000 it would take 441 seconds (~7.5 Minutes) to process 2.5M rows. \n\nFor batch size of 50, it would take 12,380 seconds (~206 Minutes ~ 3.43 Hours). \nAdding 206 Minutes + 15 Minutes (loading data) + 20 Minutes (Training and preprocessing) = 241 Minutes, which is far less than 9 hours (540 Minutes).\n\nSo why is my notebook timing-out? Is it ran on different hardware during submission or am I missing something ? If anyone experimented with such a thing, please provide us with your input!",
      "votes": null
    },
    {
      "id": "1068783",
      "postDate": "11/03/2020 18:30:00",
      "content": "<p>According to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">this</a> there are lots of very small batches.</p>",
      "rawMarkdown": "According to [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282) there are lots of very small batches.",
      "votes": null
    },
    {
      "id": "1068832",
      "postDate": "11/03/2020 19:24:50",
      "content": "<p>Thanks for the answer. Which library you think is the fastest for small batches processing ? </p>",
      "rawMarkdown": "Thanks for the answer. Which library you think is the fastest for small batches processing ?",
      "votes": null
    },
    {
      "id": "1069029",
      "postDate": "11/04/2020 01:58:28",
      "content": "<p>Check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377#1066724\" target=\"_blank\">this response</a>. The whole idea is to move away from pandas.</p>",
      "rawMarkdown": "Check [this response](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377#1066724). The whole idea is to move away from pandas.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1068783,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "11/03/2020 18:30:00",
      "content": "<p>According to <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282\" target=\"_blank\">this</a> there are lots of very small batches.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1068832,
          "author_name": "abdessalemboukil",
          "author_url": "",
          "post_date": "11/03/2020 19:24:50",
          "content": "<p>Thanks for the answer. Which library you think is the fastest for small batches processing ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1069029,
          "author_name": "rohanrao",
          "author_url": "",
          "post_date": "11/04/2020 01:58:28",
          "content": "<p>Check <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377#1066724\" target=\"_blank\">this response</a>. The whole idea is to move away from pandas.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1068770": "Running my submission notebook (includes training, features generation and inference ) takes approximately 15, 20 minutes at max. However when I submit it, it times-out. \n\nSo in an attempt to understand why, I came across this [thread that talks about how we should optimize code for small batches. ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192779#1058493)\n\nThis made me attempt to guess which batch size private test set is using during inference. To do this, I created a simulation: Preprocessing and predicting values of a dataframe of 2,500,000 rows using different batch sizes, and calculating the time it takes in total for every batch size.\n\nThe result was this (x: Batch or chunk size, y: time in seconds):\n![](https://i.imgur.com/9DqvPj9.png)\nAs you see, the smaller the batch, the more time in seconds preprocessing and inference takes. For a batch size of 10,000 it would take 441 seconds (~7.5 Minutes) to process 2.5M rows. \n\nFor batch size of 50, it would take 12,380 seconds (~206 Minutes ~ 3.43 Hours). \nAdding 206 Minutes + 15 Minutes (loading data) + 20 Minutes (Training and preprocessing) = 241 Minutes, which is far less than 9 hours (540 Minutes).\n\nSo why is my notebook timing-out? Is it ran on different hardware during submission or am I missing something ? If anyone experimented with such a thing, please provide us with your input!",
    "1068783": "According to [this](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124#1056282) there are lots of very small batches.",
    "1068832": "Thanks for the answer. Which library you think is the fastest for small batches processing ?",
    "1069029": "Check [this response](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/194377#1066724). The whole idea is to move away from pandas."
  },
  "source": "meta"
}