{
  "id": 122993,
  "title": "Inference within 20 minutes",
  "url": "/competitions/bengaliai-cv19/discussion/122993",
  "author_name": "Peter",
  "post_date": "2019-12-24T01:44:04.182000",
  "votes": 79,
  "comment_count": 22,
  "views": 0,
  "content": "<p>I spent a little time optimizing my inference code, and the running time reduced from ~2 hours to ~20 minutes.</p>\n\n<p>I am using Pytorch's dataset/dataloader, but you can apply this optimization anyway.</p>\n\n<p>So the issue is with this code in my custom dataset:</p>\n\n<p><code>\ndef __getitem__(self, idx):\n    # Using pd.DataFrame is the problem!!!\n    img = self.data.iloc[idx, 1:].values.reshape(HEIGHT, WIDTH)\n</code></p>\n\n<p>If you convert your datastore to Numpy array, the process will ~30x faster.</p>\n\n<p>```\nclass BengaliParquetDataset(Dataset):</p>\n\n<pre><code>def __init__(self, parquet_file, transform=None):\n\n    self.data = pd.read_parquet(parquet_file)\n    self.data = self.data.iloc[:, 1:].values\n\n    [...more code...]\n\ndef __getitem__(self, idx):\n    img = self.data[idx, :].reshape(HEIGHT, WIDTH)\n\n    [...more code...]\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": 701863,
      "postDate": "2019-12-24T01:44:04.183Z",
      "content": "<p>I spent a little time optimizing my inference code, and the running time reduced from ~2 hours to ~20 minutes.</p>\n\n<p>I am using Pytorch's dataset/dataloader, but you can apply this optimization anyway.</p>\n\n<p>So the issue is with this code in my custom dataset:</p>\n\n<p><code>\ndef __getitem__(self, idx):\n    # Using pd.DataFrame is the problem!!!\n    img = self.data.iloc[idx, 1:].values.reshape(HEIGHT, WIDTH)\n</code></p>\n\n<p>If you convert your datastore to Numpy array, the process will ~30x faster.</p>\n\n<p>```\nclass BengaliParquetDataset(Dataset):</p>\n\n<pre><code>def __init__(self, parquet_file, transform=None):\n\n    self.data = pd.read_parquet(parquet_file)\n    self.data = self.data.iloc[:, 1:].values\n\n    [...more code...]\n\ndef __getitem__(self, idx):\n    img = self.data[idx, :].reshape(HEIGHT, WIDTH)\n\n    [...more code...]\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I spent a little time optimizing my inference code, and the running time reduced from ~2 hours to ~20 minutes.\n\nI am using Pytorch's dataset/dataloader, but you can apply this optimization anyway.\n\nSo the issue is with this code in my custom dataset:\n\n```\ndef __getitem__(self, idx):\n    # Using pd.DataFrame is the problem!!!\n    img = self.data.iloc[idx, 1:].values.reshape(HEIGHT, WIDTH)\n```\n\nIf you convert your datastore to Numpy array, the process will ~30x faster.\n\n```\nclass BengaliParquetDataset(Dataset):\n\n    def __init__(self, parquet_file, transform=None):\n\n        self.data = pd.read_parquet(parquet_file)\n        self.data = self.data.iloc[:, 1:].values\n\n        [...more code...]\n\n    def __getitem__(self, idx):\n        img = self.data[idx, :].reshape(HEIGHT, WIDTH)\n\n        [...more code...]\n```\n\n",
      "votes": 78
    },
    {
      "id": 704910,
      "postDate": "2019-12-28T06:18:48.583Z",
      "content": "<p>To get to about 10 minutes, you can convert the numpy array to pytorch tensor and use TensorDataset.</p>",
      "rawMarkdown": "To get to about 10 minutes, you can convert the numpy array to pytorch tensor and use TensorDataset.",
      "votes": 3
    },
    {
      "id": 702377,
      "postDate": "2019-12-24T16:05:42.007Z",
      "content": "<p>how insightful ! upvoted . thanks for sharing !</p>",
      "rawMarkdown": "how insightful ! upvoted . thanks for sharing !",
      "votes": 3
    },
    {
      "id": 704093,
      "postDate": "2019-12-27T03:34:44.080Z",
      "content": "<p>you are the best :) </p>",
      "rawMarkdown": "you are the best :) ",
      "votes": 1
    },
    {
      "id": 704106,
      "postDate": "2019-12-27T03:59:11.203Z",
      "content": "<p>There is another major benefit , that is not highlighted by you <a href=\"/pestipeti\">@pestipeti</a>  . I changed the code for training . One epoch used to take 41 Minutes for full data , now its 11 minutes :D .</p>",
      "rawMarkdown": "There is another major benefit , that is not highlighted by you @pestipeti  . I changed the code for training . One epoch used to take 41 Minutes for full data , now its 11 minutes :D .",
      "votes": 2,
      "replies": [
        {
          "id": 714774,
          "postDate": "2020-01-09T18:22:27.240Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> is your 11 minutes per epoch time on full data after unfreezing all the weights?</p>",
          "rawMarkdown": "@phoenix9032 is your 11 minutes per epoch time on full data after unfreezing all the weights?"
        }
      ]
    },
    {
      "id": 772319,
      "postDate": "2020-03-15T10:38:37.103Z",
      "content": "<p>To say that this is  a life saver would be an understatement. Thank you very much !</p>",
      "rawMarkdown": "To say that this is  a life saver would be an understatement. Thank you very much !"
    },
    {
      "id": 731480,
      "postDate": "2020-01-28T17:23:18.173Z",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thanks a lot. This saves me from continuous \"Notebook Exceeded Allowed Compute\" nightmares.</p>",
      "rawMarkdown": "@pestipeti Thanks a lot. This saves me from continuous \"Notebook Exceeded Allowed Compute\" nightmares."
    },
    {
      "id": 730045,
      "postDate": "2020-01-27T03:08:00.103Z",
      "content": "<p>I had same issue, but i fixed it after discovering this thread. Thank you a lot !!!</p>",
      "rawMarkdown": "I had same issue, but i fixed it after discovering this thread. Thank you a lot !!!"
    },
    {
      "id": 717083,
      "postDate": "2020-01-12T17:45:26.993Z",
      "content": "<p>I've also tried to reshape whole dataframe, but got kernel restarted. Any luck with it?</p>",
      "rawMarkdown": "I've also tried to reshape whole dataframe, but got kernel restarted. Any luck with it?"
    },
    {
      "id": 706365,
      "postDate": "2019-12-30T09:42:38.513Z",
      "content": "<p>Thanks for sharing Peter 👍\nI am able to resolve my \"Notebook Exceeded Allowed Compute\" error using your suggestion.</p>\n\n<p>Right now I am facing this \"Submission CSV Not Found\". <br>\nAny suggestion?</p>",
      "rawMarkdown": "Thanks for sharing Peter 👍\nI am able to resolve my \"Notebook Exceeded Allowed Compute\" error using your suggestion.\n\nRight now I am facing this \"Submission CSV Not Found\".  \nAny suggestion?"
    },
    {
      "id": 706264,
      "postDate": "2019-12-30T06:48:56.110Z",
      "content": "<p>What's the parquet_file? numpy? <a href=\"/pestipeti\">@pestipeti</a> </p>",
      "rawMarkdown": "What's the parquet_file? numpy? @pestipeti "
    },
    {
      "id": 704345,
      "postDate": "2019-12-27T11:07:00.297Z",
      "content": "<p>Does the data fit into kaggle kernels RAM? <a href=\"/pestipeti\">@pestipeti</a></p>\n\n<p>I mean, I'm doing nearly the same as you advised but in Keras, and when i try to convert the whole dataframe into a numpy array i get memory error (while training).</p>\n\n<p>Now, assuming the test set is \"roughly the same size of the training set\", I assume that my submission will raise a memory error if doing as you suggested, isn't it?</p>",
      "rawMarkdown": "Does the data fit into kaggle kernels RAM? @pestipeti\n\nI mean, I'm doing nearly the same as you advised but in Keras, and when i try to convert the whole dataframe into a numpy array i get memory error (while training).\n\nNow, assuming the test set is \"roughly the same size of the training set\", I assume that my submission will raise a memory error if doing as you suggested, isn't it?",
      "replies": [
        {
          "id": 704356,
          "postDate": "2019-12-27T11:20:43.593Z",
          "content": "<p>I processed by one parquet/df at a time. </p>",
          "rawMarkdown": "I processed by one parquet/df at a time. ",
          "votes": 1
        },
        {
          "id": 704371,
          "postDate": "2019-12-27T11:55:20.713Z",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> are you talking about processing each df one by one during training time ? I have faced a problem that half way through the epoch the RAM gets full and epoch gets stuck . But that happened during old way of loading also in Kaggle Kernel .however I have not faced that issue in colab . Weird .</p>",
          "rawMarkdown": "@pestipeti are you talking about processing each df one by one during training time ? I have faced a problem that half way through the epoch the RAM gets full and epoch gets stuck . But that happened during old way of loading also in Kaggle Kernel .however I have not faced that issue in colab . Weird .",
          "votes": 1
        },
        {
          "id": 704423,
          "postDate": "2019-12-27T13:13:50.357Z",
          "content": "<p>I train locally and I have enough ram, so I did not optimize my training process for memory. During inference I load one df, convert to numpy and predict. </p>",
          "rawMarkdown": "I train locally and I have enough ram, so I did not optimize my training process for memory. During inference I load one df, convert to numpy and predict. ",
          "votes": 1
        },
        {
          "id": 704912,
          "postDate": "2019-12-28T06:21:25.323Z",
          "content": "<p>Yes it will fit, if you keep your dtype to byte. Then convert to float just a batch of data.</p>",
          "rawMarkdown": "Yes it will fit, if you keep your dtype to byte. Then convert to float just a batch of data."
        }
      ]
    },
    {
      "id": 715207,
      "postDate": "2020-01-10T08:52:52.660Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 702934,
      "postDate": "2019-12-25T11:28:10.540Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 769706,
      "postDate": "2020-03-12T06:55:52.340Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    },
    {
      "id": 740362,
      "postDate": "2020-02-09T10:18:13.277Z",
      "content": "<p>Thanks, now I can add more models!!!</p>",
      "rawMarkdown": "Thanks, now I can add more models!!!"
    },
    {
      "id": 728902,
      "postDate": "2020-01-25T13:04:08.210Z",
      "content": "<p>this is a great help - thank you!</p>",
      "rawMarkdown": "this is a great help - thank you!"
    },
    {
      "id": 716799,
      "postDate": "2020-01-12T09:00:12.647Z",
      "content": "<p>very useful, thanks your share!</p>",
      "rawMarkdown": "very useful, thanks your share!"
    }
  ],
  "comments": [
    {
      "id": 704910,
      "author_name": "ibraheemmoosa",
      "author_url": "",
      "post_date": "2019-12-28T06:18:48.583000",
      "content": "<p>To get to about 10 minutes, you can convert the numpy array to pytorch tensor and use TensorDataset.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 702377,
      "author_name": "ravi tanwar",
      "author_url": "",
      "post_date": "2019-12-24T16:05:42.007000",
      "content": "<p>how insightful ! upvoted . thanks for sharing !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 704093,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-12-27T03:34:44.080000",
      "content": "<p>you are the best :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 704106,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2019-12-27T03:59:11.203000",
      "content": "<p>There is another major benefit , that is not highlighted by you <a href=\"/pestipeti\">@pestipeti</a>  . I changed the code for training . One epoch used to take 41 Minutes for full data , now its 11 minutes :D .</p>",
      "votes": 2,
      "replies": [
        {
          "id": 714774,
          "author_name": "timetraveller",
          "author_url": "",
          "post_date": "2020-01-09T18:22:27.240000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> is your 11 minutes per epoch time on full data after unfreezing all the weights?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 772319,
      "author_name": "jwelliav",
      "author_url": "",
      "post_date": "2020-03-15T10:38:37.103000",
      "content": "<p>To say that this is  a life saver would be an understatement. Thank you very much !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 731480,
      "author_name": "Gaurav Yadav",
      "author_url": "",
      "post_date": "2020-01-28T17:23:18.173000",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a> Thanks a lot. This saves me from continuous \"Notebook Exceeded Allowed Compute\" nightmares.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 730045,
      "author_name": "damien.lee",
      "author_url": "",
      "post_date": "2020-01-27T03:08:00.103000",
      "content": "<p>I had same issue, but i fixed it after discovering this thread. Thank you a lot !!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 717083,
      "author_name": "Vladimir Smirnov",
      "author_url": "",
      "post_date": "2020-01-12T17:45:26.993000",
      "content": "<p>I've also tried to reshape whole dataframe, but got kernel restarted. Any luck with it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706365,
      "author_name": "Manish Nayak",
      "author_url": "",
      "post_date": "2019-12-30T09:42:38.513000",
      "content": "<p>Thanks for sharing Peter 👍\nI am able to resolve my \"Notebook Exceeded Allowed Compute\" error using your suggestion.</p>\n\n<p>Right now I am facing this \"Submission CSV Not Found\". <br>\nAny suggestion?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706264,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2019-12-30T06:48:56.110000",
      "content": "<p>What's the parquet_file? numpy? <a href=\"/pestipeti\">@pestipeti</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 704345,
      "author_name": "LAZCoder",
      "author_url": "",
      "post_date": "2019-12-27T11:07:00.297000",
      "content": "<p>Does the data fit into kaggle kernels RAM? <a href=\"/pestipeti\">@pestipeti</a></p>\n\n<p>I mean, I'm doing nearly the same as you advised but in Keras, and when i try to convert the whole dataframe into a numpy array i get memory error (while training).</p>\n\n<p>Now, assuming the test set is \"roughly the same size of the training set\", I assume that my submission will raise a memory error if doing as you suggested, isn't it?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 704356,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2019-12-27T11:20:43.593000",
          "content": "<p>I processed by one parquet/df at a time. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704371,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2019-12-27T11:55:20.713000",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> are you talking about processing each df one by one during training time ? I have faced a problem that half way through the epoch the RAM gets full and epoch gets stuck . But that happened during old way of loading also in Kaggle Kernel .however I have not faced that issue in colab . Weird .</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704423,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2019-12-27T13:13:50.357000",
          "content": "<p>I train locally and I have enough ram, so I did not optimize my training process for memory. During inference I load one df, convert to numpy and predict. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 704912,
          "author_name": "ibraheemmoosa",
          "author_url": "",
          "post_date": "2019-12-28T06:21:25.323000",
          "content": "<p>Yes it will fit, if you keep your dtype to byte. Then convert to float just a batch of data.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 715207,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-10T08:52:52.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 702934,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-25T11:28:10.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 769706,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-12T06:55:52.340000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 740362,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-02-09T10:18:13.277000",
      "content": "<p>Thanks, now I can add more models!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728902,
      "author_name": "james hubbard",
      "author_url": "",
      "post_date": "2020-01-25T13:04:08.210000",
      "content": "<p>this is a great help - thank you!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 716799,
      "author_name": "mayfive",
      "author_url": "",
      "post_date": "2020-01-12T09:00:12.647000",
      "content": "<p>very useful, thanks your share!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "701863": "I spent a little time optimizing my inference code, and the running time reduced from ~2 hours to ~20 minutes.\n\nI am using Pytorch's dataset/dataloader, but you can apply this optimization anyway.\n\nSo the issue is with this code in my custom dataset:\n\n```\ndef __getitem__(self, idx):\n    # Using pd.DataFrame is the problem!!!\n    img = self.data.iloc[idx, 1:].values.reshape(HEIGHT, WIDTH)\n```\n\nIf you convert your datastore to Numpy array, the process will ~30x faster.\n\n```\nclass BengaliParquetDataset(Dataset):\n\n    def __init__(self, parquet_file, transform=None):\n\n        self.data = pd.read_parquet(parquet_file)\n        self.data = self.data.iloc[:, 1:].values\n\n        [...more code...]\n\n    def __getitem__(self, idx):\n        img = self.data[idx, :].reshape(HEIGHT, WIDTH)\n\n        [...more code...]\n```\n\n",
    "704910": "To get to about 10 minutes, you can convert the numpy array to pytorch tensor and use TensorDataset.",
    "702377": "how insightful ! upvoted . thanks for sharing !",
    "704093": "you are the best :) ",
    "704106": "There is another major benefit , that is not highlighted by you @pestipeti  . I changed the code for training . One epoch used to take 41 Minutes for full data , now its 11 minutes :D .",
    "772319": "To say that this is  a life saver would be an understatement. Thank you very much !",
    "731480": "@pestipeti Thanks a lot. This saves me from continuous \"Notebook Exceeded Allowed Compute\" nightmares.",
    "730045": "I had same issue, but i fixed it after discovering this thread. Thank you a lot !!!",
    "717083": "I've also tried to reshape whole dataframe, but got kernel restarted. Any luck with it?",
    "706365": "Thanks for sharing Peter 👍\nI am able to resolve my \"Notebook Exceeded Allowed Compute\" error using your suggestion.\n\nRight now I am facing this \"Submission CSV Not Found\".  \nAny suggestion?",
    "706264": "What's the parquet_file? numpy? @pestipeti ",
    "704345": "Does the data fit into kaggle kernels RAM? @pestipeti\n\nI mean, I'm doing nearly the same as you advised but in Keras, and when i try to convert the whole dataframe into a numpy array i get memory error (while training).\n\nNow, assuming the test set is \"roughly the same size of the training set\", I assume that my submission will raise a memory error if doing as you suggested, isn't it?",
    "715207": "",
    "702934": "",
    "769706": "Thanks a lot!",
    "740362": "Thanks, now I can add more models!!!",
    "728902": "this is a great help - thank you!",
    "716799": "very useful, thanks your share!"
  }
}