{
  "id": 73007,
  "title": "Disk Bottlenecks during training",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73007",
  "author_name": "",
  "post_date": "2018-11-29T01:56:30.481020600Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'm trying to diagnose slow performance during training despite minimal CPU/GPU usage. It seems like my HDD is the bottleneck. I suppose one option is to transfer my data to my smaller SSD and see if that helps.</p>\n\n<p>However, I did notice on my resource monitor that python is only my third highest total B/sec disk activity... the highest is MsMpEng.exe, which seems to belong to windows defender. This process is accessing every image in my training folder, presumably scanning each image as python is loading it. I guess I could disable windows defender but that seems extreme. Has anyone had or solved this issue?</p>",
  "messages": [
    {
      "id": "429539",
      "postDate": "11/29/2018 01:56:30",
      "content": "<p>I'm trying to diagnose slow performance during training despite minimal CPU/GPU usage. It seems like my HDD is the bottleneck. I suppose one option is to transfer my data to my smaller SSD and see if that helps.</p>\n\n<p>However, I did notice on my resource monitor that python is only my third highest total B/sec disk activity... the highest is MsMpEng.exe, which seems to belong to windows defender. This process is accessing every image in my training folder, presumably scanning each image as python is loading it. I guess I could disable windows defender but that seems extreme. Has anyone had or solved this issue?</p>",
      "rawMarkdown": "I'm trying to diagnose slow performance during training despite minimal CPU/GPU usage. It seems like my HDD is the bottleneck. I suppose one option is to transfer my data to my smaller SSD and see if that helps.\n\nHowever, I did notice on my resource monitor that python is only my third highest total B/sec disk activity... the highest is MsMpEng.exe, which seems to belong to windows defender. This process is accessing every image in my training folder, presumably scanning each image as python is loading it. I guess I could disable windows defender but that seems extreme. Has anyone had or solved this issue?",
      "votes": null
    },
    {
      "id": "429660",
      "postDate": "11/29/2018 06:30:48",
      "content": "<p>Sorry can't help you with the Windows issue, but I just want to confirm that HDD is not an option... Good luck!</p>",
      "rawMarkdown": "Sorry can't help you with the Windows issue, but I just want to confirm that HDD is not an option... Good luck!",
      "votes": null
    },
    {
      "id": "429796",
      "postDate": "11/29/2018 10:49:51",
      "content": "<p>Tell Windows Defender to exclude your image directory: <a href=\"https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security\">https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security</a></p>",
      "rawMarkdown": "Tell Windows Defender to exclude your image directory: https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security",
      "votes": null
    },
    {
      "id": "430003",
      "postDate": "11/29/2018 16:43:29",
      "content": "<p>I wrote a script to reduce the amount of disk access during training.  Try that and see if it helps.</p>\n\n<p><a href=\"https://www.kaggle.com/robertkag/fast-image-loading\">https://www.kaggle.com/robertkag/fast-image-loading</a></p>",
      "rawMarkdown": "I wrote a script to reduce the amount of disk access during training.  Try that and see if it helps.\n\nhttps://www.kaggle.com/robertkag/fast-image-loading",
      "votes": null
    },
    {
      "id": "430038",
      "postDate": "11/29/2018 17:55:54",
      "content": "<p>Disk is definitely a bottleneck for me. Right now my training is running at a continuous 100MB/sec from and SSD. If I freeze the larger convolution layers, I can pin the SSD to 250MB/sec. I've got a new drive coming in the mail to help with this, a NVME SSD.</p>",
      "rawMarkdown": "Disk is definitely a bottleneck for me. Right now my training is running at a continuous 100MB/sec from and SSD. If I freeze the larger convolution layers, I can pin the SSD to 250MB/sec. I've got a new drive coming in the mail to help with this, a NVME SSD.",
      "votes": null
    },
    {
      "id": "430871",
      "postDate": "12/01/2018 04:50:04",
      "content": "<p>Highly recommend trying to reformat the data to hdf5 or tfrecord in order to load it in faster. It took me from 1200s epochs to 700s on the 512x512 images. I also have the benefit of having a smasung 970 evo, but still it makes a huge difference. </p>",
      "rawMarkdown": "Highly recommend trying to reformat the data to hdf5 or tfrecord in order to load it in faster. It took me from 1200s epochs to 700s on the 512x512 images. I also have the benefit of having a smasung 970 evo, but still it makes a huge difference.",
      "votes": null
    },
    {
      "id": "431197",
      "postDate": "12/01/2018 20:05:59",
      "content": "<p>I'll have to give that a try and see if that helps. I was using a samsung 840 pro, now upgraded to a HP EX900 NVMe drive that is amazingly fast so far.</p>",
      "rawMarkdown": "I'll have to give that a try and see if that helps. I was using a samsung 840 pro, now upgraded to a HP EX900 NVMe drive that is amazingly fast so far.",
      "votes": null
    },
    {
      "id": "432017",
      "postDate": "12/03/2018 09:09:08",
      "content": "<p>I am using PIL.Image.open instead of opencv or skimage and it is a lot faster.</p>",
      "rawMarkdown": "I am using PIL.Image.open instead of opencv or skimage and it is a lot faster.",
      "votes": null
    },
    {
      "id": "432018",
      "postDate": "12/03/2018 09:09:42",
      "content": "<p>How did you convert to HDF5? Do you mean to compress the image folder to HDF5? Thanks.</p>",
      "rawMarkdown": "How did you convert to HDF5? Do you mean to compress the image folder to HDF5? Thanks.",
      "votes": null
    },
    {
      "id": "432364",
      "postDate": "12/03/2018 19:14:38",
      "content": "<p>I would use h5py, save that to a single file:</p>\n\n<pre>import h5py\nf = h5py.file('somefile.h5','w')\n</pre>\n\n<p>Then for each image:</p>\n\n<pre>img = cv2.imread('path to img',cv2.IMREAD_UNCHANGED)\ndset = f.create_dataset(\"image id/name or whatever\", img)\n</pre>",
      "rawMarkdown": "I would use h5py, save that to a single file:\n<pre>import h5py\nf = h5py.file('somefile.h5','w')\n</pre>\n\nThen for each image:\n<pre>img = cv2.imread('path to img',cv2.IMREAD_UNCHANGED)\ndset = f.create_dataset(\"image id/name or whatever\", img)\n</pre>",
      "votes": null
    },
    {
      "id": "433132",
      "postDate": "12/04/2018 18:37:49",
      "content": "<p>I'll post a simple kernel later today to help you guys. Dont think I will be further pursuing this competition so dont want this work to go to waste. </p>",
      "rawMarkdown": "I'll post a simple kernel later today to help you guys. Dont think I will be further pursuing this competition so dont want this work to go to waste.",
      "votes": null
    },
    {
      "id": "436371",
      "postDate": "12/10/2018 07:35:00",
      "content": "<p><a href=\"/ryches\">@ryches</a>, I finally came back to this one as I just upgraded my video card and became CPU limited on the image loading. I'm taking my generator that converts everything to the right format and storing them as float16 with this code:</p>\n\n<pre>import h5py\nimport gc\nfrom tqdm import tqdm_notebook\ngc.collect()\nwith h5py.File('K:/data/hpa/512.hdf5', 'w') as f:\n    for (imgs,imgids) in tqdm_notebook(valid_gen):\n        for (img,imgid) in tqdm_notebook(zip(imgs,imgids)):\n            try:\n                f.create_dataset(imgid, data=img, shape=img.shape, maxshape=img.shape, compression='lzf', dtype='float16')\n            except:\n                print(\"Error on id: \" + str(imgid))\n        del imgs\n        del imgids\n        gc.collect()\n</pre>\n\n<p>While that is running I adapted a generator example to load from the saved file. Haven't tested this part yet though...</p>\n\n<pre>class DataGenerator(keras.utils.Sequence):\n    'Generates data for Keras'\n    def __init__(self, in_df, y_col=\"target_vec_float\", batch_size=16,shuffle=True):\n        'Initialization'\n        self.in_df = in_df\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.f = h5py.File('K:/data/hpa/512.hdf5', 'r')\n        self.on_epoch_end()\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(self.in_df.shape[0] / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        # Generate data\n        X, y = self.__data_generation(indexes)\n        return X, y\n\n    def on_epoch_end(self):\n        'Updates indexes after each epoch'\n        self.indexes = np.arange(self.in_df.shape[0])\n        if self.shuffle == True:\n            np.random.shuffle(self.indexes)\n\n    def __data_generation(self, indices):\n        'Generates data containing batch_size samples' # X : (n_samples, *dim, n_channels)\n        # Initialization\n        batch_data = np.empty((self.batch_size, desired_height, desired_width, nb_channels), dtype=np.float32)\n        batch_labels = []\n\n        # Generate data\n        for j, idx in enumerate(next_batch):\n            img_id = data.iloc[idx]['Id']\n\n            img = self.f[img_id]\n            label = data.iloc[idx][y_col]\n\n            batch_data[j] = img\n            batch_labels.append(label)\n\n        batch_labels = np.array(batch_labels).astype(np.float32)\n        return batch_data, batch_labels\n</pre>",
      "rawMarkdown": "ryches, I finally came back to this one as I just upgraded my video card and became CPU limited on the image loading. I'm taking my generator that converts everything to the right format and storing them as float16 with this code:\n<pre>import h5py\nimport gc\nfrom tqdm import tqdm_notebook\ngc.collect()\nwith h5py.File('K:/data/hpa/512.hdf5', 'w') as f:\n    for (imgs,imgids) in tqdm_notebook(valid_gen):\n        for (img,imgid) in tqdm_notebook(zip(imgs,imgids)):\n            try:\n                f.create_dataset(imgid, data=img, shape=img.shape, maxshape=img.shape, compression='lzf', dtype='float16')\n            except:\n                print(\"Error on id: \" + str(imgid))\n        del imgs\n        del imgids\n        gc.collect()\n</pre>\n\nWhile that is running I adapted a generator example to load from the saved file. Haven't tested this part yet though...\n\n<pre>class DataGenerator(keras.utils.Sequence):\n    'Generates data for Keras'\n    def __init__(self, in_df, y_col=\"target_vec_float\", batch_size=16,shuffle=True):\n        'Initialization'\n        self.in_df = in_df\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.f = h5py.File('K:/data/hpa/512.hdf5', 'r')\n        self.on_epoch_end()\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(self.in_df.shape[0] / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        # Generate data\n        X, y = self.__data_generation(indexes)\n        return X, y\n\n    def on_epoch_end(self):\n        'Updates indexes after each epoch'\n        self.indexes = np.arange(self.in_df.shape[0])\n        if self.shuffle == True:\n            np.random.shuffle(self.indexes)\n\n    def __data_generation(self, indices):\n        'Generates data containing batch_size samples' # X : (n_samples, *dim, n_channels)\n        # Initialization\n        batch_data = np.empty((self.batch_size, desired_height, desired_width, nb_channels), dtype=np.float32)\n        batch_labels = []\n\n        # Generate data\n        for j, idx in enumerate(next_batch):\n            img_id = data.iloc[idx]['Id']\n\n            img = self.f[img_id]\n            label = data.iloc[idx][y_col]\n\n            batch_data[j] = img\n            batch_labels.append(label)\n \n        batch_labels = np.array(batch_labels).astype(np.float32)\n        return batch_data, batch_labels\n</pre>",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 429660,
      "author_name": "stecasasso",
      "author_url": "",
      "post_date": "11/29/2018 06:30:48",
      "content": "<p>Sorry can't help you with the Windows issue, but I just want to confirm that HDD is not an option... Good luck!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 429796,
      "author_name": "raojx1",
      "author_url": "",
      "post_date": "11/29/2018 10:49:51",
      "content": "<p>Tell Windows Defender to exclude your image directory: <a href=\"https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security\">https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 430003,
      "author_name": "robertkag",
      "author_url": "",
      "post_date": "11/29/2018 16:43:29",
      "content": "<p>I wrote a script to reduce the amount of disk access during training.  Try that and see if it helps.</p>\n\n<p><a href=\"https://www.kaggle.com/robertkag/fast-image-loading\">https://www.kaggle.com/robertkag/fast-image-loading</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 430038,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/29/2018 17:55:54",
      "content": "<p>Disk is definitely a bottleneck for me. Right now my training is running at a continuous 100MB/sec from and SSD. If I freeze the larger convolution layers, I can pin the SSD to 250MB/sec. I've got a new drive coming in the mail to help with this, a NVME SSD.</p>",
      "votes": null,
      "replies": [
        {
          "id": 430871,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/01/2018 04:50:04",
          "content": "<p>Highly recommend trying to reformat the data to hdf5 or tfrecord in order to load it in faster. It took me from 1200s epochs to 700s on the 512x512 images. I also have the benefit of having a smasung 970 evo, but still it makes a huge difference. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 431197,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/01/2018 20:05:59",
          "content": "<p>I'll have to give that a try and see if that helps. I was using a samsung 840 pro, now upgraded to a HP EX900 NVMe drive that is amazingly fast so far.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432018,
          "author_name": "lpachuong",
          "author_url": "",
          "post_date": "12/03/2018 09:09:42",
          "content": "<p>How did you convert to HDF5? Do you mean to compress the image folder to HDF5? Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 432364,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/03/2018 19:14:38",
          "content": "<p>I would use h5py, save that to a single file:</p>\n\n<pre>import h5py\nf = h5py.file('somefile.h5','w')\n</pre>\n\n<p>Then for each image:</p>\n\n<pre>img = cv2.imread('path to img',cv2.IMREAD_UNCHANGED)\ndset = f.create_dataset(\"image id/name or whatever\", img)\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 433132,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/04/2018 18:37:49",
          "content": "<p>I'll post a simple kernel later today to help you guys. Dont think I will be further pursuing this competition so dont want this work to go to waste. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436371,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/10/2018 07:35:00",
          "content": "<p><a href=\"/ryches\">@ryches</a>, I finally came back to this one as I just upgraded my video card and became CPU limited on the image loading. I'm taking my generator that converts everything to the right format and storing them as float16 with this code:</p>\n\n<pre>import h5py\nimport gc\nfrom tqdm import tqdm_notebook\ngc.collect()\nwith h5py.File('K:/data/hpa/512.hdf5', 'w') as f:\n    for (imgs,imgids) in tqdm_notebook(valid_gen):\n        for (img,imgid) in tqdm_notebook(zip(imgs,imgids)):\n            try:\n                f.create_dataset(imgid, data=img, shape=img.shape, maxshape=img.shape, compression='lzf', dtype='float16')\n            except:\n                print(\"Error on id: \" + str(imgid))\n        del imgs\n        del imgids\n        gc.collect()\n</pre>\n\n<p>While that is running I adapted a generator example to load from the saved file. Haven't tested this part yet though...</p>\n\n<pre>class DataGenerator(keras.utils.Sequence):\n    'Generates data for Keras'\n    def __init__(self, in_df, y_col=\"target_vec_float\", batch_size=16,shuffle=True):\n        'Initialization'\n        self.in_df = in_df\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.f = h5py.File('K:/data/hpa/512.hdf5', 'r')\n        self.on_epoch_end()\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(self.in_df.shape[0] / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        # Generate data\n        X, y = self.__data_generation(indexes)\n        return X, y\n\n    def on_epoch_end(self):\n        'Updates indexes after each epoch'\n        self.indexes = np.arange(self.in_df.shape[0])\n        if self.shuffle == True:\n            np.random.shuffle(self.indexes)\n\n    def __data_generation(self, indices):\n        'Generates data containing batch_size samples' # X : (n_samples, *dim, n_channels)\n        # Initialization\n        batch_data = np.empty((self.batch_size, desired_height, desired_width, nb_channels), dtype=np.float32)\n        batch_labels = []\n\n        # Generate data\n        for j, idx in enumerate(next_batch):\n            img_id = data.iloc[idx]['Id']\n\n            img = self.f[img_id]\n            label = data.iloc[idx][y_col]\n\n            batch_data[j] = img\n            batch_labels.append(label)\n\n        batch_labels = np.array(batch_labels).astype(np.float32)\n        return batch_data, batch_labels\n</pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 432017,
      "author_name": "lpachuong",
      "author_url": "",
      "post_date": "12/03/2018 09:09:08",
      "content": "<p>I am using PIL.Image.open instead of opencv or skimage and it is a lot faster.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "429539": "I'm trying to diagnose slow performance during training despite minimal CPU/GPU usage. It seems like my HDD is the bottleneck. I suppose one option is to transfer my data to my smaller SSD and see if that helps.\n\nHowever, I did notice on my resource monitor that python is only my third highest total B/sec disk activity... the highest is MsMpEng.exe, which seems to belong to windows defender. This process is accessing every image in my training folder, presumably scanning each image as python is loading it. I guess I could disable windows defender but that seems extreme. Has anyone had or solved this issue?",
    "429660": "Sorry can't help you with the Windows issue, but I just want to confirm that HDD is not an option... Good luck!",
    "429796": "Tell Windows Defender to exclude your image directory: https://support.microsoft.com/en-us/help/4028485/windows-10-add-an-exclusion-to-windows-security",
    "430003": "I wrote a script to reduce the amount of disk access during training.  Try that and see if it helps.\n\nhttps://www.kaggle.com/robertkag/fast-image-loading",
    "430038": "Disk is definitely a bottleneck for me. Right now my training is running at a continuous 100MB/sec from and SSD. If I freeze the larger convolution layers, I can pin the SSD to 250MB/sec. I've got a new drive coming in the mail to help with this, a NVME SSD.",
    "430871": "Highly recommend trying to reformat the data to hdf5 or tfrecord in order to load it in faster. It took me from 1200s epochs to 700s on the 512x512 images. I also have the benefit of having a smasung 970 evo, but still it makes a huge difference.",
    "431197": "I'll have to give that a try and see if that helps. I was using a samsung 840 pro, now upgraded to a HP EX900 NVMe drive that is amazingly fast so far.",
    "432017": "I am using PIL.Image.open instead of opencv or skimage and it is a lot faster.",
    "432018": "How did you convert to HDF5? Do you mean to compress the image folder to HDF5? Thanks.",
    "432364": "I would use h5py, save that to a single file:\n<pre>import h5py\nf = h5py.file('somefile.h5','w')\n</pre>\n\nThen for each image:\n<pre>img = cv2.imread('path to img',cv2.IMREAD_UNCHANGED)\ndset = f.create_dataset(\"image id/name or whatever\", img)\n</pre>",
    "433132": "I'll post a simple kernel later today to help you guys. Dont think I will be further pursuing this competition so dont want this work to go to waste.",
    "436371": "ryches, I finally came back to this one as I just upgraded my video card and became CPU limited on the image loading. I'm taking my generator that converts everything to the right format and storing them as float16 with this code:\n<pre>import h5py\nimport gc\nfrom tqdm import tqdm_notebook\ngc.collect()\nwith h5py.File('K:/data/hpa/512.hdf5', 'w') as f:\n    for (imgs,imgids) in tqdm_notebook(valid_gen):\n        for (img,imgid) in tqdm_notebook(zip(imgs,imgids)):\n            try:\n                f.create_dataset(imgid, data=img, shape=img.shape, maxshape=img.shape, compression='lzf', dtype='float16')\n            except:\n                print(\"Error on id: \" + str(imgid))\n        del imgs\n        del imgids\n        gc.collect()\n</pre>\n\nWhile that is running I adapted a generator example to load from the saved file. Haven't tested this part yet though...\n\n<pre>class DataGenerator(keras.utils.Sequence):\n    'Generates data for Keras'\n    def __init__(self, in_df, y_col=\"target_vec_float\", batch_size=16,shuffle=True):\n        'Initialization'\n        self.in_df = in_df\n        self.batch_size = batch_size\n        self.shuffle = shuffle\n        self.f = h5py.File('K:/data/hpa/512.hdf5', 'r')\n        self.on_epoch_end()\n\n    def __len__(self):\n        'Denotes the number of batches per epoch'\n        return int(np.floor(self.in_df.shape[0] / self.batch_size))\n\n    def __getitem__(self, index):\n        'Generate one batch of data'\n        # Generate indexes of the batch\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        # Generate data\n        X, y = self.__data_generation(indexes)\n        return X, y\n\n    def on_epoch_end(self):\n        'Updates indexes after each epoch'\n        self.indexes = np.arange(self.in_df.shape[0])\n        if self.shuffle == True:\n            np.random.shuffle(self.indexes)\n\n    def __data_generation(self, indices):\n        'Generates data containing batch_size samples' # X : (n_samples, *dim, n_channels)\n        # Initialization\n        batch_data = np.empty((self.batch_size, desired_height, desired_width, nb_channels), dtype=np.float32)\n        batch_labels = []\n\n        # Generate data\n        for j, idx in enumerate(next_batch):\n            img_id = data.iloc[idx]['Id']\n\n            img = self.f[img_id]\n            label = data.iloc[idx][y_col]\n\n            batch_data[j] = img\n            batch_labels.append(label)\n \n        batch_labels = np.array(batch_labels).astype(np.float32)\n        return batch_data, batch_labels\n</pre>"
  },
  "source": "meta"
}