{
  "id": 126667,
  "title": "Keras data generator doesn't work in multiprocessing.",
  "url": "/competitions/bengaliai-cv19/discussion/126667",
  "author_name": "",
  "post_date": "2020-01-19T08:31:43.693836300Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Since I could not put all the data in memory, I wrote my own data generator.\n(update: I fixed some lines but it still don`t works.)</p>\n\n<p>```\nclass data_generator(Sequence):\n    def <strong>init</strong>(self,\n                 y_root,\n                 y_vowel,\n                 y_consonant,\n                 transformer,\n                 idxs = None,\n                 batch_size=16,\n                 image_shape=(64, 64, 1),\n                 random_seed=2434,\n                 shuffle=True):\n        self.y_root = y_root\n        self.y_vowel = y_vowel\n        self.y_consonant = y_consonant\n        np.random.seed(random_seed)\n        pyrandom.seed(random_seed)</p>\n\n<pre><code>    if idxs is None:\n        self.idxs = np.arange(len(y_root))\n    else:\n        self.idxs = np.array(idxs)\n    if shuffle:\n        np.random.shuffle(self.idxs)\n    self.batch_size = batch_size\n    self.image_shape = image_shape\n    self.transformer = transformer\n    self.shuffle = True\n\ndef __len__(self):\n    return -(-len(self.idxs) // self.batch_size)\n\n\ndef __getitem__(self, i):\n    batch_idxs = self.idxs[self.batch_size*i: self.batch_size*(i+1)]\n    batch_images = np.zeros((len(batch_idxs), self.image_shape[0], self.image_shape[1], 1), dtype=np.float32)\n    batch_labels = {}\n    batch_labels[\"root\"] = self.y_root[batch_idxs]\n    batch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\n    batch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n\n    for i, batch_idx in enumerate(batch_idxs):\n        batch_images[i] = self.augment(self.load_image(i))\n\n    return batch_images, batch_labels\n\ndef load_image(self, i):\n    return np.load(f\"./input/{i}.npy\")\n\ndef augment(self, image):\n    return self.transformer.transform(image)\n\ndef on_epoch_end(self):\n    if self.shuffle:\n        np.random.shuffle(self.idxs)\ndef on_epoch_end(self):\n    if shuffle:\n        np.random.shuffle(self.idxs)\n</code></pre>\n\n<p>```\nAnd this is used as below</p>\n\n<p><code>\nhistory = model.fit_generator(dev_generator,\n                              validation_data = val_generator,\n                              epochs = epochs,\n                              use_multiprocessing=True,\n                              workers=4)\n</code>\nWhen use_multiprocessing = False it works, but when use_multiprocessing = True it raise BrokenPipeError. \nI think there is a cause around pickle, but I have not been able to identify it yet. Please let me know if you have any ideas.</p>",
  "messages": [
    {
      "id": "722875",
      "postDate": "01/19/2020 08:31:43",
      "content": "<p>Since I could not put all the data in memory, I wrote my own data generator.\n(update: I fixed some lines but it still don`t works.)</p>\n\n<p>```\nclass data_generator(Sequence):\n    def <strong>init</strong>(self,\n                 y_root,\n                 y_vowel,\n                 y_consonant,\n                 transformer,\n                 idxs = None,\n                 batch_size=16,\n                 image_shape=(64, 64, 1),\n                 random_seed=2434,\n                 shuffle=True):\n        self.y_root = y_root\n        self.y_vowel = y_vowel\n        self.y_consonant = y_consonant\n        np.random.seed(random_seed)\n        pyrandom.seed(random_seed)</p>\n\n<pre><code>    if idxs is None:\n        self.idxs = np.arange(len(y_root))\n    else:\n        self.idxs = np.array(idxs)\n    if shuffle:\n        np.random.shuffle(self.idxs)\n    self.batch_size = batch_size\n    self.image_shape = image_shape\n    self.transformer = transformer\n    self.shuffle = True\n\ndef __len__(self):\n    return -(-len(self.idxs) // self.batch_size)\n\n\ndef __getitem__(self, i):\n    batch_idxs = self.idxs[self.batch_size*i: self.batch_size*(i+1)]\n    batch_images = np.zeros((len(batch_idxs), self.image_shape[0], self.image_shape[1], 1), dtype=np.float32)\n    batch_labels = {}\n    batch_labels[\"root\"] = self.y_root[batch_idxs]\n    batch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\n    batch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n\n    for i, batch_idx in enumerate(batch_idxs):\n        batch_images[i] = self.augment(self.load_image(i))\n\n    return batch_images, batch_labels\n\ndef load_image(self, i):\n    return np.load(f\"./input/{i}.npy\")\n\ndef augment(self, image):\n    return self.transformer.transform(image)\n\ndef on_epoch_end(self):\n    if self.shuffle:\n        np.random.shuffle(self.idxs)\ndef on_epoch_end(self):\n    if shuffle:\n        np.random.shuffle(self.idxs)\n</code></pre>\n\n<p>```\nAnd this is used as below</p>\n\n<p><code>\nhistory = model.fit_generator(dev_generator,\n                              validation_data = val_generator,\n                              epochs = epochs,\n                              use_multiprocessing=True,\n                              workers=4)\n</code>\nWhen use_multiprocessing = False it works, but when use_multiprocessing = True it raise BrokenPipeError. \nI think there is a cause around pickle, but I have not been able to identify it yet. Please let me know if you have any ideas.</p>",
      "rawMarkdown": "Since I could not put all the data in memory, I wrote my own data generator.\n(update: I fixed some lines but it still don`t works.)\n\n```\nclass data_generator(Sequence):\n    def __init__(self,\n                 y_root,\n                 y_vowel,\n                 y_consonant,\n                 transformer,\n                 idxs = None,\n                 batch_size=16,\n                 image_shape=(64, 64, 1),\n                 random_seed=2434,\n                 shuffle=True):\n        self.y_root = y_root\n        self.y_vowel = y_vowel\n        self.y_consonant = y_consonant\n        np.random.seed(random_seed)\n        pyrandom.seed(random_seed)\n        \n        if idxs is None:\n            self.idxs = np.arange(len(y_root))\n        else:\n            self.idxs = np.array(idxs)\n        if shuffle:\n            np.random.shuffle(self.idxs)\n        self.batch_size = batch_size\n        self.image_shape = image_shape\n        self.transformer = transformer\n        self.shuffle = True\n        \n    def __len__(self):\n        return -(-len(self.idxs) // self.batch_size)\n    \n\n    def __getitem__(self, i):\n        batch_idxs = self.idxs[self.batch_size*i: self.batch_size*(i+1)]\n        batch_images = np.zeros((len(batch_idxs), self.image_shape[0], self.image_shape[1], 1), dtype=np.float32)\n        batch_labels = {}\n        batch_labels[\"root\"] = self.y_root[batch_idxs]\n        batch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\n        batch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n        \n        for i, batch_idx in enumerate(batch_idxs):\n            batch_images[i] = self.augment(self.load_image(i))\n\n        return batch_images, batch_labels\n\n    def load_image(self, i):\n        return np.load(f\"./input/{i}.npy\")\n    \n    def augment(self, image):\n        return self.transformer.transform(image)\n\n    def on_epoch_end(self):\n        if self.shuffle:\n            np.random.shuffle(self.idxs)\n    def on_epoch_end(self):\n        if shuffle:\n            np.random.shuffle(self.idxs)\n```\nAnd this is used as below\n\n```\nhistory = model.fit_generator(dev_generator,\n                              validation_data = val_generator,\n                              epochs = epochs,\n                              use_multiprocessing=True,\n                              workers=4)\n```\nWhen use\\_multiprocessing = False it works, but when use\\_multiprocessing = True it raise BrokenPipeError. \nI think there is a cause around pickle, but I have not been able to identify it yet. Please let me know if you have any ideas.",
      "votes": null
    },
    {
      "id": "722897",
      "postDate": "01/19/2020 09:27:31",
      "content": "<p>What a strange way of getting the number of steps:</p>\n\n<p><code>return -(-len(self.idxs) // self.batch_size)</code></p>\n\n<p>I mean why 2 minus and not simply <code>len(self.idxs)</code> // self.batch_size?`</p>\n\n<p>Anyway, Don't know if it is the cause of the problem but when you do</p>\n\n<p><code>self.idxs = np.arange(len(X))</code></p>\n\n<p>the variable X in the generator is never declared (this is not a problem when you run your notebook but probably it will be when you call fit_generator).</p>\n\n<p>Finally, my advice is to not use dataframe in keras data generator as it could be non-pickable. Try to convert all your dataframes / dictionaries into numpy arrays or lists and try if it works</p>",
      "rawMarkdown": "What a strange way of getting the number of steps:\n\n`return -(-len(self.idxs) // self.batch_size)`\n\nI mean why 2 minus and not simply `len(self.idxs)` // self.batch_size?`\n\nAnyway, Don't know if it is the cause of the problem but when you do\n\n`self.idxs = np.arange(len(X))`\n\nthe variable X in the generator is never declared (this is not a problem when you run your notebook but probably it will be when you call fit_generator).\n\nFinally, my advice is to not use dataframe in keras data generator as it could be non-pickable. Try to convert all your dataframes / dictionaries into numpy arrays or lists and try if it works",
      "votes": null
    },
    {
      "id": "722898",
      "postDate": "01/19/2020 09:28:54",
      "content": "<p>GPU on??\nIf so, you can use only two workers..</p>",
      "rawMarkdown": "GPU on??\nIf so, you can use only two workers..",
      "votes": null
    },
    {
      "id": "722917",
      "postDate": "01/19/2020 09:54:00",
      "content": "<p>I run this code in my local environment.</p>",
      "rawMarkdown": "I run this code in my local environment.",
      "votes": null
    },
    {
      "id": "722925",
      "postDate": "01/19/2020 10:10:46",
      "content": "<p><code>-(-len(self.idxs) // self.batch_size)</code>\n is equivalent to\n<code>ceil(len(self.idxs) / self.batch_size)</code></p>\n\n<p>And I fixed \n<code>self.idxs = np.arange(len(X))</code>\nto\n<code>self.idxs = np.arange(len(y_root))</code>\nbut it still doesn`t works...</p>",
      "rawMarkdown": "`-(-len(self.idxs) // self.batch_size)`\n is equivalent to\n`ceil(len(self.idxs) / self.batch_size)`\n\nAnd I fixed \n`self.idxs = np.arange(len(X))`\nto\n`self.idxs = np.arange(len(y_root))`\nbut it still doesn`t works...",
      "votes": null
    },
    {
      "id": "722974",
      "postDate": "01/19/2020 11:07:30",
      "content": "<p>You can use multiprocessing.cpu_count() to get the number of workers from the machine</p>",
      "rawMarkdown": "You can use multiprocessing.cpu_count() to get the number of workers from the machine",
      "votes": null
    },
    {
      "id": "722977",
      "postDate": "01/19/2020 11:09:04",
      "content": "<p>Try to use a list instead of a dictionary here\n<code>\nbatch_labels = {}\nbatch_labels[\"root\"] = self.y_root[batch_idxs]\nbatch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\nbatch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n</code></p>\n\n<p>I implemented getitem in this way:</p>\n\n<p>```\ndef <strong>getitem</strong>(self, step):</p>\n\n<pre><code>   current_images = self.images[step * self.batch_size : \\\n                          (step + 1) * self.batch_size]\n\n    y1 = self.y1[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    y2 = self.y2[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    y3 = self.y3[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    X, y1, y2, y3 = self.__generate_batch(current_images,\n                                          y1, y2, y3,\n                                          step)\n  return X, [y1, y2, y3]\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Try to use a list instead of a dictionary here\n```\nbatch_labels = {}\nbatch_labels[\"root\"] = self.y_root[batch_idxs]\nbatch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\nbatch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n```\n\nI implemented getitem in this way:\n\n```\ndef __getitem__(self, step):\n        \n       current_images = self.images[step * self.batch_size : \\\n                              (step + 1) * self.batch_size]\n\n        y1 = self.y1[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n        \n        y2 = self.y2[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n        \n        y3 = self.y3[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n\n        X, y1, y2, y3 = self.__generate_batch(current_images,\n                                              y1, y2, y3,\n                                              step)\n      return X, [y1, y2, y3]\n\n```",
      "votes": null
    },
    {
      "id": "722994",
      "postDate": "01/19/2020 11:24:26",
      "content": "<p>t looks like an error caused by windows. keras data_generator multiprocessing does not seem to support windows. Thank you for helping with the investigation. I will give up this method and look for another method.</p>",
      "rawMarkdown": "t looks like an error caused by windows. keras data_generator multiprocessing does not seem to support windows. Thank you for helping with the investigation. I will give up this method and look for another method.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722897,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "01/19/2020 09:27:31",
      "content": "<p>What a strange way of getting the number of steps:</p>\n\n<p><code>return -(-len(self.idxs) // self.batch_size)</code></p>\n\n<p>I mean why 2 minus and not simply <code>len(self.idxs)</code> // self.batch_size?`</p>\n\n<p>Anyway, Don't know if it is the cause of the problem but when you do</p>\n\n<p><code>self.idxs = np.arange(len(X))</code></p>\n\n<p>the variable X in the generator is never declared (this is not a problem when you run your notebook but probably it will be when you call fit_generator).</p>\n\n<p>Finally, my advice is to not use dataframe in keras data generator as it could be non-pickable. Try to convert all your dataframes / dictionaries into numpy arrays or lists and try if it works</p>",
      "votes": null,
      "replies": [
        {
          "id": 722925,
          "author_name": "nadare",
          "author_url": "",
          "post_date": "01/19/2020 10:10:46",
          "content": "<p><code>-(-len(self.idxs) // self.batch_size)</code>\n is equivalent to\n<code>ceil(len(self.idxs) / self.batch_size)</code></p>\n\n<p>And I fixed \n<code>self.idxs = np.arange(len(X))</code>\nto\n<code>self.idxs = np.arange(len(y_root))</code>\nbut it still doesn`t works...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722977,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "01/19/2020 11:09:04",
          "content": "<p>Try to use a list instead of a dictionary here\n<code>\nbatch_labels = {}\nbatch_labels[\"root\"] = self.y_root[batch_idxs]\nbatch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\nbatch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n</code></p>\n\n<p>I implemented getitem in this way:</p>\n\n<p>```\ndef <strong>getitem</strong>(self, step):</p>\n\n<pre><code>   current_images = self.images[step * self.batch_size : \\\n                          (step + 1) * self.batch_size]\n\n    y1 = self.y1[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    y2 = self.y2[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    y3 = self.y3[step * self.batch_size : \\\n                 (step + 1) * self.batch_size]\n\n    X, y1, y2, y3 = self.__generate_batch(current_images,\n                                          y1, y2, y3,\n                                          step)\n  return X, [y1, y2, y3]\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722994,
          "author_name": "nadare",
          "author_url": "",
          "post_date": "01/19/2020 11:24:26",
          "content": "<p>t looks like an error caused by windows. keras data_generator multiprocessing does not seem to support windows. Thank you for helping with the investigation. I will give up this method and look for another method.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 722898,
      "author_name": "inoueu1",
      "author_url": "",
      "post_date": "01/19/2020 09:28:54",
      "content": "<p>GPU on??\nIf so, you can use only two workers..</p>",
      "votes": null,
      "replies": [
        {
          "id": 722917,
          "author_name": "nadare",
          "author_url": "",
          "post_date": "01/19/2020 09:54:00",
          "content": "<p>I run this code in my local environment.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 722974,
          "author_name": "lazcoder",
          "author_url": "",
          "post_date": "01/19/2020 11:07:30",
          "content": "<p>You can use multiprocessing.cpu_count() to get the number of workers from the machine</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "722875": "Since I could not put all the data in memory, I wrote my own data generator.\n(update: I fixed some lines but it still don`t works.)\n\n```\nclass data_generator(Sequence):\n    def __init__(self,\n                 y_root,\n                 y_vowel,\n                 y_consonant,\n                 transformer,\n                 idxs = None,\n                 batch_size=16,\n                 image_shape=(64, 64, 1),\n                 random_seed=2434,\n                 shuffle=True):\n        self.y_root = y_root\n        self.y_vowel = y_vowel\n        self.y_consonant = y_consonant\n        np.random.seed(random_seed)\n        pyrandom.seed(random_seed)\n        \n        if idxs is None:\n            self.idxs = np.arange(len(y_root))\n        else:\n            self.idxs = np.array(idxs)\n        if shuffle:\n            np.random.shuffle(self.idxs)\n        self.batch_size = batch_size\n        self.image_shape = image_shape\n        self.transformer = transformer\n        self.shuffle = True\n        \n    def __len__(self):\n        return -(-len(self.idxs) // self.batch_size)\n    \n\n    def __getitem__(self, i):\n        batch_idxs = self.idxs[self.batch_size*i: self.batch_size*(i+1)]\n        batch_images = np.zeros((len(batch_idxs), self.image_shape[0], self.image_shape[1], 1), dtype=np.float32)\n        batch_labels = {}\n        batch_labels[\"root\"] = self.y_root[batch_idxs]\n        batch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\n        batch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n        \n        for i, batch_idx in enumerate(batch_idxs):\n            batch_images[i] = self.augment(self.load_image(i))\n\n        return batch_images, batch_labels\n\n    def load_image(self, i):\n        return np.load(f\"./input/{i}.npy\")\n    \n    def augment(self, image):\n        return self.transformer.transform(image)\n\n    def on_epoch_end(self):\n        if self.shuffle:\n            np.random.shuffle(self.idxs)\n    def on_epoch_end(self):\n        if shuffle:\n            np.random.shuffle(self.idxs)\n```\nAnd this is used as below\n\n```\nhistory = model.fit_generator(dev_generator,\n                              validation_data = val_generator,\n                              epochs = epochs,\n                              use_multiprocessing=True,\n                              workers=4)\n```\nWhen use\\_multiprocessing = False it works, but when use\\_multiprocessing = True it raise BrokenPipeError. \nI think there is a cause around pickle, but I have not been able to identify it yet. Please let me know if you have any ideas.",
    "722897": "What a strange way of getting the number of steps:\n\n`return -(-len(self.idxs) // self.batch_size)`\n\nI mean why 2 minus and not simply `len(self.idxs)` // self.batch_size?`\n\nAnyway, Don't know if it is the cause of the problem but when you do\n\n`self.idxs = np.arange(len(X))`\n\nthe variable X in the generator is never declared (this is not a problem when you run your notebook but probably it will be when you call fit_generator).\n\nFinally, my advice is to not use dataframe in keras data generator as it could be non-pickable. Try to convert all your dataframes / dictionaries into numpy arrays or lists and try if it works",
    "722898": "GPU on??\nIf so, you can use only two workers..",
    "722917": "I run this code in my local environment.",
    "722925": "`-(-len(self.idxs) // self.batch_size)`\n is equivalent to\n`ceil(len(self.idxs) / self.batch_size)`\n\nAnd I fixed \n`self.idxs = np.arange(len(X))`\nto\n`self.idxs = np.arange(len(y_root))`\nbut it still doesn`t works...",
    "722974": "You can use multiprocessing.cpu_count() to get the number of workers from the machine",
    "722977": "Try to use a list instead of a dictionary here\n```\nbatch_labels = {}\nbatch_labels[\"root\"] = self.y_root[batch_idxs]\nbatch_labels[\"vowel\"] = self.y_vowel[batch_idxs]\nbatch_labels[\"consonant\"] = self.y_consonant[batch_idxs]\n```\n\nI implemented getitem in this way:\n\n```\ndef __getitem__(self, step):\n        \n       current_images = self.images[step * self.batch_size : \\\n                              (step + 1) * self.batch_size]\n\n        y1 = self.y1[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n        \n        y2 = self.y2[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n        \n        y3 = self.y3[step * self.batch_size : \\\n                     (step + 1) * self.batch_size]\n\n        X, y1, y2, y3 = self.__generate_batch(current_images,\n                                              y1, y2, y3,\n                                              step)\n      return X, [y1, y2, y3]\n\n```",
    "722994": "t looks like an error caused by windows. keras data_generator multiprocessing does not seem to support windows. Thank you for helping with the investigation. I will give up this method and look for another method."
  },
  "source": "meta"
}