{
  "id": 211269,
  "title": "Code for tf.keras.utils.Sequence Prediction.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211269",
  "author_name": "",
  "post_date": "2021-01-14T13:12:41.798734700Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p><code>\ntest_df = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')\ntrain_df = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')\nDIM =300\nimport cv2\n</code></p>\n<p><code>class DataGenerator(tf.keras.utils.Sequence):\n    def __init__(self, path, list_IDs, labels, batch_size, img_size, img_channel):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.indexes = np.arange(len(self.list_IDs))</code><br>\n`</p>\n<p><code>def __len__(self):\n        return int(np.floor(len(self.list_IDs)/self.batch_size))\n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n</code></p>\n<pre><code>def __data_generation(self, list_IDs_temp):\n    X = np.empty((self.batch_size, self.img_size, self.img_size, self.img_channel)) #use with caution. Faster than np.zeros\n    #Return a new array of given shape and type, without initializing entries.\n    # Generates a numpy array size (1, 300, 300, 1) with\n\n\n    y = np.empty((self.batch_size, 5), dtype=int)  #generates a numpy array size (1 by 5)\n\n    for i, ID in enumerate(list_IDs_temp): #for all path id get int and ID #enumerate is like pandas iterrows. You can set the start too.\n        img = cv2.imread(self.path+ID) #read image\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) #convert from RGB to BGR since cv2 reads in RGB.\n        img = cv2.resize(img, (self.img_size, self.img_size)) #resize\n        X[i, ] = img\n        y[i, ] = self.labels[i]\n    X = X.astype('float32')\n    X /= 255.\n    return X, y\ndef preprocess_image(image_path, desired_size=DIM):\n    img = cv2.imread(image_path) #get image\n    img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR) #convert to RGB as cv2 reads in BGR\n    im = cv2.resize(img, (desired_size,desired_size), interpolation = cv2.INTER_AREA) #resize \n    im = np.array(im) #convert to a numpy array\n    return im\n</code></pre>\n<p><code>\ntest_generator = DataGenerator('../input/cassava-leaf-disease-classification/'+'train_images/',\n                               #test_df['image_id'],\n                               #test_df['label'], \n                               #test for many images\n                               train_df['image_id'],\n                               train_df['label'],\n                               1, \n                               DIM,\n                               3)\n</code></p>\n<h4>time a piece of code.</h4>\n<p><code>import time\nstart = time.process_time()\n</code></p>\n<h4>predict on test images.</h4>\n<p><code>\npreds = []\nprobabilities = []\npreds = model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)\nprobabilities.append(preds)\n</code><br>\n<code>\npredictions = np.array(probabilities).mean(0).argmax(-1)\nprint(predictions)\n</code></p>\n<h4>The code</h4>\n<p><code>print('Prediction time : ',time.process_time() - start)\n</code></p>",
  "messages": [
    {
      "id": "1152812",
      "postDate": "01/14/2021 13:12:41",
      "content": "<p><code>\ntest_df = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')\ntrain_df = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')\nDIM =300\nimport cv2\n</code></p>\n<p><code>class DataGenerator(tf.keras.utils.Sequence):\n    def __init__(self, path, list_IDs, labels, batch_size, img_size, img_channel):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.indexes = np.arange(len(self.list_IDs))</code><br>\n`</p>\n<p><code>def __len__(self):\n        return int(np.floor(len(self.list_IDs)/self.batch_size))\n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n</code></p>\n<pre><code>def __data_generation(self, list_IDs_temp):\n    X = np.empty((self.batch_size, self.img_size, self.img_size, self.img_channel)) #use with caution. Faster than np.zeros\n    #Return a new array of given shape and type, without initializing entries.\n    # Generates a numpy array size (1, 300, 300, 1) with\n\n\n    y = np.empty((self.batch_size, 5), dtype=int)  #generates a numpy array size (1 by 5)\n\n    for i, ID in enumerate(list_IDs_temp): #for all path id get int and ID #enumerate is like pandas iterrows. You can set the start too.\n        img = cv2.imread(self.path+ID) #read image\n        img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) #convert from RGB to BGR since cv2 reads in RGB.\n        img = cv2.resize(img, (self.img_size, self.img_size)) #resize\n        X[i, ] = img\n        y[i, ] = self.labels[i]\n    X = X.astype('float32')\n    X /= 255.\n    return X, y\ndef preprocess_image(image_path, desired_size=DIM):\n    img = cv2.imread(image_path) #get image\n    img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR) #convert to RGB as cv2 reads in BGR\n    im = cv2.resize(img, (desired_size,desired_size), interpolation = cv2.INTER_AREA) #resize \n    im = np.array(im) #convert to a numpy array\n    return im\n</code></pre>\n<p><code>\ntest_generator = DataGenerator('../input/cassava-leaf-disease-classification/'+'train_images/',\n                               #test_df['image_id'],\n                               #test_df['label'], \n                               #test for many images\n                               train_df['image_id'],\n                               train_df['label'],\n                               1, \n                               DIM,\n                               3)\n</code></p>\n<h4>time a piece of code.</h4>\n<p><code>import time\nstart = time.process_time()\n</code></p>\n<h4>predict on test images.</h4>\n<p><code>\npreds = []\nprobabilities = []\npreds = model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)\nprobabilities.append(preds)\n</code><br>\n<code>\npredictions = np.array(probabilities).mean(0).argmax(-1)\nprint(predictions)\n</code></p>\n<h4>The code</h4>\n<p><code>print('Prediction time : ',time.process_time() - start)\n</code></p>",
      "rawMarkdown": "`\ntest_df = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')\ntrain_df = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')\nDIM =300\nimport cv2\n`\n\n`class DataGenerator(tf.keras.utils.Sequence):\n    def __init__(self, path, list_IDs, labels, batch_size, img_size, img_channel):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.indexes = np.arange(len(self.list_IDs))`\n`\n\n`    def __len__(self):\n        return int(np.floor(len(self.list_IDs)/self.batch_size))\n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n`\n    \n    def __data_generation(self, list_IDs_temp):\n        X = np.empty((self.batch_size, self.img_size, self.img_size, self.img_channel)) #use with caution. Faster than np.zeros\n        #Return a new array of given shape and type, without initializing entries.\n        # Generates a numpy array size (1, 300, 300, 1) with\n        \n\n        y = np.empty((self.batch_size, 5), dtype=int)  #generates a numpy array size (1 by 5)\n\n        for i, ID in enumerate(list_IDs_temp): #for all path id get int and ID #enumerate is like pandas iterrows. You can set the start too.\n            img = cv2.imread(self.path+ID) #read image\n            img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) #convert from RGB to BGR since cv2 reads in RGB.\n            img = cv2.resize(img, (self.img_size, self.img_size)) #resize\n            X[i, ] = img\n            y[i, ] = self.labels[i]\n        X = X.astype('float32')\n        X /= 255.\n        return X, y\n    def preprocess_image(image_path, desired_size=DIM):\n        img = cv2.imread(image_path) #get image\n        img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR) #convert to RGB as cv2 reads in BGR\n        im = cv2.resize(img, (desired_size,desired_size), interpolation = cv2.INTER_AREA) #resize \n        im = np.array(im) #convert to a numpy array\n        return im\n`\ntest_generator = DataGenerator('../input/cassava-leaf-disease-classification/'+'train_images/',\n                               #test_df['image_id'],\n                               #test_df['label'], \n                               #test for many images\n                               train_df['image_id'],\n                               train_df['label'],\n                               1, \n                               DIM,\n                               3)\n`\n#### time a piece of code.\n`import time\nstart = time.process_time()\n`\n#### predict on test images.\n`\npreds = []\nprobabilities = []\npreds = model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)\nprobabilities.append(preds)\n`\n`\npredictions = np.array(probabilities).mean(0).argmax(-1)\nprint(predictions)\n`\n\n#### The code\n`print('Prediction time : ',time.process_time() - start)\n`",
      "votes": null
    },
    {
      "id": "1153221",
      "postDate": "01/14/2021 17:36:51",
      "content": "<p>Question ? <br>\nWhen I predict the above way on train images  (just for testing purposes) I get all predictions as 4 ie <code>21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714</code></p>\n<p>When I run a for loop like this <br>\n<code>import time</code><br>\n<code>images_test = []</code><br>\n<code>preds_try = []</code></p>\n<h1>load all images</h1>\n<p><code>start_load = time.process_time()</code><br>\n<code>train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')</code><br>\n<code>for img in train_images:</code><br>\n<code>img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)</code><br>\n<code>img = img.resize(size)</code><br>\n<code>img = np.expand_dims(img, axis=0)</code><br>\n<code>preds_try.extend(model.predict(img).argmax(axis = 1))</code><br>\n<code>print('Prediction time : ',time.process_time() - start_load)</code></p>\n<p>I get the desired output. What could be the cause of this inconsistency . Based on the second the issue isn't the model, it is the generator.</p>",
      "rawMarkdown": "Question ? \nWhen I predict the above way on train images  (just for testing purposes) I get all predictions as 4 ie `21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714`\n\nWhen I run a for loop like this \n`import time`\n`images_test = []`\n`preds_try = []`\n# load all images\n`start_load = time.process_time()`\n`train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')`\n`for img in train_images:`\n`    img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)`\n`    img = img.resize(size)`\n`    img = np.expand_dims(img, axis=0)`\n`    preds_try.extend(model.predict(img).argmax(axis = 1))`\n`print('Prediction time : ',time.process_time() - start_load)`\n \nI get the desired output. What could be the cause of this inconsistency . Based on the second the issue isn't the model, it is the generator.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1153221,
      "author_name": "hypnotu",
      "author_url": "",
      "post_date": "01/14/2021 17:36:51",
      "content": "<p>Question ? <br>\nWhen I predict the above way on train images  (just for testing purposes) I get all predictions as 4 ie <code>21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714</code></p>\n<p>When I run a for loop like this <br>\n<code>import time</code><br>\n<code>images_test = []</code><br>\n<code>preds_try = []</code></p>\n<h1>load all images</h1>\n<p><code>start_load = time.process_time()</code><br>\n<code>train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')</code><br>\n<code>for img in train_images:</code><br>\n<code>img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)</code><br>\n<code>img = img.resize(size)</code><br>\n<code>img = np.expand_dims(img, axis=0)</code><br>\n<code>preds_try.extend(model.predict(img).argmax(axis = 1))</code><br>\n<code>print('Prediction time : ',time.process_time() - start_load)</code></p>\n<p>I get the desired output. What could be the cause of this inconsistency . Based on the second the issue isn't the model, it is the generator.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1152812": "`\ntest_df = pd.read_csv('../input/cassava-leaf-disease-classification/sample_submission.csv')\ntrain_df = pd.read_csv('../input/cassava-leaf-disease-classification/train.csv')\nDIM =300\nimport cv2\n`\n\n`class DataGenerator(tf.keras.utils.Sequence):\n    def __init__(self, path, list_IDs, labels, batch_size, img_size, img_channel):\n        self.path = path\n        self.list_IDs = list_IDs\n        self.labels = labels\n        self.batch_size = batch_size\n        self.img_size = img_size\n        self.img_channel = img_channel\n        self.indexes = np.arange(len(self.list_IDs))`\n`\n\n`    def __len__(self):\n        return int(np.floor(len(self.list_IDs)/self.batch_size))\n    def __getitem__(self, index):\n        indexes = self.indexes[index*self.batch_size:(index+1)*self.batch_size]\n        list_IDs_temp = [self.list_IDs[k] for k in indexes]\n        X, y = self.__data_generation(list_IDs_temp)\n        return X, y\n`\n    \n    def __data_generation(self, list_IDs_temp):\n        X = np.empty((self.batch_size, self.img_size, self.img_size, self.img_channel)) #use with caution. Faster than np.zeros\n        #Return a new array of given shape and type, without initializing entries.\n        # Generates a numpy array size (1, 300, 300, 1) with\n        \n\n        y = np.empty((self.batch_size, 5), dtype=int)  #generates a numpy array size (1 by 5)\n\n        for i, ID in enumerate(list_IDs_temp): #for all path id get int and ID #enumerate is like pandas iterrows. You can set the start too.\n            img = cv2.imread(self.path+ID) #read image\n            img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB) #convert from RGB to BGR since cv2 reads in RGB.\n            img = cv2.resize(img, (self.img_size, self.img_size)) #resize\n            X[i, ] = img\n            y[i, ] = self.labels[i]\n        X = X.astype('float32')\n        X /= 255.\n        return X, y\n    def preprocess_image(image_path, desired_size=DIM):\n        img = cv2.imread(image_path) #get image\n        img = cv2.cvtColor(img, cv2.COLOR_RGB2BGR) #convert to RGB as cv2 reads in BGR\n        im = cv2.resize(img, (desired_size,desired_size), interpolation = cv2.INTER_AREA) #resize \n        im = np.array(im) #convert to a numpy array\n        return im\n`\ntest_generator = DataGenerator('../input/cassava-leaf-disease-classification/'+'train_images/',\n                               #test_df['image_id'],\n                               #test_df['label'], \n                               #test for many images\n                               train_df['image_id'],\n                               train_df['label'],\n                               1, \n                               DIM,\n                               3)\n`\n#### time a piece of code.\n`import time\nstart = time.process_time()\n`\n#### predict on test images.\n`\npreds = []\nprobabilities = []\npreds = model.predict(test_generator, verbose=1, use_multiprocessing=True, workers=4)\nprobabilities.append(preds)\n`\n`\npredictions = np.array(probabilities).mean(0).argmax(-1)\nprint(predictions)\n`\n\n#### The code\n`print('Prediction time : ',time.process_time() - start)\n`",
    "1153221": "Question ? \nWhen I predict the above way on train images  (just for testing purposes) I get all predictions as 4 ie `21397/21397 [==============================] - 764s 36ms/step\n[4 4 4 ... 4 4 4]\nPrediction time :  762.911825714`\n\nWhen I run a for loop like this \n`import time`\n`images_test = []`\n`preds_try = []`\n# load all images\n`start_load = time.process_time()`\n`train_images = os.listdir('../input/cassava-leaf-disease-classification/'+'train_images/')`\n`for img in train_images:`\n`    img = Image.open('../input/cassava-leaf-disease-classification/'+'train_images/' + img)`\n`    img = img.resize(size)`\n`    img = np.expand_dims(img, axis=0)`\n`    preds_try.extend(model.predict(img).argmax(axis = 1))`\n`print('Prediction time : ',time.process_time() - start_load)`\n \nI get the desired output. What could be the cause of this inconsistency . Based on the second the issue isn't the model, it is the generator."
  },
  "source": "meta"
}