{
  "id": 225010,
  "title": "How can we load data with multiple labels using `flow_from_dataframe`?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/225010",
  "author_name": "vinesmsuic",
  "post_date": "2021-03-10T14:16:07.597000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi, I am using the following code:</p>\n<pre><code>train_datagen = tf.keras.preprocessing.image.ImageDataGenerator(\n                                     validation_split = 0.2,\n                                     preprocessing_function = None,\n                                     zoom_range = 0.2,\n                                     cval = 0.2,\n                                     horizontal_flip = True,\n                                     vertical_flip = True,\n                                     fill_mode = 'nearest',\n                                     shear_range = 0.2,\n                                     height_shift_range = 0.2,\n                                     width_shift_range = 0.2)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df,\n                                                    directory = train_dir,\n                                                    subset = \"training\",\n                                                    x_col = \"path\",\n                                                    y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                    target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                    batch_size=BATCH_SIZE,\n                                                    class_mode='multi_output')\n\nvalidation_datagen = tf.keras.preprocessing.image.ImageDataGenerator(validation_split = 0.2)\n\n\nvalidation_generator = validation_datagen.flow_from_dataframe(train_df,\n                                                                directory = train_dir,\n                                                                subset = \"validation\",\n                                                                x_col = \"path\",\n                                                                y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                                target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                                batch_size=BATCH_SIZE,\n                                                                class_mode='multi_output')\n</code></pre>\n<p>It outputs that it can locate the images.</p>\n<p>But when I try to fit it using <code>binary_crossentropy</code> as loss function, it gives error : <code>ValueError: logits and labels must have the same shape ((None, 11) vs (None, 1))</code>.</p>\n<p>My last layer was <code>tf.keras.layers.Dense(11, activation='sigmoid')</code>.</p>\n<p>I doubt if I am doing it wrong in the <code>y_col</code> parameter of <code>flow_from_dataframe</code> function.</p>",
  "messages": [
    {
      "id": 1233571,
      "postDate": "2021-03-10T14:16:07.597Z",
      "content": "<p>Hi, I am using the following code:</p>\n<pre><code>train_datagen = tf.keras.preprocessing.image.ImageDataGenerator(\n                                     validation_split = 0.2,\n                                     preprocessing_function = None,\n                                     zoom_range = 0.2,\n                                     cval = 0.2,\n                                     horizontal_flip = True,\n                                     vertical_flip = True,\n                                     fill_mode = 'nearest',\n                                     shear_range = 0.2,\n                                     height_shift_range = 0.2,\n                                     width_shift_range = 0.2)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df,\n                                                    directory = train_dir,\n                                                    subset = \"training\",\n                                                    x_col = \"path\",\n                                                    y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                    target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                    batch_size=BATCH_SIZE,\n                                                    class_mode='multi_output')\n\nvalidation_datagen = tf.keras.preprocessing.image.ImageDataGenerator(validation_split = 0.2)\n\n\nvalidation_generator = validation_datagen.flow_from_dataframe(train_df,\n                                                                directory = train_dir,\n                                                                subset = \"validation\",\n                                                                x_col = \"path\",\n                                                                y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                                target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                                batch_size=BATCH_SIZE,\n                                                                class_mode='multi_output')\n</code></pre>\n<p>It outputs that it can locate the images.</p>\n<p>But when I try to fit it using <code>binary_crossentropy</code> as loss function, it gives error : <code>ValueError: logits and labels must have the same shape ((None, 11) vs (None, 1))</code>.</p>\n<p>My last layer was <code>tf.keras.layers.Dense(11, activation='sigmoid')</code>.</p>\n<p>I doubt if I am doing it wrong in the <code>y_col</code> parameter of <code>flow_from_dataframe</code> function.</p>",
      "rawMarkdown": "Hi, I am using the following code:\n```python\ntrain_datagen = tf.keras.preprocessing.image.ImageDataGenerator(\n                                     validation_split = 0.2,\n                                     preprocessing_function = None,\n                                     zoom_range = 0.2,\n                                     cval = 0.2,\n                                     horizontal_flip = True,\n                                     vertical_flip = True,\n                                     fill_mode = 'nearest',\n                                     shear_range = 0.2,\n                                     height_shift_range = 0.2,\n                                     width_shift_range = 0.2)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df,\n                                                    directory = train_dir,\n                                                    subset = \"training\",\n                                                    x_col = \"path\",\n                                                    y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                    target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                    batch_size=BATCH_SIZE,\n                                                    class_mode='multi_output')\n\nvalidation_datagen = tf.keras.preprocessing.image.ImageDataGenerator(validation_split = 0.2)\n\n\nvalidation_generator = validation_datagen.flow_from_dataframe(train_df,\n                                                                directory = train_dir,\n                                                                subset = \"validation\",\n                                                                x_col = \"path\",\n                                                                y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                                target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                                batch_size=BATCH_SIZE,\n                                                                class_mode='multi_output')\n```\nIt outputs that it can locate the images.\n\nBut when I try to fit it using `binary_crossentropy` as loss function, it gives error : `ValueError: logits and labels must have the same shape ((None, 11) vs (None, 1))`.\n\nMy last layer was `tf.keras.layers.Dense(11, activation='sigmoid')`.\n\nI doubt if I am doing it wrong in the `y_col` parameter of `flow_from_dataframe` function.",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1233571": "Hi, I am using the following code:\n```python\ntrain_datagen = tf.keras.preprocessing.image.ImageDataGenerator(\n                                     validation_split = 0.2,\n                                     preprocessing_function = None,\n                                     zoom_range = 0.2,\n                                     cval = 0.2,\n                                     horizontal_flip = True,\n                                     vertical_flip = True,\n                                     fill_mode = 'nearest',\n                                     shear_range = 0.2,\n                                     height_shift_range = 0.2,\n                                     width_shift_range = 0.2)\n\ntrain_generator = train_datagen.flow_from_dataframe(train_df,\n                                                    directory = train_dir,\n                                                    subset = \"training\",\n                                                    x_col = \"path\",\n                                                    y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                    target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                    batch_size=BATCH_SIZE,\n                                                    class_mode='multi_output')\n\nvalidation_datagen = tf.keras.preprocessing.image.ImageDataGenerator(validation_split = 0.2)\n\n\nvalidation_generator = validation_datagen.flow_from_dataframe(train_df,\n                                                                directory = train_dir,\n                                                                subset = \"validation\",\n                                                                x_col = \"path\",\n                                                                y_col = ['ETT - Abnormal', 'ETT - Borderline', 'ETT - Normal', 'NGT - Abnormal', 'NGT - Borderline', 'NGT - Incompletely Imaged', 'NGT - Normal', 'CVC - Abnormal', 'CVC - Borderline', 'CVC - Normal', 'Swan Ganz Catheter Present'],\n                                                                target_size=(TARGET_SIZE, TARGET_SIZE),\n                                                                batch_size=BATCH_SIZE,\n                                                                class_mode='multi_output')\n```\nIt outputs that it can locate the images.\n\nBut when I try to fit it using `binary_crossentropy` as loss function, it gives error : `ValueError: logits and labels must have the same shape ((None, 11) vs (None, 1))`.\n\nMy last layer was `tf.keras.layers.Dense(11, activation='sigmoid')`.\n\nI doubt if I am doing it wrong in the `y_col` parameter of `flow_from_dataframe` function."
  }
}