{
  "id": 306437,
  "title": "Images without any channel?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/306437",
  "author_name": "",
  "post_date": "2022-02-09T11:38:31.530756100Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Should I arbitrarily make channel for those 2000 images that does not have any channel?</p>\n<p>As I've explored through images of the dataset, I discovered that 2000 images out of 51033 train-set images do not have any channels.</p>\n<p>After resizing the image to 512 x 512, I expected images to be [h, w, c] = [512, 512, 3] which hold true for most of the images. However, in few grayscale images, it seems like it lacks channel as [h, w] = [512, 512] not as [h, w, c] = [512, 512, 1]. </p>\n<p>Currently, I am dropping those from the train-set. <br>\nWhat would you guys choose?</p>",
  "messages": [
    {
      "id": "1682796",
      "postDate": "02/09/2022 11:38:31",
      "content": "<p>Should I arbitrarily make channel for those 2000 images that does not have any channel?</p>\n<p>As I've explored through images of the dataset, I discovered that 2000 images out of 51033 train-set images do not have any channels.</p>\n<p>After resizing the image to 512 x 512, I expected images to be [h, w, c] = [512, 512, 3] which hold true for most of the images. However, in few grayscale images, it seems like it lacks channel as [h, w] = [512, 512] not as [h, w, c] = [512, 512, 1]. </p>\n<p>Currently, I am dropping those from the train-set. <br>\nWhat would you guys choose?</p>",
      "rawMarkdown": "Should I arbitrarily make channel for those 2000 images that does not have any channel?\n\nAs I've explored through images of the dataset, I discovered that 2000 images out of 51033 train-set images do not have any channels.\n\nAfter resizing the image to 512 x 512, I expected images to be [h, w, c] = [512, 512, 3] which hold true for most of the images. However, in few grayscale images, it seems like it lacks channel as [h, w] = [512, 512] not as [h, w, c] = [512, 512, 1]. \n\nCurrently, I am dropping those from the train-set. \nWhat would you guys choose?",
      "votes": null
    },
    {
      "id": "1682924",
      "postDate": "02/09/2022 13:21:43",
      "content": "<p>Hi, I noticed this as well. However, it doesn't mean that they don't have any channel, they just have 1 channel.<br>\nThis is because those images are grayscale images, so they don't have the 3 color channels but just the 1 channel of grayscale values.<br>\nYou can still use them, however. Technically, you just have to stack the 1 channel you have 2 times on top of it. </p>\n<p>Assuming numpy arrays you can do:</p>\n<pre><code>import numpy as np\nfrom skimage import io\n\nimage = io.imread(img_path)\nif len(image.shape) == 2:\n    image = np.dstack((image,)*3)\n</code></pre>",
      "rawMarkdown": "Hi, I noticed this as well. However, it doesn't mean that they don't have any channel, they just have 1 channel.\nThis is because those images are grayscale images, so they don't have the 3 color channels but just the 1 channel of grayscale values.\nYou can still use them, however. Technically, you just have to stack the 1 channel you have 2 times on top of it. \n\nAssuming numpy arrays you can do:\n```python\nimport numpy as np\nfrom skimage import io\n\nimage = io.imread(img_path)\nif len(image.shape) == 2:\n    image = np.dstack((image,)*3)\n```",
      "votes": null
    },
    {
      "id": "1683338",
      "postDate": "02/09/2022 18:32:15",
      "content": "<p>Like <a href=\"https://www.kaggle.com/gordonbee\" target=\"_blank\">@gordonbee</a>, I would use these grayscale (single channel) images. I would also add b&amp;w conversion to the data augmentations used in your training. Among other benefits, your model should better handle any grayscale images in the test set. </p>",
      "rawMarkdown": "Like @gordonbee, I would use these grayscale (single channel) images. I would also add b&w conversion to the data augmentations used in your training. Among other benefits, your model should better handle any grayscale images in the test set.",
      "votes": null
    },
    {
      "id": "1684265",
      "postDate": "02/10/2022 11:24:23",
      "content": "<p>Even though I extracted metadata such as red, green, blue channel statistics and shapes from images, I noticed there are grayscale images after seeing this topic. There are 2366 grayscale images in the entire dataset. 2033 of them are in the training set and 333 of them are in the test set.</p>\n<p>I initially didn't notice or get any errors because I was using OpenCV. By default OpenCV reads images in BGR format with 3 channels and I was converting them to RGB. You can safely do</p>\n<pre><code>image = cv2.imread('path/to/image')\nimage = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n</code></pre>\n<p>I don't think you should do anything special to them at all.</p>",
      "rawMarkdown": "Even though I extracted metadata such as red, green, blue channel statistics and shapes from images, I noticed there are grayscale images after seeing this topic. There are 2366 grayscale images in the entire dataset. 2033 of them are in the training set and 333 of them are in the test set.\n\nI initially didn't notice or get any errors because I was using OpenCV. By default OpenCV reads images in BGR format with 3 channels and I was converting them to RGB. You can safely do\n\n```\nimage = cv2.imread('path/to/image')\nimage = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n```\n\nI don't think you should do anything special to them at all.",
      "votes": null
    },
    {
      "id": "1684478",
      "postDate": "02/10/2022 14:17:41",
      "content": "<p>Thank you all! I've now am capable of using all the images in the dataset 🤗</p>",
      "rawMarkdown": "Thank you all! I've now am capable of using all the images in the dataset 🤗",
      "votes": null
    },
    {
      "id": "1685170",
      "postDate": "02/11/2022 04:45:30",
      "content": "<p>You might want to analyze the composition of these single channel images. Check if there are other samples for the same individual before you drop them. </p>",
      "rawMarkdown": "You might want to analyze the composition of these single channel images. Check if there are other samples for the same individual before you drop them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1682924,
      "author_name": "gordonbee",
      "author_url": "",
      "post_date": "02/09/2022 13:21:43",
      "content": "<p>Hi, I noticed this as well. However, it doesn't mean that they don't have any channel, they just have 1 channel.<br>\nThis is because those images are grayscale images, so they don't have the 3 color channels but just the 1 channel of grayscale values.<br>\nYou can still use them, however. Technically, you just have to stack the 1 channel you have 2 times on top of it. </p>\n<p>Assuming numpy arrays you can do:</p>\n<pre><code>import numpy as np\nfrom skimage import io\n\nimage = io.imread(img_path)\nif len(image.shape) == 2:\n    image = np.dstack((image,)*3)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1683338,
      "author_name": "rturley",
      "author_url": "",
      "post_date": "02/09/2022 18:32:15",
      "content": "<p>Like <a href=\"https://www.kaggle.com/gordonbee\" target=\"_blank\">@gordonbee</a>, I would use these grayscale (single channel) images. I would also add b&amp;w conversion to the data augmentations used in your training. Among other benefits, your model should better handle any grayscale images in the test set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1684265,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/10/2022 11:24:23",
      "content": "<p>Even though I extracted metadata such as red, green, blue channel statistics and shapes from images, I noticed there are grayscale images after seeing this topic. There are 2366 grayscale images in the entire dataset. 2033 of them are in the training set and 333 of them are in the test set.</p>\n<p>I initially didn't notice or get any errors because I was using OpenCV. By default OpenCV reads images in BGR format with 3 channels and I was converting them to RGB. You can safely do</p>\n<pre><code>image = cv2.imread('path/to/image')\nimage = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n</code></pre>\n<p>I don't think you should do anything special to them at all.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1684478,
      "author_name": "snoop2head",
      "author_url": "",
      "post_date": "02/10/2022 14:17:41",
      "content": "<p>Thank you all! I've now am capable of using all the images in the dataset 🤗</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1685170,
      "author_name": "bsridatta",
      "author_url": "",
      "post_date": "02/11/2022 04:45:30",
      "content": "<p>You might want to analyze the composition of these single channel images. Check if there are other samples for the same individual before you drop them. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1682796": "Should I arbitrarily make channel for those 2000 images that does not have any channel?\n\nAs I've explored through images of the dataset, I discovered that 2000 images out of 51033 train-set images do not have any channels.\n\nAfter resizing the image to 512 x 512, I expected images to be [h, w, c] = [512, 512, 3] which hold true for most of the images. However, in few grayscale images, it seems like it lacks channel as [h, w] = [512, 512] not as [h, w, c] = [512, 512, 1]. \n\nCurrently, I am dropping those from the train-set. \nWhat would you guys choose?",
    "1682924": "Hi, I noticed this as well. However, it doesn't mean that they don't have any channel, they just have 1 channel.\nThis is because those images are grayscale images, so they don't have the 3 color channels but just the 1 channel of grayscale values.\nYou can still use them, however. Technically, you just have to stack the 1 channel you have 2 times on top of it. \n\nAssuming numpy arrays you can do:\n```python\nimport numpy as np\nfrom skimage import io\n\nimage = io.imread(img_path)\nif len(image.shape) == 2:\n    image = np.dstack((image,)*3)\n```",
    "1683338": "Like @gordonbee, I would use these grayscale (single channel) images. I would also add b&w conversion to the data augmentations used in your training. Among other benefits, your model should better handle any grayscale images in the test set.",
    "1684265": "Even though I extracted metadata such as red, green, blue channel statistics and shapes from images, I noticed there are grayscale images after seeing this topic. There are 2366 grayscale images in the entire dataset. 2033 of them are in the training set and 333 of them are in the test set.\n\nI initially didn't notice or get any errors because I was using OpenCV. By default OpenCV reads images in BGR format with 3 channels and I was converting them to RGB. You can safely do\n\n```\nimage = cv2.imread('path/to/image')\nimage = cv2.cvtColor(image, cv2.COLOR_BGR2RGB)\n```\n\nI don't think you should do anything special to them at all.",
    "1684478": "Thank you all! I've now am capable of using all the images in the dataset 🤗",
    "1685170": "You might want to analyze the composition of these single channel images. Check if there are other samples for the same individual before you drop them."
  },
  "source": "meta"
}