{
  "id": 313547,
  "title": "How do you use the masks provided in the dataset?",
  "url": "/competitions/hotel-id-to-combat-human-trafficking-2022-fgvc9/discussion/313547",
  "author_name": "",
  "post_date": "2022-03-17T17:59:18.236938700Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I am new to image classification and was wondering how to use the image mask with the images in the given dataset. I do understand they might be used for the areas to cover it up but how do we relate these?</p>",
  "messages": [
    {
      "id": "1726177",
      "postDate": "03/17/2022 17:59:18",
      "content": "<p>I am new to image classification and was wondering how to use the image mask with the images in the given dataset. I do understand they might be used for the areas to cover it up but how do we relate these?</p>",
      "rawMarkdown": "I am new to image classification and was wondering how to use the image mask with the images in the given dataset. I do understand they might be used for the areas to cover it up but how do we relate these?",
      "votes": null
    },
    {
      "id": "1727477",
      "postDate": "03/18/2022 02:15:51",
      "content": "<p>From a fellow beginner, I would suggest incorporating some layers for edge detection.  The image mask (I'm understanding that to mean the redacted personal identifiers in the image of \"the blanked our area\") seems consistently large enough to detect and adjust for.</p>\n<p>Now as for how to train those layers, I don't know.  But that's where I'd start in dealing with the \"blank\" pixel data.</p>",
      "rawMarkdown": "From a fellow beginner, I would suggest incorporating some layers for edge detection.  The image mask (I'm understanding that to mean the redacted personal identifiers in the image of \"the blanked our area\") seems consistently large enough to detect and adjust for.\n\nNow as for how to train those layers, I don't know.  But that's where I'd start in dealing with the \"blank\" pixel data.",
      "votes": null
    },
    {
      "id": "1727536",
      "postDate": "03/18/2022 03:26:47",
      "content": "<p><a href=\"https://www.kaggle.com/sliderulemath\" target=\"_blank\">@sliderulemath</a> Ah! I see. You mean we need to merge both the images and generate the model. So my understanding is that we need to find similar files and superimpose them to generate a new image and that will help me match the input with the output. Seems like a milti-class problem. </p>",
      "rawMarkdown": "sliderulemath Ah! I see. You mean we need to merge both the images and generate the model. So my understanding is that we need to find similar files and superimpose them to generate a new image and that will help me match the input with the output. Seems like a milti-class problem.",
      "votes": null
    },
    {
      "id": "1732931",
      "postDate": "03/23/2022 21:07:16",
      "content": "<p>I didn't really know how to match the masks with the images too. I decided to apply a random mask to each image. To do so I resize the mask to match the image shape.</p>",
      "rawMarkdown": "I didn't really know how to match the masks with the images too. I decided to apply a random mask to each image. To do so I resize the mask to match the image shape.",
      "votes": null
    },
    {
      "id": "1732937",
      "postDate": "03/23/2022 21:18:57",
      "content": "<p>The mask files are available for your convenience -- if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG. You can use these for processing of the query images in case it's easier than detecting the mask from the JPG, and also may also use any of the masks for whatever purpose in your training (as some other comments have suggested). You aren't required to use the PNG masks for anything -- like I said, they're just there for convenience.</p>",
      "rawMarkdown": "The mask files are available for your convenience -- if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG. You can use these for processing of the query images in case it's easier than detecting the mask from the JPG, and also may also use any of the masks for whatever purpose in your training (as some other comments have suggested). You aren't required to use the PNG masks for anything -- like I said, they're just there for convenience.",
      "votes": null
    },
    {
      "id": "1734481",
      "postDate": "03/25/2022 11:44:44",
      "content": "<blockquote>\n  <p>if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG</p>\n</blockquote>\n<p>But the train image names have format 000000000.jpg while mask names are 00000.png, there are no masks that match any images in training dataset. Even if we prepand the extra 0s to mask name it doesn't seem to match either (the resolution of images is different). I checked it in <a href=\"https://www.kaggle.com/code/michaln/masks-and-occlusions\" target=\"_blank\">this</a> notebook and couldn't find any matches.</p>\n<p>Are you sure that provided masks are related to the data in training dataset or did I just miss something?</p>",
      "rawMarkdown": "> if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG\n\nBut the train image names have format 000000000.jpg while mask names are 00000.png, there are no masks that match any images in training dataset. Even if we prepand the extra 0s to mask name it doesn't seem to match either (the resolution of images is different). I checked it in [this](https://www.kaggle.com/code/michaln/masks-and-occlusions) notebook and couldn't find any matches.\n\nAre you sure that provided masks are related to the data in training dataset or did I just miss something?",
      "votes": null
    },
    {
      "id": "1734655",
      "postDate": "03/25/2022 14:33:21",
      "content": "<p>The nomenclature is a bit confusing, but the train_masks apply to the (unseen) test images.  So hidden test will have a 00000.jpg that will go with the train_mask 00000.png that you can see.  At least this is my understanding.</p>",
      "rawMarkdown": "The nomenclature is a bit confusing, but the train_masks apply to the (unseen) test images.  So hidden test will have a 00000.jpg that will go with the train_mask 00000.png that you can see.  At least this is my understanding.",
      "votes": null
    },
    {
      "id": "1734687",
      "postDate": "03/25/2022 15:00:41",
      "content": "<p>Oh! I have figured out the source of the confusion. There was a mixup on the host end -- the \"train_masks\" folder should be named \"test_masks\" (I've asked the kaggle team to update this). There are no training masks provided. This matches the real world setting, where the \"test\" images (from investigations) have occlusions in the region of the image where the victim is located.</p>\n<p>Training images, on the other hand, are not (by default) occluded. Competitors may choose to include occlusions in their training process, but we do not dictate that (or any other approach). If a competitor chose to incorporate masks, they could either generate their own, or repurpose the ones that match the test images (resizing them as necessary).</p>\n<p>Thanks for the heads up on the train_masks issue -- I'm really sorry we didn't catch that sooner!</p>",
      "rawMarkdown": "Oh! I have figured out the source of the confusion. There was a mixup on the host end -- the \"train_masks\" folder should be named \"test_masks\" (I've asked the kaggle team to update this). There are no training masks provided. This matches the real world setting, where the \"test\" images (from investigations) have occlusions in the region of the image where the victim is located.\n\nTraining images, on the other hand, are not (by default) occluded. Competitors may choose to include occlusions in their training process, but we do not dictate that (or any other approach). If a competitor chose to incorporate masks, they could either generate their own, or repurpose the ones that match the test images (resizing them as necessary).\n\nThanks for the heads up on the train_masks issue -- I'm really sorry we didn't catch that sooner!",
      "votes": null
    },
    {
      "id": "1758111",
      "postDate": "04/17/2022 10:24:22",
      "content": "<p>I discovered a strange behavior of test masks, but I don't know if it's a problem of my code or a problem on the test masks.<br>\nI just added a double check that reapply the mask for every test images in the Pytorch dataset. I expect that if I reapply the mask the image will be unchanged. This is the code of the <strong>getitem</strong> function:</p>\n<pre><code>def __getitem__(self, index):\n    path = self.sample[index]\n    img = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB)\n    # ------------   Double check code  ----------- // Start\n    # get the mask fileneame replacing jpg with png\n    mask_filename = os.path.basename(path).replace(\".jpg\", \".png\")\n    # Get the global path of mask (mask_root defined globally)\n    mask_path = os.path.join(mask_root, mask_filename)\n    if os.path.exists(mask_path):\n        # read mask\n        mask = cv2.imread(mask_path)\n        # Transform to GRAY to get the non-zero points\n        mask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\n        # Just re-apply the mask values\n        img[mask_g&gt;0] = mask[mask_g&gt;0]\n   # ------------   Double check code  ----------- // End\n\n    img = cv2.resize(img, (test_input_size, test_input_size), interpolation=cv2.INTER_AREA)\n    img = self.transform(image=img)[\"image\"]\n    return img, path\n</code></pre>\n<p>After this double check I expect the exact same performance as the <strong>getitem</strong> without the \"Double check code\", but I obtained a drop of 0.1 in mAP. Let me know if it's my fault (probably) or a problem with masks.</p>",
      "rawMarkdown": "I discovered a strange behavior of test masks, but I don't know if it's a problem of my code or a problem on the test masks.\nI just added a double check that reapply the mask for every test images in the Pytorch dataset. I expect that if I reapply the mask the image will be unchanged. This is the code of the __getitem__ function:\n\n    def __getitem__(self, index):\n        path = self.sample[index]\n        img = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB)\n        # ------------   Double check code  ----------- // Start\n        # get the mask fileneame replacing jpg with png\n        mask_filename = os.path.basename(path).replace(\".jpg\", \".png\")\n        # Get the global path of mask (mask_root defined globally)\n        mask_path = os.path.join(mask_root, mask_filename)\n        if os.path.exists(mask_path):\n            # read mask\n            mask = cv2.imread(mask_path)\n            # Transform to GRAY to get the non-zero points\n            mask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\n            # Just re-apply the mask values\n            img[mask_g>0] = mask[mask_g>0]\n       # ------------   Double check code  ----------- // End\n        \n        img = cv2.resize(img, (test_input_size, test_input_size), interpolation=cv2.INTER_AREA)\n        img = self.transform(image=img)[\"image\"]\n        return img, path\n\nAfter this double check I expect the exact same performance as the __getitem__ without the \"Double check code\", but I obtained a drop of 0.1 in mAP. Let me know if it's my fault (probably) or a problem with masks.",
      "votes": null
    },
    {
      "id": "1758135",
      "postDate": "04/17/2022 11:06:56",
      "content": "<p>You are applying the mask to the wrong color channel.<br>\nEither do cv2.COLOR_BGR2RGB conversion for the mask before you convert it to gray (don't ask my why it works) or apply it manually to each channel (1 for red channel, 0s for others).</p>\n<p>Code to visualize</p>\n<pre><code>import cv2\nimport matplotlib.pyplot as plt\n\ndef show_im_with_channels(image, title=None):\n    channels = [\"Red\", \"Green\", \"Blue\"]\n\n    fig, ax = plt.subplots(1, 4, figsize=(22,6))\n    ax[0].imshow(image)\n    ax[0].set_ylabel(title)\n    if len(image.shape) &gt; 2 and image.shape[2] &gt; 1:\n        for i in range(3):\n            ax[i+1].imshow(image[:, :, i])\n            ax[i+1].set_title(channels[i])\n\nimg = cv2.cvtColor(cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_images/100206/000034116.jpg\"), cv2.COLOR_BGR2RGB)\nmask = cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_masks/00000.png\")\nmask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\nmasked_image = img.copy()\nmasked_image[mask_g&gt;0] = mask[mask_g&gt;0]\n\nshow_im_with_channels(img, \"Original image\")\nshow_im_with_channels(masked_image, \"Masked image\")\n\nshow_im_with_channels(mask, \"Mask\")\nshow_im_with_channels(mask_g, \"Gray mask\")\n\ntest_image = cv2.cvtColor(cv2.imread('../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/abc.jpg'), cv2.COLOR_BGR2RGB)\nshow_im_with_channels(test_image, \"Test image\")\n</code></pre>",
      "rawMarkdown": "You are applying the mask to the wrong color channel.\nEither do cv2.COLOR_BGR2RGB conversion for the mask before you convert it to gray (don't ask my why it works) or apply it manually to each channel (1 for red channel, 0s for others).\n\nCode to visualize\n```\nimport cv2\nimport matplotlib.pyplot as plt\n\ndef show_im_with_channels(image, title=None):\n    channels = [\"Red\", \"Green\", \"Blue\"]\n    \n    fig, ax = plt.subplots(1, 4, figsize=(22,6))\n    ax[0].imshow(image)\n    ax[0].set_ylabel(title)\n    if len(image.shape) > 2 and image.shape[2] > 1:\n        for i in range(3):\n            ax[i+1].imshow(image[:, :, i])\n            ax[i+1].set_title(channels[i])\n\nimg = cv2.cvtColor(cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_images/100206/000034116.jpg\"), cv2.COLOR_BGR2RGB)\nmask = cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_masks/00000.png\")\nmask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\nmasked_image = img.copy()\nmasked_image[mask_g>0] = mask[mask_g>0]\n\nshow_im_with_channels(img, \"Original image\")\nshow_im_with_channels(masked_image, \"Masked image\")\n\nshow_im_with_channels(mask, \"Mask\")\nshow_im_with_channels(mask_g, \"Gray mask\")\n\ntest_image = cv2.cvtColor(cv2.imread('../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/abc.jpg'), cv2.COLOR_BGR2RGB)\nshow_im_with_channels(test_image, \"Test image\")\n```",
      "votes": null
    },
    {
      "id": "1758140",
      "postDate": "04/17/2022 11:10:23",
      "content": "<p>You are right! Thank you</p>",
      "rawMarkdown": "You are right! Thank you",
      "votes": null
    },
    {
      "id": "1758292",
      "postDate": "04/17/2022 14:33:53",
      "content": "<p>Ok, I fixed the BGR bug and I founded that the score has decreased by 0.03 mAP (better than before but not equal to the original one)! It's like the test images have different masks (maybe something different than the provided bounding boxes). I don't know of course if I am wrong, but I'm trying to understand better the problem, I am a bit confused.🤔</p>",
      "rawMarkdown": "Ok, I fixed the BGR bug and I founded that the score has decreased by 0.03 mAP (better than before but not equal to the original one)! It's like the test images have different masks (maybe something different than the provided bounding boxes). I don't know of course if I am wrong, but I'm trying to understand better the problem, I am a bit confused.🤔",
      "votes": null
    },
    {
      "id": "1803282",
      "postDate": "05/27/2022 16:25:33",
      "content": "<p>Have you figured it out? I checked the test picture again and I think because of the jpg compression the occlusion is not as perfect as we would expect it. The the main body of occlusion has values [254, 0, 0] and on the edges the values are even less clear. So I think when you recreate the occlusion using mask the result is not the same as the original image and can lead to different score.</p>",
      "rawMarkdown": "Have you figured it out? I checked the test picture again and I think because of the jpg compression the occlusion is not as perfect as we would expect it. The the main body of occlusion has values [254, 0, 0] and on the edges the values are even less clear. So I think when you recreate the occlusion using mask the result is not the same as the original image and can lead to different score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1727477,
      "author_name": "sliderulemath",
      "author_url": "",
      "post_date": "03/18/2022 02:15:51",
      "content": "<p>From a fellow beginner, I would suggest incorporating some layers for edge detection.  The image mask (I'm understanding that to mean the redacted personal identifiers in the image of \"the blanked our area\") seems consistently large enough to detect and adjust for.</p>\n<p>Now as for how to train those layers, I don't know.  But that's where I'd start in dealing with the \"blank\" pixel data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1727536,
          "author_name": "shivkumarganesh",
          "author_url": "",
          "post_date": "03/18/2022 03:26:47",
          "content": "<p><a href=\"https://www.kaggle.com/sliderulemath\" target=\"_blank\">@sliderulemath</a> Ah! I see. You mean we need to merge both the images and generate the model. So my understanding is that we need to find similar files and superimpose them to generate a new image and that will help me match the input with the output. Seems like a milti-class problem. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1732931,
      "author_name": "laurentpoyet",
      "author_url": "",
      "post_date": "03/23/2022 21:07:16",
      "content": "<p>I didn't really know how to match the masks with the images too. I decided to apply a random mask to each image. To do so I resize the mask to match the image shape.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1732937,
      "author_name": "abbystylianou",
      "author_url": "",
      "post_date": "03/23/2022 21:18:57",
      "content": "<p>The mask files are available for your convenience -- if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG. You can use these for processing of the query images in case it's easier than detecting the mask from the JPG, and also may also use any of the masks for whatever purpose in your training (as some other comments have suggested). You aren't required to use the PNG masks for anything -- like I said, they're just there for convenience.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1734481,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "03/25/2022 11:44:44",
          "content": "<blockquote>\n  <p>if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG</p>\n</blockquote>\n<p>But the train image names have format 000000000.jpg while mask names are 00000.png, there are no masks that match any images in training dataset. Even if we prepand the extra 0s to mask name it doesn't seem to match either (the resolution of images is different). I checked it in <a href=\"https://www.kaggle.com/code/michaln/masks-and-occlusions\" target=\"_blank\">this</a> notebook and couldn't find any matches.</p>\n<p>Are you sure that provided masks are related to the data in training dataset or did I just miss something?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1734655,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "03/25/2022 14:33:21",
          "content": "<p>The nomenclature is a bit confusing, but the train_masks apply to the (unseen) test images.  So hidden test will have a 00000.jpg that will go with the train_mask 00000.png that you can see.  At least this is my understanding.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1734687,
          "author_name": "abbystylianou",
          "author_url": "",
          "post_date": "03/25/2022 15:00:41",
          "content": "<p>Oh! I have figured out the source of the confusion. There was a mixup on the host end -- the \"train_masks\" folder should be named \"test_masks\" (I've asked the kaggle team to update this). There are no training masks provided. This matches the real world setting, where the \"test\" images (from investigations) have occlusions in the region of the image where the victim is located.</p>\n<p>Training images, on the other hand, are not (by default) occluded. Competitors may choose to include occlusions in their training process, but we do not dictate that (or any other approach). If a competitor chose to incorporate masks, they could either generate their own, or repurpose the ones that match the test images (resizing them as necessary).</p>\n<p>Thanks for the heads up on the train_masks issue -- I'm really sorry we didn't catch that sooner!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1758111,
      "author_name": "alenic",
      "author_url": "",
      "post_date": "04/17/2022 10:24:22",
      "content": "<p>I discovered a strange behavior of test masks, but I don't know if it's a problem of my code or a problem on the test masks.<br>\nI just added a double check that reapply the mask for every test images in the Pytorch dataset. I expect that if I reapply the mask the image will be unchanged. This is the code of the <strong>getitem</strong> function:</p>\n<pre><code>def __getitem__(self, index):\n    path = self.sample[index]\n    img = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB)\n    # ------------   Double check code  ----------- // Start\n    # get the mask fileneame replacing jpg with png\n    mask_filename = os.path.basename(path).replace(\".jpg\", \".png\")\n    # Get the global path of mask (mask_root defined globally)\n    mask_path = os.path.join(mask_root, mask_filename)\n    if os.path.exists(mask_path):\n        # read mask\n        mask = cv2.imread(mask_path)\n        # Transform to GRAY to get the non-zero points\n        mask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\n        # Just re-apply the mask values\n        img[mask_g&gt;0] = mask[mask_g&gt;0]\n   # ------------   Double check code  ----------- // End\n\n    img = cv2.resize(img, (test_input_size, test_input_size), interpolation=cv2.INTER_AREA)\n    img = self.transform(image=img)[\"image\"]\n    return img, path\n</code></pre>\n<p>After this double check I expect the exact same performance as the <strong>getitem</strong> without the \"Double check code\", but I obtained a drop of 0.1 in mAP. Let me know if it's my fault (probably) or a problem with masks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1758135,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "04/17/2022 11:06:56",
          "content": "<p>You are applying the mask to the wrong color channel.<br>\nEither do cv2.COLOR_BGR2RGB conversion for the mask before you convert it to gray (don't ask my why it works) or apply it manually to each channel (1 for red channel, 0s for others).</p>\n<p>Code to visualize</p>\n<pre><code>import cv2\nimport matplotlib.pyplot as plt\n\ndef show_im_with_channels(image, title=None):\n    channels = [\"Red\", \"Green\", \"Blue\"]\n\n    fig, ax = plt.subplots(1, 4, figsize=(22,6))\n    ax[0].imshow(image)\n    ax[0].set_ylabel(title)\n    if len(image.shape) &gt; 2 and image.shape[2] &gt; 1:\n        for i in range(3):\n            ax[i+1].imshow(image[:, :, i])\n            ax[i+1].set_title(channels[i])\n\nimg = cv2.cvtColor(cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_images/100206/000034116.jpg\"), cv2.COLOR_BGR2RGB)\nmask = cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_masks/00000.png\")\nmask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\nmasked_image = img.copy()\nmasked_image[mask_g&gt;0] = mask[mask_g&gt;0]\n\nshow_im_with_channels(img, \"Original image\")\nshow_im_with_channels(masked_image, \"Masked image\")\n\nshow_im_with_channels(mask, \"Mask\")\nshow_im_with_channels(mask_g, \"Gray mask\")\n\ntest_image = cv2.cvtColor(cv2.imread('../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/abc.jpg'), cv2.COLOR_BGR2RGB)\nshow_im_with_channels(test_image, \"Test image\")\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1758140,
          "author_name": "alenic",
          "author_url": "",
          "post_date": "04/17/2022 11:10:23",
          "content": "<p>You are right! Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1758292,
          "author_name": "alenic",
          "author_url": "",
          "post_date": "04/17/2022 14:33:53",
          "content": "<p>Ok, I fixed the BGR bug and I founded that the score has decreased by 0.03 mAP (better than before but not equal to the original one)! It's like the test images have different masks (maybe something different than the provided bounding boxes). I don't know of course if I am wrong, but I'm trying to understand better the problem, I am a bit confused.🤔</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1803282,
          "author_name": "michaln",
          "author_url": "",
          "post_date": "05/27/2022 16:25:33",
          "content": "<p>Have you figured it out? I checked the test picture again and I think because of the jpg compression the occlusion is not as perfect as we would expect it. The the main body of occlusion has values [254, 0, 0] and on the edges the values are even less clear. So I think when you recreate the occlusion using mask the result is not the same as the original image and can lead to different score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1726177": "I am new to image classification and was wondering how to use the image mask with the images in the given dataset. I do understand they might be used for the areas to cover it up but how do we relate these?",
    "1727477": "From a fellow beginner, I would suggest incorporating some layers for edge detection.  The image mask (I'm understanding that to mean the redacted personal identifiers in the image of \"the blanked our area\") seems consistently large enough to detect and adjust for.\n\nNow as for how to train those layers, I don't know.  But that's where I'd start in dealing with the \"blank\" pixel data.",
    "1727536": "sliderulemath Ah! I see. You mean we need to merge both the images and generate the model. So my understanding is that we need to find similar files and superimpose them to generate a new image and that will help me match the input with the output. Seems like a milti-class problem.",
    "1732931": "I didn't really know how to match the masks with the images too. I decided to apply a random mask to each image. To do so I resize the mask to match the image shape.",
    "1732937": "The mask files are available for your convenience -- if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG. You can use these for processing of the query images in case it's easier than detecting the mask from the JPG, and also may also use any of the masks for whatever purpose in your training (as some other comments have suggested). You aren't required to use the PNG masks for anything -- like I said, they're just there for convenience.",
    "1734481": "> if there's a query image called 0001.jpg and a mask called 0001.png, the mask is simply a PNG that includes the exact mask from the query JPG\n\nBut the train image names have format 000000000.jpg while mask names are 00000.png, there are no masks that match any images in training dataset. Even if we prepand the extra 0s to mask name it doesn't seem to match either (the resolution of images is different). I checked it in [this](https://www.kaggle.com/code/michaln/masks-and-occlusions) notebook and couldn't find any matches.\n\nAre you sure that provided masks are related to the data in training dataset or did I just miss something?",
    "1734655": "The nomenclature is a bit confusing, but the train_masks apply to the (unseen) test images.  So hidden test will have a 00000.jpg that will go with the train_mask 00000.png that you can see.  At least this is my understanding.",
    "1734687": "Oh! I have figured out the source of the confusion. There was a mixup on the host end -- the \"train_masks\" folder should be named \"test_masks\" (I've asked the kaggle team to update this). There are no training masks provided. This matches the real world setting, where the \"test\" images (from investigations) have occlusions in the region of the image where the victim is located.\n\nTraining images, on the other hand, are not (by default) occluded. Competitors may choose to include occlusions in their training process, but we do not dictate that (or any other approach). If a competitor chose to incorporate masks, they could either generate their own, or repurpose the ones that match the test images (resizing them as necessary).\n\nThanks for the heads up on the train_masks issue -- I'm really sorry we didn't catch that sooner!",
    "1758111": "I discovered a strange behavior of test masks, but I don't know if it's a problem of my code or a problem on the test masks.\nI just added a double check that reapply the mask for every test images in the Pytorch dataset. I expect that if I reapply the mask the image will be unchanged. This is the code of the __getitem__ function:\n\n    def __getitem__(self, index):\n        path = self.sample[index]\n        img = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB)\n        # ------------   Double check code  ----------- // Start\n        # get the mask fileneame replacing jpg with png\n        mask_filename = os.path.basename(path).replace(\".jpg\", \".png\")\n        # Get the global path of mask (mask_root defined globally)\n        mask_path = os.path.join(mask_root, mask_filename)\n        if os.path.exists(mask_path):\n            # read mask\n            mask = cv2.imread(mask_path)\n            # Transform to GRAY to get the non-zero points\n            mask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\n            # Just re-apply the mask values\n            img[mask_g>0] = mask[mask_g>0]\n       # ------------   Double check code  ----------- // End\n        \n        img = cv2.resize(img, (test_input_size, test_input_size), interpolation=cv2.INTER_AREA)\n        img = self.transform(image=img)[\"image\"]\n        return img, path\n\nAfter this double check I expect the exact same performance as the __getitem__ without the \"Double check code\", but I obtained a drop of 0.1 in mAP. Let me know if it's my fault (probably) or a problem with masks.",
    "1758135": "You are applying the mask to the wrong color channel.\nEither do cv2.COLOR_BGR2RGB conversion for the mask before you convert it to gray (don't ask my why it works) or apply it manually to each channel (1 for red channel, 0s for others).\n\nCode to visualize\n```\nimport cv2\nimport matplotlib.pyplot as plt\n\ndef show_im_with_channels(image, title=None):\n    channels = [\"Red\", \"Green\", \"Blue\"]\n    \n    fig, ax = plt.subplots(1, 4, figsize=(22,6))\n    ax[0].imshow(image)\n    ax[0].set_ylabel(title)\n    if len(image.shape) > 2 and image.shape[2] > 1:\n        for i in range(3):\n            ax[i+1].imshow(image[:, :, i])\n            ax[i+1].set_title(channels[i])\n\nimg = cv2.cvtColor(cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_images/100206/000034116.jpg\"), cv2.COLOR_BGR2RGB)\nmask = cv2.imread(\"../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/train_masks/00000.png\")\nmask_g = cv2.cvtColor(mask, cv2.COLOR_RGB2GRAY)\nmasked_image = img.copy()\nmasked_image[mask_g>0] = mask[mask_g>0]\n\nshow_im_with_channels(img, \"Original image\")\nshow_im_with_channels(masked_image, \"Masked image\")\n\nshow_im_with_channels(mask, \"Mask\")\nshow_im_with_channels(mask_g, \"Gray mask\")\n\ntest_image = cv2.cvtColor(cv2.imread('../input/hotel-id-to-combat-human-trafficking-2022-fgvc9/test_images/abc.jpg'), cv2.COLOR_BGR2RGB)\nshow_im_with_channels(test_image, \"Test image\")\n```",
    "1758140": "You are right! Thank you",
    "1758292": "Ok, I fixed the BGR bug and I founded that the score has decreased by 0.03 mAP (better than before but not equal to the original one)! It's like the test images have different masks (maybe something different than the provided bounding boxes). I don't know of course if I am wrong, but I'm trying to understand better the problem, I am a bit confused.🤔",
    "1803282": "Have you figured it out? I checked the test picture again and I think because of the jpg compression the occlusion is not as perfect as we would expect it. The the main body of occlusion has values [254, 0, 0] and on the edges the values are even less clear. So I think when you recreate the occlusion using mask the result is not the same as the original image and can lead to different score."
  },
  "source": "meta"
}