{
  "id": 304104,
  "title": "Yolov5 feature maps - see how network understand our starfishes …",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/304104",
  "author_name": "",
  "post_date": "2022-01-30T21:32:49.811965200Z",
  "votes": 38,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Today was day of funny experiments. I used yolov5 and generated feature maps and Grad-CAM (I will publish soon). Feature maps are generated by applying Filters or Feature detectors to the input image. Feature map visualization will provide insight into the internal representations for specific input. Layer 17 has 256 feature maps, layer 20 has 512 and layer 23 has 1024 feature maps. So there are almost 1800 feature maps that you can look at in just the YOLOv5l output layers (for a single input image).</p>\n<p>Can you see starfish? 😄😍 </p>\n<p>Yolo5 architecture (source: <a href=\"https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444\" target=\"_blank\">https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444</a>)<br>\n<img src=\"https://i.ibb.co/R73NCwY/yolo5-architecture.png\" alt=\"Architecture\"></p>\n<p>Detection<br>\n<img src=\"https://i.ibb.co/fSWtVGq/115-EEE61-B4-C9-4-B47-93-B4-B3087954-F5-DF.jpg\" alt=\"map\"></p>\n<p>Feature map - section 23 of yolov5<br>\n<img src=\"https://i.ibb.co/FnfbCPk/9473-CE31-8-B9-D-4-F8-A-A286-97-DDA3355-E2-D.jpg\" alt=\"Section 23\"></p>\n<p><img src=\"https://i.ibb.co/c605xW6/image0.jpg\" alt=\"f1\"></p>\n<p><img src=\"https://i.ibb.co/KXvJTxm/image1.jpg\" alt=\"f1\"></p>\n<p>Feature map - section 17 of yolov5. Below you can see starfish and its features.<br>\n<img src=\"https://i.ibb.co/2FfFPhC/5-CC2-BC31-DA30-48-A2-BFAA-03877-AF98994.png\" alt=\"Section 17\"></p>",
  "messages": [
    {
      "id": "1669679",
      "postDate": "01/30/2022 21:32:49",
      "content": "<p>Today was day of funny experiments. I used yolov5 and generated feature maps and Grad-CAM (I will publish soon). Feature maps are generated by applying Filters or Feature detectors to the input image. Feature map visualization will provide insight into the internal representations for specific input. Layer 17 has 256 feature maps, layer 20 has 512 and layer 23 has 1024 feature maps. So there are almost 1800 feature maps that you can look at in just the YOLOv5l output layers (for a single input image).</p>\n<p>Can you see starfish? 😄😍 </p>\n<p>Yolo5 architecture (source: <a href=\"https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444\" target=\"_blank\">https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444</a>)<br>\n<img src=\"https://i.ibb.co/R73NCwY/yolo5-architecture.png\" alt=\"Architecture\"></p>\n<p>Detection<br>\n<img src=\"https://i.ibb.co/fSWtVGq/115-EEE61-B4-C9-4-B47-93-B4-B3087954-F5-DF.jpg\" alt=\"map\"></p>\n<p>Feature map - section 23 of yolov5<br>\n<img src=\"https://i.ibb.co/FnfbCPk/9473-CE31-8-B9-D-4-F8-A-A286-97-DDA3355-E2-D.jpg\" alt=\"Section 23\"></p>\n<p><img src=\"https://i.ibb.co/c605xW6/image0.jpg\" alt=\"f1\"></p>\n<p><img src=\"https://i.ibb.co/KXvJTxm/image1.jpg\" alt=\"f1\"></p>\n<p>Feature map - section 17 of yolov5. Below you can see starfish and its features.<br>\n<img src=\"https://i.ibb.co/2FfFPhC/5-CC2-BC31-DA30-48-A2-BFAA-03877-AF98994.png\" alt=\"Section 17\"></p>",
      "rawMarkdown": "Today was day of funny experiments. I used yolov5 and generated feature maps and Grad-CAM (I will publish soon). Feature maps are generated by applying Filters or Feature detectors to the input image. Feature map visualization will provide insight into the internal representations for specific input. Layer 17 has 256 feature maps, layer 20 has 512 and layer 23 has 1024 feature maps. So there are almost 1800 feature maps that you can look at in just the YOLOv5l output layers (for a single input image).\n\n Can you see starfish? 😄😍 \n\nYolo5 architecture (source: https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444)\n![Architecture](https://i.ibb.co/R73NCwY/yolo5-architecture.png)\n\nDetection\n![map](https://i.ibb.co/fSWtVGq/115-EEE61-B4-C9-4-B47-93-B4-B3087954-F5-DF.jpg)\n\nFeature map - section 23 of yolov5\n![Section 23](https://i.ibb.co/FnfbCPk/9473-CE31-8-B9-D-4-F8-A-A286-97-DDA3355-E2-D.jpg)\n\n![f1](https://i.ibb.co/c605xW6/image0.jpg)\n\n![f1](https://i.ibb.co/KXvJTxm/image1.jpg)\n\nFeature map - section 17 of yolov5. Below you can see starfish and its features.\n![Section 17](https://i.ibb.co/2FfFPhC/5-CC2-BC31-DA30-48-A2-BFAA-03877-AF98994.png)",
      "votes": null
    },
    {
      "id": "1670155",
      "postDate": "01/31/2022 09:04:02",
      "content": "<p>Model explanation is important to create ethical AI. I have used the saliency maps but Grad-cam looks like more advanced method </p>",
      "rawMarkdown": "Model explanation is important to create ethical AI. I have used the saliency maps but Grad-cam looks like more advanced method",
      "votes": null
    },
    {
      "id": "1670337",
      "postDate": "01/31/2022 12:40:19",
      "content": "<p>Now I see starfishes everywhere 🤔</p>",
      "rawMarkdown": "Now I see starfishes everywhere 🤔",
      "votes": null
    },
    {
      "id": "1670354",
      "postDate": "01/31/2022 12:57:18",
      "content": "<p>Don't worry, I have the same :)</p>",
      "rawMarkdown": "Don't worry, I have the same :)",
      "votes": null
    },
    {
      "id": "1670668",
      "postDate": "01/31/2022 18:27:10",
      "content": "<p>Cool<br>\nI like this analytics</p>",
      "rawMarkdown": "Cool\nI like this analytics",
      "votes": null
    },
    {
      "id": "1671282",
      "postDate": "02/01/2022 10:31:29",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> , I never understood how to interpret these feature maps.</p>\n<p>By the way, do you know how to interpret these feature maps with regards to classification (between different objects) ? Not applicable to this competition because there is only one object to classify, but for multi-object classification (e.g. between cats / dogs / humans in COCO dataset).</p>",
      "rawMarkdown": "Thanks @remekkinas , I never understood how to interpret these feature maps.\n\nBy the way, do you know how to interpret these feature maps with regards to classification (between different objects) ? Not applicable to this competition because there is only one object to classify, but for multi-object classification (e.g. between cats / dogs / humans in COCO dataset).",
      "votes": null
    },
    {
      "id": "1671416",
      "postDate": "02/01/2022 12:52:26",
      "content": "<p>Great question. Unfortunately I am not able to answer your question because I have not tested such cases (I used yolov5 for multiclass object detection but not generated feature maps). I will do it in my next project.</p>",
      "rawMarkdown": "Great question. Unfortunately I am not able to answer your question because I have not tested such cases (I used yolov5 for multiclass object detection but not generated feature maps). I will do it in my next project.",
      "votes": null
    },
    {
      "id": "1671438",
      "postDate": "02/01/2022 13:07:15",
      "content": "<p>we are interested to know which pixels (in the input image) contribute to the decision. e,g, the objectness map may be correct for COTS (i.e. it correctly localized the center object). but if the contribution is from a nearby rock (rather than the starfish itself), we may suspect it doesn't generalise well.</p>\n<p>(other example of wrong contributions, an image is classified as fish because of blue ocean, cow because it sits on green grass … all these cannot be reflected on eval meric. but these errors can be revealed by CAM map)</p>\n<p>in image classification, there is GAP pooling. i.e. decision comes from pooled region.</p>\n<p>in object detection, there is no pooling (for one stage methods like yolo. but there is roi pooling in faster rcnn), so CAM map is harder to generate.</p>\n<p>transformer has attention map, so it is easier to visualised.</p>",
      "rawMarkdown": "we are interested to know which pixels (in the input image) contribute to the decision. e,g, the objectness map may be correct for COTS (i.e. it correctly localized the center object). but if the contribution is from a nearby rock (rather than the starfish itself), we may suspect it doesn't generalise well.\n\n(other example of wrong contributions, an image is classified as fish because of blue ocean, cow because it sits on green grass ... all these cannot be reflected on eval meric. but these errors can be revealed by CAM map)\n\nin image classification, there is GAP pooling. i.e. decision comes from pooled region.\n\n\nin object detection, there is no pooling (for one stage methods like yolo. but there is roi pooling in faster rcnn), so CAM map is harder to generate.\n\ntransformer has attention map, so it is easier to visualised.",
      "votes": null
    },
    {
      "id": "1671458",
      "postDate": "02/01/2022 13:17:35",
      "content": "<p>this is cam map i have for cots image classifier (224x224)<br>\n<img src=\"https://i.ibb.co/bv3vFvL/Selection-022.png\" alt=\"https://i.ibb.co/bv3vFvL/Selection-022.png\"><br>\n<img src=\"https://i.ibb.co/vcbjzK7/Selection-023.png\" alt=\"https://i.ibb.co/vcbjzK7/Selection-023.png\"></p>\n<p>actually there are many mistakes. this prove that there is not enough samples to train a classifier</p>\n<p>model i used is : <a href=\"https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\" target=\"_blank\">https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md</a></p>\n<hr>\n<p>you may want to check this:<br>\n<a href=\"https://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf\" target=\"_blank\">https://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf</a></p>",
      "rawMarkdown": "this is cam map i have for cots image classifier (224x224)\n![https://i.ibb.co/bv3vFvL/Selection-022.png](https://i.ibb.co/bv3vFvL/Selection-022.png)\n![https://i.ibb.co/vcbjzK7/Selection-023.png](https://i.ibb.co/vcbjzK7/Selection-023.png)\n\nactually there are many mistakes. this prove that there is not enough samples to train a classifier\n\nmodel i used is : https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\n\n---\n\nyou may want to check this:\nhttps://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1670155,
      "author_name": "marcinstasko",
      "author_url": "",
      "post_date": "01/31/2022 09:04:02",
      "content": "<p>Model explanation is important to create ethical AI. I have used the saliency maps but Grad-cam looks like more advanced method </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1670337,
      "author_name": "robsonsan",
      "author_url": "",
      "post_date": "01/31/2022 12:40:19",
      "content": "<p>Now I see starfishes everywhere 🤔</p>",
      "votes": null,
      "replies": [
        {
          "id": 1670354,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "01/31/2022 12:57:18",
          "content": "<p>Don't worry, I have the same :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1670668,
      "author_name": "nicksergievskiy",
      "author_url": "",
      "post_date": "01/31/2022 18:27:10",
      "content": "<p>Cool<br>\nI like this analytics</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1671282,
      "author_name": "alexchwong",
      "author_url": "",
      "post_date": "02/01/2022 10:31:29",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> , I never understood how to interpret these feature maps.</p>\n<p>By the way, do you know how to interpret these feature maps with regards to classification (between different objects) ? Not applicable to this competition because there is only one object to classify, but for multi-object classification (e.g. between cats / dogs / humans in COCO dataset).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1671416,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/01/2022 12:52:26",
          "content": "<p>Great question. Unfortunately I am not able to answer your question because I have not tested such cases (I used yolov5 for multiclass object detection but not generated feature maps). I will do it in my next project.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1671438,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/01/2022 13:07:15",
          "content": "<p>we are interested to know which pixels (in the input image) contribute to the decision. e,g, the objectness map may be correct for COTS (i.e. it correctly localized the center object). but if the contribution is from a nearby rock (rather than the starfish itself), we may suspect it doesn't generalise well.</p>\n<p>(other example of wrong contributions, an image is classified as fish because of blue ocean, cow because it sits on green grass … all these cannot be reflected on eval meric. but these errors can be revealed by CAM map)</p>\n<p>in image classification, there is GAP pooling. i.e. decision comes from pooled region.</p>\n<p>in object detection, there is no pooling (for one stage methods like yolo. but there is roi pooling in faster rcnn), so CAM map is harder to generate.</p>\n<p>transformer has attention map, so it is easier to visualised.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1671458,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/01/2022 13:17:35",
          "content": "<p>this is cam map i have for cots image classifier (224x224)<br>\n<img src=\"https://i.ibb.co/bv3vFvL/Selection-022.png\" alt=\"https://i.ibb.co/bv3vFvL/Selection-022.png\"><br>\n<img src=\"https://i.ibb.co/vcbjzK7/Selection-023.png\" alt=\"https://i.ibb.co/vcbjzK7/Selection-023.png\"></p>\n<p>actually there are many mistakes. this prove that there is not enough samples to train a classifier</p>\n<p>model i used is : <a href=\"https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\" target=\"_blank\">https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md</a></p>\n<hr>\n<p>you may want to check this:<br>\n<a href=\"https://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf\" target=\"_blank\">https://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1669679": "Today was day of funny experiments. I used yolov5 and generated feature maps and Grad-CAM (I will publish soon). Feature maps are generated by applying Filters or Feature detectors to the input image. Feature map visualization will provide insight into the internal representations for specific input. Layer 17 has 256 feature maps, layer 20 has 512 and layer 23 has 1024 feature maps. So there are almost 1800 feature maps that you can look at in just the YOLOv5l output layers (for a single input image).\n\n Can you see starfish? 😄😍 \n\nYolo5 architecture (source: https://github.com/ultralytics/yolov5/issues/280#issuecomment-1000948444)\n![Architecture](https://i.ibb.co/R73NCwY/yolo5-architecture.png)\n\nDetection\n![map](https://i.ibb.co/fSWtVGq/115-EEE61-B4-C9-4-B47-93-B4-B3087954-F5-DF.jpg)\n\nFeature map - section 23 of yolov5\n![Section 23](https://i.ibb.co/FnfbCPk/9473-CE31-8-B9-D-4-F8-A-A286-97-DDA3355-E2-D.jpg)\n\n![f1](https://i.ibb.co/c605xW6/image0.jpg)\n\n![f1](https://i.ibb.co/KXvJTxm/image1.jpg)\n\nFeature map - section 17 of yolov5. Below you can see starfish and its features.\n![Section 17](https://i.ibb.co/2FfFPhC/5-CC2-BC31-DA30-48-A2-BFAA-03877-AF98994.png)",
    "1670155": "Model explanation is important to create ethical AI. I have used the saliency maps but Grad-cam looks like more advanced method",
    "1670337": "Now I see starfishes everywhere 🤔",
    "1670354": "Don't worry, I have the same :)",
    "1670668": "Cool\nI like this analytics",
    "1671282": "Thanks @remekkinas , I never understood how to interpret these feature maps.\n\nBy the way, do you know how to interpret these feature maps with regards to classification (between different objects) ? Not applicable to this competition because there is only one object to classify, but for multi-object classification (e.g. between cats / dogs / humans in COCO dataset).",
    "1671416": "Great question. Unfortunately I am not able to answer your question because I have not tested such cases (I used yolov5 for multiclass object detection but not generated feature maps). I will do it in my next project.",
    "1671438": "we are interested to know which pixels (in the input image) contribute to the decision. e,g, the objectness map may be correct for COTS (i.e. it correctly localized the center object). but if the contribution is from a nearby rock (rather than the starfish itself), we may suspect it doesn't generalise well.\n\n(other example of wrong contributions, an image is classified as fish because of blue ocean, cow because it sits on green grass ... all these cannot be reflected on eval meric. but these errors can be revealed by CAM map)\n\nin image classification, there is GAP pooling. i.e. decision comes from pooled region.\n\n\nin object detection, there is no pooling (for one stage methods like yolo. but there is roi pooling in faster rcnn), so CAM map is harder to generate.\n\ntransformer has attention map, so it is easier to visualised.",
    "1671458": "this is cam map i have for cots image classifier (224x224)\n![https://i.ibb.co/bv3vFvL/Selection-022.png](https://i.ibb.co/bv3vFvL/Selection-022.png)\n![https://i.ibb.co/vcbjzK7/Selection-023.png](https://i.ibb.co/vcbjzK7/Selection-023.png)\n\nactually there are many mistakes. this prove that there is not enough samples to train a classifier\n\nmodel i used is : https://github.com/facebookresearch/deit/blob/main/README_patchconvnet.md\n\n---\n\nyou may want to check this:\nhttps://openresearch-repository.anu.edu.au/bitstream/1885/232544/1/CNN_based_small_object_detection_and_visualization_with_feature_activation_map_edit.pdf"
  },
  "source": "meta"
}