{
  "id": 268314,
  "title": "SETI Grad Cam Images and Explanation.",
  "url": "/competitions/seti-breakthrough-listen/discussion/268314",
  "author_name": "Chris Deotte",
  "post_date": "2021-08-26T21:14:01.322000",
  "votes": 51,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>What is SETI Target=1?</h1>\n<p>Specifically what part of each positive image makes it a <code>target=1</code>? One way to investigate this is to use Grad Cam (also called class activation maps).</p>\n<h1>What is Grad Cam?</h1>\n<p>Many people don't realize that a pretrained CNN backbone outputs segmentation masks. That's right, a CNN backbone outputs segmentation masks. </p>\n<p>When we input a train image of size <code>512x512</code> into an EfficientNet, we get 1024 segmentation masks of dimension 16x16 output. (The outputs are 32x smaller). Each segmentation mask has been trained to find certain patterns. </p>\n<p>In the example below we input a 512x512 image and get output of segmentation masks that locate <code>red trapezoids</code>, <code>green triangles</code>, <code>orange hexagons</code>, and <code>blue circles</code>. The segmentation mask contains a one in every location of the original image that had this pattern and a zero otherwise. (More explanation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">here</a>)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam.png\" alt=\"\"></p>\n<p>Next we apply <code>global average pooling</code> which turns the 1024 segmentation masks into 1024 numbers. Each number is the average of 1s and 0s present in the segmentation mask. Therefore if a pattern does not exist, then the global average pooling is zero. And if the pattern is everywhere we would see an average pooling of one.</p>\n<p>Lastly we apply a dense layer with 1 sigmoid output. This dense layer multiplies each of the 1024 numbers by a weight coefficient <code>w1, w2, ..., w1023, w1024</code>. The model will use weight equal 1 when a pattern is important for <code>target=1</code> and weight equal 0 when a pattern is not important.</p>\n<p><strong>Grad cam is just the display of the important segmentation masks all combined into one overlapping image</strong>. We find this grad cam image, by inserting a train or test image into the CNN. We then display all the segmentation masks that are important as determined by the final dense layers weights.</p>\n<h1>Grad Cam Examples</h1>\n<p>In the plots below. The image on the left is the original \"on\" cadence spectrogram. (i.e. <code>np.vstack( img[::2] )</code>). The next image to the right is the same image but with emboss and histogram equalization to make it easier for our human eye to see. The next image is the \"off\" cadence image. And finally, the image on the right is the grad cam from a silver medal trained model. Lastly we take the contour of this grad cam and superimpose it on top of the second image.</p>\n<h1>Starter Notebook</h1>\n<p>I published a starter notebook here<br>\n<a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780</a><br>\n(and older discussion <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/116058\" target=\"_blank\">here</a> and older notebook <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\" target=\"_blank\">here</a> from Cloud comp)</p>\n<h1>SETI Grad Cam - Train OOF Examples</h1>\n<p>The examples below are 100% generated by code with <strong>no</strong> human labeling! The model makes a prediction and then circles what caused it to predict <code>target=1</code>.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_1.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_2.png\" alt=\"\"></p>\n<h1>SETI Grad Cam - Test Prediction Examples</h1>\n<p>The examples below are 100% generated by code with <strong>no</strong> human labeling! The model makes a prediction and then circles what caused it to predict <code>target=1</code>.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_4.png\" alt=\"\"></p>\n<h1>Applications</h1>\n<p>Using grad cam, we can see what makes a particular image a <code>target=1</code>. This helps us improve our model to better detect these patterns. It also helps us understand our models and demystifies the \"black box\". It's similar to boosted trees (XGB, LGBM, etc) feature importance.</p>",
  "messages": [
    {
      "id": 1492072,
      "postDate": "2021-08-26T21:14:01.323Z",
      "content": "<h1>What is SETI Target=1?</h1>\n<p>Specifically what part of each positive image makes it a <code>target=1</code>? One way to investigate this is to use Grad Cam (also called class activation maps).</p>\n<h1>What is Grad Cam?</h1>\n<p>Many people don't realize that a pretrained CNN backbone outputs segmentation masks. That's right, a CNN backbone outputs segmentation masks. </p>\n<p>When we input a train image of size <code>512x512</code> into an EfficientNet, we get 1024 segmentation masks of dimension 16x16 output. (The outputs are 32x smaller). Each segmentation mask has been trained to find certain patterns. </p>\n<p>In the example below we input a 512x512 image and get output of segmentation masks that locate <code>red trapezoids</code>, <code>green triangles</code>, <code>orange hexagons</code>, and <code>blue circles</code>. The segmentation mask contains a one in every location of the original image that had this pattern and a zero otherwise. (More explanation <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\" target=\"_blank\">here</a>)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam.png\" alt=\"\"></p>\n<p>Next we apply <code>global average pooling</code> which turns the 1024 segmentation masks into 1024 numbers. Each number is the average of 1s and 0s present in the segmentation mask. Therefore if a pattern does not exist, then the global average pooling is zero. And if the pattern is everywhere we would see an average pooling of one.</p>\n<p>Lastly we apply a dense layer with 1 sigmoid output. This dense layer multiplies each of the 1024 numbers by a weight coefficient <code>w1, w2, ..., w1023, w1024</code>. The model will use weight equal 1 when a pattern is important for <code>target=1</code> and weight equal 0 when a pattern is not important.</p>\n<p><strong>Grad cam is just the display of the important segmentation masks all combined into one overlapping image</strong>. We find this grad cam image, by inserting a train or test image into the CNN. We then display all the segmentation masks that are important as determined by the final dense layers weights.</p>\n<h1>Grad Cam Examples</h1>\n<p>In the plots below. The image on the left is the original \"on\" cadence spectrogram. (i.e. <code>np.vstack( img[::2] )</code>). The next image to the right is the same image but with emboss and histogram equalization to make it easier for our human eye to see. The next image is the \"off\" cadence image. And finally, the image on the right is the grad cam from a silver medal trained model. Lastly we take the contour of this grad cam and superimpose it on top of the second image.</p>\n<h1>Starter Notebook</h1>\n<p>I published a starter notebook here<br>\n<a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780</a><br>\n(and older discussion <a href=\"https://www.kaggle.com/c/understanding_cloud_organization/discussion/116058\" target=\"_blank\">here</a> and older notebook <a href=\"https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\" target=\"_blank\">here</a> from Cloud comp)</p>\n<h1>SETI Grad Cam - Train OOF Examples</h1>\n<p>The examples below are 100% generated by code with <strong>no</strong> human labeling! The model makes a prediction and then circles what caused it to predict <code>target=1</code>.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_1.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_2.png\" alt=\"\"></p>\n<h1>SETI Grad Cam - Test Prediction Examples</h1>\n<p>The examples below are 100% generated by code with <strong>no</strong> human labeling! The model makes a prediction and then circles what caused it to predict <code>target=1</code>.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png\" alt=\"\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_4.png\" alt=\"\"></p>\n<h1>Applications</h1>\n<p>Using grad cam, we can see what makes a particular image a <code>target=1</code>. This helps us improve our model to better detect these patterns. It also helps us understand our models and demystifies the \"black box\". It's similar to boosted trees (XGB, LGBM, etc) feature importance.</p>",
      "rawMarkdown": "# What is SETI Target=1?\nSpecifically what part of each positive image makes it a `target=1`? One way to investigate this is to use Grad Cam (also called class activation maps).\n# What is Grad Cam?\nMany people don't realize that a pretrained CNN backbone outputs segmentation masks. That's right, a CNN backbone outputs segmentation masks. \n\nWhen we input a train image of size `512x512` into an EfficientNet, we get 1024 segmentation masks of dimension 16x16 output. (The outputs are 32x smaller). Each segmentation mask has been trained to find certain patterns. \n\nIn the example below we input a 512x512 image and get output of segmentation masks that locate `red trapezoids`, `green triangles`, `orange hexagons`, and `blue circles`. The segmentation mask contains a one in every location of the original image that had this pattern and a zero otherwise. (More explanation [here][2])\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam.png)\n\nNext we apply `global average pooling` which turns the 1024 segmentation masks into 1024 numbers. Each number is the average of 1s and 0s present in the segmentation mask. Therefore if a pattern does not exist, then the global average pooling is zero. And if the pattern is everywhere we would see an average pooling of one.\n\nLastly we apply a dense layer with 1 sigmoid output. This dense layer multiplies each of the 1024 numbers by a weight coefficient `w1, w2, ..., w1023, w1024`. The model will use weight equal 1 when a pattern is important for `target=1` and weight equal 0 when a pattern is not important.\n\n**Grad cam is just the display of the important segmentation masks all combined into one overlapping image**. We find this grad cam image, by inserting a train or test image into the CNN. We then display all the segmentation masks that are important as determined by the final dense layers weights.\n\n\n# Grad Cam Examples\nIn the plots below. The image on the left is the original \"on\" cadence spectrogram. (i.e. `np.vstack( img[::2] )`). The next image to the right is the same image but with emboss and histogram equalization to make it easier for our human eye to see. The next image is the \"off\" cadence image. And finally, the image on the right is the grad cam from a silver medal trained model. Lastly we take the contour of this grad cam and superimpose it on top of the second image.\n\n# Starter Notebook\nI published a starter notebook here\nhttps://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\n(and older discussion [here][3] and older notebook [here][1] from Cloud comp)\n\n# SETI Grad Cam - Train OOF Examples\nThe examples below are 100% generated by code with **no** human labeling! The model makes a prediction and then circles what caused it to predict `target=1`.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_1.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_2.png)\n\n# SETI Grad Cam - Test Prediction Examples\nThe examples below are 100% generated by code with **no** human labeling! The model makes a prediction and then circles what caused it to predict `target=1`.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_4.png)\n\n# Applications\nUsing grad cam, we can see what makes a particular image a `target=1`. This helps us improve our model to better detect these patterns. It also helps us understand our models and demystifies the \"black box\". It's similar to boosted trees (XGB, LGBM, etc) feature importance.\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[3]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/116058",
      "votes": 51
    },
    {
      "id": 1496929,
      "postDate": "2021-08-30T18:17:09.730Z",
      "content": "<p>Really cool =)) , Thank u for sharing <a href=\"https://www.kaggle.com/cdeoette\" target=\"_blank\">@cdeoette</a> </p>",
      "rawMarkdown": "Really cool =)) , Thank u for sharing @cdeoette ",
      "votes": 2
    },
    {
      "id": 1494262,
      "postDate": "2021-08-28T14:30:38.653Z",
      "content": "<p>Probably my favorite explainability method when working with image neural networks. Thanks for the clear explanation and examples. 👌</p>",
      "rawMarkdown": "Probably my favorite explainability method when working with image neural networks. Thanks for the clear explanation and examples. 👌",
      "votes": 2
    },
    {
      "id": 1495701,
      "postDate": "2021-08-29T17:44:49.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1492840,
      "postDate": "2021-08-27T13:09:53.460Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 1492973,
          "postDate": "2021-08-27T14:59:30.907Z",
          "content": "<p>Great question. In this competition, i used grad cam (after the comp ended) to improve my models and make sure they were learning the correct thing. The CV LB gap implied that models were overfitting train and learning the wrong thing.</p>\n<p>In this competition, models were learning the background and not the foreground (signal). This is analogous to training a model to classify dogs and cats. In all the training images, we put dogs in cars and cats in boats. Then our models will predict dog when they see a car. And cat when they see a boat.</p>\n<p>Then for test images, we put dogs in boats and cats in cars. Our model makes all the wrong predictions on test data because our model sees a car and predicts dog when there is a cat present. And our model will have a large CV LB gap.</p>\n<p>By using grad cam, we can make sure our model is learning the signal and not a spurious correlation (i.e. the background). If our model is incorrectly learning a spurious correlation, then we need to determine how to remove the background. In this competition, the 1st place team removed the background by subtracting the background <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a>. Another technique is aggressive mixup <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a> with <code>alpha &gt; 3</code> and max target (i.e. <code>y_new = max(y_one, y_two)</code>.</p>",
          "rawMarkdown": "Great question. In this competition, i used grad cam (after the comp ended) to improve my models and make sure they were learning the correct thing. The CV LB gap implied that models were overfitting train and learning the wrong thing.\n\nIn this competition, models were learning the background and not the foreground (signal). This is analogous to training a model to classify dogs and cats. In all the training images, we put dogs in cars and cats in boats. Then our models will predict dog when they see a car. And cat when they see a boat.\n\nThen for test images, we put dogs in boats and cats in cars. Our model makes all the wrong predictions on test data because our model sees a car and predicts dog when there is a cat present. And our model will have a large CV LB gap.\n\nBy using grad cam, we can make sure our model is learning the signal and not a spurious correlation (i.e. the background). If our model is incorrectly learning a spurious correlation, then we need to determine how to remove the background. In this competition, the 1st place team removed the background by subtracting the background [here][1]. Another technique is aggressive mixup [here][2] with `alpha > 3` and max target (i.e. `y_new = max(y_one, y_two)`.\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\n[2]: https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780",
          "votes": 3
        }
      ]
    },
    {
      "id": 1495376,
      "postDate": "2021-08-29T14:11:07.273Z",
      "content": "<p>Awesome, thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
      "rawMarkdown": "Awesome, thanks for sharing @cdeotte !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1496929,
      "author_name": "fireflies",
      "author_url": "",
      "post_date": "2021-08-30T18:17:09.730000",
      "content": "<p>Really cool =)) , Thank u for sharing <a href=\"https://www.kaggle.com/cdeoette\" target=\"_blank\">@cdeoette</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1494262,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2021-08-28T14:30:38.653000",
      "content": "<p>Probably my favorite explainability method when working with image neural networks. Thanks for the clear explanation and examples. 👌</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1495701,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-29T17:44:49.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1492840,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-08-27T13:09:53.460000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 1492973,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-27T14:59:30.907000",
          "content": "<p>Great question. In this competition, i used grad cam (after the comp ended) to improve my models and make sure they were learning the correct thing. The CV LB gap implied that models were overfitting train and learning the wrong thing.</p>\n<p>In this competition, models were learning the background and not the foreground (signal). This is analogous to training a model to classify dogs and cats. In all the training images, we put dogs in cars and cats in boats. Then our models will predict dog when they see a car. And cat when they see a boat.</p>\n<p>Then for test images, we put dogs in boats and cats in cars. Our model makes all the wrong predictions on test data because our model sees a car and predicts dog when there is a cat present. And our model will have a large CV LB gap.</p>\n<p>By using grad cam, we can make sure our model is learning the signal and not a spurious correlation (i.e. the background). If our model is incorrectly learning a spurious correlation, then we need to determine how to remove the background. In this competition, the 1st place team removed the background by subtracting the background <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a>. Another technique is aggressive mixup <a href=\"https://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\" target=\"_blank\">here</a> with <code>alpha &gt; 3</code> and max target (i.e. <code>y_new = max(y_one, y_two)</code>.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1495376,
      "author_name": "Old Monk",
      "author_url": "",
      "post_date": "2021-08-29T14:11:07.273000",
      "content": "<p>Awesome, thanks for sharing <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1492072": "# What is SETI Target=1?\nSpecifically what part of each positive image makes it a `target=1`? One way to investigate this is to use Grad Cam (also called class activation maps).\n# What is Grad Cam?\nMany people don't realize that a pretrained CNN backbone outputs segmentation masks. That's right, a CNN backbone outputs segmentation masks. \n\nWhen we input a train image of size `512x512` into an EfficientNet, we get 1024 segmentation masks of dimension 16x16 output. (The outputs are 32x smaller). Each segmentation mask has been trained to find certain patterns. \n\nIn the example below we input a 512x512 image and get output of segmentation masks that locate `red trapezoids`, `green triangles`, `orange hexagons`, and `blue circles`. The segmentation mask contains a one in every location of the original image that had this pattern and a zero otherwise. (More explanation [here][2])\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam.png)\n\nNext we apply `global average pooling` which turns the 1024 segmentation masks into 1024 numbers. Each number is the average of 1s and 0s present in the segmentation mask. Therefore if a pattern does not exist, then the global average pooling is zero. And if the pattern is everywhere we would see an average pooling of one.\n\nLastly we apply a dense layer with 1 sigmoid output. This dense layer multiplies each of the 1024 numbers by a weight coefficient `w1, w2, ..., w1023, w1024`. The model will use weight equal 1 when a pattern is important for `target=1` and weight equal 0 when a pattern is not important.\n\n**Grad cam is just the display of the important segmentation masks all combined into one overlapping image**. We find this grad cam image, by inserting a train or test image into the CNN. We then display all the segmentation masks that are important as determined by the final dense layers weights.\n\n\n# Grad Cam Examples\nIn the plots below. The image on the left is the original \"on\" cadence spectrogram. (i.e. `np.vstack( img[::2] )`). The next image to the right is the same image but with emboss and histogram equalization to make it easier for our human eye to see. The next image is the \"off\" cadence image. And finally, the image on the right is the grad cam from a silver medal trained model. Lastly we take the contour of this grad cam and superimpose it on top of the second image.\n\n# Starter Notebook\nI published a starter notebook here\nhttps://www.kaggle.com/cdeotte/silver-medal-with-grad-cam-lb-0-780\n(and older discussion [here][3] and older notebook [here][1] from Cloud comp)\n\n# SETI Grad Cam - Train OOF Examples\nThe examples below are 100% generated by code with **no** human labeling! The model makes a prediction and then circles what caused it to predict `target=1`.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_1.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_2.png)\n\n# SETI Grad Cam - Test Prediction Examples\nThe examples below are 100% generated by code with **no** human labeling! The model makes a prediction and then circles what caused it to predict `target=1`.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_3.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/grad_cam_4.png)\n\n# Applications\nUsing grad cam, we can see what makes a particular image a `target=1`. This helps us improve our model to better detect these patterns. It also helps us understand our models and demystifies the \"black box\". It's similar to boosted trees (XGB, LGBM, etc) feature importance.\n\n[1]: https://www.kaggle.com/cdeotte/unsupervised-masks-cv-0-60\n[2]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160147\n[3]: https://www.kaggle.com/c/understanding_cloud_organization/discussion/116058",
    "1496929": "Really cool =)) , Thank u for sharing @cdeoette ",
    "1494262": "Probably my favorite explainability method when working with image neural networks. Thanks for the clear explanation and examples. 👌",
    "1495701": "",
    "1492840": "",
    "1495376": "Awesome, thanks for sharing @cdeotte !"
  }
}