{
  "id": 21994,
  "title": "heat map of CNN output",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/21994",
  "author_name": "hengck23",
  "post_date": "2016-07-01T16:16:53.263000",
  "votes": 29,
  "comment_count": 40,
  "views": 8405,
  "content": "<p>I am using the method CAM-VGG16 <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\">1</a> + finetune on the driver train images. Here, I generate the heatmap of the CNN response on the evaluation images. I also give the top 5 scores. This gives you an idea why your network is not performing well. Hope it helps!</p>\n\n<p>See attachment image.\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\" alt=\"enter image description here\">\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\">1</a> of <a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a> </p>",
  "messages": [
    {
      "id": 125696,
      "postDate": "2016-07-01T16:16:53.263Z",
      "content": "<p>I am using the method CAM-VGG16 <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\">1</a> + finetune on the driver train images. Here, I generate the heatmap of the CNN response on the evaluation images. I also give the top 5 scores. This gives you an idea why your network is not performing well. Hope it helps!</p>\n\n<p>See attachment image.\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\" alt=\"enter image description here\">\n<a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/125696/4524/Clipboard04.jpg\">1</a> of <a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a> </p>",
      "rawMarkdown": "I am using the method CAM-VGG16 [1] + finetune on the driver train images. Here, I generate the heatmap of the CNN response on the evaluation images. I also give the top 5 scores. This gives you an idea why your network is not performing well. Hope it helps!\n\nSee attachment image.\n  ![enter image description here][1]\n[1] of http://cnnlocalization.csail.mit.edu/ \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/125696/4524/Clipboard04.jpg",
      "votes": 26
    },
    {
      "id": 126399,
      "postDate": "2016-07-08T06:36:48.523Z",
      "content": "<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>",
      "rawMarkdown": "Nice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras",
      "votes": 3
    },
    {
      "id": 128873,
      "postDate": "2016-07-25T01:12:22.690Z",
      "content": "<p>@ ZFTurbo</p>\n\n<p>Actually it does compute all 10 classes at the same time.\nThe label is only used to plot the map corresponding to the true class of the image.</p>",
      "rawMarkdown": "@ ZFTurbo\r\n\r\nActually it does compute all 10 classes at the same time.\r\nThe label is only used to plot the map corresponding to the true class of the image.",
      "votes": 1
    },
    {
      "id": 128865,
      "postDate": "2016-07-24T23:22:45.753Z",
      "content": "<p>I only output the heatmap of the max class</p>\n\n<p>[quote=ZFTurbo;128855]</p>\n\n<p><strong><em>tmain</em></strong> - thanks for the script. I'm trying it now. Your function requires &quot;label&quot; as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like &quot;Heng CherKeng&quot;'s?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "I only output the heatmap of the max class\r\n\r\n[quote=ZFTurbo;128855]\r\n\r\n***tmain*** - thanks for the script. I'm trying it now. Your function requires \"label\" as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like \"Heng CherKeng\"'s?\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 126077,
      "postDate": "2016-07-06T01:32:14.187Z",
      "content": "<p>googlenet CAM + vgg16 CAM results</p>",
      "rawMarkdown": "googlenet CAM + vgg16 CAM results",
      "votes": 1
    },
    {
      "id": 125855,
      "postDate": "2016-07-03T14:43:07.427Z",
      "content": "<p>@DavidGbodiOdaibo\nHere are two way to use the CAM. I think this is going to work!</p>\n\n<ol>\n<li><p>one problem is that the training loss decreases too fast due to lack of training data. We can prevent this by making confusing train samples. </p>\n\n<ul><li>confuse_sample = train_sample + random_crop_of_another_image_of_different_class</li></ul></li>\n</ol>\n\n<p>by controlling how we mix the samples, we  make problem more difficult and the network can train &quot;infinitely&quot;. CAM can be use to guide the cropping of image region</p>\n\n<ol start=\"2\">\n<li>Instead of just learning the label (and assume we have the ground truth class activation region), use CAM learn both the region and label.  This prevent the network to learn the wrong features.</li>\n</ol>",
      "rawMarkdown": "@DavidGbodiOdaibo\r\nHere are two way to use the CAM. I think this is going to work!\r\n\r\n1. one problem is that the training loss decreases too fast due to lack of training data. We can prevent this by making confusing train samples. \r\n\r\n - confuse_sample = train_sample + random_crop_of_another_image_of_different_class\r\n\r\nby controlling how we mix the samples, we  make problem more difficult and the network can train \"infinitely\". CAM can be use to guide the cropping of image region\r\n\r\n\r\n\r\n2. Instead of just learning the label (and assume we have the ground truth class activation region), use CAM learn both the region and label.  This prevent the network to learn the wrong features.\r\n",
      "votes": 1
    },
    {
      "id": 125762,
      "postDate": "2016-07-02T08:32:42.580Z",
      "content": "<p>@DavidGbodiOdaibo\nSee attchment!</p>\n\n<p>[quote=DavidGbodiOdaibo;125698]</p>\n\n<p>This is really cool, I would like to see more heat maps for the &quot;hair and makeup samples&quot;,  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "@DavidGbodiOdaibo\r\nSee attchment!\r\n\r\n[quote=DavidGbodiOdaibo;125698]\r\n\r\nThis is really cool, I would like to see more heat maps for the \"hair and makeup samples\",  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.\r\n\r\n[/quote]\r\n",
      "votes": 1
    },
    {
      "id": 125891,
      "postDate": "2016-07-04T04:41:26.690Z",
      "content": "<p>The idea of using synthetic train samples works very well. This slow down convergence and prevent overfitting as well. Here is updated results on single googlenet-CAM:</p>\n\n<ul>\n<li>single crop: 0.28831</li>\n<li>mult icrop (different scale):0.26392</li>\n</ul>",
      "rawMarkdown": "The idea of using synthetic train samples works very well. This slow down convergence and prevent overfitting as well. Here is updated results on single googlenet-CAM:\r\n\r\n - single crop: 0.28831\r\n - mult icrop (different scale):0.26392\r\n",
      "votes": 2
    },
    {
      "id": 1137355,
      "postDate": "2021-01-03T21:28:29.180Z",
      "content": "<p><a href=\"https://www.kaggle.com/thibmain\" target=\"_blank\">@thibmain</a> thanks much for sharing this amazing Keras implementation. I am trying to use your code and I got the following error:</p>\n<p><code>KeyError: \"Can't open attribute (can't locate attribute: 'nb_layers')\"</code></p>\n<p>What could be the issue? Is it related to Keras? How can I solve this issue? Many thanks in advance.</p>",
      "rawMarkdown": "@thibmain thanks much for sharing this amazing Keras implementation. I am trying to use your code and I got the following error:\n\n`KeyError: \"Can't open attribute (can't locate attribute: 'nb_layers')\"`\n\nWhat could be the issue? Is it related to Keras? How can I solve this issue? Many thanks in advance."
    },
    {
      "id": 129278,
      "postDate": "2016-07-28T08:19:23.960Z",
      "content": "<p>@Ferris</p>\n\n<p>I'm not using the script, but i'm still doing the mean pixel subtraction &amp; float conversion. I only tried adding dropout to the CAM model after the flatten layer not the VGG-16 one above, I'm just doing that by iterating through the layers and setting weights to the pre-trained weights as required.</p>\n\n<p>In any case I'm giving up on this for now, I've tried all reasonable ranges of learning rates and have tried fine-tuning using a variety of different methods as well as heavy data augmentation, i'm not able to get any lower than a validation loss of 0.90 using these pre-trained models.</p>",
      "rawMarkdown": "@Ferris\r\n\r\nI'm not using the script, but i'm still doing the mean pixel subtraction & float conversion. I only tried adding dropout to the CAM model after the flatten layer not the VGG-16 one above, I'm just doing that by iterating through the layers and setting weights to the pre-trained weights as required.\r\n\r\nIn any case I'm giving up on this for now, I've tried all reasonable ranges of learning rates and have tried fine-tuning using a variety of different methods as well as heavy data augmentation, i'm not able to get any lower than a validation loss of 0.90 using these pre-trained models."
    },
    {
      "id": 129262,
      "postDate": "2016-07-28T03:30:38.720Z",
      "content": "<p>@tetmin</p>\n\n<p>Have you modified the mean pixel subtraction or float conversion part? (assuming you are using JiaoDong's script)</p>\n\n<p>If not, how did you add dropout layer after flatten layer? Here's the 'normal way',</p>\n\n<p>```</p>\n\n<pre><code>model.add(Flatten())\nmodel.add(Dense(4096, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(4096, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(1000, activation='softmax'))\n\nmodel.load_weights('../input/vgg16_weights.h5')\n\nmodel.layers.pop()  # Get rid of the classification layer\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\n\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "@tetmin\r\n\r\nHave you modified the mean pixel subtraction or float conversion part? (assuming you are using JiaoDong's script)\r\n\r\nIf not, how did you add dropout layer after flatten layer? Here's the 'normal way',\r\n\r\n\r\n```\r\n\r\n    model.add(Flatten())\r\n    model.add(Dense(4096, activation='relu'))\r\n    model.add(Dropout(0.5))\r\n    model.add(Dense(4096, activation='relu'))\r\n    model.add(Dropout(0.5))\r\n    model.add(Dense(1000, activation='softmax'))\r\n\r\n    model.load_weights('../input/vgg16_weights.h5')\r\n\r\n    model.layers.pop()  # Get rid of the classification layer\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n\r\n    model.add(Dense(10, activation='softmax'))\r\n```"
    },
    {
      "id": 129229,
      "postDate": "2016-07-27T17:26:33.683Z",
      "content": "<p>@tmain, thanks for that. I've now tried a range of learning rates from 1e-4 to 1e-7 with both ADAM and SGD as well as adding dropout after the Flatten layer and testing with batch normalisation. The best i've managed to get is a validation loss of 0.9 before the network overfits. I'm also using heavy image augmentation (scale, shift, shear, rotate) as suggested by some others.</p>\n\n<p>Were people really able to get losses of around 0.3 with what i've mentioned above or am I missing something? I'm just trying to get to the baseline of what others have reported quickly before experimenting further.</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "@tmain, thanks for that. I've now tried a range of learning rates from 1e-4 to 1e-7 with both ADAM and SGD as well as adding dropout after the Flatten layer and testing with batch normalisation. The best i've managed to get is a validation loss of 0.9 before the network overfits. I'm also using heavy image augmentation (scale, shift, shear, rotate) as suggested by some others.\r\n\r\nWere people really able to get losses of around 0.3 with what i've mentioned above or am I missing something? I'm just trying to get to the baseline of what others have reported quickly before experimenting further.\r\n\r\nThanks"
    },
    {
      "id": 129185,
      "postDate": "2016-07-27T11:53:35.620Z",
      "content": "<p>@tetmin</p>\n\n<p>Usually a batch size of 16 or 32. </p>",
      "rawMarkdown": "@tetmin\r\n\r\nUsually a batch size of 16 or 32. "
    },
    {
      "id": 129184,
      "postDate": "2016-07-27T11:48:34.757Z",
      "content": "<p>@tmain, thanks for that. What batch size were you using with that learning rate out of interest?</p>\n\n<p>For anyone trying to use this script with Tensorflow as the Keras backend, it can be sped up 2.5x by doing the following:\n- Change dimension ordering from th to tf\n- Transpose all the pre-trained weights of the convolutional layers so that they are in the usual Tensorflow order\n- Delete all of the ZeroPadding layers and add 'same' padding to the Conv2D layers</p>",
      "rawMarkdown": "@tmain, thanks for that. What batch size were you using with that learning rate out of interest?\r\n\r\nFor anyone trying to use this script with Tensorflow as the Keras backend, it can be sped up 2.5x by doing the following:\r\n- Change dimension ordering from th to tf\r\n- Transpose all the pre-trained weights of the convolutional layers so that they are in the usual Tensorflow order\r\n- Delete all of the ZeroPadding layers and add 'same' padding to the Conv2D layers"
    },
    {
      "id": 129141,
      "postDate": "2016-07-27T02:14:49.343Z",
      "content": "<p>@tetmin</p>\n\n<p>Normally I use Adam with standard parameters and learning rate of 1E-5.</p>",
      "rawMarkdown": "@tetmin\r\n\r\nNormally I use Adam with standard parameters and learning rate of 1E-5."
    },
    {
      "id": 129135,
      "postDate": "2016-07-27T00:14:47.063Z",
      "content": "<p>@tetmin</p>\n\n<p>I had the same problem.</p>\n\n<p>I added BN layer before the added convolution layer to the pre trained model. That helped the network to converge.</p>",
      "rawMarkdown": "@tetmin\r\n\r\nI had the same problem.\r\n\r\nI added BN layer before the added convolution layer to the pre trained model. That helped the network to converge."
    },
    {
      "id": 129096,
      "postDate": "2016-07-26T17:12:12.233Z",
      "content": "<p>What learning rate are people using to retrain this, with the default settings in tmain's github link this does not seem to converge, I get a huge loss of 14.70. A single epoch is also taking around 7000s on an AWS g2 instance, does that sound about right?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "What learning rate are people using to retrain this, with the default settings in tmain's github link this does not seem to converge, I get a huge loss of 14.70. A single epoch is also taking around 7000s on an AWS g2 instance, does that sound about right?\r\n\r\nThanks"
    },
    {
      "id": 129037,
      "postDate": "2016-07-26T03:31:04.413Z",
      "content": "<p>very nice work!</p>\n\n<p>[quote=tmain;126399]</p>\n\n<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "very nice work!\r\n\r\n[quote=tmain;126399]\r\n\r\nNice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 128899,
      "postDate": "2016-07-25T06:31:07.057Z",
      "content": "<p><strong>Heng CherKeng, tmain</strong>, thanks. This makes sense.</p>",
      "rawMarkdown": "**Heng CherKeng, tmain**, thanks. This makes sense."
    },
    {
      "id": 128855,
      "postDate": "2016-07-24T20:13:50.520Z",
      "content": "<p><strong><em>tmain</em></strong> - thanks for the script. I'm trying it now. Your function requires &quot;label&quot; as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like &quot;Heng CherKeng&quot;'s?</p>",
      "rawMarkdown": "***tmain*** - thanks for the script. I'm trying it now. Your function requires \"label\" as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like \"Heng CherKeng\"'s?"
    },
    {
      "id": 128734,
      "postDate": "2016-07-23T08:24:41.250Z",
      "content": "<p>Hello Heng,</p>\n\n<p>I have a question about how you choose which epoch to use for testing/submission. The attached figure is for learning with Googlenet. Around epoch 6 i get the best loss3 on validation and i tend to use model from that snapshot. Is that the right way to do it?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hello Heng,\r\n\r\nI have a question about how you choose which epoch to use for testing/submission. The attached figure is for learning with Googlenet. Around epoch 6 i get the best loss3 on validation and i tend to use model from that snapshot. Is that the right way to do it?\r\n\r\nThanks"
    },
    {
      "id": 128699,
      "postDate": "2016-07-22T20:06:35.223Z",
      "content": "<p>&quot;downscaling and upscaling the images and then cropping/padding&quot; is correct. But later i found that using ensemble of models of single original crops gives the best results. But it has to depend on how much perturbation of scale you used in training, it is trial and error kind of work.</p>\n\n<p>[quote=Rishab Gargeya;128680]</p>\n\n<p>@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say &quot;center crop, 0.8&quot; and &quot;center crop, 1.2&quot;. Are you downscaling and upscaling the images and then cropping/padding?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "\"downscaling and upscaling the images and then cropping/padding\" is correct. But later i found that using ensemble of models of single original crops gives the best results. But it has to depend on how much perturbation of scale you used in training, it is trial and error kind of work.\r\n\r\n[quote=Rishab Gargeya;128680]\r\n\r\n@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say \"center crop, 0.8\" and \"center crop, 1.2\". Are you downscaling and upscaling the images and then cropping/padding?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 128680,
      "postDate": "2016-07-22T14:56:13.530Z",
      "content": "<p>@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say &quot;center crop, 0.8&quot; and &quot;center crop, 1.2&quot;. Are you downscaling and upscaling the images and then cropping/padding?</p>",
      "rawMarkdown": "@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say \"center crop, 0.8\" and \"center crop, 1.2\". Are you downscaling and upscaling the images and then cropping/padding?"
    },
    {
      "id": 127915,
      "postDate": "2016-07-16T07:44:11.873Z",
      "content": "<p>I am not sure what you mean by that. The procedure I followed was:</p>\n\n<ul>\n<li>Use pre-trained VGG16 weights and train (cf. JiaoDong's  keras post). Save the weights of the newly trained VGGCAM model.</li>\n<li>Re-use these trained weights to get the Activation Map </li>\n</ul>",
      "rawMarkdown": "I am not sure what you mean by that. The procedure I followed was:\r\n\r\n- Use pre-trained VGG16 weights and train (cf. JiaoDong's  keras post). Save the weights of the newly trained VGGCAM model.\r\n- Re-use these trained weights to get the Activation Map "
    },
    {
      "id": 127644,
      "postDate": "2016-07-14T19:51:37.920Z",
      "content": "<p>How did you combine those models? </p>\n\n<p>[quote=tmain;126399]</p>\n\n<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "How did you combine those models? \r\n\r\n\r\n[quote=tmain;126399]\r\n\r\nNice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 126427,
      "postDate": "2016-07-08T11:50:51.223Z",
      "content": "<p>@tmain nice work!</p>",
      "rawMarkdown": "@tmain nice work!"
    },
    {
      "id": 126069,
      "postDate": "2016-07-05T23:47:39.360Z",
      "content": "<p>It is always interesting to see results on out-of-sample data .... see attachment.\nIt test my system on some random images download from the internet. </p>",
      "rawMarkdown": "It is always interesting to see results on out-of-sample data .... see attachment.\r\nIt test my system on some random images download from the internet. "
    },
    {
      "id": 126023,
      "postDate": "2016-07-05T15:49:25.183Z",
      "content": "<p>Thanks for the idea :)\nSaw this and thought we can play with it too <a href=\"http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv\">http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv</a></p>",
      "rawMarkdown": "Thanks for the idea :)\r\nSaw this and thought we can play with it too http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv"
    },
    {
      "id": 125953,
      "postDate": "2016-07-04T23:29:31.890Z",
      "content": "<p>@shenzhenwei\nhere are the details.</p>\n\n<ul>\n<li>croping and mixing patches  for training.\nThis my idea. You can simply put random patches and get results by trial and error. The closest paper that apply this is &quot;Frankenstein: Learning Deep Face Representations using Small Data&quot;</li>\n<li><p>generating deep feature results. Let the network be:</p>\n\n<p>conv ---&gt; ave  ---&gt; ip</p></li>\n</ul>\n\n<p>for one sample, say the dims = C x H X W of the data are:  conv = 1024 x 14 x 14, ave = 1024 x1 x1, ip = 10 x 1 x 1. The weights of inner products are: weights = 10 x 1024. to create heat_map of the k-class (out of 10 classes), use heat_map_k = sum_c { conv_c * weight_k_c }. c runs from 1 to 1024. heat_map_k = 1x14x14. upscale heat_map_k to input image size to overlay results.</p>",
      "rawMarkdown": "@shenzhenwei\r\nhere are the details.\r\n\r\n - croping and mixing patches  for training.\r\n This my idea. You can simply put random patches and get results by trial and error. The closest paper that apply this is \"Frankenstein: Learning Deep Face Representations using Small Data\"\r\n - generating deep feature results. Let the network be:\r\n\r\n\r\n conv ---> ave  ---> ip\r\n\r\nfor one sample, say the dims = C x H X W of the data are:  conv = 1024 x 14 x 14, ave = 1024 x1 x1, ip = 10 x 1 x 1. The weights of inner products are: weights = 10 x 1024. to create heat_map of the k-class (out of 10 classes), use heat_map_k = sum_c { conv_c * weight_k_c }. c runs from 1 to 1024. heat_map_k = 1x14x14. upscale heat_map_k to input image size to overlay results.\r\n\r\n"
    },
    {
      "id": 125885,
      "postDate": "2016-07-04T02:22:01.060Z",
      "content": "<p>This is cool.  Thank you for your sharing.</p>",
      "rawMarkdown": "This is cool.  Thank you for your sharing."
    },
    {
      "id": 125853,
      "postDate": "2016-07-03T12:59:46.607Z",
      "content": "<p>upsampling the  14x14 output is only for visualisation (the heat map is 14x14).  The upsampling is simply any image resizing of the heat map to the input image size, so that we can overlay the heatmap onto the image.</p>\n\n<p>You do not need to do upsampling if you only want to get the class score. see also the attachment for the network structure.</p>\n\n<p>[quote=DavidGbodiOdaibo;125852]</p>\n\n<p>is anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?</p>\n\n<p>[/quote]</p>",
      "rawMarkdown": "upsampling the  14x14 output is only for visualisation (the heat map is 14x14).  The upsampling is simply any image resizing of the heat map to the input image size, so that we can overlay the heatmap onto the image.\r\n\r\nYou do not need to do upsampling if you only want to get the class score. see also the attachment for the network structure.\r\n\r\n\r\n\r\n[quote=DavidGbodiOdaibo;125852]\r\n\r\nis anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?\r\n\r\n[/quote]\r\n"
    },
    {
      "id": 125852,
      "postDate": "2016-07-03T12:55:03.270Z",
      "content": "<p>is anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?</p>",
      "rawMarkdown": "is anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?"
    },
    {
      "id": 125842,
      "postDate": "2016-07-03T11:11:35.450Z",
      "content": "<p>These CAM models are good. Single googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. Please see attachment. I am not sure if the results comes from the CAM structure or my training method. Here are the training details:</p>\n\n<ol>\n<li><p>split of train and validation set. Divide train images into  groups of {driver,action}.\n   random select {driver,action} for training and validation.</p></li>\n<li><p>use 320x240 as image size into the CAM network</p></li>\n<li><p>apply aggressive augmentation. I use scale change of +/- 0.2, displacement of random shift up to 0.2*MIN(width,height). Further, zero out a random block of 60x60 in each train image to simulate occlusion.</p></li>\n</ol>",
      "rawMarkdown": "These CAM models are good. Single googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. Please see attachment. I am not sure if the results comes from the CAM structure or my training method. Here are the training details:\r\n\r\n 1.  split of train and validation set. Divide train images into  groups of {driver,action}.\r\n       random select {driver,action} for training and validation.\r\n\r\n 2.  use 320x240 as image size into the CAM network\r\n\r\n 3. apply aggressive augmentation. I use scale change of +/- 0.2, displacement of random shift up to 0.2*MIN(width,height). Further, zero out a random block of 60x60 in each train image to simulate occlusion.\r\n\r\n "
    },
    {
      "id": 125698,
      "postDate": "2016-07-01T16:28:20.300Z",
      "content": "<p>This is really cool, I would like to see more heat maps for the &quot;hair and makeup samples&quot;,  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.</p>",
      "rawMarkdown": "This is really cool, I would like to see more heat maps for the \"hair and makeup samples\",  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down."
    },
    {
      "id": 125697,
      "postDate": "2016-07-01T16:21:11.740Z",
      "content": "<p>It looks <em>awesome</em>.</p>",
      "rawMarkdown": "It looks *awesome*."
    },
    {
      "id": 175242,
      "postDate": "2017-04-14T14:44:54.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 125960,
      "postDate": "2016-07-05T01:49:11.907Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 125930,
      "postDate": "2016-07-04T16:16:00.020Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 175148,
          "postDate": "2017-04-14T01:36:08.247Z",
          "content": "<p>@shenzhenwei saw one iccv 2017 paper tody: <a href=\"https://arxiv.org/pdf/1704.04232.pdf\">https://arxiv.org/pdf/1704.04232.pdf</a>\n\"Our key idea is to hide patches in a training image randomly, forcing the\nnetwork to seek other relevant parts when the most dis-criminative part is hidden\"</p>\n\n<p>\"Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-supervised Object and Action Localization\"- Krishna Kumar Singh and Yong Jae Lee (University of California, Davis), ICCV 2017</p>",
          "rawMarkdown": "@shenzhenwei saw one iccv 2017 paper tody: https://arxiv.org/pdf/1704.04232.pdf\n\"Our key idea is to hide patches in a training image randomly, forcing the\nnetwork to seek other relevant parts when the most dis-criminative part is hidden\"\n\n\"Hide-and-Seek: Forcing a Network to be Meticulous for Weakly-supervised Object and Action Localization\"- Krishna Kumar Singh and Yong Jae Lee (University of California, Davis), ICCV 2017"
        }
      ]
    },
    {
      "id": 125854,
      "postDate": "2016-07-03T13:46:16.307Z",
      "content": "<p>@Heng CherKeng  thanks</p>",
      "rawMarkdown": "@Heng CherKeng  thanks"
    },
    {
      "id": 125739,
      "postDate": "2016-07-02T03:51:05.133Z",
      "content": "<p>Interesting! Thanks for your sharing!</p>",
      "rawMarkdown": "Interesting! Thanks for your sharing!"
    }
  ],
  "comments": [
    {
      "id": 126399,
      "author_name": "tmain",
      "author_url": "",
      "post_date": "2016-07-08T06:36:48.523000",
      "content": "<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 128873,
      "author_name": "tmain",
      "author_url": "",
      "post_date": "2016-07-25T01:12:22.690000",
      "content": "<p>@ ZFTurbo</p>\n\n<p>Actually it does compute all 10 classes at the same time.\nThe label is only used to plot the map corresponding to the true class of the image.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 128865,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-24T23:22:45.753000",
      "content": "<p>I only output the heatmap of the max class</p>\n\n<p>[quote=ZFTurbo;128855]</p>\n\n<p><strong><em>tmain</em></strong> - thanks for the script. I'm trying it now. Your function requires &quot;label&quot; as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like &quot;Heng CherKeng&quot;'s?</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 126077,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-06T01:32:14.187000",
      "content": "<p>googlenet CAM + vgg16 CAM results</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125855,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-03T14:43:07.427000",
      "content": "<p>@DavidGbodiOdaibo\nHere are two way to use the CAM. I think this is going to work!</p>\n\n<ol>\n<li><p>one problem is that the training loss decreases too fast due to lack of training data. We can prevent this by making confusing train samples. </p>\n\n<ul><li>confuse_sample = train_sample + random_crop_of_another_image_of_different_class</li></ul></li>\n</ol>\n\n<p>by controlling how we mix the samples, we  make problem more difficult and the network can train &quot;infinitely&quot;. CAM can be use to guide the cropping of image region</p>\n\n<ol start=\"2\">\n<li>Instead of just learning the label (and assume we have the ground truth class activation region), use CAM learn both the region and label.  This prevent the network to learn the wrong features.</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125762,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-02T08:32:42.580000",
      "content": "<p>@DavidGbodiOdaibo\nSee attchment!</p>\n\n<p>[quote=DavidGbodiOdaibo;125698]</p>\n\n<p>This is really cool, I would like to see more heat maps for the &quot;hair and makeup samples&quot;,  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.</p>\n\n<p>[/quote]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 125891,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-04T04:41:26.690000",
      "content": "<p>The idea of using synthetic train samples works very well. This slow down convergence and prevent overfitting as well. Here is updated results on single googlenet-CAM:</p>\n\n<ul>\n<li>single crop: 0.28831</li>\n<li>mult icrop (different scale):0.26392</li>\n</ul>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1137355,
      "author_name": "Elton Brasil da Costa",
      "author_url": "",
      "post_date": "2021-01-03T21:28:29.180000",
      "content": "<p><a href=\"https://www.kaggle.com/thibmain\" target=\"_blank\">@thibmain</a> thanks much for sharing this amazing Keras implementation. I am trying to use your code and I got the following error:</p>\n<p><code>KeyError: \"Can't open attribute (can't locate attribute: 'nb_layers')\"</code></p>\n<p>What could be the issue? Is it related to Keras? How can I solve this issue? Many thanks in advance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129278,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-07-28T08:19:23.960000",
      "content": "<p>@Ferris</p>\n\n<p>I'm not using the script, but i'm still doing the mean pixel subtraction &amp; float conversion. I only tried adding dropout to the CAM model after the flatten layer not the VGG-16 one above, I'm just doing that by iterating through the layers and setting weights to the pre-trained weights as required.</p>\n\n<p>In any case I'm giving up on this for now, I've tried all reasonable ranges of learning rates and have tried fine-tuning using a variety of different methods as well as heavy data augmentation, i'm not able to get any lower than a validation loss of 0.90 using these pre-trained models.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129262,
      "author_name": "Ferris Wu",
      "author_url": "",
      "post_date": "2016-07-28T03:30:38.720000",
      "content": "<p>@tetmin</p>\n\n<p>Have you modified the mean pixel subtraction or float conversion part? (assuming you are using JiaoDong's script)</p>\n\n<p>If not, how did you add dropout layer after flatten layer? Here's the 'normal way',</p>\n\n<p>```</p>\n\n<pre><code>model.add(Flatten())\nmodel.add(Dense(4096, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(4096, activation='relu'))\nmodel.add(Dropout(0.5))\nmodel.add(Dense(1000, activation='softmax'))\n\nmodel.load_weights('../input/vgg16_weights.h5')\n\nmodel.layers.pop()  # Get rid of the classification layer\nmodel.outputs = [model.layers[-1].output]\nmodel.layers[-1].outbound_nodes = []\n\nmodel.add(Dense(10, activation='softmax'))\n</code></pre>\n\n<p>```</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129229,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-07-27T17:26:33.683000",
      "content": "<p>@tmain, thanks for that. I've now tried a range of learning rates from 1e-4 to 1e-7 with both ADAM and SGD as well as adding dropout after the Flatten layer and testing with batch normalisation. The best i've managed to get is a validation loss of 0.9 before the network overfits. I'm also using heavy image augmentation (scale, shift, shear, rotate) as suggested by some others.</p>\n\n<p>Were people really able to get losses of around 0.3 with what i've mentioned above or am I missing something? I'm just trying to get to the baseline of what others have reported quickly before experimenting further.</p>\n\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129185,
      "author_name": "tmain",
      "author_url": "",
      "post_date": "2016-07-27T11:53:35.620000",
      "content": "<p>@tetmin</p>\n\n<p>Usually a batch size of 16 or 32. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129184,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-07-27T11:48:34.757000",
      "content": "<p>@tmain, thanks for that. What batch size were you using with that learning rate out of interest?</p>\n\n<p>For anyone trying to use this script with Tensorflow as the Keras backend, it can be sped up 2.5x by doing the following:\n- Change dimension ordering from th to tf\n- Transpose all the pre-trained weights of the convolutional layers so that they are in the usual Tensorflow order\n- Delete all of the ZeroPadding layers and add 'same' padding to the Conv2D layers</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129141,
      "author_name": "tmain",
      "author_url": "",
      "post_date": "2016-07-27T02:14:49.343000",
      "content": "<p>@tetmin</p>\n\n<p>Normally I use Adam with standard parameters and learning rate of 1E-5.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129135,
      "author_name": "kAI",
      "author_url": "",
      "post_date": "2016-07-27T00:14:47.063000",
      "content": "<p>@tetmin</p>\n\n<p>I had the same problem.</p>\n\n<p>I added BN layer before the added convolution layer to the pre trained model. That helped the network to converge.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129096,
      "author_name": "tetmin",
      "author_url": "",
      "post_date": "2016-07-26T17:12:12.233000",
      "content": "<p>What learning rate are people using to retrain this, with the default settings in tmain's github link this does not seem to converge, I get a huge loss of 14.70. A single epoch is also taking around 7000s on an AWS g2 instance, does that sound about right?</p>\n\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 129037,
      "author_name": "Vinh Nguyen",
      "author_url": "",
      "post_date": "2016-07-26T03:31:04.413000",
      "content": "<p>very nice work!</p>\n\n<p>[quote=tmain;126399]</p>\n\n<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>\n\n<p>[/quote]</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128899,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-07-25T06:31:07.057000",
      "content": "<p><strong>Heng CherKeng, tmain</strong>, thanks. This makes sense.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128855,
      "author_name": "ZFTurbo",
      "author_url": "",
      "post_date": "2016-07-24T20:13:50.520000",
      "content": "<p><strong><em>tmain</em></strong> - thanks for the script. I'm trying it now. Your function requires &quot;label&quot; as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like &quot;Heng CherKeng&quot;'s?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128734,
      "author_name": "datapool",
      "author_url": "",
      "post_date": "2016-07-23T08:24:41.250000",
      "content": "<p>Hello Heng,</p>\n\n<p>I have a question about how you choose which epoch to use for testing/submission. The attached figure is for learning with Googlenet. Around epoch 6 i get the best loss3 on validation and i tend to use model from that snapshot. Is that the right way to do it?</p>\n\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128699,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-22T20:06:35.223000",
      "content": "<p>&quot;downscaling and upscaling the images and then cropping/padding&quot; is correct. But later i found that using ensemble of models of single original crops gives the best results. But it has to depend on how much perturbation of scale you used in training, it is trial and error kind of work.</p>\n\n<p>[quote=Rishab Gargeya;128680]</p>\n\n<p>@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say &quot;center crop, 0.8&quot; and &quot;center crop, 1.2&quot;. Are you downscaling and upscaling the images and then cropping/padding?</p>\n\n<p>[/quote]</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 128680,
      "author_name": "Rishab Gargeya",
      "author_url": "",
      "post_date": "2016-07-22T14:56:13.530000",
      "content": "<p>@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say &quot;center crop, 0.8&quot; and &quot;center crop, 1.2&quot;. Are you downscaling and upscaling the images and then cropping/padding?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 127915,
      "author_name": "tmain",
      "author_url": "",
      "post_date": "2016-07-16T07:44:11.873000",
      "content": "<p>I am not sure what you mean by that. The procedure I followed was:</p>\n\n<ul>\n<li>Use pre-trained VGG16 weights and train (cf. JiaoDong's  keras post). Save the weights of the newly trained VGGCAM model.</li>\n<li>Re-use these trained weights to get the Activation Map </li>\n</ul>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 127644,
      "author_name": "Wenxin",
      "author_url": "",
      "post_date": "2016-07-14T19:51:37.920000",
      "content": "<p>How did you combine those models? </p>\n\n<p>[quote=tmain;126399]</p>\n\n<p>Nice find !</p>\n\n<p>For those interested, I have written a quick keras implementation: <a href=\"https://github.com/tdeboissiere/VGG16CAM-keras\">github link</a></p>\n\n<p>[/quote]</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126427,
      "author_name": "DavidGbodiOdaibo",
      "author_url": "",
      "post_date": "2016-07-08T11:50:51.223000",
      "content": "<p>@tmain nice work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126069,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-05T23:47:39.360000",
      "content": "<p>It is always interesting to see results on out-of-sample data .... see attachment.\nIt test my system on some random images download from the internet. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 126023,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-05T15:49:25.183000",
      "content": "<p>Thanks for the idea :)\nSaw this and thought we can play with it too <a href=\"http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv\">http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125953,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2016-07-04T23:29:31.890000",
      "content": "<p>@shenzhenwei\nhere are the details.</p>\n\n<ul>\n<li>croping and mixing patches  for training.\nThis my idea. You can simply put random patches and get results by trial and error. The closest paper that apply this is &quot;Frankenstein: Learning Deep Face Representations using Small Data&quot;</li>\n<li><p>generating deep feature results. Let the network be:</p>\n\n<p>conv ---&gt; ave  ---&gt; ip</p></li>\n</ul>\n\n<p>for one sample, say the dims = C x H X W of the data are:  conv = 1024 x 14 x 14, ave = 1024 x1 x1, ip = 10 x 1 x 1. The weights of inner products are: weights = 10 x 1024. to create heat_map of the k-class (out of 10 classes), use heat_map_k = sum_c { conv_c * weight_k_c }. c runs from 1 to 1024. heat_map_k = 1x14x14. upscale heat_map_k to input image size to overlay results.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125885,
      "author_name": "l_mahome",
      "author_url": "",
      "post_date": "2016-07-04T02:22:01.060000",
      "content": "<p>This is cool.  Thank you for your sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125853,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-03T12:59:46.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125852,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-03T12:55:03.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125842,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-03T11:11:35.450000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125698,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T16:28:20.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125697,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-01T16:21:11.740000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 175242,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-04-14T14:44:54.407000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125960,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-05T01:49:11.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125930,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-04T16:16:00.020000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 175148,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-04-14T01:36:08.247000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 125854,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-03T13:46:16.307000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 125739,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-07-02T03:51:05.133000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "125696": "I am using the method CAM-VGG16 [1] + finetune on the driver train images. Here, I generate the heatmap of the CNN response on the evaluation images. I also give the top 5 scores. This gives you an idea why your network is not performing well. Hope it helps!\n\nSee attachment image.\n  ![enter image description here][1]\n[1] of http://cnnlocalization.csail.mit.edu/ \n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/125696/4524/Clipboard04.jpg",
    "126399": "Nice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras",
    "128873": "@ ZFTurbo\r\n\r\nActually it does compute all 10 classes at the same time.\r\nThe label is only used to plot the map corresponding to the true class of the image.",
    "128865": "I only output the heatmap of the max class\r\n\r\n[quote=ZFTurbo;128855]\r\n\r\n***tmain*** - thanks for the script. I'm trying it now. Your function requires \"label\" as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like \"Heng CherKeng\"'s?\r\n\r\n[/quote]\r\n",
    "126077": "googlenet CAM + vgg16 CAM results",
    "125855": "@DavidGbodiOdaibo\r\nHere are two way to use the CAM. I think this is going to work!\r\n\r\n1. one problem is that the training loss decreases too fast due to lack of training data. We can prevent this by making confusing train samples. \r\n\r\n - confuse_sample = train_sample + random_crop_of_another_image_of_different_class\r\n\r\nby controlling how we mix the samples, we  make problem more difficult and the network can train \"infinitely\". CAM can be use to guide the cropping of image region\r\n\r\n\r\n\r\n2. Instead of just learning the label (and assume we have the ground truth class activation region), use CAM learn both the region and label.  This prevent the network to learn the wrong features.\r\n",
    "125762": "@DavidGbodiOdaibo\r\nSee attchment!\r\n\r\n[quote=DavidGbodiOdaibo;125698]\r\n\r\nThis is really cool, I would like to see more heat maps for the \"hair and makeup samples\",  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.\r\n\r\n[/quote]\r\n",
    "125891": "The idea of using synthetic train samples works very well. This slow down convergence and prevent overfitting as well. Here is updated results on single googlenet-CAM:\r\n\r\n - single crop: 0.28831\r\n - mult icrop (different scale):0.26392\r\n",
    "1137355": "@thibmain thanks much for sharing this amazing Keras implementation. I am trying to use your code and I got the following error:\n\n`KeyError: \"Can't open attribute (can't locate attribute: 'nb_layers')\"`\n\nWhat could be the issue? Is it related to Keras? How can I solve this issue? Many thanks in advance.",
    "129278": "@Ferris\r\n\r\nI'm not using the script, but i'm still doing the mean pixel subtraction & float conversion. I only tried adding dropout to the CAM model after the flatten layer not the VGG-16 one above, I'm just doing that by iterating through the layers and setting weights to the pre-trained weights as required.\r\n\r\nIn any case I'm giving up on this for now, I've tried all reasonable ranges of learning rates and have tried fine-tuning using a variety of different methods as well as heavy data augmentation, i'm not able to get any lower than a validation loss of 0.90 using these pre-trained models.",
    "129262": "@tetmin\r\n\r\nHave you modified the mean pixel subtraction or float conversion part? (assuming you are using JiaoDong's script)\r\n\r\nIf not, how did you add dropout layer after flatten layer? Here's the 'normal way',\r\n\r\n\r\n```\r\n\r\n    model.add(Flatten())\r\n    model.add(Dense(4096, activation='relu'))\r\n    model.add(Dropout(0.5))\r\n    model.add(Dense(4096, activation='relu'))\r\n    model.add(Dropout(0.5))\r\n    model.add(Dense(1000, activation='softmax'))\r\n\r\n    model.load_weights('../input/vgg16_weights.h5')\r\n\r\n    model.layers.pop()  # Get rid of the classification layer\r\n    model.outputs = [model.layers[-1].output]\r\n    model.layers[-1].outbound_nodes = []\r\n\r\n    model.add(Dense(10, activation='softmax'))\r\n```",
    "129229": "@tmain, thanks for that. I've now tried a range of learning rates from 1e-4 to 1e-7 with both ADAM and SGD as well as adding dropout after the Flatten layer and testing with batch normalisation. The best i've managed to get is a validation loss of 0.9 before the network overfits. I'm also using heavy image augmentation (scale, shift, shear, rotate) as suggested by some others.\r\n\r\nWere people really able to get losses of around 0.3 with what i've mentioned above or am I missing something? I'm just trying to get to the baseline of what others have reported quickly before experimenting further.\r\n\r\nThanks",
    "129185": "@tetmin\r\n\r\nUsually a batch size of 16 or 32. ",
    "129184": "@tmain, thanks for that. What batch size were you using with that learning rate out of interest?\r\n\r\nFor anyone trying to use this script with Tensorflow as the Keras backend, it can be sped up 2.5x by doing the following:\r\n- Change dimension ordering from th to tf\r\n- Transpose all the pre-trained weights of the convolutional layers so that they are in the usual Tensorflow order\r\n- Delete all of the ZeroPadding layers and add 'same' padding to the Conv2D layers",
    "129141": "@tetmin\r\n\r\nNormally I use Adam with standard parameters and learning rate of 1E-5.",
    "129135": "@tetmin\r\n\r\nI had the same problem.\r\n\r\nI added BN layer before the added convolution layer to the pre trained model. That helped the network to converge.",
    "129096": "What learning rate are people using to retrain this, with the default settings in tmain's github link this does not seem to converge, I get a huge loss of 14.70. A single epoch is also taking around 7000s on an AWS g2 instance, does that sound about right?\r\n\r\nThanks",
    "129037": "very nice work!\r\n\r\n[quote=tmain;126399]\r\n\r\nNice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras\r\n\r\n[/quote]\r\n",
    "128899": "**Heng CherKeng, tmain**, thanks. This makes sense.",
    "128855": "***tmain*** - thanks for the script. I'm trying it now. Your function requires \"label\" as well. So it actually create 10 images for one driver. How to achieve a heatmap for all labels at once? I mean got the pictures which looks like \"Heng CherKeng\"'s?",
    "128734": "Hello Heng,\r\n\r\nI have a question about how you choose which epoch to use for testing/submission. The attached figure is for learning with Googlenet. Around epoch 6 i get the best loss3 on validation and i tend to use model from that snapshot. Is that the right way to do it?\r\n\r\nThanks",
    "128699": "\"downscaling and upscaling the images and then cropping/padding\" is correct. But later i found that using ensemble of models of single original crops gives the best results. But it has to depend on how much perturbation of scale you used in training, it is trial and error kind of work.\r\n\r\n[quote=Rishab Gargeya;128680]\r\n\r\n@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say \"center crop, 0.8\" and \"center crop, 1.2\". Are you downscaling and upscaling the images and then cropping/padding?\r\n\r\n[/quote]\r\n",
    "128680": "@Heng Cherkeng Great work here! I'm a little confused, however, by what you mean when you say \"center crop, 0.8\" and \"center crop, 1.2\". Are you downscaling and upscaling the images and then cropping/padding?",
    "127915": "I am not sure what you mean by that. The procedure I followed was:\r\n\r\n- Use pre-trained VGG16 weights and train (cf. JiaoDong's  keras post). Save the weights of the newly trained VGGCAM model.\r\n- Re-use these trained weights to get the Activation Map ",
    "127644": "How did you combine those models? \r\n\r\n\r\n[quote=tmain;126399]\r\n\r\nNice find !\r\n\r\nFor those interested, I have written a quick keras implementation: [github link][1]\r\n\r\n\r\n  [1]: https://github.com/tdeboissiere/VGG16CAM-keras\r\n\r\n[/quote]\r\n",
    "126427": "@tmain nice work!",
    "126069": "It is always interesting to see results on out-of-sample data .... see attachment.\r\nIt test my system on some random images download from the internet. ",
    "126023": "Thanks for the idea :)\r\nSaw this and thought we can play with it too http://stackoverflow.com/questions/14063070/overlay-a-smaller-image-on-a-larger-image-python-opencv",
    "125953": "@shenzhenwei\r\nhere are the details.\r\n\r\n - croping and mixing patches  for training.\r\n This my idea. You can simply put random patches and get results by trial and error. The closest paper that apply this is \"Frankenstein: Learning Deep Face Representations using Small Data\"\r\n - generating deep feature results. Let the network be:\r\n\r\n\r\n conv ---> ave  ---> ip\r\n\r\nfor one sample, say the dims = C x H X W of the data are:  conv = 1024 x 14 x 14, ave = 1024 x1 x1, ip = 10 x 1 x 1. The weights of inner products are: weights = 10 x 1024. to create heat_map of the k-class (out of 10 classes), use heat_map_k = sum_c { conv_c * weight_k_c }. c runs from 1 to 1024. heat_map_k = 1x14x14. upscale heat_map_k to input image size to overlay results.\r\n\r\n",
    "125885": "This is cool.  Thank you for your sharing.",
    "125853": "upsampling the  14x14 output is only for visualisation (the heat map is 14x14).  The upsampling is simply any image resizing of the heat map to the input image size, so that we can overlay the heatmap onto the image.\r\n\r\nYou do not need to do upsampling if you only want to get the class score. see also the attachment for the network structure.\r\n\r\n\r\n\r\n[quote=DavidGbodiOdaibo;125852]\r\n\r\nis anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?\r\n\r\n[/quote]\r\n",
    "125852": "is anyone good with caffe/lasagne/keras? How would one implement this CAM network with  VGG in keras or lasagne. In other words how would one translate the caffe layers to layers in keras or lasagne.  The paper talks about upsampling the  14x14 output of the last conv layer, is that what the inner_product layer in the caffe model does?",
    "125842": "These CAM models are good. Single googlenet-CAM can give LB 0.38746 and single VGG-CAM give LB 0.27369. Please see attachment. I am not sure if the results comes from the CAM structure or my training method. Here are the training details:\r\n\r\n 1.  split of train and validation set. Divide train images into  groups of {driver,action}.\r\n       random select {driver,action} for training and validation.\r\n\r\n 2.  use 320x240 as image size into the CAM network\r\n\r\n 3. apply aggressive augmentation. I use scale change of +/- 0.2, displacement of random shift up to 0.2*MIN(width,height). Further, zero out a random block of 60x60 in each train image to simulate occlusion.\r\n\r\n ",
    "125698": "This is really cool, I would like to see more heat maps for the \"hair and makeup samples\",  looking at my predictions I noticed that my CNN focuses on the visor a lot and misclassifies many samples as hair and makeup whenever the visor is down.",
    "125697": "It looks *awesome*.",
    "175242": "",
    "125960": "",
    "125930": "",
    "125854": "@Heng CherKeng  thanks",
    "125739": "Interesting! Thanks for your sharing!"
  }
}