{
  "id": 70478,
  "title": "how not to overfit : attention is what you need ?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70478",
  "author_name": "hengck23",
  "post_date": "2018-11-04T11:15:35.825000",
  "votes": 65,
  "comment_count": 49,
  "views": 0,
  "content": "<p>At first look, this seems to be a multi-classification problem.</p>\n\n<p>But I think it is more suited as \"classification and weak localisation problem\".</p>\n\n<p>I note that if i visualize the network results (trained only with classification loss), most of the visualization show  rubbish results although F1 scores are high. So i think you need to give weak supervision for training</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png\" alt=\"enter image description here\"></p>",
  "messages": [
    {
      "id": 415080,
      "postDate": "2018-11-04T11:15:35.827Z",
      "content": "<p>At first look, this seems to be a multi-classification problem.</p>\n\n<p>But I think it is more suited as \"classification and weak localisation problem\".</p>\n\n<p>I note that if i visualize the network results (trained only with classification loss), most of the visualization show  rubbish results although F1 scores are high. So i think you need to give weak supervision for training</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "At first look, this seems to be a multi-classification problem.\n\nBut I think it is more suited as \"classification and weak localisation problem\".\n\nI note that if i visualize the network results (trained only with classification loss), most of the visualization show  rubbish results although F1 scores are high. So i think you need to give weak supervision for training\n\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png",
      "votes": 64
    },
    {
      "id": 416921,
      "postDate": "2018-11-07T12:48:17.343Z",
      "content": "<p>to put @Brian method into pictures:\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10629/Slide8.png\" alt=\"enter image description here\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10628/Slide9.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "to put @Brian method into pictures:\n   ![enter image description here][1]\n   ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10629/Slide8.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10628/Slide9.png",
      "votes": 10,
      "replies": [
        {
          "id": 416940,
          "postDate": "2018-11-07T13:25:59.417Z",
          "content": "<p>so using this method,  we are minimizing the combination of reconstruction loss and classification loss?</p>",
          "rawMarkdown": "so using this method,  we are minimizing the combination of reconstruction loss and classification loss?"
        },
        {
          "id": 417034,
          "postDate": "2018-11-07T16:29:48.210Z",
          "content": "<p>The top one is exactly how I am doing it. The encoder is complex, the decoder very simple. @Zhijian, yes, the model has two outputs each with a loss function. One outputs the reconstruction which I train with MSE loss, the other output the classifier with BCE+F1 loss.</p>\n\n<p>I am using Keras, here is how I set some of this up:</p>\n\n<p>Decoder:</p>\n\n<pre>def decoder_block(x,blocks=6,start_filters=512):\n    for i in range(blocks):\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = UpSampling2D()(x)\n    return x\n</pre>\n\n<p>Here is the last layer of the encoder:</p>\n\n<pre>    mpx = MaxPooling2D((2, 2), strides=(2, 2), name='b17_b6_o')(x)\n</pre>\n\n<p>Here is how the model gets built with 2 outputs:</p>\n\n<pre>    x = Dense(28, name='b17_d3')(x)\n    x = Activation('sigmoid', name='predictions')(x)\n\n    img_out = decoder_block(mpx,7)\n    img_out = Dense(3,activation='relu', name='img_out')(img_out)\n\n    # Create model\n    model = Model(img_input, [x,img_out], name='b21')\n</pre>\n\n<p>Here is how to fit a model with 2 outputs, 2 loss functions. The data generators instead of yielding X,y yield X, [y, X] which gives the target classes and original image to reconstruct:</p>\n\n<pre>loss_funcs = {\n    \"predictions\": brian_loss,\n    \"img_out\": \"mse\"\n}\nloss_weights = {\"predictions\": 1.0, \"img_out\": 10.0 } \n\nmetrics = { \"predictions\":[\n      weighted_binary_crossentropy,\n      \"acc\",\n      f1,\n      f2,\n      r_loss,\n      p_loss\n]}\n\n#train all\nfor i, layer in enumerate(model.layers):\n    model.layers[i].trainable = True\n\n# for i, layer in enumerate(model.layers):\n#     if orig_model.layers[i].name.startswith('b17'):\n#         model.layers[i].trainable = False\n\nmodel.compile(optimizer=opt,loss=loss_funcs, loss_weights=loss_weights, metrics=metrics)\nresults = model.fit_generator(train_generator,\n                              steps_per_epoch=train_df.shape[0]//BATCH_SIZE,\n                              validation_data = validation_generator,\n                              validation_steps = valid_df.shape[0]//BATCH_SIZE,\n                              epochs = 2000, \n                              callbacks=[ reduce_lr,checkpointer,tb],\n                              workers=24,\n                              max_queue_size=20)\n</pre>",
          "rawMarkdown": "The top one is exactly how I am doing it. The encoder is complex, the decoder very simple. @Zhijian, yes, the model has two outputs each with a loss function. One outputs the reconstruction which I train with MSE loss, the other output the classifier with BCE+F1 loss.\n\nI am using Keras, here is how I set some of this up:\n\nDecoder:\n<pre>def decoder_block(x,blocks=6,start_filters=512):\n    for i in range(blocks):\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = UpSampling2D()(x)\n    return x\n</pre>\n\nHere is the last layer of the encoder:\n<pre>    mpx = MaxPooling2D((2, 2), strides=(2, 2), name='b17_b6_o')(x)\n</pre>\n\nHere is how the model gets built with 2 outputs:\n<pre>    x = Dense(28, name='b17_d3')(x)\n    x = Activation('sigmoid', name='predictions')(x)\n\n    img_out = decoder_block(mpx,7)\n    img_out = Dense(3,activation='relu', name='img_out')(img_out)\n\n    # Create model\n    model = Model(img_input, [x,img_out], name='b21')\n</pre>\n\nHere is how to fit a model with 2 outputs, 2 loss functions. The data generators instead of yielding X,y yield X, [y, X] which gives the target classes and original image to reconstruct:\n<pre>loss_funcs = {\n    \"predictions\": brian_loss,\n    \"img_out\": \"mse\"\n}\nloss_weights = {\"predictions\": 1.0, \"img_out\": 10.0 } \n\nmetrics = { \"predictions\":[\n      weighted_binary_crossentropy,\n      \"acc\",\n      f1,\n      f2,\n      r_loss,\n      p_loss\n]}\n\n#train all\nfor i, layer in enumerate(model.layers):\n    model.layers[i].trainable = True\n    \n# for i, layer in enumerate(model.layers):\n#     if orig_model.layers[i].name.startswith('b17'):\n#         model.layers[i].trainable = False\n    \nmodel.compile(optimizer=opt,loss=loss_funcs, loss_weights=loss_weights, metrics=metrics)\nresults = model.fit_generator(train_generator,\n                              steps_per_epoch=train_df.shape[0]//BATCH_SIZE,\n                              validation_data = validation_generator,\n                              validation_steps = valid_df.shape[0]//BATCH_SIZE,\n                              epochs = 2000, \n                              callbacks=[ reduce_lr,checkpointer,tb],\n                              workers=24,\n                              max_queue_size=20)\n</pre>",
          "votes": 15
        },
        {
          "id": 417066,
          "postDate": "2018-11-07T17:26:18.413Z",
          "content": "<p>thanks for sharing the codes,\nnow I see the points.\nso we basically try to minimize the classification error, while keeping the relevant information as much as possible.</p>",
          "rawMarkdown": "thanks for sharing the codes,\nnow I see the points.\nso we basically try to minimize the classification error, while keeping the relevant information as much as possible."
        },
        {
          "id": 417516,
          "postDate": "2018-11-08T12:06:03.703Z",
          "content": "<p>How can you exploit the test dataset during training if the classification loss has no meaning? Do you train only the autoencoder without the supervised loss?</p>",
          "rawMarkdown": "How can you exploit the test dataset during training if the classification loss has no meaning? Do you train only the autoencoder without the supervised loss?"
        },
        {
          "id": 417694,
          "postDate": "2018-11-08T16:39:22.303Z",
          "content": "<p>First I train only test data to get the network heading in the right direction. After that I transfer the encoder and decoder weights to an autoencoder only network and train with everything. Then I go back to training and tuning with only the test set and the full network.</p>\n\n<p>The biggest issue I run into is GPU memory. I am currently trying to get a network that performs well with the higher resolution. I am trying to take the autoencoder output and feed it back into the classifcation block. Having convolutions go down then up and back down with the 1024x1024 images uses a huge amount of memory resulting in a batch size of 3 on my 4GB video card.</p>",
          "rawMarkdown": "First I train only test data to get the network heading in the right direction. After that I transfer the encoder and decoder weights to an autoencoder only network and train with everything. Then I go back to training and tuning with only the test set and the full network.\n\nThe biggest issue I run into is GPU memory. I am currently trying to get a network that performs well with the higher resolution. I am trying to take the autoencoder output and feed it back into the classifcation block. Having convolutions go down then up and back down with the 1024x1024 images uses a huge amount of memory resulting in a batch size of 3 on my 4GB video card.",
          "votes": 3
        },
        {
          "id": 417695,
          "postDate": "2018-11-08T16:44:37.750Z",
          "content": "<p>Do you mean \"train\" data?</p>",
          "rawMarkdown": "Do you mean \"train\" data?"
        },
        {
          "id": 417731,
          "postDate": "2018-11-08T17:49:10.573Z",
          "content": "<p>Be careful, BatchNorm is going to perfom poorly with 3 samples per batch.</p>",
          "rawMarkdown": "Be careful, BatchNorm is going to perfom poorly with 3 samples per batch."
        },
        {
          "id": 417772,
          "postDate": "2018-11-08T19:01:35.760Z",
          "content": "<p>@Daniel, Yes, was too early in the morning. I train with the training set first, then train autoencoder only with both training and test sets, and then train with the training set on the full network with scheduled learning rates.</p>\n\n<p>@Arm, Agree 100%. I've started running this larger network at 512x512 to see how it compares with my other 512 models. At 512x512 I can run batches of 24 on that video card and 47k images in my training set, it is taking 2300+ seconds per epoch. </p>\n\n<p>I run two experiments at once as I have 2 video cards. This one is running on the lower memory card, and I have a new version of my best model training on the faster card. That model is taking 2100 seconds per epoch.</p>\n\n<p>I'm hoping to qualify for the special prize at least, that video card sure looks nice. My faster model is scoring 5.18 on the public LB and cpu only prediction is down to ~400ms per image.</p>",
          "rawMarkdown": "@Daniel, Yes, was too early in the morning. I train with the training set first, then train autoencoder only with both training and test sets, and then train with the training set on the full network with scheduled learning rates.\n\n@Arm, Agree 100%. I've started running this larger network at 512x512 to see how it compares with my other 512 models. At 512x512 I can run batches of 24 on that video card and 47k images in my training set, it is taking 2300+ seconds per epoch. \n\nI run two experiments at once as I have 2 video cards. This one is running on the lower memory card, and I have a new version of my best model training on the faster card. That model is taking 2100 seconds per epoch.\n\nI'm hoping to qualify for the special prize at least, that video card sure looks nice. My faster model is scoring 5.18 on the public LB and cpu only prediction is down to ~400ms per image.",
          "votes": 1
        },
        {
          "id": 417903,
          "postDate": "2018-11-09T00:47:22.720Z",
          "content": "<p>What about your presence threshold? Is it far from 0.5?</p>\n\n<p>I'm having a hard time making the LB anywhere near my metric. There isn't a clear correlation.</p>",
          "rawMarkdown": "What about your presence threshold? Is it far from 0.5?\n\nI'm having a hard time making the LB anywhere near my metric. There isn't a clear correlation."
        },
        {
          "id": 417909,
          "postDate": "2018-11-09T01:02:38.733Z",
          "content": "<p>Different for each class. I run a search of thresholds and apply the lowest threshold with the best score. Locally I get much higher values than  public LB. Here are the last thresholds I used, I doubt it is too useful, it changes every training session: </p>\n\n<pre>[0.84716478 0.3        0.69826551 0.3        0.46891261 0.71547698\n 0.35203469 0.3        0.3        0.3        0.3        0.3\n 0.57338225 0.3        0.3        0.3        0.47331554 0.3\n 0.56697799 0.3        0.3        0.8903936  0.35883923 0.82715143\n 0.76230821 0.78992662 0.67384923 0.3       ]\n</pre>",
          "rawMarkdown": "Different for each class. I run a search of thresholds and apply the lowest threshold with the best score. Locally I get much higher values than  public LB. Here are the last thresholds I used, I doubt it is too useful, it changes every training session: \n<pre>[0.84716478 0.3        0.69826551 0.3        0.46891261 0.71547698\n 0.35203469 0.3        0.3        0.3        0.3        0.3\n 0.57338225 0.3        0.3        0.3        0.47331554 0.3\n 0.56697799 0.3        0.3        0.8903936  0.35883923 0.82715143\n 0.76230821 0.78992662 0.67384923 0.3       ]\n</pre>",
          "votes": 2
        },
        {
          "id": 418008,
          "postDate": "2018-11-09T05:22:40.807Z",
          "content": "<p>One thing I thought to mention, during training I am using a static 0.5 threshold. Before submitting I predict on the entire training set and use those predictions to come up with the best thresholds for the test set. Getting the thresholds right gave me a 0.02 gain in LB score.</p>",
          "rawMarkdown": "One thing I thought to mention, during training I am using a static 0.5 threshold. Before submitting I predict on the entire training set and use those predictions to come up with the best thresholds for the test set. Getting the thresholds right gave me a 0.02 gain in LB score.",
          "votes": 2
        },
        {
          "id": 420040,
          "postDate": "2018-11-13T01:01:22.650Z",
          "content": "<p>@Brian Thanks for sharing your work! I've learned a lot from this! I have one question though - you mention you train the Autoencoder using both the train and test sets? I was wondering if we're allowed to do that, or I've misinterpreted what was mentioned in this thread about using test features here -<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68665#404486\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68665#404486</a>. </p>\n\n<p>It would be great if we could use test features as well!</p>",
          "rawMarkdown": "@Brian Thanks for sharing your work! I've learned a lot from this! I have one question though - you mention you train the Autoencoder using both the train and test sets? I was wondering if we're allowed to do that, or I've misinterpreted what was mentioned in this thread about using test features here -https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68665#404486. \n\nIt would be great if we could use test features as well!"
        },
        {
          "id": 420128,
          "postDate": "2018-11-13T05:23:57.933Z",
          "content": "<p>It does look like that the first answer to the question there prohibits this. I don't see why it would be different than any other external data.</p>\n\n<p>In the end, I've abandoned this approach mainly due to the performance issues. It didn't seem to gain too much over methods like dropout, normalization, and weight regularization.</p>",
          "rawMarkdown": "It does look like that the first answer to the question there prohibits this. I don't see why it would be different than any other external data.\n\nIn the end, I've abandoned this approach mainly due to the performance issues. It didn't seem to gain too much over methods like dropout, normalization, and weight regularization.",
          "votes": 1
        },
        {
          "id": 421754,
          "postDate": "2018-11-15T11:21:14.317Z",
          "content": "<p>@Heng CherKeng One question regarding the first model on your image. Why does test set go into decoder (the upper branch of the model)? Shouldn't it go just to classifier instead? Brian says below, he removes decoder for predictions.</p>",
          "rawMarkdown": "@Heng CherKeng One question regarding the first model on your image. Why does test set go into decoder (the upper branch of the model)? Shouldn't it go just to classifier instead? Brian says below, he removes decoder for predictions."
        },
        {
          "id": 422314,
          "postDate": "2018-11-16T03:56:44.117Z",
          "content": "<p>During training the test data can be used with the autoencoder only portion and all of the data at the start, then classification added later. I thought the idea of an autoencoder was neat so I wanted to see what I could do with it.</p>",
          "rawMarkdown": "During training the test data can be used with the autoencoder only portion and all of the data at the start, then classification added later. I thought the idea of an autoencoder was neat so I wanted to see what I could do with it.",
          "votes": 3
        },
        {
          "id": 422561,
          "postDate": "2018-11-16T12:03:22.050Z",
          "content": "<p>People have won previous competitions using autoencoders like this (considering the test data as well).   </p>\n\n<p>I'm not sure about the rules and their wording, but why not accept it?</p>",
          "rawMarkdown": "People have won previous competitions using autoencoders like this (considering the test data as well).   \n\nI'm not sure about the rules and their wording, but why not accept it?"
        }
      ]
    },
    {
      "id": 416765,
      "postDate": "2018-11-07T08:50:40.210Z",
      "content": "<p>The way I have an attention mechanism generates images like attached. I posted these in the other visualization thread too. </p>\n\n<p>The network starts with the an encoder. The encoder then connects to a decoder and is trained as an autoencoder.  Additionally the classification head is connected to the same point the decoder is, trained with BCE+F1. After training the weights are saved and the decoder is removed from the network. </p>\n\n<p><img src=\"http://brians.network/images/c4.png\" alt=\"\">\n<img src=\"http://brians.network/images/c1.png\" alt=\"\">\n<img src=\"http://brians.network/images/c2.png\" alt=\"\">\n<img src=\"http://brians.network/images/c3.png\" alt=\"\"></p>",
      "rawMarkdown": "The way I have an attention mechanism generates images like attached. I posted these in the other visualization thread too. \n\nThe network starts with the an encoder. The encoder then connects to a decoder and is trained as an autoencoder.  Additionally the classification head is connected to the same point the decoder is, trained with BCE+F1. After training the weights are saved and the decoder is removed from the network. \n\n![](http://brians.network/images/c4.png)\n![](http://brians.network/images/c1.png)\n![](http://brians.network/images/c2.png)\n![](http://brians.network/images/c3.png)",
      "votes": 10,
      "replies": [
        {
          "id": 416779,
          "postDate": "2018-11-07T09:16:53.447Z",
          "content": "<p>Cool :)</p>\n\n<p>Which loss are you using for the autoencoder?</p>",
          "rawMarkdown": "Cool :)\n\nWhich loss are you using for the autoencoder?"
        },
        {
          "id": 416782,
          "postDate": "2018-11-07T09:19:15.923Z",
          "content": "<p>Relu activation with mean squared error loss.</p>",
          "rawMarkdown": "Relu activation with mean squared error loss."
        },
        {
          "id": 416914,
          "postDate": "2018-11-07T12:33:46.890Z",
          "content": "<p>So, instead of predicting on the original image, you are actualy predicting the classes from a compact representation of the image. Also, for the autoencoder part, are you training w.r.t. to the original 4-channel input, or a modified version that gives more supervision to the classifier (like 28*s*s green mask)?</p>",
          "rawMarkdown": "So, instead of predicting on the original image, you are actualy predicting the classes from a compact representation of the image. Also, for the autoencoder part, are you training w.r.t. to the original 4-channel input, or a modified version that gives more supervision to the classifier (like 28*s*s green mask)?"
        },
        {
          "id": 417060,
          "postDate": "2018-11-07T17:09:04.947Z",
          "content": "<p>Currently the autoencoder is training to the same as the input. I'm using 1024x1024x3 with RGB filters, discarding yellow. Using a modified version sounds like a good idea, the decoder can be changed to output one channel and then use only the green for training.</p>\n\n<p>My goal when I put this together was to use the decoding as a way to force the encoder to learn a 'better' representation by including enough information to reconstruct the image. It does help but also increases training time. </p>\n\n<p>Another thought I had was to take the image output and input it back into another encoder that the classifier then can use. If I did that though, I wouldn't be able to reduce the model after training.</p>",
          "rawMarkdown": "Currently the autoencoder is training to the same as the input. I'm using 1024x1024x3 with RGB filters, discarding yellow. Using a modified version sounds like a good idea, the decoder can be changed to output one channel and then use only the green for training.\n\nMy goal when I put this together was to use the decoding as a way to force the encoder to learn a 'better' representation by including enough information to reconstruct the image. It does help but also increases training time. \n\nAnother thought I had was to take the image output and input it back into another encoder that the classifier then can use. If I did that though, I wouldn't be able to reduce the model after training.",
          "votes": 2
        },
        {
          "id": 417087,
          "postDate": "2018-11-07T18:02:26.673Z",
          "content": "<p>You are telling us too much! I really hope that you win. Thank you for all your great ideas.</p>",
          "rawMarkdown": "You are telling us too much! I really hope that you win. Thank you for all your great ideas."
        },
        {
          "id": 417702,
          "postDate": "2018-11-08T16:54:05.373Z",
          "content": "<p>I've got my fingers crossed. The contest is long from over. I think my current advantage is only that I started early. With image contests it reminds me more of an arms race for processing power. I only spend an hour or two a day on this at most. My work takes up most the time, and I usually have to wait for a few days to run enough epochs to see a result.</p>",
          "rawMarkdown": "I've got my fingers crossed. The contest is long from over. I think my current advantage is only that I started early. With image contests it reminds me more of an arms race for processing power. I only spend an hour or two a day on this at most. My work takes up most the time, and I usually have to wait for a few days to run enough epochs to see a result."
        },
        {
          "id": 417714,
          "postDate": "2018-11-08T17:14:13.593Z",
          "content": "<p>Update on the results.... Relu is not the way to go. After 20 epochs the image output layer died, outputting all zeros. Now running with linear output.</p>",
          "rawMarkdown": "Update on the results.... Relu is not the way to go. After 20 epochs the image output layer died, outputting all zeros. Now running with linear output."
        },
        {
          "id": 417726,
          "postDate": "2018-11-08T17:38:39.457Z",
          "content": "<p>Just a suggestion...  (given than with your current LB position I doubt you need any... ;) )</p>\n\n<p>Why don't you test your whole pipeline with smaller images and reserve the full 1024 later on? It is not enough to get meaninful data?</p>",
          "rawMarkdown": "Just a suggestion...  (given than with your current LB position I doubt you need any... ;) )\n\nWhy don't you test your whole pipeline with smaller images and reserve the full 1024 later on? It is not enough to get meaninful data?"
        },
        {
          "id": 418000,
          "postDate": "2018-11-09T05:08:43.140Z",
          "content": "<p>I'm at about that point now. My current LB score is from the 512x512 R G and B images. I tried scaling the model and weights to 1024 and didn't have much success on that front. Right now I have this complicated model running at 512x512 and my newest model design running at 1024x1024.  My fast model overfits quite a bit but trains quickly. The model that includes the autoencoder is not showing signs of overfitting, but training is much slower.  <strong>Edit: It's only much slower when I forget to set the learning rate back up</strong></p>\n\n<p>Both of these have been running most of the day and were started at the same time. The model processing 1024x1024 images is making good progress, the changes I made seem to be working. It is still overfitting and the validation loss is slow to improve. Batch size 48:</p>\n\n<pre>1179/1179 [==============================] - 2109s 2s/step - loss: 2.5490 - weighted_binary_crossentropy: 0.4317 - acc: 0.6400 - f1: 0.5625 - f2: 0.7676 - r_loss: 0.1661 - p_loss: 0.3144 - val_loss: 4.0923 - val_weighted_binary_crossentropy: 0.7664 - val_acc: 0.5022 - val_f1: 0.3895 - val_f2: 0.6534 - val_r_loss: 0.2660 - val_p_loss: 0.4270\n</pre>\n\n<p>This model includes the autoencoder, running at 512x512. No signs of overfitting yet, the validation loss is better than the training loss still. Even though it's been running the same time it is not even close to even an attempt at the public LB.  Every epoch the best validation is surpassed, and there is not a large difference between train and validation. Batch size 24 and less training images:</p>\n\n<pre>1995/1995 [==============================] - 2322s 1s/step - loss: 9.6247 - predictions_loss: 6.4718 - img_out_loss: 0.6306 - predictions_weighted_binary_crossentropy: 0.5607 - predictions_acc: 0.1541 - predictions_f1: 0.0989 - predictions_f2: 0.2290 - predictions_r_loss: 0.5546 - predictions_p_loss: 0.8593 - val_loss: 9.3711 - val_predictions_loss: 6.4072 - val_img_out_loss: 0.5928 - val_predictions_weighted_binary_crossentropy: 0.5488 - val_predictions_acc: 0.2066 - val_predictions_f1: 0.1102 - val_predictions_f2: 0.2764 - val_predictions_r_loss: 0.4176 - val_predictions_p_loss: 0.8676\n</pre>",
          "rawMarkdown": "I'm at about that point now. My current LB score is from the 512x512 R G and B images. I tried scaling the model and weights to 1024 and didn't have much success on that front. Right now I have this complicated model running at 512x512 and my newest model design running at 1024x1024.  My fast model overfits quite a bit but trains quickly. The model that includes the autoencoder is not showing signs of overfitting, but training is much slower.  **Edit: It's only much slower when I forget to set the learning rate back up**\n\nBoth of these have been running most of the day and were started at the same time. The model processing 1024x1024 images is making good progress, the changes I made seem to be working. It is still overfitting and the validation loss is slow to improve. Batch size 48:\n<pre>1179/1179 [==============================] - 2109s 2s/step - loss: 2.5490 - weighted_binary_crossentropy: 0.4317 - acc: 0.6400 - f1: 0.5625 - f2: 0.7676 - r_loss: 0.1661 - p_loss: 0.3144 - val_loss: 4.0923 - val_weighted_binary_crossentropy: 0.7664 - val_acc: 0.5022 - val_f1: 0.3895 - val_f2: 0.6534 - val_r_loss: 0.2660 - val_p_loss: 0.4270\n</pre>\n\nThis model includes the autoencoder, running at 512x512. No signs of overfitting yet, the validation loss is better than the training loss still. Even though it's been running the same time it is not even close to even an attempt at the public LB.  Every epoch the best validation is surpassed, and there is not a large difference between train and validation. Batch size 24 and less training images:\n\n<pre>1995/1995 [==============================] - 2322s 1s/step - loss: 9.6247 - predictions_loss: 6.4718 - img_out_loss: 0.6306 - predictions_weighted_binary_crossentropy: 0.5607 - predictions_acc: 0.1541 - predictions_f1: 0.0989 - predictions_f2: 0.2290 - predictions_r_loss: 0.5546 - predictions_p_loss: 0.8593 - val_loss: 9.3711 - val_predictions_loss: 6.4072 - val_img_out_loss: 0.5928 - val_predictions_weighted_binary_crossentropy: 0.5488 - val_predictions_acc: 0.2066 - val_predictions_f1: 0.1102 - val_predictions_f2: 0.2764 - val_predictions_r_loss: 0.4176 - val_predictions_p_loss: 0.8676\n</pre>"
        },
        {
          "id": 419488,
          "postDate": "2018-11-12T03:52:59.303Z",
          "content": "<p>I've dropped this model for now, it ended up overfitting in the end anyways. The new model from this weekend is looking good so far using dropout and batch normalization.</p>",
          "rawMarkdown": "I've dropped this model for now, it ended up overfitting in the end anyways. The new model from this weekend is looking good so far using dropout and batch normalization.",
          "votes": 3
        }
      ]
    },
    {
      "id": 415944,
      "postDate": "2018-11-06T00:09:33.170Z",
      "content": "<p>code for CAM (class activation map) visualisation. This what the trained CNN thinks \"the target class would look like\".</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10615/2dfacb0e-bbad-11e8-b2ba-ac1f6b6435d0_07.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10617/5b67cf00-bbb3-11e8-b2ba-ac1f6b6435d0_02.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10618/5c08e04a-bbac-11e8-b2ba-ac1f6b6435d0_11.png\" alt=\"enter image description here\">\nsee also: \n<a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994</a></p>",
      "rawMarkdown": "code for CAM (class activation map) visualisation. This what the trained CNN thinks \"the target class would look like\".\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n  ![enter image description here][3]\nsee also: \nhttp://cnnlocalization.csail.mit.edu/\n\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10615/2dfacb0e-bbad-11e8-b2ba-ac1f6b6435d0_07.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10617/5b67cf00-bbb3-11e8-b2ba-ac1f6b6435d0_02.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10618/5c08e04a-bbac-11e8-b2ba-ac1f6b6435d0_11.png",
      "votes": 9
    },
    {
      "id": 415082,
      "postDate": "2018-11-04T11:24:08.657Z",
      "content": "<p>you can google for more papers. </p>\n\n<p>keywords: image classification, weak localisation, attention.</p>\n\n<p>But do note that our case is slightly different because we do have the ground truth attention mask (the green channel, which is the antibody marker). </p>\n\n<p>maybe we do not need to predict the attention mask, we may be able to use the green as input attention directly.</p>\n\n<hr>\n\n<p>reference:</p>\n\n<p>Residual Attention Network for Image Classification - Fei Wang, cvpr 2017</p>\n\n<p><a href=\"https://arxiv.org/abs/1704.06904\">https://arxiv.org/abs/1704.06904</a></p>\n\n<p><a href=\"https://www.youtube.com/watch?v=Deq1BGTHIPA\">https://www.youtube.com/watch?v=Deq1BGTHIPA</a></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821003-bcfb4216-76ee-11e7-947b-95feafcc34f2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821007-bf3fd8fc-76ee-11e7-916a-7ba989d94058.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821012-c46cbcbe-76ee-11e7-9a70-b07138c8c34b.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "you can google for more papers. \n\nkeywords: image classification, weak localisation, attention.\n\nBut do note that our case is slightly different because we do have the ground truth attention mask (the green channel, which is the antibody marker). \n\nmaybe we do not need to predict the attention mask, we may be able to use the green as input attention directly.\n\n----\nreference:\n\nResidual Attention Network for Image Classification - Fei Wang, cvpr 2017\n\nhttps://arxiv.org/abs/1704.06904\n\nhttps://www.youtube.com/watch?v=Deq1BGTHIPA\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n  \n  ![enter image description here][3]\n\n\n  [1]: https://user-images.githubusercontent.com/9496670/28821003-bcfb4216-76ee-11e7-947b-95feafcc34f2.png\n  [2]: https://user-images.githubusercontent.com/9496670/28821007-bf3fd8fc-76ee-11e7-916a-7ba989d94058.png\n  [3]: https://user-images.githubusercontent.com/9496670/28821012-c46cbcbe-76ee-11e7-9a70-b07138c8c34b.png",
      "votes": 9,
      "replies": [
        {
          "id": 416919,
          "postDate": "2018-11-07T12:39:28.063Z",
          "content": "<p>Thanks for sharing :)</p>",
          "rawMarkdown": "Thanks for sharing :)"
        }
      ]
    },
    {
      "id": 416787,
      "postDate": "2018-11-07T09:26:30.417Z",
      "content": "<p>Tell Me Where to Look: Guided Attention Inference Network-cvpr 1028</p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf</a></p>\n\n<p>imagine that you can label just a few pixel region of the target class, and the learning system will automatically expand it to  all pixel region of the same class.</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/15357902/37857732-a34e056c-2f38-11e8-982d-59c4299de981.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Tell Me Where to Look: Guided Attention Inference Network-cvpr 1028\n\nhttp://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf\n\nimagine that you can label just a few pixel region of the target class, and the learning system will automatically expand it to  all pixel region of the same class.\n\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://user-images.githubusercontent.com/15357902/37857732-a34e056c-2f38-11e8-982d-59c4299de981.png",
      "votes": 7,
      "replies": [
        {
          "id": 416846,
          "postDate": "2018-11-07T11:14:54.437Z",
          "content": "<p>This is awesome.</p>\n\n<p>I get wondered with how things can be so \"how couldn't I think about this\"?</p>",
          "rawMarkdown": "This is awesome.\n\nI get wondered with how things can be so \"how couldn't I think about this\"?",
          "votes": 1
        }
      ]
    },
    {
      "id": 415739,
      "postDate": "2018-11-05T15:26:47.893Z",
      "content": "<p>some of my localisation results:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415739/10613/localisation_results.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "some of my localisation results:\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415739/10613/localisation_results.png",
      "votes": 6
    },
    {
      "id": 415253,
      "postDate": "2018-11-04T18:55:20.270Z",
      "content": "<p>Also according to some threads, there are label noises in both training and test sets. A way to combat this might be mixup:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a></p>",
      "rawMarkdown": "Also according to some threads, there are label noises in both training and test sets. A way to combat this might be mixup:\n\nhttps://arxiv.org/pdf/1710.09412.pdf",
      "votes": 5,
      "replies": [
        {
          "id": 415801,
          "postDate": "2018-11-05T18:03:09.143Z",
          "content": "<p>Is this what I think it is?</p>\n\n<p>We sum two images together, in different proportions, then we sum their labels in the same proportions and train the sum as one sample?</p>",
          "rawMarkdown": "Is this what I think it is?\n\nWe sum two images together, in different proportions, then we sum their labels in the same proportions and train the sum as one sample?",
          "votes": 2
        },
        {
          "id": 421977,
          "postDate": "2018-11-15T16:18:42.197Z",
          "content": "<p>Yup, exactly this.\nit helped me to beat overfitting a llittle in the Ship Detection Challenge</p>",
          "rawMarkdown": "Yup, exactly this.\nit helped me to beat overfitting a llittle in the Ship Detection Challenge"
        }
      ]
    },
    {
      "id": 415385,
      "postDate": "2018-11-05T02:49:40.797Z",
      "content": "<p>related paper:\n<a href=\"https://arxiv.org/pdf/1811.00871.pdf\">https://arxiv.org/pdf/1811.00871.pdf</a></p>\n\n<p>Classification of Findings with Localized Lesions in Fundoscopic Images using a Regionally Guided CNN\n- Jaemin Son</p>\n\n<p>a perfect example of similar task: attention for localisation+classification:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10609/network_localisation.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10610/results_localisation.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "related paper:\nhttps://arxiv.org/pdf/1811.00871.pdf\n\nClassification of Findings with Localized Lesions in Fundoscopic Images using a Regionally Guided CNN\n- Jaemin Son\n \na perfect example of similar task: attention for localisation+classification:\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10609/network_localisation.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10610/results_localisation.png",
      "votes": 4,
      "replies": [
        {
          "id": 415390,
          "postDate": "2018-11-05T03:15:44.727Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 416176,
      "postDate": "2018-11-06T10:15:29.287Z",
      "content": "<p>related topics: multiple instance learning:\n<a href=\"https://academic.oup.com/bioinformatics/article/32/12/i52/2288769\">https://academic.oup.com/bioinformatics/article/32/12/i52/2288769</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10622/mil-2.png\" alt=\"enter image description here\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10621/mil-1.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "related topics: multiple instance learning:\nhttps://academic.oup.com/bioinformatics/article/32/12/i52/2288769\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10622/mil-2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10621/mil-1.png",
      "votes": 5
    },
    {
      "id": 415613,
      "postDate": "2018-11-05T11:45:46.237Z",
      "content": "<p>this is how i understand the problem setup for this challenge. please let me know if i am wrong. thanks!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415613/10608/setup.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "this is how i understand the problem setup for this challenge. please let me know if i am wrong. thanks!\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415613/10608/setup.png",
      "votes": 3,
      "replies": [
        {
          "id": 415938,
          "postDate": "2018-11-05T23:45:16.430Z",
          "content": "<p>This how I see it also(the picture of the interpretation of the data with stains, etc)</p>",
          "rawMarkdown": "This how I see it also(the picture of the interpretation of the data with stains, etc)"
        }
      ]
    },
    {
      "id": 415977,
      "postDate": "2018-11-06T01:55:08.527Z",
      "content": "<p><a href=\"https://arxiv.org/pdf/1809.08264.pdf\">https://arxiv.org/pdf/1809.08264.pdf</a></p>\n\n<p>Global Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification</p>\n\n<p>In this work, we first tackle the problem of simultaneous pixel-level localization and image-level classification\nwith only image-level labels for fully convolutional network training. We investigate the global pooling method\nwhich plays a vital role in this task. Classical global max pooling and average pooling methods are hard to indicate\nthe precise regions of objects. Therefore, we revisit the global weighted average pooling (GWAP) method for this\ntask and propose the class-agnostic GWAP module and the class-specific GWAP module in this paper. We evaluate\nthe classification and pixel-level localization ability on the ILSVRC benchmark dataset</p>",
      "rawMarkdown": "https://arxiv.org/pdf/1809.08264.pdf\n\nGlobal Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification\n\n\nIn this work, we first tackle the problem of simultaneous pixel-level localization and image-level classification\nwith only image-level labels for fully convolutional network training. We investigate the global pooling method\nwhich plays a vital role in this task. Classical global max pooling and average pooling methods are hard to indicate\nthe precise regions of objects. Therefore, we revisit the global weighted average pooling (GWAP) method for this\ntask and propose the class-agnostic GWAP module and the class-specific GWAP module in this paper. We evaluate\nthe classification and pixel-level localization ability on the ILSVRC benchmark dataset\n"
    },
    {
      "id": 416140,
      "postDate": "2018-11-06T08:54:09.887Z",
      "content": "<p>Note that:</p>\n\n<ol>\n<li><p>if you create new \"train image = red+blue+yellow+empty_green\", this can be negative class (all classes are absent)</p></li>\n<li><p>train with both  new \"train image = red+blue+yellow+green\" and original  \"train image = red+blue+yellow+empty_green\" can help the network to focus on the correct region.</p></li>\n</ol>",
      "rawMarkdown": "Note that:\n\n1.  if you create new \"train image = red+blue+yellow+empty_green\", this can be negative class (all classes are absent)\n\n2. train with both  new \"train image = red+blue+yellow+green\" and original  \"train image = red+blue+yellow+empty_green\" can help the network to focus on the correct region.\n",
      "votes": 1
    },
    {
      "id": 449499,
      "postDate": "2019-01-03T08:34:40.117Z",
      "content": "<p>@Heng CherKeng Hi, could you please share how did you threshold green channel for mask preparation. Thanks</p>",
      "rawMarkdown": "@Heng CherKeng Hi, could you please share how did you threshold green channel for mask preparation. Thanks"
    },
    {
      "id": 436412,
      "postDate": "2018-12-10T08:45:37.840Z",
      "content": "<p>Hey guys,\nWhat annotation tool are you using for labelling the images?</p>",
      "rawMarkdown": "Hey guys,\nWhat annotation tool are you using for labelling the images?"
    },
    {
      "id": 415467,
      "postDate": "2018-11-05T07:23:07.993Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "replies": [
        {
          "id": 423349,
          "postDate": "2018-11-18T01:22:04.113Z",
          "content": "<p>Wow, this is superb sharing, lots of things to digest...\nMany thanks.</p>",
          "rawMarkdown": "Wow, this is superb sharing, lots of things to digest...\nMany thanks."
        }
      ]
    },
    {
      "id": 449536,
      "postDate": "2019-01-03T10:11:03.807Z",
      "content": "<p>Thanks for your sharing. 谢谢</p>",
      "rawMarkdown": "Thanks for your sharing. 谢谢"
    }
  ],
  "comments": [
    {
      "id": 416921,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-07T12:48:17.343000",
      "content": "<p>to put @Brian method into pictures:\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10629/Slide8.png\" alt=\"enter image description here\">\n   <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10628/Slide9.png\" alt=\"enter image description here\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 416940,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2018-11-07T13:25:59.417000",
          "content": "<p>so using this method,  we are minimizing the combination of reconstruction loss and classification loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417034,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-07T16:29:48.210000",
          "content": "<p>The top one is exactly how I am doing it. The encoder is complex, the decoder very simple. @Zhijian, yes, the model has two outputs each with a loss function. One outputs the reconstruction which I train with MSE loss, the other output the classifier with BCE+F1 loss.</p>\n\n<p>I am using Keras, here is how I set some of this up:</p>\n\n<p>Decoder:</p>\n\n<pre>def decoder_block(x,blocks=6,start_filters=512):\n    for i in range(blocks):\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = Conv2D(start_filters//(2**i),(3,3), activation='relu', padding='same')(x)\n        x = UpSampling2D()(x)\n    return x\n</pre>\n\n<p>Here is the last layer of the encoder:</p>\n\n<pre>    mpx = MaxPooling2D((2, 2), strides=(2, 2), name='b17_b6_o')(x)\n</pre>\n\n<p>Here is how the model gets built with 2 outputs:</p>\n\n<pre>    x = Dense(28, name='b17_d3')(x)\n    x = Activation('sigmoid', name='predictions')(x)\n\n    img_out = decoder_block(mpx,7)\n    img_out = Dense(3,activation='relu', name='img_out')(img_out)\n\n    # Create model\n    model = Model(img_input, [x,img_out], name='b21')\n</pre>\n\n<p>Here is how to fit a model with 2 outputs, 2 loss functions. The data generators instead of yielding X,y yield X, [y, X] which gives the target classes and original image to reconstruct:</p>\n\n<pre>loss_funcs = {\n    \"predictions\": brian_loss,\n    \"img_out\": \"mse\"\n}\nloss_weights = {\"predictions\": 1.0, \"img_out\": 10.0 } \n\nmetrics = { \"predictions\":[\n      weighted_binary_crossentropy,\n      \"acc\",\n      f1,\n      f2,\n      r_loss,\n      p_loss\n]}\n\n#train all\nfor i, layer in enumerate(model.layers):\n    model.layers[i].trainable = True\n\n# for i, layer in enumerate(model.layers):\n#     if orig_model.layers[i].name.startswith('b17'):\n#         model.layers[i].trainable = False\n\nmodel.compile(optimizer=opt,loss=loss_funcs, loss_weights=loss_weights, metrics=metrics)\nresults = model.fit_generator(train_generator,\n                              steps_per_epoch=train_df.shape[0]//BATCH_SIZE,\n                              validation_data = validation_generator,\n                              validation_steps = valid_df.shape[0]//BATCH_SIZE,\n                              epochs = 2000, \n                              callbacks=[ reduce_lr,checkpointer,tb],\n                              workers=24,\n                              max_queue_size=20)\n</pre>",
          "votes": 15,
          "replies": []
        },
        {
          "id": 417066,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2018-11-07T17:26:18.413000",
          "content": "<p>thanks for sharing the codes,\nnow I see the points.\nso we basically try to minimize the classification error, while keeping the relevant information as much as possible.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417516,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-11-08T12:06:03.703000",
          "content": "<p>How can you exploit the test dataset during training if the classification loss has no meaning? Do you train only the autoencoder without the supervised loss?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417694,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-08T16:39:22.303000",
          "content": "<p>First I train only test data to get the network heading in the right direction. After that I transfer the encoder and decoder weights to an autoencoder only network and train with everything. Then I go back to training and tuning with only the test set and the full network.</p>\n\n<p>The biggest issue I run into is GPU memory. I am currently trying to get a network that performs well with the higher resolution. I am trying to take the autoencoder output and feed it back into the classifcation block. Having convolutions go down then up and back down with the 1024x1024 images uses a huge amount of memory resulting in a batch size of 3 on my 4GB video card.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 417695,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-08T16:44:37.750000",
          "content": "<p>Do you mean \"train\" data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417731,
          "author_name": "FlYM",
          "author_url": "",
          "post_date": "2018-11-08T17:49:10.573000",
          "content": "<p>Be careful, BatchNorm is going to perfom poorly with 3 samples per batch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417772,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-08T19:01:35.760000",
          "content": "<p>@Daniel, Yes, was too early in the morning. I train with the training set first, then train autoencoder only with both training and test sets, and then train with the training set on the full network with scheduled learning rates.</p>\n\n<p>@Arm, Agree 100%. I've started running this larger network at 512x512 to see how it compares with my other 512 models. At 512x512 I can run batches of 24 on that video card and 47k images in my training set, it is taking 2300+ seconds per epoch. </p>\n\n<p>I run two experiments at once as I have 2 video cards. This one is running on the lower memory card, and I have a new version of my best model training on the faster card. That model is taking 2100 seconds per epoch.</p>\n\n<p>I'm hoping to qualify for the special prize at least, that video card sure looks nice. My faster model is scoring 5.18 on the public LB and cpu only prediction is down to ~400ms per image.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 417903,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-09T00:47:22.720000",
          "content": "<p>What about your presence threshold? Is it far from 0.5?</p>\n\n<p>I'm having a hard time making the LB anywhere near my metric. There isn't a clear correlation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417909,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-09T01:02:38.733000",
          "content": "<p>Different for each class. I run a search of thresholds and apply the lowest threshold with the best score. Locally I get much higher values than  public LB. Here are the last thresholds I used, I doubt it is too useful, it changes every training session: </p>\n\n<pre>[0.84716478 0.3        0.69826551 0.3        0.46891261 0.71547698\n 0.35203469 0.3        0.3        0.3        0.3        0.3\n 0.57338225 0.3        0.3        0.3        0.47331554 0.3\n 0.56697799 0.3        0.3        0.8903936  0.35883923 0.82715143\n 0.76230821 0.78992662 0.67384923 0.3       ]\n</pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 418008,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-09T05:22:40.807000",
          "content": "<p>One thing I thought to mention, during training I am using a static 0.5 threshold. Before submitting I predict on the entire training set and use those predictions to come up with the best thresholds for the test set. Getting the thresholds right gave me a 0.02 gain in LB score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 420040,
          "author_name": "arao",
          "author_url": "",
          "post_date": "2018-11-13T01:01:22.650000",
          "content": "<p>@Brian Thanks for sharing your work! I've learned a lot from this! I have one question though - you mention you train the Autoencoder using both the train and test sets? I was wondering if we're allowed to do that, or I've misinterpreted what was mentioned in this thread about using test features here -<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68665#404486\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68665#404486</a>. </p>\n\n<p>It would be great if we could use test features as well!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 420128,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-13T05:23:57.933000",
          "content": "<p>It does look like that the first answer to the question there prohibits this. I don't see why it would be different than any other external data.</p>\n\n<p>In the end, I've abandoned this approach mainly due to the performance issues. It didn't seem to gain too much over methods like dropout, normalization, and weight regularization.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 421754,
          "author_name": "Michal Haltuf",
          "author_url": "",
          "post_date": "2018-11-15T11:21:14.317000",
          "content": "<p>@Heng CherKeng One question regarding the first model on your image. Why does test set go into decoder (the upper branch of the model)? Shouldn't it go just to classifier instead? Brian says below, he removes decoder for predictions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422314,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-16T03:56:44.117000",
          "content": "<p>During training the test data can be used with the autoencoder only portion and all of the data at the start, then classification added later. I thought the idea of an autoencoder was neat so I wanted to see what I could do with it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 422561,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-16T12:03:22.050000",
          "content": "<p>People have won previous competitions using autoencoders like this (considering the test data as well).   </p>\n\n<p>I'm not sure about the rules and their wording, but why not accept it?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416765,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-11-07T08:50:40.210000",
      "content": "<p>The way I have an attention mechanism generates images like attached. I posted these in the other visualization thread too. </p>\n\n<p>The network starts with the an encoder. The encoder then connects to a decoder and is trained as an autoencoder.  Additionally the classification head is connected to the same point the decoder is, trained with BCE+F1. After training the weights are saved and the decoder is removed from the network. </p>\n\n<p><img src=\"http://brians.network/images/c4.png\" alt=\"\">\n<img src=\"http://brians.network/images/c1.png\" alt=\"\">\n<img src=\"http://brians.network/images/c2.png\" alt=\"\">\n<img src=\"http://brians.network/images/c3.png\" alt=\"\"></p>",
      "votes": 10,
      "replies": [
        {
          "id": 416779,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-07T09:16:53.447000",
          "content": "<p>Cool :)</p>\n\n<p>Which loss are you using for the autoencoder?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416782,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-07T09:19:15.923000",
          "content": "<p>Relu activation with mean squared error loss.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 416914,
          "author_name": "Peiyuan Liao",
          "author_url": "",
          "post_date": "2018-11-07T12:33:46.890000",
          "content": "<p>So, instead of predicting on the original image, you are actualy predicting the classes from a compact representation of the image. Also, for the autoencoder part, are you training w.r.t. to the original 4-channel input, or a modified version that gives more supervision to the classifier (like 28*s*s green mask)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417060,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-07T17:09:04.947000",
          "content": "<p>Currently the autoencoder is training to the same as the input. I'm using 1024x1024x3 with RGB filters, discarding yellow. Using a modified version sounds like a good idea, the decoder can be changed to output one channel and then use only the green for training.</p>\n\n<p>My goal when I put this together was to use the decoding as a way to force the encoder to learn a 'better' representation by including enough information to reconstruct the image. It does help but also increases training time. </p>\n\n<p>Another thought I had was to take the image output and input it back into another encoder that the classifier then can use. If I did that though, I wouldn't be able to reduce the model after training.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 417087,
          "author_name": "FlYM",
          "author_url": "",
          "post_date": "2018-11-07T18:02:26.673000",
          "content": "<p>You are telling us too much! I really hope that you win. Thank you for all your great ideas.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417702,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-08T16:54:05.373000",
          "content": "<p>I've got my fingers crossed. The contest is long from over. I think my current advantage is only that I started early. With image contests it reminds me more of an arms race for processing power. I only spend an hour or two a day on this at most. My work takes up most the time, and I usually have to wait for a few days to run enough epochs to see a result.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417714,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-08T17:14:13.593000",
          "content": "<p>Update on the results.... Relu is not the way to go. After 20 epochs the image output layer died, outputting all zeros. Now running with linear output.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 417726,
          "author_name": "Thundo",
          "author_url": "",
          "post_date": "2018-11-08T17:38:39.457000",
          "content": "<p>Just a suggestion...  (given than with your current LB position I doubt you need any... ;) )</p>\n\n<p>Why don't you test your whole pipeline with smaller images and reserve the full 1024 later on? It is not enough to get meaninful data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 418000,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-09T05:08:43.140000",
          "content": "<p>I'm at about that point now. My current LB score is from the 512x512 R G and B images. I tried scaling the model and weights to 1024 and didn't have much success on that front. Right now I have this complicated model running at 512x512 and my newest model design running at 1024x1024.  My fast model overfits quite a bit but trains quickly. The model that includes the autoencoder is not showing signs of overfitting, but training is much slower.  <strong>Edit: It's only much slower when I forget to set the learning rate back up</strong></p>\n\n<p>Both of these have been running most of the day and were started at the same time. The model processing 1024x1024 images is making good progress, the changes I made seem to be working. It is still overfitting and the validation loss is slow to improve. Batch size 48:</p>\n\n<pre>1179/1179 [==============================] - 2109s 2s/step - loss: 2.5490 - weighted_binary_crossentropy: 0.4317 - acc: 0.6400 - f1: 0.5625 - f2: 0.7676 - r_loss: 0.1661 - p_loss: 0.3144 - val_loss: 4.0923 - val_weighted_binary_crossentropy: 0.7664 - val_acc: 0.5022 - val_f1: 0.3895 - val_f2: 0.6534 - val_r_loss: 0.2660 - val_p_loss: 0.4270\n</pre>\n\n<p>This model includes the autoencoder, running at 512x512. No signs of overfitting yet, the validation loss is better than the training loss still. Even though it's been running the same time it is not even close to even an attempt at the public LB.  Every epoch the best validation is surpassed, and there is not a large difference between train and validation. Batch size 24 and less training images:</p>\n\n<pre>1995/1995 [==============================] - 2322s 1s/step - loss: 9.6247 - predictions_loss: 6.4718 - img_out_loss: 0.6306 - predictions_weighted_binary_crossentropy: 0.5607 - predictions_acc: 0.1541 - predictions_f1: 0.0989 - predictions_f2: 0.2290 - predictions_r_loss: 0.5546 - predictions_p_loss: 0.8593 - val_loss: 9.3711 - val_predictions_loss: 6.4072 - val_img_out_loss: 0.5928 - val_predictions_weighted_binary_crossentropy: 0.5488 - val_predictions_acc: 0.2066 - val_predictions_f1: 0.1102 - val_predictions_f2: 0.2764 - val_predictions_r_loss: 0.4176 - val_predictions_p_loss: 0.8676\n</pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 419488,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-12T03:52:59.303000",
          "content": "<p>I've dropped this model for now, it ended up overfitting in the end anyways. The new model from this weekend is looking good so far using dropout and batch normalization.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 415944,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T00:09:33.170000",
      "content": "<p>code for CAM (class activation map) visualisation. This what the trained CNN thinks \"the target class would look like\".</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10615/2dfacb0e-bbad-11e8-b2ba-ac1f6b6435d0_07.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10617/5b67cf00-bbb3-11e8-b2ba-ac1f6b6435d0_02.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10618/5c08e04a-bbac-11e8-b2ba-ac1f6b6435d0_11.png\" alt=\"enter image description here\">\nsee also: \n<a href=\"http://cnnlocalization.csail.mit.edu/\">http://cnnlocalization.csail.mit.edu/</a></p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994</a></p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 415082,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-04T11:24:08.657000",
      "content": "<p>you can google for more papers. </p>\n\n<p>keywords: image classification, weak localisation, attention.</p>\n\n<p>But do note that our case is slightly different because we do have the ground truth attention mask (the green channel, which is the antibody marker). </p>\n\n<p>maybe we do not need to predict the attention mask, we may be able to use the green as input attention directly.</p>\n\n<hr>\n\n<p>reference:</p>\n\n<p>Residual Attention Network for Image Classification - Fei Wang, cvpr 2017</p>\n\n<p><a href=\"https://arxiv.org/abs/1704.06904\">https://arxiv.org/abs/1704.06904</a></p>\n\n<p><a href=\"https://www.youtube.com/watch?v=Deq1BGTHIPA\">https://www.youtube.com/watch?v=Deq1BGTHIPA</a></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821003-bcfb4216-76ee-11e7-947b-95feafcc34f2.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821007-bf3fd8fc-76ee-11e7-916a-7ba989d94058.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://user-images.githubusercontent.com/9496670/28821012-c46cbcbe-76ee-11e7-9a70-b07138c8c34b.png\" alt=\"enter image description here\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 416919,
          "author_name": "Kaushik Perika",
          "author_url": "",
          "post_date": "2018-11-07T12:39:28.063000",
          "content": "<p>Thanks for sharing :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 416787,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-07T09:26:30.417000",
      "content": "<p>Tell Me Where to Look: Guided Attention Inference Network-cvpr 1028</p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf\">http://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf</a></p>\n\n<p>imagine that you can label just a few pixel region of the target class, and the learning system will automatically expand it to  all pixel region of the same class.</p>\n\n<p><img src=\"https://user-images.githubusercontent.com/15357902/37857732-a34e056c-2f38-11e8-982d-59c4299de981.png\" alt=\"enter image description here\"></p>",
      "votes": 7,
      "replies": [
        {
          "id": 416846,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-07T11:14:54.437000",
          "content": "<p>This is awesome.</p>\n\n<p>I get wondered with how things can be so \"how couldn't I think about this\"?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 415739,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-05T15:26:47.893000",
      "content": "<p>some of my localisation results:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415739/10613/localisation_results.png\" alt=\"enter image description here\"></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 415253,
      "author_name": "Peiyuan Liao",
      "author_url": "",
      "post_date": "2018-11-04T18:55:20.270000",
      "content": "<p>Also according to some threads, there are label noises in both training and test sets. A way to combat this might be mixup:</p>\n\n<p><a href=\"https://arxiv.org/pdf/1710.09412.pdf\">https://arxiv.org/pdf/1710.09412.pdf</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 415801,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-05T18:03:09.143000",
          "content": "<p>Is this what I think it is?</p>\n\n<p>We sum two images together, in different proportions, then we sum their labels in the same proportions and train the sum as one sample?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 421977,
          "author_name": "Borys Tymchenko",
          "author_url": "",
          "post_date": "2018-11-15T16:18:42.197000",
          "content": "<p>Yup, exactly this.\nit helped me to beat overfitting a llittle in the Ship Detection Challenge</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415385,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-05T02:49:40.797000",
      "content": "<p>related paper:\n<a href=\"https://arxiv.org/pdf/1811.00871.pdf\">https://arxiv.org/pdf/1811.00871.pdf</a></p>\n\n<p>Classification of Findings with Localized Lesions in Fundoscopic Images using a Regionally Guided CNN\n- Jaemin Son</p>\n\n<p>a perfect example of similar task: attention for localisation+classification:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10609/network_localisation.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10610/results_localisation.png\" alt=\"enter image description here\"></p>",
      "votes": 4,
      "replies": [
        {
          "id": 415390,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-05T03:15:44.727000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 416176,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T10:15:29.287000",
      "content": "<p>related topics: multiple instance learning:\n<a href=\"https://academic.oup.com/bioinformatics/article/32/12/i52/2288769\">https://academic.oup.com/bioinformatics/article/32/12/i52/2288769</a></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10622/mil-2.png\" alt=\"enter image description here\">\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10621/mil-1.png\" alt=\"enter image description here\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 415613,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-05T11:45:46.237000",
      "content": "<p>this is how i understand the problem setup for this challenge. please let me know if i am wrong. thanks!</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415613/10608/setup.png\" alt=\"enter image description here\"></p>",
      "votes": 3,
      "replies": [
        {
          "id": 415938,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-11-05T23:45:16.430000",
          "content": "<p>This how I see it also(the picture of the interpretation of the data with stains, etc)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 415977,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T01:55:08.527000",
      "content": "<p><a href=\"https://arxiv.org/pdf/1809.08264.pdf\">https://arxiv.org/pdf/1809.08264.pdf</a></p>\n\n<p>Global Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification</p>\n\n<p>In this work, we first tackle the problem of simultaneous pixel-level localization and image-level classification\nwith only image-level labels for fully convolutional network training. We investigate the global pooling method\nwhich plays a vital role in this task. Classical global max pooling and average pooling methods are hard to indicate\nthe precise regions of objects. Therefore, we revisit the global weighted average pooling (GWAP) method for this\ntask and propose the class-agnostic GWAP module and the class-specific GWAP module in this paper. We evaluate\nthe classification and pixel-level localization ability on the ILSVRC benchmark dataset</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 416140,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-11-06T08:54:09.887000",
      "content": "<p>Note that:</p>\n\n<ol>\n<li><p>if you create new \"train image = red+blue+yellow+empty_green\", this can be negative class (all classes are absent)</p></li>\n<li><p>train with both  new \"train image = red+blue+yellow+green\" and original  \"train image = red+blue+yellow+empty_green\" can help the network to focus on the correct region.</p></li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 449499,
      "author_name": "Md Adilur Rahim",
      "author_url": "",
      "post_date": "2019-01-03T08:34:40.117000",
      "content": "<p>@Heng CherKeng Hi, could you please share how did you threshold green channel for mask preparation. Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 436412,
      "author_name": "It's_me :)",
      "author_url": "",
      "post_date": "2018-12-10T08:45:37.840000",
      "content": "<p>Hey guys,\nWhat annotation tool are you using for labelling the images?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 415467,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2018-11-05T07:23:07.993000",
      "content": "<p>Thank you for sharing</p>",
      "votes": 0,
      "replies": [
        {
          "id": 423349,
          "author_name": "Zungmann",
          "author_url": "",
          "post_date": "2018-11-18T01:22:04.113000",
          "content": "<p>Wow, this is superb sharing, lots of things to digest...\nMany thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 449536,
      "author_name": "BenChur",
      "author_url": "",
      "post_date": "2019-01-03T10:11:03.807000",
      "content": "<p>Thanks for your sharing. 谢谢</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "415080": "At first look, this seems to be a multi-classification problem.\n\nBut I think it is more suited as \"classification and weak localisation problem\".\n\nI note that if i visualize the network results (trained only with classification loss), most of the visualization show  rubbish results although F1 scores are high. So i think you need to give weak supervision for training\n\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png",
    "416921": "to put @Brian method into pictures:\n   ![enter image description here][1]\n   ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10629/Slide8.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/416921/10628/Slide9.png",
    "416765": "The way I have an attention mechanism generates images like attached. I posted these in the other visualization thread too. \n\nThe network starts with the an encoder. The encoder then connects to a decoder and is trained as an autoencoder.  Additionally the classification head is connected to the same point the decoder is, trained with BCE+F1. After training the weights are saved and the decoder is removed from the network. \n\n![](http://brians.network/images/c4.png)\n![](http://brians.network/images/c1.png)\n![](http://brians.network/images/c2.png)\n![](http://brians.network/images/c3.png)",
    "415944": "code for CAM (class activation map) visualisation. This what the trained CNN thinks \"the target class would look like\".\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n  ![enter image description here][3]\nsee also: \nhttp://cnnlocalization.csail.mit.edu/\n\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/discussion/21994\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10615/2dfacb0e-bbad-11e8-b2ba-ac1f6b6435d0_07.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10617/5b67cf00-bbb3-11e8-b2ba-ac1f6b6435d0_02.png\n  [3]: https://storage.googleapis.com/kaggle-forum-message-attachments/415944/10618/5c08e04a-bbac-11e8-b2ba-ac1f6b6435d0_11.png",
    "415082": "you can google for more papers. \n\nkeywords: image classification, weak localisation, attention.\n\nBut do note that our case is slightly different because we do have the ground truth attention mask (the green channel, which is the antibody marker). \n\nmaybe we do not need to predict the attention mask, we may be able to use the green as input attention directly.\n\n----\nreference:\n\nResidual Attention Network for Image Classification - Fei Wang, cvpr 2017\n\nhttps://arxiv.org/abs/1704.06904\n\nhttps://www.youtube.com/watch?v=Deq1BGTHIPA\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n  \n  ![enter image description here][3]\n\n\n  [1]: https://user-images.githubusercontent.com/9496670/28821003-bcfb4216-76ee-11e7-947b-95feafcc34f2.png\n  [2]: https://user-images.githubusercontent.com/9496670/28821007-bf3fd8fc-76ee-11e7-916a-7ba989d94058.png\n  [3]: https://user-images.githubusercontent.com/9496670/28821012-c46cbcbe-76ee-11e7-9a70-b07138c8c34b.png",
    "416787": "Tell Me Where to Look: Guided Attention Inference Network-cvpr 1028\n\nhttp://openaccess.thecvf.com/content_cvpr_2018/papers/Li_Tell_Me_Where_CVPR_2018_paper.pdf\n\nimagine that you can label just a few pixel region of the target class, and the learning system will automatically expand it to  all pixel region of the same class.\n\n\n\n  ![enter image description here][1]\n\n\n  [1]: https://user-images.githubusercontent.com/15357902/37857732-a34e056c-2f38-11e8-982d-59c4299de981.png",
    "415739": "some of my localisation results:\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415739/10613/localisation_results.png",
    "415253": "Also according to some threads, there are label noises in both training and test sets. A way to combat this might be mixup:\n\nhttps://arxiv.org/pdf/1710.09412.pdf",
    "415385": "related paper:\nhttps://arxiv.org/pdf/1811.00871.pdf\n\nClassification of Findings with Localized Lesions in Fundoscopic Images using a Regionally Guided CNN\n- Jaemin Son\n \na perfect example of similar task: attention for localisation+classification:\n\n  ![enter image description here][1]\n\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10609/network_localisation.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415385/10610/results_localisation.png",
    "416176": "related topics: multiple instance learning:\nhttps://academic.oup.com/bioinformatics/article/32/12/i52/2288769\n\n  ![enter image description here][1]\n  ![enter image description here][2]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10622/mil-2.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/416176/10621/mil-1.png",
    "415613": "this is how i understand the problem setup for this challenge. please let me know if i am wrong. thanks!\n\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/415613/10608/setup.png",
    "415977": "https://arxiv.org/pdf/1809.08264.pdf\n\nGlobal Weighted Average Pooling Bridges Pixel-level Localization and Image-level Classification\n\n\nIn this work, we first tackle the problem of simultaneous pixel-level localization and image-level classification\nwith only image-level labels for fully convolutional network training. We investigate the global pooling method\nwhich plays a vital role in this task. Classical global max pooling and average pooling methods are hard to indicate\nthe precise regions of objects. Therefore, we revisit the global weighted average pooling (GWAP) method for this\ntask and propose the class-agnostic GWAP module and the class-specific GWAP module in this paper. We evaluate\nthe classification and pixel-level localization ability on the ILSVRC benchmark dataset\n",
    "416140": "Note that:\n\n1.  if you create new \"train image = red+blue+yellow+empty_green\", this can be negative class (all classes are absent)\n\n2. train with both  new \"train image = red+blue+yellow+green\" and original  \"train image = red+blue+yellow+empty_green\" can help the network to focus on the correct region.\n",
    "449499": "@Heng CherKeng Hi, could you please share how did you threshold green channel for mask preparation. Thanks",
    "436412": "Hey guys,\nWhat annotation tool are you using for labelling the images?",
    "415467": "Thank you for sharing",
    "449536": "Thanks for your sharing. 谢谢"
  }
}