{
  "id": 77325,
  "title": "12th place solution",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/arm-12th-place-solution",
  "author_name": "",
  "post_date": "2019-01-16T15:21:32.643Z",
  "votes": 38,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hey kagglers, you know what they say: Be careful of overfitting, rare classes and bestfitting! Hah.</p>\n\n<p>Jokes apart, congratulations to all the winners, this was a really tough competition. Thanks to the Kaggle organization and Human Protein Atlas for providing such a great challenge and also thanks to @brian @heng and @tilii for all the useful posts, I have learnt a lot from you. </p>\n\n<p>I am going to briefly detail my solution:</p>\n\n<h2>Hardware</h2>\n\n<p>First I had a 980Ti GPU, 2600k CPU and 14GB RAM. Four weeks ago I bought a 1080Ti.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Removing duplicates was essential for improving your validation set and consequently for finding  your class thresholds. To do so I compared image channels using the imagehash library proposed by  @tilii</p>\n\n<h2>Base model</h2>\n\n<p>ResNet50 trained with float16 precision, RGB,  1024x1024, lots of data augmentation, \n  weight decay, sgd, gradient clipping, bce loss and multi-scale resolution images. All in PyTorch.\n  Best single model [0.628 public, 0.553 private]</p>\n\n<h2>Gamma correction</h2>\n\n<p>Due to the dark exposure of some images, half of my models were trained with an extra initial layer to correct the gamma. The goal was to find the best gamma per channel for all the images. I tried to make a layer that extracted a custom gamma per image taking into account the image as a context but without success. I tried things like  i) using statistical image values as features and ii) applying a small convolution to the image then fc; all of them very unstable. The approach that worked best was the simplest: learn independent gammas (y_1 , y_2 , y_3) per channel and only use one parameter initialized at 1 without an activation. This reduced my global val_loss and increased my internal F1_score of some classes.\n  <img src=\"https://i.postimg.cc/rmh5h58T/Captura-de-pantalla-2019-01-11-a-les-12-34-16.png\" alt=\"Gamma correction formula\">\nAmusingly all the learnt gammas were quite similar, in the range [0.6-0.65]. Meaning that the images were quite cleared. I guess dark images benefitted from this layers but others - specially noisy images - don't, that's why I trained also networks without it.</p>\n\n<h2>What should I trust?</h2>\n\n<p>Is my model overfitting? \nThis was the big question of all the competition. It was hard answering it for the rare classes but for the other ones I found a pattern that I believe was quite useful. When comparing models, most of the times the models that were overfitting had the largest F1-Score in a zone where Precision was extremely dominating vs Recall.  Example in the following first figure. My hypothesis is that such models are learning things that are only present in the training set (similar cell/protein types, microscope features, who knows) and then become very confident in some samples increasing the Precision. When this happened I lowered thresholds or I discarded the model and choose another one (second figure).\n<img src=\"https://i.postimg.cc/SR3cW84K/Captura-de-pantalla-2019-01-11-a-les-13-11-08.png\" alt=\"F1 curves\"></p>\n\n<h2>Model variety</h2>\n\n<p>I trained 11 models, 2 DenseNets121 and 9 ResNets50. They were fast to train and provided the best results in my val_set.  Each model had different settings such as:\nExcluding data leak, including it, change model seed, using yellow images , different class balance, full sized images, crops...</p>\n\n<h2>Rare classes</h2>\n\n<p>I did not know how to tackle them properly and I guess my fall has been mainly due to these classes (It's a pity that we cannot see F1-Score per class). I was about to try some Few Shot Learning approaches, but the image size was a limiting factor. Apart from this, I assume that the data leak eliminated most of the rare samples so detecting even just one of them was crucial (maybe too much). I hope I can learn from the top solutions about this.</p>\n\n<h2>Ensembling</h2>\n\n<p>I tried Logistic Regression, Averaging, Stacking and XGBoosting. In abundant classes it worked, but with others it overfitted and I thought it was not reliable. In the end I did the following:</p>\n\n<p><em>Expert models</em>\nI decided which model were the best for specific classes, for instance: Yellow image model for classes 6, 7;   cropped image model for large classes such as 0, 25, full for 16; etc. These decisions were made taking into account hpa web info data, F1-curves, validation data, and public leaderboard.</p>\n\n<p><em>One vote ensembling with top predictions</em>\nApart from expert models I increased the recall of my predictions by adding the top predictions from other models. Most of my models scored similar F1-scores per class, so it made sense. I increased the thresholds of non-expert models using this heuristic formula T_opt = T_opt + (1 - T_opt) *0.3 and included them as in one vote is enough.</p>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>Training a Multi Head Attention Module (Transformer paper) from the predictions of crops of an image in order to focus the attention in the spatial domain. </li>\n<li>Training GapNet style/Feature Pyramid with ResNet18 as backbone did not provide any benefits in my setup (I guess I did not play enough with it because @Dieter was able to make it work).</li>\n<li>Training with 2048x2048 images, I guess batch size was too small and the extra resolution was not that important for the majority of the samples.</li>\n</ul>",
  "messages": [
    {
      "id": "454372",
      "postDate": "01/11/2019 14:45:30",
      "content": "<p>Hey kagglers, you know what they say: Be careful of overfitting, rare classes and bestfitting! Hah.</p>\n\n<p>Jokes apart, congratulations to all the winners, this was a really tough competition. Thanks to the Kaggle organization and Human Protein Atlas for providing such a great challenge and also thanks to @brian @heng and @tilii for all the useful posts, I have learnt a lot from you. </p>\n\n<p>I am going to briefly detail my solution:</p>\n\n<h2>Hardware</h2>\n\n<p>First I had a 980Ti GPU, 2600k CPU and 14GB RAM. Four weeks ago I bought a 1080Ti.</p>\n\n<h2>Preprocessing</h2>\n\n<p>Removing duplicates was essential for improving your validation set and consequently for finding  your class thresholds. To do so I compared image channels using the imagehash library proposed by  @tilii</p>\n\n<h2>Base model</h2>\n\n<p>ResNet50 trained with float16 precision, RGB,  1024x1024, lots of data augmentation, \n  weight decay, sgd, gradient clipping, bce loss and multi-scale resolution images. All in PyTorch.\n  Best single model [0.628 public, 0.553 private]</p>\n\n<h2>Gamma correction</h2>\n\n<p>Due to the dark exposure of some images, half of my models were trained with an extra initial layer to correct the gamma. The goal was to find the best gamma per channel for all the images. I tried to make a layer that extracted a custom gamma per image taking into account the image as a context but without success. I tried things like  i) using statistical image values as features and ii) applying a small convolution to the image then fc; all of them very unstable. The approach that worked best was the simplest: learn independent gammas (y_1 , y_2 , y_3) per channel and only use one parameter initialized at 1 without an activation. This reduced my global val_loss and increased my internal F1_score of some classes.\n  <img src=\"https://i.postimg.cc/rmh5h58T/Captura-de-pantalla-2019-01-11-a-les-12-34-16.png\" alt=\"Gamma correction formula\">\nAmusingly all the learnt gammas were quite similar, in the range [0.6-0.65]. Meaning that the images were quite cleared. I guess dark images benefitted from this layers but others - specially noisy images - don't, that's why I trained also networks without it.</p>\n\n<h2>What should I trust?</h2>\n\n<p>Is my model overfitting? \nThis was the big question of all the competition. It was hard answering it for the rare classes but for the other ones I found a pattern that I believe was quite useful. When comparing models, most of the times the models that were overfitting had the largest F1-Score in a zone where Precision was extremely dominating vs Recall.  Example in the following first figure. My hypothesis is that such models are learning things that are only present in the training set (similar cell/protein types, microscope features, who knows) and then become very confident in some samples increasing the Precision. When this happened I lowered thresholds or I discarded the model and choose another one (second figure).\n<img src=\"https://i.postimg.cc/SR3cW84K/Captura-de-pantalla-2019-01-11-a-les-13-11-08.png\" alt=\"F1 curves\"></p>\n\n<h2>Model variety</h2>\n\n<p>I trained 11 models, 2 DenseNets121 and 9 ResNets50. They were fast to train and provided the best results in my val_set.  Each model had different settings such as:\nExcluding data leak, including it, change model seed, using yellow images , different class balance, full sized images, crops...</p>\n\n<h2>Rare classes</h2>\n\n<p>I did not know how to tackle them properly and I guess my fall has been mainly due to these classes (It's a pity that we cannot see F1-Score per class). I was about to try some Few Shot Learning approaches, but the image size was a limiting factor. Apart from this, I assume that the data leak eliminated most of the rare samples so detecting even just one of them was crucial (maybe too much). I hope I can learn from the top solutions about this.</p>\n\n<h2>Ensembling</h2>\n\n<p>I tried Logistic Regression, Averaging, Stacking and XGBoosting. In abundant classes it worked, but with others it overfitted and I thought it was not reliable. In the end I did the following:</p>\n\n<p><em>Expert models</em>\nI decided which model were the best for specific classes, for instance: Yellow image model for classes 6, 7;   cropped image model for large classes such as 0, 25, full for 16; etc. These decisions were made taking into account hpa web info data, F1-curves, validation data, and public leaderboard.</p>\n\n<p><em>One vote ensembling with top predictions</em>\nApart from expert models I increased the recall of my predictions by adding the top predictions from other models. Most of my models scored similar F1-scores per class, so it made sense. I increased the thresholds of non-expert models using this heuristic formula T_opt = T_opt + (1 - T_opt) *0.3 and included them as in one vote is enough.</p>\n\n<h2>Things that did not work</h2>\n\n<ul>\n<li>Training a Multi Head Attention Module (Transformer paper) from the predictions of crops of an image in order to focus the attention in the spatial domain. </li>\n<li>Training GapNet style/Feature Pyramid with ResNet18 as backbone did not provide any benefits in my setup (I guess I did not play enough with it because @Dieter was able to make it work).</li>\n<li>Training with 2048x2048 images, I guess batch size was too small and the extra resolution was not that important for the majority of the samples.</li>\n</ul>",
      "rawMarkdown": "Hey kagglers, you know what they say: Be careful of overfitting, rare classes and bestfitting! Hah.\n \nJokes apart, congratulations to all the winners, this was a really tough competition. Thanks to the Kaggle organization and Human Protein Atlas for providing such a great challenge and also thanks to @brian @heng and @tilii for all the useful posts, I have learnt a lot from you. \n\nI am going to briefly detail my solution:\n\n\n##Hardware##\n  First I had a 980Ti GPU, 2600k CPU and 14GB RAM. Four weeks ago I bought a 1080Ti.\n\n##Preprocessing##\n  Removing duplicates was essential for improving your validation set and consequently for finding  your class thresholds. To do so I compared image channels using the imagehash library proposed by  @tilii\n\n\n##Base model##\n  ResNet50 trained with float16 precision, RGB,  1024x1024, lots of data augmentation, \n  weight decay, sgd, gradient clipping, bce loss and multi-scale resolution images. All in PyTorch.\n  Best single model [0.628 public, 0.553 private]\n\n##Gamma correction##\n Due to the dark exposure of some images, half of my models were trained with an extra initial layer to correct the gamma. The goal was to find the best gamma per channel for all the images. I tried to make a layer that extracted a custom gamma per image taking into account the image as a context but without success. I tried things like  i) using statistical image values as features and ii) applying a small convolution to the image then fc; all of them very unstable. The approach that worked best was the simplest: learn independent gammas (y_1 , y_2 , y_3) per channel and only use one parameter initialized at 1 without an activation. This reduced my global val_loss and increased my internal F1_score of some classes.\n  ![Gamma correction formula][1]\nAmusingly all the learnt gammas were quite similar, in the range [0.6-0.65]. Meaning that the images were quite cleared. I guess dark images benefitted from this layers but others - specially noisy images - don't, that's why I trained also networks without it.\n\n\n\n##What should I trust?##\nIs my model overfitting? \nThis was the big question of all the competition. It was hard answering it for the rare classes but for the other ones I found a pattern that I believe was quite useful. When comparing models, most of the times the models that were overfitting had the largest F1-Score in a zone where Precision was extremely dominating vs Recall.  Example in the following first figure. My hypothesis is that such models are learning things that are only present in the training set (similar cell/protein types, microscope features, who knows) and then become very confident in some samples increasing the Precision. When this happened I lowered thresholds or I discarded the model and choose another one (second figure).\n![F1 curves][2]\n\n##Model variety##\nI trained 11 models, 2 DenseNets121 and 9 ResNets50. They were fast to train and provided the best results in my val_set.  Each model had different settings such as:\nExcluding data leak, including it, change model seed, using yellow images , different class balance, full sized images, crops...\n\n##Rare classes##\nI did not know how to tackle them properly and I guess my fall has been mainly due to these classes (It's a pity that we cannot see F1-Score per class). I was about to try some Few Shot Learning approaches, but the image size was a limiting factor. Apart from this, I assume that the data leak eliminated most of the rare samples so detecting even just one of them was crucial (maybe too much). I hope I can learn from the top solutions about this.\n\n##Ensembling##\nI tried Logistic Regression, Averaging, Stacking and XGBoosting. In abundant classes it worked, but with others it overfitted and I thought it was not reliable. In the end I did the following:\n\n*Expert models*\nI decided which model were the best for specific classes, for instance: Yellow image model for classes 6, 7;   cropped image model for large classes such as 0, 25, full for 16; etc. These decisions were made taking into account hpa web info data, F1-curves, validation data, and public leaderboard.\n\n*One vote ensembling with top predictions*\nApart from expert models I increased the recall of my predictions by adding the top predictions from other models. Most of my models scored similar F1-scores per class, so it made sense. I increased the thresholds of non-expert models using this heuristic formula T_opt = T_opt + (1 - T_opt) *0.3 and included them as in one vote is enough.\n\n\n##Things that did not work##\n- Training a Multi Head Attention Module (Transformer paper) from the predictions of crops of an image in order to focus the attention in the spatial domain. \n-  Training GapNet style/Feature Pyramid with ResNet18 as backbone did not provide any benefits in my setup (I guess I did not play enough with it because @Dieter was able to make it work).\n- Training with 2048x2048 images, I guess batch size was too small and the extra resolution was not that important for the majority of the samples.\n\n  [1]: https://i.postimg.cc/rmh5h58T/Captura-de-pantalla-2019-01-11-a-les-12-34-16.png\n  [2]: https://i.postimg.cc/SR3cW84K/Captura-de-pantalla-2019-01-11-a-les-13-11-08.png",
      "votes": null
    },
    {
      "id": "454380",
      "postDate": "01/11/2019 14:56:37",
      "content": "<blockquote>It's a pity that we cannot see F1-Score per class</blockquote>\n\n<p>You can probe it by submitting your .csv file with some class and without it. Difference between two scores will tell you about binary f1 score of each class.</p>",
      "rawMarkdown": "<blockquote>It's a pity that we cannot see F1-Score per class</blockquote>\n\nYou can probe it by submitting your .csv file with some class and without it. Difference between two scores will tell you about binary f1 score of each class.",
      "votes": null
    },
    {
      "id": "454384",
      "postDate": "01/11/2019 15:02:22",
      "content": "<p>Great idea. Even though what I would like is to compare my F1-Scores per class with the ones of top winners.  I guess when top winners upload their code I can generate their submission and extract theirs as well. Nice.</p>",
      "rawMarkdown": "Great idea. Even though what I would like is to compare my F1-Scores per class with the ones of top winners.  I guess when top winners upload their code I can generate their submission and extract theirs as well. Nice.",
      "votes": null
    },
    {
      "id": "454625",
      "postDate": "01/11/2019 22:33:16",
      "content": "<p>Thanks for sharing. We were quite curious about how you were able to achieve your LB score. Congratulations to your solo-gold. </p>",
      "rawMarkdown": "Thanks for sharing. We were quite curious about how you were able to achieve your LB score. Congratulations to your solo-gold.",
      "votes": null
    },
    {
      "id": "454705",
      "postDate": "01/12/2019 03:42:06",
      "content": "<p>This is Amazing! You able to chieve you score by a 1080Ti! I am quite curious about how long does it take to train a 1024*1024 model? \nI fail to train a 1024 * 1024 because I only have a 1080Ti. Before I read your solution, I guess it is due to too small batch size. You can train it better so I want to see if there is a problem with my optimization method or if I have not trained enough time.</p>",
      "rawMarkdown": "This is Amazing! You able to chieve you score by a 1080Ti! I am quite curious about how long does it take to train a 1024*1024 model? \nI fail to train a 1024 * 1024 because I only have a 1080Ti. Before I read your solution, I guess it is due to too small batch size. You can train it better so I want to see if there is a problem with my optimization method or if I have not trained enough time.",
      "votes": null
    },
    {
      "id": "454827",
      "postDate": "01/12/2019 10:10:10",
      "content": "<p>Two days. The key is using float16 in Pytorch, if you do that you can double the size of your batch.</p>",
      "rawMarkdown": "Two days. The key is using float16 in Pytorch, if you do that you can double the size of your batch.",
      "votes": null
    },
    {
      "id": "455067",
      "postDate": "01/12/2019 21:40:57",
      "content": "<p>I'm getting the impression you ran with comp data only, am I incorrect? Congrats on the score, textbook model diversity!</p>",
      "rawMarkdown": "I'm getting the impression you ran with comp data only, am I incorrect? Congrats on the score, textbook model diversity!",
      "votes": null
    },
    {
      "id": "455085",
      "postDate": "01/12/2019 23:46:22",
      "content": "<p>I had some models with only comp data because it was cleaner. Nevertheless, for rare classes I trained with HPA leaks.</p>",
      "rawMarkdown": "I had some models with only comp data because it was cleaner. Nevertheless, for rare classes I trained with HPA leaks.",
      "votes": null
    },
    {
      "id": "455123",
      "postDate": "01/13/2019 02:56:07",
      "content": "<p>Thanks！</p>",
      "rawMarkdown": "Thanks！",
      "votes": null
    },
    {
      "id": "455462",
      "postDate": "01/14/2019 00:46:17",
      "content": "<p>Great stuff, man and thanks for the float 16 tip - I cannot wait to go and test it myself! I have one question for you, how exactly did you learn the gammas (what kind of layer) and what do you mean by \"only use one parameter initialized at 1 without an activation\"?</p>",
      "rawMarkdown": "Great stuff, man and thanks for the float 16 tip - I cannot wait to go and test it myself! I have one question for you, how exactly did you learn the gammas (what kind of layer) and what do you mean by \"only use one parameter initialized at 1 without an activation\"?",
      "votes": null
    },
    {
      "id": "456231",
      "postDate": "01/15/2019 11:42:31",
      "content": "<p>I made a new layer in PyTorch. It's just a layer that raises each channel to a specific scalar value.</p>",
      "rawMarkdown": "I made a new layer in PyTorch. It's just a layer that raises each channel to a specific scalar value.",
      "votes": null
    },
    {
      "id": "456482",
      "postDate": "01/15/2019 22:55:58",
      "content": "<p>That's the obvious way to do it and I felt dump right after asking the question... ;-) Thanks for your answer though and I hope some of that pain has started to fade away! Personally, I am really impressed by your performance.</p>",
      "rawMarkdown": "That's the obvious way to do it and I felt dump right after asking the question... ;-) Thanks for your answer though and I hope some of that pain has started to fade away! Personally, I am really impressed by your performance.",
      "votes": null
    },
    {
      "id": "456578",
      "postDate": "01/16/2019 05:09:37",
      "content": "<p>3rd place  used gradient accumulation. gonna use that later to test how it works, after I finish some other tests</p>",
      "rawMarkdown": "3rd place  used gradient accumulation. gonna use that later to test how it works, after I finish some other tests",
      "votes": null
    },
    {
      "id": "456579",
      "postDate": "01/16/2019 05:15:40",
      "content": "<p>If there are no batch norms, it works out-of-box in PyTorch. But batch norms require custom handling.</p>",
      "rawMarkdown": "If there are no batch norms, it works out-of-box in PyTorch. But batch norms require custom handling.",
      "votes": null
    },
    {
      "id": "456690",
      "postDate": "01/16/2019 09:58:44",
      "content": "<p>Thanks for your sharing. It's a tough competition, I have just gotten 0.1+ score.</p>",
      "rawMarkdown": "Thanks for your sharing. It's a tough competition, I have just gotten 0.1+ score.",
      "votes": null
    },
    {
      "id": "458245",
      "postDate": "01/19/2019 07:03:47",
      "content": "<p>Thanks for sharing this. I was stuck on a loop as I didn't know what else I could have done to increase the accuracy. Thanks for the post mate!</p>",
      "rawMarkdown": "Thanks for sharing this. I was stuck on a loop as I didn't know what else I could have done to increase the accuracy. Thanks for the post mate!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454380,
      "author_name": "hokmund",
      "author_url": "",
      "post_date": "01/11/2019 14:56:37",
      "content": "<blockquote>It's a pity that we cannot see F1-Score per class</blockquote>\n\n<p>You can probe it by submitting your .csv file with some class and without it. Difference between two scores will tell you about binary f1 score of each class.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454384,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "01/11/2019 15:02:22",
          "content": "<p>Great idea. Even though what I would like is to compare my F1-Scores per class with the ones of top winners.  I guess when top winners upload their code I can generate their submission and extract theirs as well. Nice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 454625,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "01/11/2019 22:33:16",
      "content": "<p>Thanks for sharing. We were quite curious about how you were able to achieve your LB score. Congratulations to your solo-gold. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454705,
      "author_name": "lbyg1994",
      "author_url": "",
      "post_date": "01/12/2019 03:42:06",
      "content": "<p>This is Amazing! You able to chieve you score by a 1080Ti! I am quite curious about how long does it take to train a 1024*1024 model? \nI fail to train a 1024 * 1024 because I only have a 1080Ti. Before I read your solution, I guess it is due to too small batch size. You can train it better so I want to see if there is a problem with my optimization method or if I have not trained enough time.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454827,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "01/12/2019 10:10:10",
          "content": "<p>Two days. The key is using float16 in Pytorch, if you do that you can double the size of your batch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 455123,
          "author_name": "lbyg1994",
          "author_url": "",
          "post_date": "01/13/2019 02:56:07",
          "content": "<p>Thanks！</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 456578,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "01/16/2019 05:09:37",
          "content": "<p>3rd place  used gradient accumulation. gonna use that later to test how it works, after I finish some other tests</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 456579,
          "author_name": "hokmund",
          "author_url": "",
          "post_date": "01/16/2019 05:15:40",
          "content": "<p>If there are no batch norms, it works out-of-box in PyTorch. But batch norms require custom handling.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 455067,
      "author_name": "gabrielhae",
      "author_url": "",
      "post_date": "01/12/2019 21:40:57",
      "content": "<p>I'm getting the impression you ran with comp data only, am I incorrect? Congrats on the score, textbook model diversity!</p>",
      "votes": null,
      "replies": [
        {
          "id": 455085,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "01/12/2019 23:46:22",
          "content": "<p>I had some models with only comp data because it was cleaner. Nevertheless, for rare classes I trained with HPA leaks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 455462,
      "author_name": "sventleman",
      "author_url": "",
      "post_date": "01/14/2019 00:46:17",
      "content": "<p>Great stuff, man and thanks for the float 16 tip - I cannot wait to go and test it myself! I have one question for you, how exactly did you learn the gammas (what kind of layer) and what do you mean by \"only use one parameter initialized at 1 without an activation\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 456231,
          "author_name": "arnaurm",
          "author_url": "",
          "post_date": "01/15/2019 11:42:31",
          "content": "<p>I made a new layer in PyTorch. It's just a layer that raises each channel to a specific scalar value.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 456482,
          "author_name": "sventleman",
          "author_url": "",
          "post_date": "01/15/2019 22:55:58",
          "content": "<p>That's the obvious way to do it and I felt dump right after asking the question... ;-) Thanks for your answer though and I hope some of that pain has started to fade away! Personally, I am really impressed by your performance.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456690,
      "author_name": "benchur",
      "author_url": "",
      "post_date": "01/16/2019 09:58:44",
      "content": "<p>Thanks for your sharing. It's a tough competition, I have just gotten 0.1+ score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 458245,
      "author_name": "nikhilpandey360",
      "author_url": "",
      "post_date": "01/19/2019 07:03:47",
      "content": "<p>Thanks for sharing this. I was stuck on a loop as I didn't know what else I could have done to increase the accuracy. Thanks for the post mate!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454372": "Hey kagglers, you know what they say: Be careful of overfitting, rare classes and bestfitting! Hah.\n \nJokes apart, congratulations to all the winners, this was a really tough competition. Thanks to the Kaggle organization and Human Protein Atlas for providing such a great challenge and also thanks to @brian @heng and @tilii for all the useful posts, I have learnt a lot from you. \n\nI am going to briefly detail my solution:\n\n\n##Hardware##\n  First I had a 980Ti GPU, 2600k CPU and 14GB RAM. Four weeks ago I bought a 1080Ti.\n\n##Preprocessing##\n  Removing duplicates was essential for improving your validation set and consequently for finding  your class thresholds. To do so I compared image channels using the imagehash library proposed by  @tilii\n\n\n##Base model##\n  ResNet50 trained with float16 precision, RGB,  1024x1024, lots of data augmentation, \n  weight decay, sgd, gradient clipping, bce loss and multi-scale resolution images. All in PyTorch.\n  Best single model [0.628 public, 0.553 private]\n\n##Gamma correction##\n Due to the dark exposure of some images, half of my models were trained with an extra initial layer to correct the gamma. The goal was to find the best gamma per channel for all the images. I tried to make a layer that extracted a custom gamma per image taking into account the image as a context but without success. I tried things like  i) using statistical image values as features and ii) applying a small convolution to the image then fc; all of them very unstable. The approach that worked best was the simplest: learn independent gammas (y_1 , y_2 , y_3) per channel and only use one parameter initialized at 1 without an activation. This reduced my global val_loss and increased my internal F1_score of some classes.\n  ![Gamma correction formula][1]\nAmusingly all the learnt gammas were quite similar, in the range [0.6-0.65]. Meaning that the images were quite cleared. I guess dark images benefitted from this layers but others - specially noisy images - don't, that's why I trained also networks without it.\n\n\n\n##What should I trust?##\nIs my model overfitting? \nThis was the big question of all the competition. It was hard answering it for the rare classes but for the other ones I found a pattern that I believe was quite useful. When comparing models, most of the times the models that were overfitting had the largest F1-Score in a zone where Precision was extremely dominating vs Recall.  Example in the following first figure. My hypothesis is that such models are learning things that are only present in the training set (similar cell/protein types, microscope features, who knows) and then become very confident in some samples increasing the Precision. When this happened I lowered thresholds or I discarded the model and choose another one (second figure).\n![F1 curves][2]\n\n##Model variety##\nI trained 11 models, 2 DenseNets121 and 9 ResNets50. They were fast to train and provided the best results in my val_set.  Each model had different settings such as:\nExcluding data leak, including it, change model seed, using yellow images , different class balance, full sized images, crops...\n\n##Rare classes##\nI did not know how to tackle them properly and I guess my fall has been mainly due to these classes (It's a pity that we cannot see F1-Score per class). I was about to try some Few Shot Learning approaches, but the image size was a limiting factor. Apart from this, I assume that the data leak eliminated most of the rare samples so detecting even just one of them was crucial (maybe too much). I hope I can learn from the top solutions about this.\n\n##Ensembling##\nI tried Logistic Regression, Averaging, Stacking and XGBoosting. In abundant classes it worked, but with others it overfitted and I thought it was not reliable. In the end I did the following:\n\n*Expert models*\nI decided which model were the best for specific classes, for instance: Yellow image model for classes 6, 7;   cropped image model for large classes such as 0, 25, full for 16; etc. These decisions were made taking into account hpa web info data, F1-curves, validation data, and public leaderboard.\n\n*One vote ensembling with top predictions*\nApart from expert models I increased the recall of my predictions by adding the top predictions from other models. Most of my models scored similar F1-scores per class, so it made sense. I increased the thresholds of non-expert models using this heuristic formula T_opt = T_opt + (1 - T_opt) *0.3 and included them as in one vote is enough.\n\n\n##Things that did not work##\n- Training a Multi Head Attention Module (Transformer paper) from the predictions of crops of an image in order to focus the attention in the spatial domain. \n-  Training GapNet style/Feature Pyramid with ResNet18 as backbone did not provide any benefits in my setup (I guess I did not play enough with it because @Dieter was able to make it work).\n- Training with 2048x2048 images, I guess batch size was too small and the extra resolution was not that important for the majority of the samples.\n\n  [1]: https://i.postimg.cc/rmh5h58T/Captura-de-pantalla-2019-01-11-a-les-12-34-16.png\n  [2]: https://i.postimg.cc/SR3cW84K/Captura-de-pantalla-2019-01-11-a-les-13-11-08.png",
    "454380": "<blockquote>It's a pity that we cannot see F1-Score per class</blockquote>\n\nYou can probe it by submitting your .csv file with some class and without it. Difference between two scores will tell you about binary f1 score of each class.",
    "454384": "Great idea. Even though what I would like is to compare my F1-Scores per class with the ones of top winners.  I guess when top winners upload their code I can generate their submission and extract theirs as well. Nice.",
    "454625": "Thanks for sharing. We were quite curious about how you were able to achieve your LB score. Congratulations to your solo-gold.",
    "454705": "This is Amazing! You able to chieve you score by a 1080Ti! I am quite curious about how long does it take to train a 1024*1024 model? \nI fail to train a 1024 * 1024 because I only have a 1080Ti. Before I read your solution, I guess it is due to too small batch size. You can train it better so I want to see if there is a problem with my optimization method or if I have not trained enough time.",
    "454827": "Two days. The key is using float16 in Pytorch, if you do that you can double the size of your batch.",
    "455067": "I'm getting the impression you ran with comp data only, am I incorrect? Congrats on the score, textbook model diversity!",
    "455085": "I had some models with only comp data because it was cleaner. Nevertheless, for rare classes I trained with HPA leaks.",
    "455123": "Thanks！",
    "455462": "Great stuff, man and thanks for the float 16 tip - I cannot wait to go and test it myself! I have one question for you, how exactly did you learn the gammas (what kind of layer) and what do you mean by \"only use one parameter initialized at 1 without an activation\"?",
    "456231": "I made a new layer in PyTorch. It's just a layer that raises each channel to a specific scalar value.",
    "456482": "That's the obvious way to do it and I felt dump right after asking the question... ;-) Thanks for your answer though and I hope some of that pain has started to fade away! Personally, I am really impressed by your performance.",
    "456578": "3rd place  used gradient accumulation. gonna use that later to test how it works, after I finish some other tests",
    "456579": "If there are no batch norms, it works out-of-box in PyTorch. But batch norms require custom handling.",
    "456690": "Thanks for your sharing. It's a tough competition, I have just gotten 0.1+ score.",
    "458245": "Thanks for sharing this. I was stuck on a loop as I didn't know what else I could have done to increase the accuracy. Thanks for the post mate!"
  },
  "source": "meta"
}