{
  "id": 77335,
  "title": "73th solution, only resnet18, with pytorch code",
  "url": "/competitions/human-protein-atlas-image-classification/writeups/cowboy-bebop-73th-solution-only-resnet18-with-pyto",
  "author_name": "",
  "post_date": "2019-01-13T16:05:46.397Z",
  "votes": 20,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I attend this competition to test my config-based pytorch pipeline borrowed from FAIR's <a href=\"https://github.com/facebookresearch/maskrcnn-benchmark/tree/master/maskrcnn_benchmark\">maskrcnn-benchmark</a>, which has been proven to be a nice architecture for fast prototyping using config files while maintaining simple &amp; clear reusable and easy-to-scale code base. I'm going to share some tools which might be helpful to build cv projects by pytorch 1.0.</p>\n\n<p>I did a lot of explorations like different networks, loss function, optimizer, lr schedulers etc. but it turns out <strong>a simple 4-Fold resnet18</strong> ensemble could achieve <strong>0.530</strong> (I failed to select my best submission). The key to train a good-to-go single model(about <strong>0.580</strong> LB) needs:</p>\n\n<ol>\n<li>Train valid set split using <a href=\"https://github.com/trent-b/iterative-stratification\">Multilabel Stratification</a>, you can find my implementation <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L32\">here</a></li>\n<li><strong>Weighted sampler</strong> to tackle unbalanced data in each batch, which could by achieved by <code>torch.utils.data.WeightedRandomSampler</code>, the weights could be generated by <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L79\">this method here</a></li>\n<li><strong>Train augmentation</strong> by random crop and resize. you could find the code <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/data/transforms/build.py#L5\">here</a>, I'd say <a href=\"https://github.com/aleju/imgaug\">imgaug</a> is really a useful tool for image augmentation. I use 288~448 crop size at a step of 32 then resize to 512</li>\n<li><strong>External data</strong>. remember to use correct preprocessing to match mean/std between train/external. this has been mentioned in the external data thread. I'd like to even do a histogram match between train and extra if I had time</li>\n<li><strong>Macro F1 loss</strong>. It seems that Weighted BCE works fine as well, but I'm using macro f1 directly since it has better performance than BCE in my experiments. you can find my pytorch implementation <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/model/loss.py#L25\">here</a>, which was inspired by <a href=\"https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric\">this kernel</a></li>\n<li>After all these preparations above, I used a simple <strong>resnet-18 with global average pooling + global max pooling concat</strong> as the final pooling layer. but merely using avgpool works fine in my experiments(only slight loss on performance in my observation). I also tried res34, res50, bn-inception, gapnet but none of them outperformanced my res18. you can find all my implementations <a href=\"https://github.com/shawnau/kaggle-HPA/tree/master/dl_backbone/model/base\">here</a>. Pretrainedmodels accelerates training, but training from scratch should achieve similar performance with enough epochs.</li>\n</ol>\n\n<hr>\n\n<ol>\n<li><strong>TTA</strong> didnt help, but no harm as well. I didn't try to submit with tta due to limited time. I just put  my implementations <a href=\"https://github.com/shawnau/kaggle-HPA/blob/2f58e4b7a4739b29f74e988c4b554774fdff1cd4/dl_backbone/data/transforms/build.py#L74\">here</a> for refer</li>\n<li>RGB have similar performance like RGBY, I just ensemble RGB with RGBY, but it's not a must for achieving 0.530 PB. It seems <strong>Y channel is not quite useful</strong> like many other kagglers reported</li>\n<li><strong>Threshold optimize</strong> always hurt my performance....I think its due to inconsistent of external+train dataset and test set. I used a brute-force method to pick hand-made thresholds for top-5 frequent classes, which boost public LB for 0.005 but not working in private LB.</li>\n</ol>\n\n<p>All the configurations of different settings(network, loss function, lr_scheduler, optimizer, sampler, tta, etc) could easily been achieved by <a href=\"https://github.com/shawnau/kaggle-HPA/tree/master/tools/config\">different config files here</a>. you could see how convenient to build config-based experiment pipeline.</p>\n\n<p>Finally, thanks kaggle for this competition, hope my code might be helpful to the kagglers who fights kaggle by pytorch. I almost switched to fastai at the beginning of the competition seeing <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">lafoss' kernel</a>....it's just too easy to use haha</p>\n\n<p><a href=\"https://github.com/shawnau/kaggle-HPA\">My contest code is here</a></p>",
  "messages": [
    {
      "id": "454424",
      "postDate": "01/11/2019 16:13:19",
      "content": "<p>I attend this competition to test my config-based pytorch pipeline borrowed from FAIR's <a href=\"https://github.com/facebookresearch/maskrcnn-benchmark/tree/master/maskrcnn_benchmark\">maskrcnn-benchmark</a>, which has been proven to be a nice architecture for fast prototyping using config files while maintaining simple &amp; clear reusable and easy-to-scale code base. I'm going to share some tools which might be helpful to build cv projects by pytorch 1.0.</p>\n\n<p>I did a lot of explorations like different networks, loss function, optimizer, lr schedulers etc. but it turns out <strong>a simple 4-Fold resnet18</strong> ensemble could achieve <strong>0.530</strong> (I failed to select my best submission). The key to train a good-to-go single model(about <strong>0.580</strong> LB) needs:</p>\n\n<ol>\n<li>Train valid set split using <a href=\"https://github.com/trent-b/iterative-stratification\">Multilabel Stratification</a>, you can find my implementation <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L32\">here</a></li>\n<li><strong>Weighted sampler</strong> to tackle unbalanced data in each batch, which could by achieved by <code>torch.utils.data.WeightedRandomSampler</code>, the weights could be generated by <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L79\">this method here</a></li>\n<li><strong>Train augmentation</strong> by random crop and resize. you could find the code <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/data/transforms/build.py#L5\">here</a>, I'd say <a href=\"https://github.com/aleju/imgaug\">imgaug</a> is really a useful tool for image augmentation. I use 288~448 crop size at a step of 32 then resize to 512</li>\n<li><strong>External data</strong>. remember to use correct preprocessing to match mean/std between train/external. this has been mentioned in the external data thread. I'd like to even do a histogram match between train and extra if I had time</li>\n<li><strong>Macro F1 loss</strong>. It seems that Weighted BCE works fine as well, but I'm using macro f1 directly since it has better performance than BCE in my experiments. you can find my pytorch implementation <a href=\"https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/model/loss.py#L25\">here</a>, which was inspired by <a href=\"https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric\">this kernel</a></li>\n<li>After all these preparations above, I used a simple <strong>resnet-18 with global average pooling + global max pooling concat</strong> as the final pooling layer. but merely using avgpool works fine in my experiments(only slight loss on performance in my observation). I also tried res34, res50, bn-inception, gapnet but none of them outperformanced my res18. you can find all my implementations <a href=\"https://github.com/shawnau/kaggle-HPA/tree/master/dl_backbone/model/base\">here</a>. Pretrainedmodels accelerates training, but training from scratch should achieve similar performance with enough epochs.</li>\n</ol>\n\n<hr>\n\n<ol>\n<li><strong>TTA</strong> didnt help, but no harm as well. I didn't try to submit with tta due to limited time. I just put  my implementations <a href=\"https://github.com/shawnau/kaggle-HPA/blob/2f58e4b7a4739b29f74e988c4b554774fdff1cd4/dl_backbone/data/transforms/build.py#L74\">here</a> for refer</li>\n<li>RGB have similar performance like RGBY, I just ensemble RGB with RGBY, but it's not a must for achieving 0.530 PB. It seems <strong>Y channel is not quite useful</strong> like many other kagglers reported</li>\n<li><strong>Threshold optimize</strong> always hurt my performance....I think its due to inconsistent of external+train dataset and test set. I used a brute-force method to pick hand-made thresholds for top-5 frequent classes, which boost public LB for 0.005 but not working in private LB.</li>\n</ol>\n\n<p>All the configurations of different settings(network, loss function, lr_scheduler, optimizer, sampler, tta, etc) could easily been achieved by <a href=\"https://github.com/shawnau/kaggle-HPA/tree/master/tools/config\">different config files here</a>. you could see how convenient to build config-based experiment pipeline.</p>\n\n<p>Finally, thanks kaggle for this competition, hope my code might be helpful to the kagglers who fights kaggle by pytorch. I almost switched to fastai at the beginning of the competition seeing <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">lafoss' kernel</a>....it's just too easy to use haha</p>\n\n<p><a href=\"https://github.com/shawnau/kaggle-HPA\">My contest code is here</a></p>",
      "rawMarkdown": "I attend this competition to test my config-based pytorch pipeline borrowed from FAIR's [maskrcnn-benchmark][1], which has been proven to be a nice architecture for fast prototyping using config files while maintaining simple &amp; clear reusable and easy-to-scale code base. I'm going to share some tools which might be helpful to build cv projects by pytorch 1.0.\n\nI did a lot of explorations like different networks, loss function, optimizer, lr schedulers etc. but it turns out **a simple 4-Fold resnet18** ensemble could achieve **0.530** (I failed to select my best submission). The key to train a good-to-go single model(about **0.580** LB) needs:\n\n1. Train valid set split using [Multilabel Stratification](https://github.com/trent-b/iterative-stratification), you can find my implementation [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L32)\n2. **Weighted sampler** to tackle unbalanced data in each batch, which could by achieved by `torch.utils.data.WeightedRandomSampler`, the weights could be generated by [this method here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L79)\n3. **Train augmentation** by random crop and resize. you could find the code [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/data/transforms/build.py#L5), I'd say [imgaug](https://github.com/aleju/imgaug) is really a useful tool for image augmentation. I use 288~448 crop size at a step of 32 then resize to 512\n4. **External data**. remember to use correct preprocessing to match mean/std between train/external. this has been mentioned in the external data thread. I'd like to even do a histogram match between train and extra if I had time\n5. **Macro F1 loss**. It seems that Weighted BCE works fine as well, but I'm using macro f1 directly since it has better performance than BCE in my experiments. you can find my pytorch implementation [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/model/loss.py#L25), which was inspired by [this kernel](https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric)\n6. After all these preparations above, I used a simple **resnet-18 with global average pooling + global max pooling concat** as the final pooling layer. but merely using avgpool works fine in my experiments(only slight loss on performance in my observation). I also tried res34, res50, bn-inception, gapnet but none of them outperformanced my res18. you can find all my implementations [here](https://github.com/shawnau/kaggle-HPA/tree/master/dl_backbone/model/base). Pretrainedmodels accelerates training, but training from scratch should achieve similar performance with enough epochs.\n\n---\n\n7. **TTA** didnt help, but no harm as well. I didn't try to submit with tta due to limited time. I just put  my implementations [here](https://github.com/shawnau/kaggle-HPA/blob/2f58e4b7a4739b29f74e988c4b554774fdff1cd4/dl_backbone/data/transforms/build.py#L74) for refer\n8. RGB have similar performance like RGBY, I just ensemble RGB with RGBY, but it's not a must for achieving 0.530 PB. It seems **Y channel is not quite useful** like many other kagglers reported\n9. **Threshold optimize** always hurt my performance....I think its due to inconsistent of external+train dataset and test set. I used a brute-force method to pick hand-made thresholds for top-5 frequent classes, which boost public LB for 0.005 but not working in private LB.\n\nAll the configurations of different settings(network, loss function, lr_scheduler, optimizer, sampler, tta, etc) could easily been achieved by [different config files here](https://github.com/shawnau/kaggle-HPA/tree/master/tools/config). you could see how convenient to build config-based experiment pipeline.\n\nFinally, thanks kaggle for this competition, hope my code might be helpful to the kagglers who fights kaggle by pytorch. I almost switched to fastai at the beginning of the competition seeing [lafoss' kernel](https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb)....it's just too easy to use haha\n\n[My contest code is here](https://github.com/shawnau/kaggle-HPA)\n\n  [1]: https://github.com/facebookresearch/maskrcnn-benchmark/tree/master/maskrcnn_benchmark",
      "votes": null
    },
    {
      "id": "454442",
      "postDate": "01/11/2019 16:38:49",
      "content": "<p>I also experienced that global average + global max pooling and avgpool were similar and slightly better to use global average + global max pooling. It's quite shock for me to learn resnet18 which has low capacity was quite enough for this competition after reading some solution overviews. Thank you for the code and explanation.</p>",
      "rawMarkdown": "I also experienced that global average + global max pooling and avgpool were similar and slightly better to use global average + global max pooling. It's quite shock for me to learn resnet18 which has low capacity was quite enough for this competition after reading some solution overviews. Thank you for the code and explanation.",
      "votes": null
    },
    {
      "id": "454458",
      "postDate": "01/11/2019 16:59:51",
      "content": "<p>thanks for sharing! btw, in my experiments, removing dropouts led to overfitting, although it is said that using dropout with bachnorm could be harmful...I still kept it in my codes. (I forgot the source of this rumor</p>",
      "rawMarkdown": "thanks for sharing! btw, in my experiments, removing dropouts led to overfitting, although it is said that using dropout with bachnorm could be harmful...I still kept it in my codes. (I forgot the source of this rumor",
      "votes": null
    },
    {
      "id": "455246",
      "postDate": "01/13/2019 11:55:46",
      "content": "<p>Thanks for publishing your solution. I use a similar pipeline with json files, it really helps with keeping track of experiments. Though, I was never able to achieve the performance you got, not even with more complex networks so I'll probably have a close look at your code and see where I went wrong.</p>\n\n<p>I have a question about the cropping; do you change the target when you crop? Because some of the proteins can be quite localized and if you crop out that region of the image the target is no longer true. </p>",
      "rawMarkdown": "Thanks for publishing your solution. I use a similar pipeline with json files, it really helps with keeping track of experiments. Though, I was never able to achieve the performance you got, not even with more complex networks so I'll probably have a close look at your code and see where I went wrong.\n\nI have a question about the cropping; do you change the target when you crop? Because some of the proteins can be quite localized and if you crop out that region of the image the target is no longer true.",
      "votes": null
    },
    {
      "id": "455329",
      "postDate": "01/13/2019 16:02:14",
      "content": "<p>Good question. I was concerning about this issue when using crop and resize as augmentation as well, so I just used a straightforward method: shrink(-16 pixel) the crop size step by step, to see the performance of each experiment. I stopped at 256*256, it do have better performance on my local cv. But I didn't try smaller crops due to limited resource...thank you!</p>",
      "rawMarkdown": "Good question. I was concerning about this issue when using crop and resize as augmentation as well, so I just used a straightforward method: shrink(-16 pixel) the crop size step by step, to see the performance of each experiment. I stopped at 256*256, it do have better performance on my local cv. But I didn't try smaller crops due to limited resource...thank you!",
      "votes": null
    },
    {
      "id": "456287",
      "postDate": "01/15/2019 13:45:18",
      "content": "<p>Nice solution，quite helpful for beginners like me.</p>\n\n<p>I see Focal Loss in your code, and it is quite different from the focal loss in papers.</p>\n\n<p>So I am very confused and need some help about it.</p>\n\n<p>Could you give me some advice or information?</p>\n\n<p>Thanks!!!!!!!</p>",
      "rawMarkdown": "Nice solution，quite helpful for beginners like me.\n\nI see Focal Loss in your code, and it is quite different from the focal loss in papers.\n\nSo I am very confused and need some help about it.\n\nCould you give me some advice or information?\n\nThanks!!!!!!!",
      "votes": null
    },
    {
      "id": "531068",
      "postDate": "05/14/2019 08:22:56",
      "content": "<p>Hi, thanks for your sharing, it's helpful for a beginner like me. \nI have some questions about Weight Sampler and Weighted BCELoss. From my point, when I use Weight Sampler, the real proportion and weight of the input data will be changed to make them more balanced. So the weight used by the Weighted BCELoss will be not corrected anymore. So I think, we can't use them simultaneously, is my opinion correct, and is there any method to solve this problem?</p>",
      "rawMarkdown": "Hi, thanks for your sharing, it's helpful for a beginner like me. \nI have some questions about Weight Sampler and Weighted BCELoss. From my point, when I use Weight Sampler, the real proportion and weight of the input data will be changed to make them more balanced. So the weight used by the Weighted BCELoss will be not corrected anymore. So I think, we can't use them simultaneously, is my opinion correct, and is there any method to solve this problem?",
      "votes": null
    },
    {
      "id": "534020",
      "postDate": "05/20/2019 13:14:28",
      "content": "<p>thanks for the reply! I agree with you. Using both will break the balance for the loss. I think using one of them is the correct choice, but maybe the weights could be adjusted more.</p>",
      "rawMarkdown": "thanks for the reply! I agree with you. Using both will break the balance for the loss. I think using one of them is the correct choice, but maybe the weights could be adjusted more.",
      "votes": null
    },
    {
      "id": "535501",
      "postDate": "05/23/2019 03:37:07",
      "content": "<p>Hi, thanks for sharing your solution.</p>\n\n<p>Is it ok to use weighted sampler and macro f1 loss simultaneously? It seems that macro f1 loss is sort of similar to weighted BCELoss because it treats all classes in the same way.</p>",
      "rawMarkdown": "Hi, thanks for sharing your solution.\n\nIs it ok to use weighted sampler and macro f1 loss simultaneously? It seems that macro f1 loss is sort of similar to weighted BCELoss because it treats all classes in the same way.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454442,
      "author_name": "soonhwankwon",
      "author_url": "",
      "post_date": "01/11/2019 16:38:49",
      "content": "<p>I also experienced that global average + global max pooling and avgpool were similar and slightly better to use global average + global max pooling. It's quite shock for me to learn resnet18 which has low capacity was quite enough for this competition after reading some solution overviews. Thank you for the code and explanation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 454458,
          "author_name": "atm584",
          "author_url": "",
          "post_date": "01/11/2019 16:59:51",
          "content": "<p>thanks for sharing! btw, in my experiments, removing dropouts led to overfitting, although it is said that using dropout with bachnorm could be harmful...I still kept it in my codes. (I forgot the source of this rumor</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 455246,
      "author_name": "dr1t10",
      "author_url": "",
      "post_date": "01/13/2019 11:55:46",
      "content": "<p>Thanks for publishing your solution. I use a similar pipeline with json files, it really helps with keeping track of experiments. Though, I was never able to achieve the performance you got, not even with more complex networks so I'll probably have a close look at your code and see where I went wrong.</p>\n\n<p>I have a question about the cropping; do you change the target when you crop? Because some of the proteins can be quite localized and if you crop out that region of the image the target is no longer true. </p>",
      "votes": null,
      "replies": [
        {
          "id": 455329,
          "author_name": "atm584",
          "author_url": "",
          "post_date": "01/13/2019 16:02:14",
          "content": "<p>Good question. I was concerning about this issue when using crop and resize as augmentation as well, so I just used a straightforward method: shrink(-16 pixel) the crop size step by step, to see the performance of each experiment. I stopped at 256*256, it do have better performance on my local cv. But I didn't try smaller crops due to limited resource...thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456287,
      "author_name": "noexecution",
      "author_url": "",
      "post_date": "01/15/2019 13:45:18",
      "content": "<p>Nice solution，quite helpful for beginners like me.</p>\n\n<p>I see Focal Loss in your code, and it is quite different from the focal loss in papers.</p>\n\n<p>So I am very confused and need some help about it.</p>\n\n<p>Could you give me some advice or information?</p>\n\n<p>Thanks!!!!!!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 531068,
      "author_name": "qunyang",
      "author_url": "",
      "post_date": "05/14/2019 08:22:56",
      "content": "<p>Hi, thanks for your sharing, it's helpful for a beginner like me. \nI have some questions about Weight Sampler and Weighted BCELoss. From my point, when I use Weight Sampler, the real proportion and weight of the input data will be changed to make them more balanced. So the weight used by the Weighted BCELoss will be not corrected anymore. So I think, we can't use them simultaneously, is my opinion correct, and is there any method to solve this problem?</p>",
      "votes": null,
      "replies": [
        {
          "id": 534020,
          "author_name": "atm584",
          "author_url": "",
          "post_date": "05/20/2019 13:14:28",
          "content": "<p>thanks for the reply! I agree with you. Using both will break the balance for the loss. I think using one of them is the correct choice, but maybe the weights could be adjusted more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 535501,
          "author_name": "",
          "author_url": "",
          "post_date": "05/23/2019 03:37:07",
          "content": "<p>Hi, thanks for sharing your solution.</p>\n\n<p>Is it ok to use weighted sampler and macro f1 loss simultaneously? It seems that macro f1 loss is sort of similar to weighted BCELoss because it treats all classes in the same way.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454424": "I attend this competition to test my config-based pytorch pipeline borrowed from FAIR's [maskrcnn-benchmark][1], which has been proven to be a nice architecture for fast prototyping using config files while maintaining simple &amp; clear reusable and easy-to-scale code base. I'm going to share some tools which might be helpful to build cv projects by pytorch 1.0.\n\nI did a lot of explorations like different networks, loss function, optimizer, lr schedulers etc. but it turns out **a simple 4-Fold resnet18** ensemble could achieve **0.530** (I failed to select my best submission). The key to train a good-to-go single model(about **0.580** LB) needs:\n\n1. Train valid set split using [Multilabel Stratification](https://github.com/trent-b/iterative-stratification), you can find my implementation [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L32)\n2. **Weighted sampler** to tackle unbalanced data in each batch, which could by achieved by `torch.utils.data.WeightedRandomSampler`, the weights could be generated by [this method here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/tools/preprocess.py#L79)\n3. **Train augmentation** by random crop and resize. you could find the code [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/data/transforms/build.py#L5), I'd say [imgaug](https://github.com/aleju/imgaug) is really a useful tool for image augmentation. I use 288~448 crop size at a step of 32 then resize to 512\n4. **External data**. remember to use correct preprocessing to match mean/std between train/external. this has been mentioned in the external data thread. I'd like to even do a histogram match between train and extra if I had time\n5. **Macro F1 loss**. It seems that Weighted BCE works fine as well, but I'm using macro f1 directly since it has better performance than BCE in my experiments. you can find my pytorch implementation [here](https://github.com/shawnau/kaggle-HPA/blob/d6071ff37d5db7612b2f380323a7d48309cc13fe/dl_backbone/model/loss.py#L25), which was inspired by [this kernel](https://www.kaggle.com/rejpalcz/best-loss-function-for-f1-score-metric)\n6. After all these preparations above, I used a simple **resnet-18 with global average pooling + global max pooling concat** as the final pooling layer. but merely using avgpool works fine in my experiments(only slight loss on performance in my observation). I also tried res34, res50, bn-inception, gapnet but none of them outperformanced my res18. you can find all my implementations [here](https://github.com/shawnau/kaggle-HPA/tree/master/dl_backbone/model/base). Pretrainedmodels accelerates training, but training from scratch should achieve similar performance with enough epochs.\n\n---\n\n7. **TTA** didnt help, but no harm as well. I didn't try to submit with tta due to limited time. I just put  my implementations [here](https://github.com/shawnau/kaggle-HPA/blob/2f58e4b7a4739b29f74e988c4b554774fdff1cd4/dl_backbone/data/transforms/build.py#L74) for refer\n8. RGB have similar performance like RGBY, I just ensemble RGB with RGBY, but it's not a must for achieving 0.530 PB. It seems **Y channel is not quite useful** like many other kagglers reported\n9. **Threshold optimize** always hurt my performance....I think its due to inconsistent of external+train dataset and test set. I used a brute-force method to pick hand-made thresholds for top-5 frequent classes, which boost public LB for 0.005 but not working in private LB.\n\nAll the configurations of different settings(network, loss function, lr_scheduler, optimizer, sampler, tta, etc) could easily been achieved by [different config files here](https://github.com/shawnau/kaggle-HPA/tree/master/tools/config). you could see how convenient to build config-based experiment pipeline.\n\nFinally, thanks kaggle for this competition, hope my code might be helpful to the kagglers who fights kaggle by pytorch. I almost switched to fastai at the beginning of the competition seeing [lafoss' kernel](https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb)....it's just too easy to use haha\n\n[My contest code is here](https://github.com/shawnau/kaggle-HPA)\n\n  [1]: https://github.com/facebookresearch/maskrcnn-benchmark/tree/master/maskrcnn_benchmark",
    "454442": "I also experienced that global average + global max pooling and avgpool were similar and slightly better to use global average + global max pooling. It's quite shock for me to learn resnet18 which has low capacity was quite enough for this competition after reading some solution overviews. Thank you for the code and explanation.",
    "454458": "thanks for sharing! btw, in my experiments, removing dropouts led to overfitting, although it is said that using dropout with bachnorm could be harmful...I still kept it in my codes. (I forgot the source of this rumor",
    "455246": "Thanks for publishing your solution. I use a similar pipeline with json files, it really helps with keeping track of experiments. Though, I was never able to achieve the performance you got, not even with more complex networks so I'll probably have a close look at your code and see where I went wrong.\n\nI have a question about the cropping; do you change the target when you crop? Because some of the proteins can be quite localized and if you crop out that region of the image the target is no longer true.",
    "455329": "Good question. I was concerning about this issue when using crop and resize as augmentation as well, so I just used a straightforward method: shrink(-16 pixel) the crop size step by step, to see the performance of each experiment. I stopped at 256*256, it do have better performance on my local cv. But I didn't try smaller crops due to limited resource...thank you!",
    "456287": "Nice solution，quite helpful for beginners like me.\n\nI see Focal Loss in your code, and it is quite different from the focal loss in papers.\n\nSo I am very confused and need some help about it.\n\nCould you give me some advice or information?\n\nThanks!!!!!!!",
    "531068": "Hi, thanks for your sharing, it's helpful for a beginner like me. \nI have some questions about Weight Sampler and Weighted BCELoss. From my point, when I use Weight Sampler, the real proportion and weight of the input data will be changed to make them more balanced. So the weight used by the Weighted BCELoss will be not corrected anymore. So I think, we can't use them simultaneously, is my opinion correct, and is there any method to solve this problem?",
    "534020": "thanks for the reply! I agree with you. Using both will break the balance for the loss. I think using one of them is the correct choice, but maybe the weights could be adjusted more.",
    "535501": "Hi, thanks for sharing your solution.\n\nIs it ok to use weighted sampler and macro f1 loss simultaneously? It seems that macro f1 loss is sort of similar to weighted BCELoss because it treats all classes in the same way."
  },
  "source": "meta"
}