{
  "id": 87367,
  "title": "How I drop from 0.9805 Private LB to 0.974 (113rd Solution)",
  "url": "/competitions/histopathologic-cancer-detection/discussion/87367",
  "author_name": "Hanke Chen",
  "post_date": "2019-03-31T00:58:03.715000",
  "votes": 9,
  "comment_count": 13,
  "views": 0,
  "content": "<h1>Discoveries</h1>\n\n<ol>\n<li>Discover Duplicate: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367\">Discover Duplicate</a>  </li>\n<li>Discover Leak: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85424\">How to get LB 1.000</a>  </li>\n<li>Invent WSI Normalization: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87069\">Better Way to Normalize Slides? WSI Normalization</a>  </li>\n<li>Full Solution Summary: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367\">How I drop from 0.9805 Private LB to 0.974 (113rd Solution)</a>  </li>\n</ol>\n\n<h1>My LB 0.980 Solution</h1>\n\n<h2>Single Model</h2>\n\n<ul>\n<li>Private: - 1. 0.9785</li>\n<li>Public: 0.9753</li>\n<li>Network: SEResNeXt50 (modified)\n<ul><li>Input: 128*128 without maxpooling</li>\n<li>BN, FC, ELU with 3 layers as a classifier</li>\n<li>cat maxpooling and avg pooling</li></ul></li>\n<li>Optimizer: SGD w/ momentum 0.9 (Adam is bad for me)</li>\n<li>LR Scheduler: Triangle lr with Decrease on Plateau</li>\n<li>Fold: 5 fold ensemble</li>\n<li>TTA: 16</li>\n<li>Validation Time Augmentation: 4</li>\n<li>Augmentation: from Kernels + 20% center focus by padding</li>\n<li>Dropout FC: 0.9</li>\n<li>Dropout CNN: 0.2</li>\n<li>Loss: BCE</li>\n<li>Activation: ELU with KaimingInit</li>\n<li>Training: Freeze 3 epoch, train to 6th epoch is the best</li>\n<li>Tricks:\n<ul><li>Snapshot Ensemble (about 2 each fold)</li>\n<li>Shakeup simulation</li>\n<li>8x scheduled Augmentation for faster converge</li>\n<li>Create a stable CV based on WSI</li>\n<li>Max and Min lr based on the Paper (1 cycle + gradient warmup)</li>\n<li>Faster loading with preprocessing images into .npy</li></ul></li>\n</ul>\n\n<h2>Pseudo labeling (with all probability)</h2>\n\n<p>I did not have time to train 5 fold pseudo\nBut here are the scores with 4 tta ensemble:\n(They are generally better +0.008 on both public and private)\nFold, Private, Public\nFold1: 0.9765, 0.9763\nFold2: 0.9787, 0.9668</p>\n\n<h2>Ensemble</h2>\n\n<p>Simple weighted average ensemble with 1 fold NasNet, DenseNet...\nGives Private 0.9805, Public 0.9780</p>\n\n<h1>Ideas that did not work</h1>\n\n<p>Jigsaw puzzle (no unified patterns found)\nAverage, Geometric, Power ensemble are about the same\nPostprocess (setting white 32*32 images in the middle to negative labels)</p>\n\n<h1>Ideas not implemented</h1>\n\n<ul>\n<li>Whole slide image normalization (no enough time + not sure if it works)</li>\n<li>Cheating with 1.000 LB (I don't think I am the first one who discovered this)</li>\n<li>WSI info from test set (I pretend I did not discover this)</li>\n<li>Distill with teacher and student</li>\n<li>Mixup</li>\n<li>Attention layers</li>\n<li>Inceptionv3, v4, X</li>\n<li>attention weighted average pooling</li>\n<li>global max pooling</li>\n<li>global std pooling</li>\n<li>gradient aggregation</li>\n<li>ImageNet pre-train (don't think it will be useful)</li>\n</ul>\n\n<h1>Self Reflection</h1>\n\n<p>The only difference between 0.980 and 0.974 is my postprocess: setting white 32*32 images in the middle to negative labels. It gives all my model a stable LB improvement 0.0001. The change is so small that I did not even think whether I should use the postprocess strategy or not.</p>\n\n<p>The mistake comes from not understanding the data: the patches are cropped from WSI. That means even if the image is empty, it may be a huge white cell inside the tumor!</p>\n\n<p>The other mistake is that I did not evaluate my postprocess through holdout set in CV. I trusted 1/5 LB too much. But looking back, I would still choose to use postprocess because the chance of getting an improvement of 0.0001 in Public LB and getting a huge drop 0.01 in Private LB is pretty low.</p>\n\n<p>See you in the next cv competition!</p>",
  "messages": [
    {
      "id": 504092,
      "postDate": "2019-03-31T00:58:03.717Z",
      "content": "<h1>Discoveries</h1>\n\n<ol>\n<li>Discover Duplicate: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367\">Discover Duplicate</a>  </li>\n<li>Discover Leak: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85424\">How to get LB 1.000</a>  </li>\n<li>Invent WSI Normalization: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87069\">Better Way to Normalize Slides? WSI Normalization</a>  </li>\n<li>Full Solution Summary: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367\">How I drop from 0.9805 Private LB to 0.974 (113rd Solution)</a>  </li>\n</ol>\n\n<h1>My LB 0.980 Solution</h1>\n\n<h2>Single Model</h2>\n\n<ul>\n<li>Private: - 1. 0.9785</li>\n<li>Public: 0.9753</li>\n<li>Network: SEResNeXt50 (modified)\n<ul><li>Input: 128*128 without maxpooling</li>\n<li>BN, FC, ELU with 3 layers as a classifier</li>\n<li>cat maxpooling and avg pooling</li></ul></li>\n<li>Optimizer: SGD w/ momentum 0.9 (Adam is bad for me)</li>\n<li>LR Scheduler: Triangle lr with Decrease on Plateau</li>\n<li>Fold: 5 fold ensemble</li>\n<li>TTA: 16</li>\n<li>Validation Time Augmentation: 4</li>\n<li>Augmentation: from Kernels + 20% center focus by padding</li>\n<li>Dropout FC: 0.9</li>\n<li>Dropout CNN: 0.2</li>\n<li>Loss: BCE</li>\n<li>Activation: ELU with KaimingInit</li>\n<li>Training: Freeze 3 epoch, train to 6th epoch is the best</li>\n<li>Tricks:\n<ul><li>Snapshot Ensemble (about 2 each fold)</li>\n<li>Shakeup simulation</li>\n<li>8x scheduled Augmentation for faster converge</li>\n<li>Create a stable CV based on WSI</li>\n<li>Max and Min lr based on the Paper (1 cycle + gradient warmup)</li>\n<li>Faster loading with preprocessing images into .npy</li></ul></li>\n</ul>\n\n<h2>Pseudo labeling (with all probability)</h2>\n\n<p>I did not have time to train 5 fold pseudo\nBut here are the scores with 4 tta ensemble:\n(They are generally better +0.008 on both public and private)\nFold, Private, Public\nFold1: 0.9765, 0.9763\nFold2: 0.9787, 0.9668</p>\n\n<h2>Ensemble</h2>\n\n<p>Simple weighted average ensemble with 1 fold NasNet, DenseNet...\nGives Private 0.9805, Public 0.9780</p>\n\n<h1>Ideas that did not work</h1>\n\n<p>Jigsaw puzzle (no unified patterns found)\nAverage, Geometric, Power ensemble are about the same\nPostprocess (setting white 32*32 images in the middle to negative labels)</p>\n\n<h1>Ideas not implemented</h1>\n\n<ul>\n<li>Whole slide image normalization (no enough time + not sure if it works)</li>\n<li>Cheating with 1.000 LB (I don't think I am the first one who discovered this)</li>\n<li>WSI info from test set (I pretend I did not discover this)</li>\n<li>Distill with teacher and student</li>\n<li>Mixup</li>\n<li>Attention layers</li>\n<li>Inceptionv3, v4, X</li>\n<li>attention weighted average pooling</li>\n<li>global max pooling</li>\n<li>global std pooling</li>\n<li>gradient aggregation</li>\n<li>ImageNet pre-train (don't think it will be useful)</li>\n</ul>\n\n<h1>Self Reflection</h1>\n\n<p>The only difference between 0.980 and 0.974 is my postprocess: setting white 32*32 images in the middle to negative labels. It gives all my model a stable LB improvement 0.0001. The change is so small that I did not even think whether I should use the postprocess strategy or not.</p>\n\n<p>The mistake comes from not understanding the data: the patches are cropped from WSI. That means even if the image is empty, it may be a huge white cell inside the tumor!</p>\n\n<p>The other mistake is that I did not evaluate my postprocess through holdout set in CV. I trusted 1/5 LB too much. But looking back, I would still choose to use postprocess because the chance of getting an improvement of 0.0001 in Public LB and getting a huge drop 0.01 in Private LB is pretty low.</p>\n\n<p>See you in the next cv competition!</p>",
      "rawMarkdown": "# Discoveries\n1. Discover Duplicate: [Discover Duplicate](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367)  \n2. Discover Leak: [How to get LB 1.000](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85424)  \n3. Invent WSI Normalization: [Better Way to Normalize Slides? WSI Normalization](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87069)  \n4. Full Solution Summary: [How I drop from 0.9805 Private LB to 0.974 (113rd Solution)](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367)  \n\n# My LB 0.980 Solution\n\n## Single Model\n- Private: - 1. 0.9785\n- Public: 0.9753\n- Network: SEResNeXt50 (modified)\n  - Input: 128*128 without maxpooling\n  - BN, FC, ELU with 3 layers as a classifier\n  - cat maxpooling and avg pooling\n- Optimizer: SGD w/ momentum 0.9 (Adam is bad for me)\n- LR Scheduler: Triangle lr with Decrease on Plateau\n- Fold: 5 fold ensemble\n- TTA: 16\n- Validation Time Augmentation: 4\n- Augmentation: from Kernels + 20% center focus by padding\n- Dropout FC: 0.9\n- Dropout CNN: 0.2\n- Loss: BCE\n- Activation: ELU with KaimingInit\n- Training: Freeze 3 epoch, train to 6th epoch is the best\n- Tricks:\n  - Snapshot Ensemble (about 2 each fold)\n  - Shakeup simulation\n  - 8x scheduled Augmentation for faster converge\n  - Create a stable CV based on WSI\n  - Max and Min lr based on the Paper (1 cycle + gradient warmup)\n  - Faster loading with preprocessing images into .npy\n\n## Pseudo labeling (with all probability)\nI did not have time to train 5 fold pseudo\nBut here are the scores with 4 tta ensemble:\n(They are generally better +0.008 on both public and private)\nFold, Private, Public\nFold1: 0.9765, 0.9763\nFold2: 0.9787, 0.9668\n\n## Ensemble\nSimple weighted average ensemble with 1 fold NasNet, DenseNet...\nGives Private 0.9805, Public 0.9780\n\n# Ideas that did not work\nJigsaw puzzle (no unified patterns found)\nAverage, Geometric, Power ensemble are about the same\nPostprocess (setting white 32*32 images in the middle to negative labels)\n\n# Ideas not implemented\n- Whole slide image normalization (no enough time + not sure if it works)\n- Cheating with 1.000 LB (I don't think I am the first one who discovered this)\n- WSI info from test set (I pretend I did not discover this)\n- Distill with teacher and student\n- Mixup\n- Attention layers\n- Inceptionv3, v4, X\n- attention weighted average pooling\n- global max pooling\n- global std pooling\n- gradient aggregation\n- ImageNet pre-train (don't think it will be useful)\n\n# Self Reflection\nThe only difference between 0.980 and 0.974 is my postprocess: setting white 32*32 images in the middle to negative labels. It gives all my model a stable LB improvement 0.0001. The change is so small that I did not even think whether I should use the postprocess strategy or not.\n\nThe mistake comes from not understanding the data: the patches are cropped from WSI. That means even if the image is empty, it may be a huge white cell inside the tumor!\n\nThe other mistake is that I did not evaluate my postprocess through holdout set in CV. I trusted 1/5 LB too much. But looking back, I would still choose to use postprocess because the chance of getting an improvement of 0.0001 in Public LB and getting a huge drop 0.01 in Private LB is pretty low.\n\nSee you in the next cv competition!",
      "votes": 9
    },
    {
      "id": 504110,
      "postDate": "2019-03-31T02:21:44.463Z",
      "content": "<p>My best model on the private lb is the one that I set as public.... I never thought it was the best, 0.9744 on public lb and 0.9791 on private lb.</p>",
      "rawMarkdown": "My best model on the private lb is the one that I set as public.... I never thought it was the best, 0.9744 on public lb and 0.9791 on private lb.",
      "votes": 7,
      "replies": [
        {
          "id": 504467,
          "postDate": "2019-03-31T16:55:20.467Z",
          "content": "<p>Crazy!</p>",
          "rawMarkdown": "Crazy!",
          "votes": 2
        }
      ]
    },
    {
      "id": 504456,
      "postDate": "2019-03-31T16:48:53.163Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": 1
    },
    {
      "id": 504103,
      "postDate": "2019-03-31T02:00:07.547Z",
      "content": "<p>Nice work. I also used ELU in my classifier. I never had much luck with se networks, seresnext in particular was so slow that I could not use it at all. My best models are all from densenet169 but senet154 and preresnet269b gave some good results, just 4-8 times longer than densenet and with lower scores. </p>",
      "rawMarkdown": "Nice work. I also used ELU in my classifier. I never had much luck with se networks, seresnext in particular was so slow that I could not use it at all. My best models are all from densenet169 but senet154 and preresnet269b gave some good results, just 4-8 times longer than densenet and with lower scores. ",
      "votes": 1,
      "replies": [
        {
          "id": 504127,
          "postDate": "2019-03-31T03:05:43.813Z",
          "content": "<p>I am using SEResNeXt50. With your <code>senet154 and preresnet269b</code>, did you experience a huge overfit? Since I discovered overfit within 8 epoch, I did not try deeper networks.</p>",
          "rawMarkdown": "I am using SEResNeXt50. With your `senet154 and preresnet269b`, did you experience a huge overfit? Since I discovered overfit within 8 epoch, I did not try deeper networks."
        }
      ]
    },
    {
      "id": 504205,
      "postDate": "2019-03-31T06:30:24.503Z",
      "content": "<p>My observations on this competition:\n1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.\n2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.\n3) More TTA helps.</p>",
      "rawMarkdown": "My observations on this competition:\n1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.\n2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.\n3) More TTA helps.",
      "votes": 2,
      "replies": [
        {
          "id": 504358,
          "postDate": "2019-03-31T13:23:10.990Z",
          "content": "<blockquote>\n  <p>1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.</p>\n</blockquote>\n\n<p>Yes. I warn myself a lot of time \"not to trust LB\", all my tricks are tested against CV except the postprocess. I learned a harsh lesson now.</p>\n\n<blockquote>\n  <p>2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.</p>\n</blockquote>\n\n<p>This question may seem dumb: How can those images (96*96) fit into your network without upscaling?</p>",
          "rawMarkdown": "&gt; 1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.\n\nYes. I warn myself a lot of time \"not to trust LB\", all my tricks are tested against CV except the postprocess. I learned a harsh lesson now.\n\n&gt; 2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.\n\nThis question may seem dumb: How can those images (96*96) fit into your network without upscaling?"
        },
        {
          "id": 504464,
          "postDate": "2019-03-31T16:52:40.657Z",
          "content": "<p>The minimum resolution for most of image networks is 32x32 that will produce 1x1 output at the last convolution layer. For 96x96 input you will have 3x3 output with the center corresponding to the region  in the original image under detection (though central pooling worked worse than concat pooling). For common network resolution like 224x224 you just have 7x7 output. If you use adaptive pooling the network works with any of above-mentioned resolutions, and the head fully connected part remains the same.</p>",
          "rawMarkdown": "The minimum resolution for most of image networks is 32x32 that will produce 1x1 output at the last convolution layer. For 96x96 input you will have 3x3 output with the center corresponding to the region  in the original image under detection (though central pooling worked worse than concat pooling). For common network resolution like 224x224 you just have 7x7 output. If you use adaptive pooling the network works with any of above-mentioned resolutions, and the head fully connected part remains the same.",
          "votes": 1
        }
      ]
    },
    {
      "id": 504179,
      "postDate": "2019-03-31T04:51:34.327Z",
      "content": "<p>Anybody got good results using stain augmentation / normalization?</p>",
      "rawMarkdown": "Anybody got good results using stain augmentation / normalization?",
      "replies": [
        {
          "id": 504390,
          "postDate": "2019-03-31T14:24:46.937Z",
          "content": "<p>I tried stain normalization twice with two different reference images and I didn't see any improvements.  I also tried mixup but it gave slightly worse results. However, I probably didn't train that enough as mixup trains much longer before starting to overfit.</p>",
          "rawMarkdown": "I tried stain normalization twice with two different reference images and I didn't see any improvements.  I also tried mixup but it gave slightly worse results. However, I probably didn't train that enough as mixup trains much longer before starting to overfit.",
          "votes": 1
        },
        {
          "id": 504468,
          "postDate": "2019-03-31T16:56:51.843Z",
          "content": "<p>I trained only one model of ensemble with randomly normalized stain images. The public score was only 0,9715 and private score 0,9723 but in the ensemble this model improved both private and public.</p>",
          "rawMarkdown": "I trained only one model of ensemble with randomly normalized stain images. The public score was only 0,9715 and private score 0,9723 but in the ensemble this model improved both private and public.",
          "votes": 1
        },
        {
          "id": 504483,
          "postDate": "2019-03-31T17:31:26.567Z",
          "content": "<p>Mixup usually slightly improved the lb score but did not make much difference for val auc, I found what worked best was training with mixup for around 10 epochs after training the unfrozen network, then removing mixup and continuing for another 5–10 epochs depending on the model. With densenets 5 epochs after mixup seemed sufficient most of the time. </p>",
          "rawMarkdown": "Mixup usually slightly improved the lb score but did not make much difference for val auc, I found what worked best was training with mixup for around 10 epochs after training the unfrozen network, then removing mixup and continuing for another 5–10 epochs depending on the model. With densenets 5 epochs after mixup seemed sufficient most of the time. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 504170,
      "postDate": "2019-03-31T04:23:33.067Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 504110,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-03-31T02:21:44.463000",
      "content": "<p>My best model on the private lb is the one that I set as public.... I never thought it was the best, 0.9744 on public lb and 0.9791 on private lb.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 504467,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-31T16:55:20.467000",
          "content": "<p>Crazy!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 504456,
      "author_name": "Dimitrij Shulkin",
      "author_url": "",
      "post_date": "2019-03-31T16:48:53.163000",
      "content": "<p>Great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 504103,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-03-31T02:00:07.547000",
      "content": "<p>Nice work. I also used ELU in my classifier. I never had much luck with se networks, seresnext in particular was so slow that I could not use it at all. My best models are all from densenet169 but senet154 and preresnet269b gave some good results, just 4-8 times longer than densenet and with lower scores. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 504127,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-03-31T03:05:43.813000",
          "content": "<p>I am using SEResNeXt50. With your <code>senet154 and preresnet269b</code>, did you experience a huge overfit? Since I discovered overfit within 8 epoch, I did not try deeper networks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 504205,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2019-03-31T06:30:24.503000",
      "content": "<p>My observations on this competition:\n1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.\n2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.\n3) More TTA helps.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 504358,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-03-31T13:23:10.990000",
          "content": "<blockquote>\n  <p>1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.</p>\n</blockquote>\n\n<p>Yes. I warn myself a lot of time \"not to trust LB\", all my tricks are tested against CV except the postprocess. I learned a harsh lesson now.</p>\n\n<blockquote>\n  <p>2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.</p>\n</blockquote>\n\n<p>This question may seem dumb: How can those images (96*96) fit into your network without upscaling?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 504464,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2019-03-31T16:52:40.657000",
          "content": "<p>The minimum resolution for most of image networks is 32x32 that will produce 1x1 output at the last convolution layer. For 96x96 input you will have 3x3 output with the center corresponding to the region  in the original image under detection (though central pooling worked worse than concat pooling). For common network resolution like 224x224 you just have 7x7 output. If you use adaptive pooling the network works with any of above-mentioned resolutions, and the head fully connected part remains the same.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 504179,
      "author_name": "Hanke Chen",
      "author_url": "",
      "post_date": "2019-03-31T04:51:34.327000",
      "content": "<p>Anybody got good results using stain augmentation / normalization?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 504390,
          "author_name": "Joni Juvonen",
          "author_url": "",
          "post_date": "2019-03-31T14:24:46.937000",
          "content": "<p>I tried stain normalization twice with two different reference images and I didn't see any improvements.  I also tried mixup but it gave slightly worse results. However, I probably didn't train that enough as mixup trains much longer before starting to overfit.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 504468,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-31T16:56:51.843000",
          "content": "<p>I trained only one model of ensemble with randomly normalized stain images. The public score was only 0,9715 and private score 0,9723 but in the ensemble this model improved both private and public.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 504483,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-03-31T17:31:26.567000",
          "content": "<p>Mixup usually slightly improved the lb score but did not make much difference for val auc, I found what worked best was training with mixup for around 10 epochs after training the unfrozen network, then removing mixup and continuing for another 5–10 epochs depending on the model. With densenets 5 epochs after mixup seemed sufficient most of the time. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 504170,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-31T04:23:33.067000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "504092": "# Discoveries\n1. Discover Duplicate: [Discover Duplicate](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367)  \n2. Discover Leak: [How to get LB 1.000](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/85424)  \n3. Invent WSI Normalization: [Better Way to Normalize Slides? WSI Normalization](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87069)  \n4. Full Solution Summary: [How I drop from 0.9805 Private LB to 0.974 (113rd Solution)](https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/87367)  \n\n# My LB 0.980 Solution\n\n## Single Model\n- Private: - 1. 0.9785\n- Public: 0.9753\n- Network: SEResNeXt50 (modified)\n  - Input: 128*128 without maxpooling\n  - BN, FC, ELU with 3 layers as a classifier\n  - cat maxpooling and avg pooling\n- Optimizer: SGD w/ momentum 0.9 (Adam is bad for me)\n- LR Scheduler: Triangle lr with Decrease on Plateau\n- Fold: 5 fold ensemble\n- TTA: 16\n- Validation Time Augmentation: 4\n- Augmentation: from Kernels + 20% center focus by padding\n- Dropout FC: 0.9\n- Dropout CNN: 0.2\n- Loss: BCE\n- Activation: ELU with KaimingInit\n- Training: Freeze 3 epoch, train to 6th epoch is the best\n- Tricks:\n  - Snapshot Ensemble (about 2 each fold)\n  - Shakeup simulation\n  - 8x scheduled Augmentation for faster converge\n  - Create a stable CV based on WSI\n  - Max and Min lr based on the Paper (1 cycle + gradient warmup)\n  - Faster loading with preprocessing images into .npy\n\n## Pseudo labeling (with all probability)\nI did not have time to train 5 fold pseudo\nBut here are the scores with 4 tta ensemble:\n(They are generally better +0.008 on both public and private)\nFold, Private, Public\nFold1: 0.9765, 0.9763\nFold2: 0.9787, 0.9668\n\n## Ensemble\nSimple weighted average ensemble with 1 fold NasNet, DenseNet...\nGives Private 0.9805, Public 0.9780\n\n# Ideas that did not work\nJigsaw puzzle (no unified patterns found)\nAverage, Geometric, Power ensemble are about the same\nPostprocess (setting white 32*32 images in the middle to negative labels)\n\n# Ideas not implemented\n- Whole slide image normalization (no enough time + not sure if it works)\n- Cheating with 1.000 LB (I don't think I am the first one who discovered this)\n- WSI info from test set (I pretend I did not discover this)\n- Distill with teacher and student\n- Mixup\n- Attention layers\n- Inceptionv3, v4, X\n- attention weighted average pooling\n- global max pooling\n- global std pooling\n- gradient aggregation\n- ImageNet pre-train (don't think it will be useful)\n\n# Self Reflection\nThe only difference between 0.980 and 0.974 is my postprocess: setting white 32*32 images in the middle to negative labels. It gives all my model a stable LB improvement 0.0001. The change is so small that I did not even think whether I should use the postprocess strategy or not.\n\nThe mistake comes from not understanding the data: the patches are cropped from WSI. That means even if the image is empty, it may be a huge white cell inside the tumor!\n\nThe other mistake is that I did not evaluate my postprocess through holdout set in CV. I trusted 1/5 LB too much. But looking back, I would still choose to use postprocess because the chance of getting an improvement of 0.0001 in Public LB and getting a huge drop 0.01 in Private LB is pretty low.\n\nSee you in the next cv competition!",
    "504110": "My best model on the private lb is the one that I set as public.... I never thought it was the best, 0.9744 on public lb and 0.9791 on private lb.",
    "504456": "Great work!",
    "504103": "Nice work. I also used ELU in my classifier. I never had much luck with se networks, seresnext in particular was so slow that I could not use it at all. My best models are all from densenet169 but senet154 and preresnet269b gave some good results, just 4-8 times longer than densenet and with lower scores. ",
    "504205": "My observations on this competition:\n1) The key to this competition was using WSI and CV. The model predictions are quite unstable, and even with 10-fold CV I expected quite high level of noise. Using public LB for validation under such conditions would be really risky idea. It is almost always better to trust CV rather than LB, but leaks should be checked first.\n2) Image upscaling really helps. I never used such trick before, but it seems to be really useful thing to do in competitions with small images. However, in this particular competition I preferred to upscale by an integer number to have the cancer detected image region multiple to 32.\n3) More TTA helps.",
    "504179": "Anybody got good results using stain augmentation / normalization?",
    "504170": ""
  }
}