{
  "id": 87400,
  "title": "24th place solution (0.9801 Private PB)",
  "url": "/competitions/histopathologic-cancer-detection/discussion/87400",
  "author_name": "Dimitrij Shulkin",
  "post_date": "2019-03-31T09:33:56.938000",
  "votes": 10,
  "comment_count": 21,
  "views": 0,
  "content": "<p>It was our first competition and we are very glad to reach 24th private (16th public). \nOur solution:\n1. Ensemble consisting of two DensNet201, Resnet50, Xception, InceptionV3 and VGG16\n2. One DenseNet201 was trained on randomly stain normalized images\n3. One Cycle Policy (SGD with cyclical learning rate and momentum over 6-7 epochs) \n3. For stain normalization images with to many white pixels were removed\n4. Images were resized to 224x224\n5. On each model was applied semi supervised technique such as pseudo labeling from the test set with different assumptions\n6. Random augmentation from imgaug package: flips, rotations, crops, saturation...\n7. 4 TTA\n8. At the end simple weighted average\n9. Only Keras\n10. Only Kaggle Kernels\n11. No external data</p>",
  "messages": [
    {
      "id": 504266,
      "postDate": "2019-03-31T09:33:56.940Z",
      "content": "<p>It was our first competition and we are very glad to reach 24th private (16th public). \nOur solution:\n1. Ensemble consisting of two DensNet201, Resnet50, Xception, InceptionV3 and VGG16\n2. One DenseNet201 was trained on randomly stain normalized images\n3. One Cycle Policy (SGD with cyclical learning rate and momentum over 6-7 epochs) \n3. For stain normalization images with to many white pixels were removed\n4. Images were resized to 224x224\n5. On each model was applied semi supervised technique such as pseudo labeling from the test set with different assumptions\n6. Random augmentation from imgaug package: flips, rotations, crops, saturation...\n7. 4 TTA\n8. At the end simple weighted average\n9. Only Keras\n10. Only Kaggle Kernels\n11. No external data</p>",
      "rawMarkdown": "It was our first competition and we are very glad to reach 24th private (16th public). \nOur solution:\n1. Ensemble consisting of two DensNet201, Resnet50, Xception, InceptionV3 and VGG16\n2. One DenseNet201 was trained on randomly stain normalized images\n3. One Cycle Policy (SGD with cyclical learning rate and momentum over 6-7 epochs) \n3. For stain normalization images with to many white pixels were removed\n4. Images were resized to 224x224\n5. On each model was applied semi supervised technique such as pseudo labeling from the test set with different assumptions\n6. Random augmentation from imgaug package: flips, rotations, crops, saturation...\n7. 4 TTA\n8. At the end simple weighted average\n9. Only Keras\n10. Only Kaggle Kernels\n11. No external data\n",
      "votes": 8
    },
    {
      "id": 504951,
      "postDate": "2019-04-01T11:14:15.347Z",
      "content": "<p>Nice job! Did you measure the significance of pseudo labeling? In other words, did you measure how much it helps? </p>",
      "rawMarkdown": "Nice job! Did you measure the significance of pseudo labeling? In other words, did you measure how much it helps? ",
      "votes": 1,
      "replies": [
        {
          "id": 505181,
          "postDate": "2019-04-01T16:51:22.673Z",
          "content": "<p>Thank you. After the final evaluation (privat score), I realized that this technique quickly leads to overfitting of public set. Thus the public score is high but the private score is low. So you have to be very careful!</p>",
          "rawMarkdown": "Thank you. After the final evaluation (privat score), I realized that this technique quickly leads to overfitting of public set. Thus the public score is high but the private score is low. So you have to be very careful!",
          "votes": 2
        },
        {
          "id": 505315,
          "postDate": "2019-04-01T21:13:30.467Z",
          "content": "<p>Pseudo labeling disrupted my original CV (due to the inability (read: I don't allow myself) to find WSI for test images). But based on the Public LB and Private LB, pseudo labeling helped a lot (both before and after tta). However, the same improvement did not reflect on my model ensembles. Maybe it is because I only had two pseudo fold vs. around 9 folds of other models. I guess if I have more folds of those, my best model can improve 0.0005 from 0.9805.</p>",
          "rawMarkdown": "Pseudo labeling disrupted my original CV (due to the inability (read: I don't allow myself) to find WSI for test images). But based on the Public LB and Private LB, pseudo labeling helped a lot (both before and after tta). However, the same improvement did not reflect on my model ensembles. Maybe it is because I only had two pseudo fold vs. around 9 folds of other models. I guess if I have more folds of those, my best model can improve 0.0005 from 0.9805.",
          "votes": 2
        }
      ]
    },
    {
      "id": 504529,
      "postDate": "2019-03-31T19:32:37.977Z",
      "content": "<p>Great job!</p>",
      "rawMarkdown": "Great job!",
      "votes": 1,
      "replies": [
        {
          "id": 504738,
          "postDate": "2019-04-01T05:16:32.503Z",
          "content": "<p>Thank you William!</p>",
          "rawMarkdown": "Thank you William!",
          "votes": 1
        }
      ]
    },
    {
      "id": 504309,
      "postDate": "2019-03-31T11:32:10.193Z",
      "content": "<p>The score of single best model performs on LB.</p>",
      "rawMarkdown": "The score of single best model performs on LB.",
      "votes": 1,
      "replies": [
        {
          "id": 504323,
          "postDate": "2019-03-31T11:56:25Z",
          "content": "<p>For example, with the best Densnet201 we got 0.980 private and 0.9824 public</p>",
          "rawMarkdown": "For example, with the best Densnet201 we got 0.980 private and 0.9824 public",
          "votes": 1
        },
        {
          "id": 504362,
          "postDate": "2019-03-31T13:29:57.873Z",
          "content": "<p>The score 0.980 on private is with <code>pseudo</code>, <code>tta</code>, and <code>one fold</code>, right? Is that one with <code>stain normalization</code> or without?</p>",
          "rawMarkdown": "The score 0.980 on private is with `pseudo`, `tta`, and `one fold`, right? Is that one with `stain normalization` or without?",
          "votes": 1
        },
        {
          "id": 504411,
          "postDate": "2019-03-31T15:18:05.177Z",
          "content": "<p>This was without one.</p>",
          "rawMarkdown": "This was without one.",
          "votes": 2
        }
      ]
    },
    {
      "id": 504268,
      "postDate": "2019-03-31T09:39:45.630Z",
      "content": "<p>Can you tell me the score of your best single model.</p>",
      "rawMarkdown": "Can you tell me the score of your best single model.",
      "votes": 1,
      "replies": [
        {
          "id": 504277,
          "postDate": "2019-03-31T10:06:22.423Z",
          "content": "<p>Do you mean the score of individual model of ensemble before averaging?</p>",
          "rawMarkdown": "Do you mean the score of individual model of ensemble before averaging?",
          "votes": 2
        },
        {
          "id": 504285,
          "postDate": "2019-03-31T10:20:00.177Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 505100,
      "postDate": "2019-04-01T15:04:48.970Z",
      "content": "<p>Great job!! I never try pseudo labeling, and in my experiments I used heavy image augmentation and then do heavy TTA not the standard TTA just use the same image augmentation with test data, then do M times TTA (every time change a seed).  I think it's like doing heavy ensemble in the testing time with same model (changing data instead of changing model) and the number of the combination of augmentation would be the limit of M.</p>",
      "rawMarkdown": "Great job!! I never try pseudo labeling, and in my experiments I used heavy image augmentation and then do heavy TTA not the standard TTA just use the same image augmentation with test data, then do M times TTA (every time change a seed).  I think it's like doing heavy ensemble in the testing time with same model (changing data instead of changing model) and the number of the combination of augmentation would be the limit of M.",
      "votes": 2,
      "replies": [
        {
          "id": 505182,
          "postDate": "2019-04-01T16:52:34.433Z",
          "content": "<p>Thank you! Do you use randomly TTA applying augmentations on test data? I didn't experiment with TTA. Just took 4 TTA (old school )))) )</p>",
          "rawMarkdown": "Thank you! Do you use randomly TTA applying augmentations on test data? I didn't experiment with TTA. Just took 4 TTA (old school )))) )",
          "votes": 2
        },
        {
          "id": 505281,
          "postDate": "2019-04-01T20:11:15.763Z",
          "content": "<p>I have done the same random transformation as we do in the training time, like this:\n<code>data_transforms = albumentations.Compose([\n    albumentations.Resize(224, 224),\n    albumentations.RandomRotate90(p=0.5),\n    albumentations.Transpose(p=0.5),\n    albumentations.Flip(p=0.5),\n    albumentations.OneOf([\n        albumentations.CLAHE(clip_limit=2), albumentations.IAASharpen(), albumentations.IAAEmboss(), \n        albumentations.RandomBrightness(), albumentations.RandomContrast(),\n        albumentations.JpegCompression(), albumentations.Blur(), albumentations.GaussNoise()], p=0.5), \n    albumentations.HueSaturationValue(p=0.5), \n    albumentations.ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n    albumentations.Normalize(),\n    AT.ToTensor()\n    ])</code>\nIn my experiments like my public kernel <a href=\"https://www.kaggle.com/jionie/tta-power-densenet169\">https://www.kaggle.com/jionie/tta-power-densenet169</a>, when you increase num_tta, the performance increases. When we do ensemble, we use same data and different models to get different predictions with high diversity, and here we use different data and same model to get different predictions (I don't check whether they have high diversity or not, but I think they have, lol).</p>",
          "rawMarkdown": "I have done the same random transformation as we do in the training time, like this:\n`data_transforms = albumentations.Compose([\n    albumentations.Resize(224, 224),\n    albumentations.RandomRotate90(p=0.5),\n    albumentations.Transpose(p=0.5),\n    albumentations.Flip(p=0.5),\n    albumentations.OneOf([\n        albumentations.CLAHE(clip_limit=2), albumentations.IAASharpen(), albumentations.IAAEmboss(), \n        albumentations.RandomBrightness(), albumentations.RandomContrast(),\n        albumentations.JpegCompression(), albumentations.Blur(), albumentations.GaussNoise()], p=0.5), \n    albumentations.HueSaturationValue(p=0.5), \n    albumentations.ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n    albumentations.Normalize(),\n    AT.ToTensor()\n    ])`\nIn my experiments like my public kernel https://www.kaggle.com/jionie/tta-power-densenet169, when you increase num_tta, the performance increases. When we do ensemble, we use same data and different models to get different predictions with high diversity, and here we use different data and same model to get different predictions (I don't check whether they have high diversity or not, but I think they have, lol).\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 512233,
      "postDate": "2019-04-10T16:34:49.253Z",
      "content": "<p>Hi Dima\nDid you try k-fold CV with any of the model?</p>",
      "rawMarkdown": "Hi Dima\nDid you try k-fold CV with any of the model?",
      "replies": [
        {
          "id": 512324,
          "postDate": "2019-04-10T17:28:22.923Z",
          "content": "<p>No, only random splits.</p>",
          "rawMarkdown": "No, only random splits.",
          "votes": 2
        },
        {
          "id": 512392,
          "postDate": "2019-04-10T18:16:46.147Z",
          "content": "<p>I wanted to make k-fold with 10 folds but decided for random Jocker )))))))</p>",
          "rawMarkdown": "I wanted to make k-fold with 10 folds but decided for random Jocker )))))))",
          "votes": 2
        }
      ]
    },
    {
      "id": 510067,
      "postDate": "2019-04-08T15:49:38.127Z",
      "content": "<p>How much images did u considered in each subset??\nWere the same images repeated again in another subset?</p>",
      "rawMarkdown": "How much images did u considered in each subset??\nWere the same images repeated again in another subset?",
      "replies": [
        {
          "id": 510544,
          "postDate": "2019-04-09T07:47:49.837Z",
          "content": "<p>Hello Aarya,\nsorry, it was a mistake. I made multiple training runs with the same train set - always with random train and validation sets.</p>",
          "rawMarkdown": "Hello Aarya,\nsorry, it was a mistake. I made multiple training runs with the same train set - always with random train and validation sets.",
          "votes": 2
        }
      ]
    },
    {
      "id": 504363,
      "postDate": "2019-03-31T13:32:02.423Z",
      "content": "<p>Cool, thank you! </p>",
      "rawMarkdown": "Cool, thank you! ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 504951,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2019-04-01T11:14:15.347000",
      "content": "<p>Nice job! Did you measure the significance of pseudo labeling? In other words, did you measure how much it helps? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 505181,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-01T16:51:22.673000",
          "content": "<p>Thank you. After the final evaluation (privat score), I realized that this technique quickly leads to overfitting of public set. Thus the public score is high but the private score is low. So you have to be very careful!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 505315,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-04-01T21:13:30.467000",
          "content": "<p>Pseudo labeling disrupted my original CV (due to the inability (read: I don't allow myself) to find WSI for test images). But based on the Public LB and Private LB, pseudo labeling helped a lot (both before and after tta). However, the same improvement did not reflect on my model ensembles. Maybe it is because I only had two pseudo fold vs. around 9 folds of other models. I guess if I have more folds of those, my best model can improve 0.0005 from 0.9805.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 504529,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-03-31T19:32:37.977000",
      "content": "<p>Great job!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 504738,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-01T05:16:32.503000",
          "content": "<p>Thank you William!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 504309,
      "author_name": "Harleys Zhang",
      "author_url": "",
      "post_date": "2019-03-31T11:32:10.193000",
      "content": "<p>The score of single best model performs on LB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 504323,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-31T11:56:25",
          "content": "<p>For example, with the best Densnet201 we got 0.980 private and 0.9824 public</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 504362,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2019-03-31T13:29:57.873000",
          "content": "<p>The score 0.980 on private is with <code>pseudo</code>, <code>tta</code>, and <code>one fold</code>, right? Is that one with <code>stain normalization</code> or without?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 504411,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-31T15:18:05.177000",
          "content": "<p>This was without one.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 504268,
      "author_name": "Harleys Zhang",
      "author_url": "",
      "post_date": "2019-03-31T09:39:45.630000",
      "content": "<p>Can you tell me the score of your best single model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 504277,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-31T10:06:22.423000",
          "content": "<p>Do you mean the score of individual model of ensemble before averaging?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 504285,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-03-31T10:20:00.177000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 505100,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-04-01T15:04:48.970000",
      "content": "<p>Great job!! I never try pseudo labeling, and in my experiments I used heavy image augmentation and then do heavy TTA not the standard TTA just use the same image augmentation with test data, then do M times TTA (every time change a seed).  I think it's like doing heavy ensemble in the testing time with same model (changing data instead of changing model) and the number of the combination of augmentation would be the limit of M.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 505182,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-01T16:52:34.433000",
          "content": "<p>Thank you! Do you use randomly TTA applying augmentations on test data? I didn't experiment with TTA. Just took 4 TTA (old school )))) )</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 505281,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-04-01T20:11:15.763000",
          "content": "<p>I have done the same random transformation as we do in the training time, like this:\n<code>data_transforms = albumentations.Compose([\n    albumentations.Resize(224, 224),\n    albumentations.RandomRotate90(p=0.5),\n    albumentations.Transpose(p=0.5),\n    albumentations.Flip(p=0.5),\n    albumentations.OneOf([\n        albumentations.CLAHE(clip_limit=2), albumentations.IAASharpen(), albumentations.IAAEmboss(), \n        albumentations.RandomBrightness(), albumentations.RandomContrast(),\n        albumentations.JpegCompression(), albumentations.Blur(), albumentations.GaussNoise()], p=0.5), \n    albumentations.HueSaturationValue(p=0.5), \n    albumentations.ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n    albumentations.Normalize(),\n    AT.ToTensor()\n    ])</code>\nIn my experiments like my public kernel <a href=\"https://www.kaggle.com/jionie/tta-power-densenet169\">https://www.kaggle.com/jionie/tta-power-densenet169</a>, when you increase num_tta, the performance increases. When we do ensemble, we use same data and different models to get different predictions with high diversity, and here we use different data and same model to get different predictions (I don't check whether they have high diversity or not, but I think they have, lol).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 512233,
      "author_name": "Aarya Patel",
      "author_url": "",
      "post_date": "2019-04-10T16:34:49.253000",
      "content": "<p>Hi Dima\nDid you try k-fold CV with any of the model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 512324,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-10T17:28:22.923000",
          "content": "<p>No, only random splits.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 512392,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-10T18:16:46.147000",
          "content": "<p>I wanted to make k-fold with 10 folds but decided for random Jocker )))))))</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 510067,
      "author_name": "Aarya Patel",
      "author_url": "",
      "post_date": "2019-04-08T15:49:38.127000",
      "content": "<p>How much images did u considered in each subset??\nWere the same images repeated again in another subset?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 510544,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-04-09T07:47:49.837000",
          "content": "<p>Hello Aarya,\nsorry, it was a mistake. I made multiple training runs with the same train set - always with random train and validation sets.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 504363,
      "author_name": "neuralartist",
      "author_url": "",
      "post_date": "2019-03-31T13:32:02.423000",
      "content": "<p>Cool, thank you! </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "504266": "It was our first competition and we are very glad to reach 24th private (16th public). \nOur solution:\n1. Ensemble consisting of two DensNet201, Resnet50, Xception, InceptionV3 and VGG16\n2. One DenseNet201 was trained on randomly stain normalized images\n3. One Cycle Policy (SGD with cyclical learning rate and momentum over 6-7 epochs) \n3. For stain normalization images with to many white pixels were removed\n4. Images were resized to 224x224\n5. On each model was applied semi supervised technique such as pseudo labeling from the test set with different assumptions\n6. Random augmentation from imgaug package: flips, rotations, crops, saturation...\n7. 4 TTA\n8. At the end simple weighted average\n9. Only Keras\n10. Only Kaggle Kernels\n11. No external data\n",
    "504951": "Nice job! Did you measure the significance of pseudo labeling? In other words, did you measure how much it helps? ",
    "504529": "Great job!",
    "504309": "The score of single best model performs on LB.",
    "504268": "Can you tell me the score of your best single model.",
    "505100": "Great job!! I never try pseudo labeling, and in my experiments I used heavy image augmentation and then do heavy TTA not the standard TTA just use the same image augmentation with test data, then do M times TTA (every time change a seed).  I think it's like doing heavy ensemble in the testing time with same model (changing data instead of changing model) and the number of the combination of augmentation would be the limit of M.",
    "512233": "Hi Dima\nDid you try k-fold CV with any of the model?",
    "510067": "How much images did u considered in each subset??\nWere the same images repeated again in another subset?",
    "504363": "Cool, thank you! "
  }
}