{
  "id": 83760,
  "title": "How to get 0.98+ AUC",
  "url": "/competitions/histopathologic-cancer-detection/discussion/83760",
  "author_name": "SM",
  "post_date": "2019-03-12T17:56:09.291000",
  "votes": 52,
  "comment_count": 58,
  "views": 0,
  "content": "<p>main issues:\n- validation, needs to be done by isolating WSIs, not random patches\n- overfitting to WSIs in train</p>\n\n<p>solutions:\n- I extracted slide ids for almost every patch, here they <a href=\"https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC\">are</a>\n- intense augmentations, dropout 3 x 90%, more Linear layers</p>\n\n<p>training:\n- image size - rescaled to model required input size \n- no cleaning\n- transformers: \n<code>\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n</code>\n- custom tail: avg and max pooling concat, multiple Linear layers with dropout and  batchnorm\n-  training:\n   - freeze all except custom tail\n   - lr * 100\n   - unfreeze all\n   - lr / 100\n   - cyclic lr (triangle)\n   - save model when score improved\n   - drop lr /2 if AUC on valid did not improve for 2-5 cycles (1 cycle is xxx iterations based on arch and bach size) and load best model\n   - lr restart *(10-100) after 4 lr drops\n- TTa predictions on all flips+rotations transforms mentioned above\n- single best model 0.975</p>\n\n<p>UPDT: forgot to mention, I use FocalLoss</p>\n\n<p><a href=\"https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9\">starter kernel</a> </p>",
  "messages": [
    {
      "id": 488582,
      "postDate": "2019-03-12T17:56:09.293Z",
      "content": "<p>main issues:\n- validation, needs to be done by isolating WSIs, not random patches\n- overfitting to WSIs in train</p>\n\n<p>solutions:\n- I extracted slide ids for almost every patch, here they <a href=\"https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC\">are</a>\n- intense augmentations, dropout 3 x 90%, more Linear layers</p>\n\n<p>training:\n- image size - rescaled to model required input size \n- no cleaning\n- transformers: \n<code>\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n</code>\n- custom tail: avg and max pooling concat, multiple Linear layers with dropout and  batchnorm\n-  training:\n   - freeze all except custom tail\n   - lr * 100\n   - unfreeze all\n   - lr / 100\n   - cyclic lr (triangle)\n   - save model when score improved\n   - drop lr /2 if AUC on valid did not improve for 2-5 cycles (1 cycle is xxx iterations based on arch and bach size) and load best model\n   - lr restart *(10-100) after 4 lr drops\n- TTa predictions on all flips+rotations transforms mentioned above\n- single best model 0.975</p>\n\n<p>UPDT: forgot to mention, I use FocalLoss</p>\n\n<p><a href=\"https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9\">starter kernel</a> </p>",
      "rawMarkdown": "main issues:\n- validation, needs to be done by isolating WSIs, not random patches\n- overfitting to WSIs in train\n\nsolutions:\n- I extracted slide ids for almost every patch, here they [are](https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC)\n- intense augmentations, dropout 3 x 90%, more Linear layers\n\ntraining:\n- image size - rescaled to model required input size \n- no cleaning\n- transformers: \n```\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n```\n- custom tail: avg and max pooling concat, multiple Linear layers with dropout and  batchnorm\n-  training:\n   - freeze all except custom tail\n   - lr * 100\n   - unfreeze all\n   - lr / 100\n   - cyclic lr (triangle)\n   - save model when score improved\n   - drop lr /2 if AUC on valid did not improve for 2-5 cycles (1 cycle is xxx iterations based on arch and bach size) and load best model\n   - lr restart *(10-100) after 4 lr drops\n- TTa predictions on all flips+rotations transforms mentioned above\n- single best model 0.975\n\nUPDT: forgot to mention, I use FocalLoss\n\n[starter kernel](https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9) ",
      "votes": 52
    },
    {
      "id": 1762477,
      "postDate": "2022-04-20T17:59:03.790Z",
      "content": "<p>I am late to the party, but this is very useful discussion. Thanks everyone!</p>",
      "rawMarkdown": "I am late to the party, but this is very useful discussion. Thanks everyone!",
      "votes": 2
    },
    {
      "id": 492571,
      "postDate": "2019-03-17T13:17:16.993Z",
      "content": "<p>Why resize training image to model required input size, rather than using average pooling to replace pooling layer? Will there be any difference in score? Why?</p>",
      "rawMarkdown": "Why resize training image to model required input size, rather than using average pooling to replace pooling layer? Will there be any difference in score? Why?",
      "votes": 3,
      "replies": [
        {
          "id": 493167,
          "postDate": "2019-03-18T11:10:37.090Z",
          "content": "<p><a href=\"/sermakarevich\">@sermakarevich</a> </p>",
          "rawMarkdown": "@sermakarevich ",
          "votes": 3
        },
        {
          "id": 493180,
          "postDate": "2019-03-18T11:38:42.210Z",
          "content": "<p>no good reason</p>",
          "rawMarkdown": "no good reason"
        }
      ]
    },
    {
      "id": 488888,
      "postDate": "2019-03-13T06:36:34.210Z",
      "content": "<p><a href=\"/ivanpan\">@ivanpan</a> \nTTA stands for test-time-augmentation. I used one unique transformer at time to extract slightly different predictions for the same image in test set. It was already described on the forum. </p>\n\n<p>@Roshan Santhosh\nI use them to create proper train/valid split without leakage. </p>\n\n<p>@Dima Shulkin \n<code>\n  (1): Sequential(\n    (0): AdaptiveConcatPool2d(\n      (average_pool): AdaptiveAvgPool2d(output_size=1)\n      (max_pool): AdaptiveMaxPool2d(output_size=1)\n    )\n    (1): Flatten()\n    (2): BatchNorm1d(3072, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (3): Dropout(p=0.8)\n    (4): Linear(in_features=3072, out_features=512, bias=True)\n    (5): ReLU(inplace)\n    (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (7): Dropout(p=0.8)\n    (8): Linear(in_features=512, out_features=256, bias=True)\n    (9): ReLU(inplace)\n    (10): BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (11): Dropout(p=0.8)\n    (12): Linear(in_features=256, out_features=1, bias=True)\n  )\n)\n</code></p>",
      "rawMarkdown": "@ivanpan \nTTA stands for test-time-augmentation. I used one unique transformer at time to extract slightly different predictions for the same image in test set. It was already described on the forum. \n\n@Roshan Santhosh\nI use them to create proper train/valid split without leakage. \n\n@Dima Shulkin \n```\n  (1): Sequential(\n    (0): AdaptiveConcatPool2d(\n      (average_pool): AdaptiveAvgPool2d(output_size=1)\n      (max_pool): AdaptiveMaxPool2d(output_size=1)\n    )\n    (1): Flatten()\n    (2): BatchNorm1d(3072, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (3): Dropout(p=0.8)\n    (4): Linear(in_features=3072, out_features=512, bias=True)\n    (5): ReLU(inplace)\n    (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (7): Dropout(p=0.8)\n    (8): Linear(in_features=512, out_features=256, bias=True)\n    (9): ReLU(inplace)\n    (10): BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (11): Dropout(p=0.8)\n    (12): Linear(in_features=256, out_features=1, bias=True)\n  )\n)\n```\n",
      "votes": 3
    },
    {
      "id": 488694,
      "postDate": "2019-03-12T22:14:50.713Z",
      "content": "<p>How exactly are you using the WSI IDs?</p>",
      "rawMarkdown": "How exactly are you using the WSI IDs?",
      "votes": 4
    },
    {
      "id": 490289,
      "postDate": "2019-03-14T13:08:58.370Z",
      "content": "<p>Weird thing: started to check WSI ids and found something strange. For instance, there is a WSI called camelyon16_train_tumor_017. And it contains 544 patches. The thing is: non of them are labeled as \"tumor\" in the competition dataset. I mean not a single one. </p>",
      "rawMarkdown": "Weird thing: started to check WSI ids and found something strange. For instance, there is a WSI called camelyon16_train_tumor_017. And it contains 544 patches. The thing is: non of them are labeled as \"tumor\" in the competition dataset. I mean not a single one. ",
      "votes": 1
    },
    {
      "id": 489142,
      "postDate": "2019-03-13T14:48:54.697Z",
      "content": "<p>Thanks for sharing WHOLE SLIDE IMAGE (WSI): A digitized histopathology glass slide that has been created on a slide scanner.  The digitized glass slide represents a high-resolution replica of the original glass that can then be manipulated through software to mimic microscope review and diagnosis. </p>",
      "rawMarkdown": "Thanks for sharing WHOLE SLIDE IMAGE (WSI): A digitized histopathology glass slide that has been created on a slide scanner.  The digitized glass slide represents a high-resolution replica of the original glass that can then be manipulated through software to mimic microscope review and diagnosis. ",
      "votes": 1
    },
    {
      "id": 489001,
      "postDate": "2019-03-13T11:06:43.937Z",
      "content": "<p>Thank you for sharing, Can you explain the mean of WSI.</p>",
      "rawMarkdown": "Thank you for sharing, Can you explain the mean of WSI.",
      "votes": 1,
      "replies": [
        {
          "id": 489003,
          "postDate": "2019-03-13T11:08:18.340Z",
          "content": "<p>Whole Slide Image</p>",
          "rawMarkdown": "Whole Slide Image",
          "votes": 1
        }
      ]
    },
    {
      "id": 488899,
      "postDate": "2019-03-13T06:53:16.743Z",
      "content": "<p>Thank you for the slide IDs!</p>\n\n<p>So in order to remove inter-WSI bias, we should partition train/validation and CV folds by WSIs. Splitting just randomly leads to too optimistic validation set estimates.</p>\n\n<p>WSI stats:\n- 27,273 training samples with unknown WSI\n- Minimum samples per WSI is 2 <code>wsi010</code>\n- Maximum samples per WSI is 3,363 <code>wsi094</code></p>\n\n<p>Did you also stratify by labels when splitting? I'm not sure how to manage that...</p>",
      "rawMarkdown": "Thank you for the slide IDs!\n\nSo in order to remove inter-WSI bias, we should partition train/validation and CV folds by WSIs. Splitting just randomly leads to too optimistic validation set estimates.\n\nWSI stats:\n- 27,273 training samples with unknown WSI\n- Minimum samples per WSI is 2 `wsi010`\n- Maximum samples per WSI is 3,363 `wsi094`\n\nDid you also stratify by labels when splitting? I'm not sure how to manage that...",
      "votes": 1,
      "replies": [
        {
          "id": 2845734,
          "postDate": "2024-05-30T16:53:52.287Z",
          "content": "<p>Hi, I extracted slide IDs for almost every patch, but I can't open the links now. Could you please resend them to me? Alternatively, you can send them to my email at ruigangge@gmail.com. Thank you!</p>",
          "rawMarkdown": "Hi, I extracted slide IDs for almost every patch, but I can't open the links now. Could you please resend them to me? Alternatively, you can send them to my email at ruigangge@gmail.com. Thank you!"
        }
      ]
    },
    {
      "id": 488877,
      "postDate": "2019-03-13T06:14:06.990Z",
      "content": "<p>Hello SM,\nthank you for your contribution! \nWhat do you mean saying \"multiple linear layers\"? Linear activation functions? \nBest\nDima</p>",
      "rawMarkdown": "Hello SM,\nthank you for your contribution! \nWhat do you mean saying \"multiple linear layers\"? Linear activation functions? \nBest\nDima",
      "votes": 1,
      "replies": [
        {
          "id": 488881,
          "postDate": "2019-03-13T06:21:50.820Z",
          "content": "<p>I think it means multiple fully connected layers.</p>",
          "rawMarkdown": "I think it means multiple fully connected layers.",
          "votes": 1
        }
      ]
    },
    {
      "id": 488596,
      "postDate": "2019-03-12T18:17:12.743Z",
      "content": "<p>What do you mean by TTA predictions on all flips+rotations mentioned above? You mean we'll have 3 images (original, and 2 for each RandomChoice) or you mean 16 (original + 15 for every mentioned transformation)?</p>",
      "rawMarkdown": "What do you mean by TTA predictions on all flips+rotations mentioned above? You mean we'll have 3 images (original, and 2 for each RandomChoice) or you mean 16 (original + 15 for every mentioned transformation)?",
      "votes": 1,
      "replies": [
        {
          "id": 488690,
          "postDate": "2019-03-12T21:52:41.973Z",
          "content": "<p>Great results, by the way. Thanks for sharing your approach. </p>",
          "rawMarkdown": "Great results, by the way. Thanks for sharing your approach. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 493049,
      "postDate": "2019-03-18T06:57:30.760Z",
      "content": "<p>Hello everyone,\ndid anybody perform stain normalization (<a href=\"https://arxiv.org/abs/1811.03815\">https://arxiv.org/abs/1811.03815</a>) ?</p>",
      "rawMarkdown": "Hello everyone,\ndid anybody perform stain normalization (https://arxiv.org/abs/1811.03815) ?",
      "votes": 2,
      "replies": [
        {
          "id": 493065,
          "postDate": "2019-03-18T07:30:23.293Z",
          "content": "<p>Hi,\nI was thinking about that not with with the approach of this paper but rather with the conventional methods. However, I could not figure it out how to choose a proper reference image. </p>",
          "rawMarkdown": "Hi,\nI was thinking about that not with with the approach of this paper but rather with the conventional methods. However, I could not figure it out how to choose a proper reference image. ",
          "votes": 1
        },
        {
          "id": 493168,
          "postDate": "2019-03-18T11:12:36.723Z",
          "content": "<p>That was also my biggest problem. I couldn't find any good approaches anywhere. I tried to select the reference image randomly. However, I also came across the other problem. The images in the train set and test set with too many white pixels could not be transformed.</p>",
          "rawMarkdown": "That was also my biggest problem. I couldn't find any good approaches anywhere. I tried to select the reference image randomly. However, I also came across the other problem. The images in the train set and test set with too many white pixels could not be transformed.",
          "votes": 2
        },
        {
          "id": 493338,
          "postDate": "2019-03-18T14:56:25.243Z",
          "content": "<p>That's true. I had the same problem with the whitish images.\nAnother solution could be using the WSI stain normalization which was used by the Camelyon 2016 challenge winner. However, as we do not have the whole slide information for this dataset, it is not possible to use that approach. </p>",
          "rawMarkdown": "That's true. I had the same problem with the whitish images.\nAnother solution could be using the WSI stain normalization which was used by the Camelyon 2016 challenge winner. However, as we do not have the whole slide information for this dataset, it is not possible to use that approach. ",
          "votes": 1
        },
        {
          "id": 493613,
          "postDate": "2019-03-18T21:42:29.733Z",
          "content": "<p>Great ideas! And thanks for bringing this to my attention. </p>\n\n<p>Your kernel on the subject is very good, it’s good to show explicitly how the transforms work. I have started to experiment with stain normalization methods but it’s too early to report anything convincing. </p>\n\n<p>For anyone who wants to try this, aside from the kernel mentioned there is a nice python package called staintools that is quite easy to use (unless you are on windows). </p>\n\n<p>I agree that deciding on a reference image to normalize the rest to is not trivial or obvious. Trial and error here should be quite a bit faster than trying to train an infogan though, but that I have no idea how to do. </p>",
          "rawMarkdown": "Great ideas! And thanks for bringing this to my attention. \n\nYour kernel on the subject is very good, it’s good to show explicitly how the transforms work. I have started to experiment with stain normalization methods but it’s too early to report anything convincing. \n\nFor anyone who wants to try this, aside from the kernel mentioned there is a nice python package called staintools that is quite easy to use (unless you are on windows). \n\nI agree that deciding on a reference image to normalize the rest to is not trivial or obvious. Trial and error here should be quite a bit faster than trying to train an infogan though, but that I have no idea how to do. ",
          "votes": 1
        },
        {
          "id": 493987,
          "postDate": "2019-03-19T10:37:11.750Z",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> Which kernel?</p>",
          "rawMarkdown": "@interneuron Which kernel?"
        },
        {
          "id": 493992,
          "postDate": "2019-03-19T10:41:20.487Z",
          "content": "<p>this one:\n<a href=\"https://www.kaggle.com/robotdreams/stain-normalization\">https://www.kaggle.com/robotdreams/stain-normalization</a></p>",
          "rawMarkdown": "this one:\nhttps://www.kaggle.com/robotdreams/stain-normalization",
          "votes": 2
        },
        {
          "id": 494197,
          "postDate": "2019-03-19T15:11:27.353Z",
          "content": "<p>Yep, that’s the one. Dima’s Kernel. Thanks masih. </p>",
          "rawMarkdown": "Yep, that’s the one. Dima’s Kernel. Thanks masih. "
        }
      ]
    },
    {
      "id": 489956,
      "postDate": "2019-03-14T07:37:15.167Z",
      "content": "<p><a href=\"/keremt\">@keremt</a> It is just matching of patches in the competition with original patches by mean pixel value. If there is only one patch in original dataset with the mean value, I used its WSI id, if multiple - I skipped it. </p>\n\n<p><a href=\"/tanlikesmath\">@tanlikesmath</a> No I did not. Resnet, densenet and some models from here <code>https://github.com/Cadene/pretrained-models.pytorch</code></p>\n\n<p><a href=\"/qitvision\">@qitvision</a> No I did not. I did not do stratified split by labels as I thought there are too many patches and random split should be good enough. Was thinking about stratified split by rounded pixel means but did not do that either.  </p>",
      "rawMarkdown": "@keremt It is just matching of patches in the competition with original patches by mean pixel value. If there is only one patch in original dataset with the mean value, I used its WSI id, if multiple - I skipped it. \n\n@tanlikesmath No I did not. Resnet, densenet and some models from here `https://github.com/Cadene/pretrained-models.pytorch`\n\n@qitvision No I did not. I did not do stratified split by labels as I thought there are too many patches and random split should be good enough. Was thinking about stratified split by rounded pixel means but did not do that either.  ",
      "votes": 2,
      "replies": [
        {
          "id": 2845742,
          "postDate": "2024-05-30T16:56:05.533Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 488663,
      "postDate": "2019-03-12T20:52:24.033Z",
      "content": "<p>Thank you! I think that I figured out that WSI file.  If I'm not mistaken, we have to replace \"camelyon16_train_tumor_066\"  with 0. </p>",
      "rawMarkdown": "Thank you! I think that I figured out that WSI file.  If I'm not mistaken, we have to replace \"camelyon16_train_tumor_066\"  with 0. ",
      "votes": 2
    },
    {
      "id": 488658,
      "postDate": "2019-03-12T20:39:36.907Z",
      "content": "<p><a href=\"/sermakarevich\">@sermakarevich</a> , thank you for sharing WSIs IDs, I'm curious how did u get them? Another question is if u tried to use all 600 GB dataset including both camelyon16 and camelyon17 for training? Though I think it isn't worth doing it for a playground competition, the result may be much better.</p>",
      "rawMarkdown": "@sermakarevich , thank you for sharing WSIs IDs, I'm curious how did u get them? Another question is if u tried to use all 600 GB dataset including both camelyon16 and camelyon17 for training? Though I think it isn't worth doing it for a playground competition, the result may be much better.",
      "votes": 2,
      "replies": [
        {
          "id": 488678,
          "postDate": "2019-03-12T21:15:38.497Z",
          "content": "<p>It`s just a playground. I was just testing a pipeline that we work on at <a href=\"https://www.lilystyle.ai/\">https://www.lilystyle.ai/</a> </p>",
          "rawMarkdown": "It`s just a playground. I was just testing a pipeline that we work on at https://www.lilystyle.ai/ ",
          "votes": 1
        },
        {
          "id": 488946,
          "postDate": "2019-03-13T09:03:33.547Z",
          "content": "<p>Very cool! I am also curious about how to get the WSI IDs :)</p>",
          "rawMarkdown": "Very cool! I am also curious about how to get the WSI IDs :)"
        }
      ]
    },
    {
      "id": 2845761,
      "postDate": "2024-05-30T17:03:55.170Z",
      "content": "<p>I hope this message finds you well. I recently came across your forum post where you shared a link for the extracted slide IDs for almost every patch. Unfortunately, the link seems to have expired. Could you please resend the information or update the link directly to my email at ruigangge@gmail.com? It would be greatly appreciated.</p>\n<p>Additionally, I would like to inquire about the slide IDs. Are all slide IDs provided officially unique, and could you please share some insights into how these IDs are obtained?</p>\n<p>Thank you for your assistance.</p>",
      "rawMarkdown": "I hope this message finds you well. I recently came across your forum post where you shared a link for the extracted slide IDs for almost every patch. Unfortunately, the link seems to have expired. Could you please resend the information or update the link directly to my email at ruigangge@gmail.com? It would be greatly appreciated.\n\nAdditionally, I would like to inquire about the slide IDs. Are all slide IDs provided officially unique, and could you please share some insights into how these IDs are obtained?\n\nThank you for your assistance."
    },
    {
      "id": 496363,
      "postDate": "2019-03-22T04:35:42.223Z",
      "content": "<p>Hi SM, many thanks for the awesome tips!!\nCould you share the code or explanation on how you extracted the slide IDs? </p>",
      "rawMarkdown": "Hi SM, many thanks for the awesome tips!!\nCould you share the code or explanation on how you extracted the slide IDs? ",
      "replies": [
        {
          "id": 499057,
          "postDate": "2019-03-24T09:11:42.393Z",
          "content": "<p>unfortunately I cant as thats a lib we write for prod at <code>www.lilystyle.ai</code>\nhowever some sketches of it you can find <a href=\"https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9\">here</a></p>",
          "rawMarkdown": "unfortunately I cant as thats a lib we write for prod at `www.lilystyle.ai`\nhowever some sketches of it you can find [here](https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9)"
        }
      ]
    },
    {
      "id": 494938,
      "postDate": "2019-03-20T12:59:25.813Z",
      "content": "<p>What is a WSI??</p>",
      "rawMarkdown": "What is a WSI??"
    },
    {
      "id": 492342,
      "postDate": "2019-03-17T05:30:10.887Z",
      "content": "<p>dataset isn't extremely imbalanced, but focal loss still provided a big advantage?</p>",
      "rawMarkdown": " dataset isn't extremely imbalanced, but focal loss still provided a big advantage?",
      "replies": [
        {
          "id": 492359,
          "postDate": "2019-03-17T06:07:12.533Z",
          "content": "<p>focal loss has not been any better than bce for me yet. </p>",
          "rawMarkdown": "focal loss has not been any better than bce for me yet. "
        },
        {
          "id": 492421,
          "postDate": "2019-03-17T09:06:03.030Z",
          "content": "<p>Same for me. I didn't see any improvement to BCE and I tried focal loss with <code>alpha=0.25, gamma=2</code>.</p>",
          "rawMarkdown": "Same for me. I didn't see any improvement to BCE and I tried focal loss with `alpha=0.25, gamma=2`."
        }
      ]
    },
    {
      "id": 491298,
      "postDate": "2019-03-15T12:48:09.380Z",
      "content": "<p>Hi SM,\nThank you! I want to know you how to divide dataset through wsi_ids? If the validation dataset accounts for 20%，the train dataset contains all wsi_ids and takes 80% of them, or 80% of wsi_ids and takes all of them?</p>",
      "rawMarkdown": "Hi SM,\nThank you! I want to know you how to divide dataset through wsi_ids? If the validation dataset accounts for 20%，the train dataset contains all wsi_ids and takes 80% of them, or 80% of wsi_ids and takes all of them?"
    },
    {
      "id": 491263,
      "postDate": "2019-03-15T12:01:56.367Z",
      "content": "<p>Hi SM,\nMy single model is only 0.9711,\nI want to know which network model you are using.</p>",
      "rawMarkdown": "Hi SM,\nMy single model is only 0.9711,\nI want to know which network model you are using.",
      "replies": [
        {
          "id": 492007,
          "postDate": "2019-03-16T14:04:01.400Z",
          "content": "<p>Do you use TTA? TTA could help my models boost from around 0.97 to 0.974.</p>",
          "rawMarkdown": "Do you use TTA? TTA could help my models boost from around 0.97 to 0.974.",
          "votes": 2
        }
      ]
    },
    {
      "id": 490513,
      "postDate": "2019-03-14T16:20:35.853Z",
      "content": "<p>Is anyone willing to share their pipeline for WSI in a kernel? </p>",
      "rawMarkdown": "Is anyone willing to share their pipeline for WSI in a kernel? ",
      "replies": [
        {
          "id": 490565,
          "postDate": "2019-03-14T17:08:42.393Z",
          "content": "<p>I'm working on it. Maybe tomorrow or the day after that, when I finish. </p>",
          "rawMarkdown": "I'm working on it. Maybe tomorrow or the day after that, when I finish. ",
          "votes": 1
        },
        {
          "id": 490608,
          "postDate": "2019-03-14T17:49:28.490Z",
          "content": "<p>Have a look at mine:\n<a href=\"https://www.kaggle.com/guntherthepenguin/fastai-v1-densenet169-with-wsi\">https://www.kaggle.com/guntherthepenguin/fastai-v1-densenet169-with-wsi</a></p>\n\n<p>Its still pretty rough on the edges but a good start I guess</p>",
          "rawMarkdown": "Have a look at mine:\nhttps://www.kaggle.com/guntherthepenguin/fastai-v1-densenet169-with-wsi\n\nIts still pretty rough on the edges but a good start I guess",
          "votes": 1
        },
        {
          "id": 490630,
          "postDate": "2019-03-14T18:10:42.060Z",
          "content": "<p>It's still running. I'll check it out once the job finish running. </p>",
          "rawMarkdown": "It's still running. I'll check it out once the job finish running. ",
          "votes": 1
        },
        {
          "id": 490658,
          "postDate": "2019-03-14T18:37:00.137Z",
          "content": "<p>Yep will take some time.\nI also added Focal Loss for good measure</p>",
          "rawMarkdown": "Yep will take some time.\nI also added Focal Loss for good measure"
        },
        {
          "id": 490722,
          "postDate": "2019-03-14T19:46:39.067Z",
          "content": "<p>Finished my version. Take a look: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a></p>\n\n<p>It generates train/cv split based on MSI. At the end of this script you will have train and cv ids as well as train and cv labels. </p>\n\n<p>It also works quite fast. </p>",
          "rawMarkdown": "Finished my version. Take a look: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\n\nIt generates train/cv split based on MSI. At the end of this script you will have train and cv ids as well as train and cv labels. \n\nIt also works quite fast. ",
          "votes": 1
        },
        {
          "id": 554506,
          "postDate": "2019-06-17T14:38:49.417Z",
          "content": "<p><a href=\"/ivanpan\">@ivanpan</a>  What does cv/cv_label means here?</p>",
          "rawMarkdown": "@ivanpan  What does cv/cv_label means here?"
        }
      ]
    },
    {
      "id": 490175,
      "postDate": "2019-03-14T10:49:03.267Z",
      "content": "<p>Am I the only one who understands the meaning of Whole Slide Image, but still does understand what is going on here?\nHow does getting WSI id's help improve the model? I'm so absolutely lost here, and I would be glad if someone could give a detailed explanation for newbies like us. </p>",
      "rawMarkdown": "Am I the only one who understands the meaning of Whole Slide Image, but still does understand what is going on here?\nHow does getting WSI id's help improve the model? I'm so absolutely lost here, and I would be glad if someone could give a detailed explanation for newbies like us. ",
      "replies": [
        {
          "id": 490188,
          "postDate": "2019-03-14T11:01:29.117Z",
          "content": "<p>Correct me if I'm wrong, but I understand it as follows: first of all, WSI is like a full image and images in our dataset are only patches (crops) of those original 'big' images. But why would getting these original big images improve the model? Because of data leakage. We should train on some WSIs, but validate on others. Otherwise, our estimation of performance of the model (that we get with CV) might (and in this case will be) too optimistic. </p>\n\n<p>So, instead of splitting to train/cv with random patches, we split using WSIs. Like, all patches from some particular WSI should either be in train or cv set. But if you split with random patches, some images from WSI will end up in train, and others (from the same WSI) in CV. Hence, data leakage, hence, our estimations are too optimistic. </p>",
          "rawMarkdown": "Correct me if I'm wrong, but I understand it as follows: first of all, WSI is like a full image and images in our dataset are only patches (crops) of those original 'big' images. But why would getting these original big images improve the model? Because of data leakage. We should train on some WSIs, but validate on others. Otherwise, our estimation of performance of the model (that we get with CV) might (and in this case will be) too optimistic. \n\nSo, instead of splitting to train/cv with random patches, we split using WSIs. Like, all patches from some particular WSI should either be in train or cv set. But if you split with random patches, some images from WSI will end up in train, and others (from the same WSI) in CV. Hence, data leakage, hence, our estimations are too optimistic. ",
          "votes": 3
        },
        {
          "id": 490194,
          "postDate": "2019-03-14T11:14:31.940Z",
          "content": "<p>Look at <a href=\"https://github.com/alexander-rakhlin/ICIAR2018\">https://github.com/alexander-rakhlin/ICIAR2018</a> </p>",
          "rawMarkdown": "Look at https://github.com/alexander-rakhlin/ICIAR2018 ",
          "votes": 1
        },
        {
          "id": 490204,
          "postDate": "2019-03-14T11:38:03.403Z",
          "content": "<p>Oh I totally understand now. So the main point of extracting the id's for each obtained WSI is to detect correlations between images belonging to the same WSI, and to have a solid validation during training.</p>\n\n<p>I have a slight reservation though. This means a particular WSI could contain patches that are both tumor positive and tumor negative; but the labels are quite definite: camelyon16train-tumor-078, camelyon16train-normal-078. This means a negative-labelled patch could still belong to one camelyon16train-tumor-106. \nIt would be great if <a href=\"/sermakarevich\">@sermakarevich</a> could explain the process behind his extraction of slide id's.</p>",
          "rawMarkdown": "Oh I totally understand now. So the main point of extracting the id's for each obtained WSI is to detect correlations between images belonging to the same WSI, and to have a solid validation during training.\n\nI have a slight reservation though. This means a particular WSI could contain patches that are both tumor positive and tumor negative; but the labels are quite definite: camelyon16train-tumor-078, camelyon16train-normal-078. This means a negative-labelled patch could still belong to one camelyon16train-tumor-106. \nIt would be great if @sermakarevich could explain the process behind his extraction of slide id's.\n",
          "votes": 1
        },
        {
          "id": 490206,
          "postDate": "2019-03-14T11:42:17.620Z",
          "content": "<p>Hm. Let me just get this straight: \"could contain\" or do contain? In other words, did you specifically check for that? </p>\n\n<p>P.S. If they do, then nice catch. I didn't think of that. </p>",
          "rawMarkdown": "Hm. Let me just get this straight: \"could contain\" or do contain? In other words, did you specifically check for that? \n\nP.S. If they do, then nice catch. I didn't think of that. ",
          "votes": 1
        },
        {
          "id": 490320,
          "postDate": "2019-03-14T13:31:11.130Z",
          "content": "<p>Thanks for the nice explanation. One thing that I still do not understand is how it is beneficial to get better score at the end for the test data/competition validation data. With this proper data separation for the internal validation sets, the AUC score for the internal validation set in each fold may drop from 99+ to lets say 96-97. But How it will give any boost at the end for test data/competition validation data?</p>",
          "rawMarkdown": "Thanks for the nice explanation. One thing that I still do not understand is how it is beneficial to get better score at the end for the test data/competition validation data. With this proper data separation for the internal validation sets, the AUC score for the internal validation set in each fold may drop from 99+ to lets say 96-97. But How it will give any boost at the end for test data/competition validation data?",
          "votes": 1
        },
        {
          "id": 490330,
          "postDate": "2019-03-14T13:37:09.463Z",
          "content": "<p>Well, I'm not exactly sure, but I think it works like this: by having a more realistic AUC score for the cv, you can build a better model. In particular, the one that generalises better. I mean a had a model that easily achieved 99.7+ AUROC on CV, but it got worse LB score than my another model that had only 99.4+ AUROC on CV. The fact that there is no correlation between CV and LB score makes it really difficult to build a model that generalises well. So, we use WSIs. </p>",
          "rawMarkdown": "Well, I'm not exactly sure, but I think it works like this: by having a more realistic AUC score for the cv, you can build a better model. In particular, the one that generalises better. I mean a had a model that easily achieved 99.7+ AUROC on CV, but it got worse LB score than my another model that had only 99.4+ AUROC on CV. The fact that there is no correlation between CV and LB score makes it really difficult to build a model that generalises well. So, we use WSIs. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 489543,
      "postDate": "2019-03-14T00:45:57.830Z",
      "content": "<p>Did you actually use a ResNet9 (I highly doubt)? I am not seeing any details regarding any cross validation... how was that performed?</p>",
      "rawMarkdown": "Did you actually use a ResNet9 (I highly doubt)? I am not seeing any details regarding any cross validation... how was that performed?"
    },
    {
      "id": 488711,
      "postDate": "2019-03-12T22:43:07.707Z",
      "content": "<p>Nice work! I thought about assembling the original images but I have not yet done so, but I am curious, it seems like it will improve the score, but I'm not sure its the sort of solution for pcam they are looking for. OTOH, I don't know the utility of the patches for medical imaging beyond giving tractable-sized images for working on, I don't know for sure but I would not expect pathologists to be working with such patches for actual diagnoses. </p>\n\n<p>Also, when you say your best single model, do you mean with or without fold averaging? </p>",
      "rawMarkdown": "Nice work! I thought about assembling the original images but I have not yet done so, but I am curious, it seems like it will improve the score, but I'm not sure its the sort of solution for pcam they are looking for. OTOH, I don't know the utility of the patches for medical imaging beyond giving tractable-sized images for working on, I don't know for sure but I would not expect pathologists to be working with such patches for actual diagnoses. \n\nAlso, when you say your best single model, do you mean with or without fold averaging? "
    },
    {
      "id": 499049,
      "postDate": "2019-03-24T08:54:42.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 490233,
      "postDate": "2019-03-14T12:20:41.900Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1762477,
      "author_name": "Alexander Y.",
      "author_url": "",
      "post_date": "2022-04-20T17:59:03.790000",
      "content": "<p>I am late to the party, but this is very useful discussion. Thanks everyone!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 492571,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "2019-03-17T13:17:16.993000",
      "content": "<p>Why resize training image to model required input size, rather than using average pooling to replace pooling layer? Will there be any difference in score? Why?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 493167,
          "author_name": "seefun",
          "author_url": "",
          "post_date": "2019-03-18T11:10:37.090000",
          "content": "<p><a href=\"/sermakarevich\">@sermakarevich</a> </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 493180,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2019-03-18T11:38:42.210000",
          "content": "<p>no good reason</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 488888,
      "author_name": "SM",
      "author_url": "",
      "post_date": "2019-03-13T06:36:34.210000",
      "content": "<p><a href=\"/ivanpan\">@ivanpan</a> \nTTA stands for test-time-augmentation. I used one unique transformer at time to extract slightly different predictions for the same image in test set. It was already described on the forum. </p>\n\n<p>@Roshan Santhosh\nI use them to create proper train/valid split without leakage. </p>\n\n<p>@Dima Shulkin \n<code>\n  (1): Sequential(\n    (0): AdaptiveConcatPool2d(\n      (average_pool): AdaptiveAvgPool2d(output_size=1)\n      (max_pool): AdaptiveMaxPool2d(output_size=1)\n    )\n    (1): Flatten()\n    (2): BatchNorm1d(3072, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (3): Dropout(p=0.8)\n    (4): Linear(in_features=3072, out_features=512, bias=True)\n    (5): ReLU(inplace)\n    (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (7): Dropout(p=0.8)\n    (8): Linear(in_features=512, out_features=256, bias=True)\n    (9): ReLU(inplace)\n    (10): BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (11): Dropout(p=0.8)\n    (12): Linear(in_features=256, out_features=1, bias=True)\n  )\n)\n</code></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 488694,
      "author_name": "Roshan Santhosh",
      "author_url": "",
      "post_date": "2019-03-12T22:14:50.713000",
      "content": "<p>How exactly are you using the WSI IDs?</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 490289,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2019-03-14T13:08:58.370000",
      "content": "<p>Weird thing: started to check WSI ids and found something strange. For instance, there is a WSI called camelyon16_train_tumor_017. And it contains 544 patches. The thing is: non of them are labeled as \"tumor\" in the competition dataset. I mean not a single one. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 489142,
      "author_name": "Taylor",
      "author_url": "",
      "post_date": "2019-03-13T14:48:54.697000",
      "content": "<p>Thanks for sharing WHOLE SLIDE IMAGE (WSI): A digitized histopathology glass slide that has been created on a slide scanner.  The digitized glass slide represents a high-resolution replica of the original glass that can then be manipulated through software to mimic microscope review and diagnosis. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 489001,
      "author_name": "zxyu",
      "author_url": "",
      "post_date": "2019-03-13T11:06:43.937000",
      "content": "<p>Thank you for sharing, Can you explain the mean of WSI.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 489003,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-13T11:08:18.340000",
          "content": "<p>Whole Slide Image</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 488899,
      "author_name": "Joni Juvonen",
      "author_url": "",
      "post_date": "2019-03-13T06:53:16.743000",
      "content": "<p>Thank you for the slide IDs!</p>\n\n<p>So in order to remove inter-WSI bias, we should partition train/validation and CV folds by WSIs. Splitting just randomly leads to too optimistic validation set estimates.</p>\n\n<p>WSI stats:\n- 27,273 training samples with unknown WSI\n- Minimum samples per WSI is 2 <code>wsi010</code>\n- Maximum samples per WSI is 3,363 <code>wsi094</code></p>\n\n<p>Did you also stratify by labels when splitting? I'm not sure how to manage that...</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2845734,
          "author_name": "ruigangge",
          "author_url": "",
          "post_date": "2024-05-30T16:53:52.287000",
          "content": "<p>Hi, I extracted slide IDs for almost every patch, but I can't open the links now. Could you please resend them to me? Alternatively, you can send them to my email at ruigangge@gmail.com. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 488877,
      "author_name": "Dimitrij Shulkin",
      "author_url": "",
      "post_date": "2019-03-13T06:14:06.990000",
      "content": "<p>Hello SM,\nthank you for your contribution! \nWhat do you mean saying \"multiple linear layers\"? Linear activation functions? \nBest\nDima</p>",
      "votes": 1,
      "replies": [
        {
          "id": 488881,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-13T06:21:50.820000",
          "content": "<p>I think it means multiple fully connected layers.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 488596,
      "author_name": "Ivan Panshin",
      "author_url": "",
      "post_date": "2019-03-12T18:17:12.743000",
      "content": "<p>What do you mean by TTA predictions on all flips+rotations mentioned above? You mean we'll have 3 images (original, and 2 for each RandomChoice) or you mean 16 (original + 15 for every mentioned transformation)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 488690,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-12T21:52:41.973000",
          "content": "<p>Great results, by the way. Thanks for sharing your approach. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 493049,
      "author_name": "Dimitrij Shulkin",
      "author_url": "",
      "post_date": "2019-03-18T06:57:30.760000",
      "content": "<p>Hello everyone,\ndid anybody perform stain normalization (<a href=\"https://arxiv.org/abs/1811.03815\">https://arxiv.org/abs/1811.03815</a>) ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 493065,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2019-03-18T07:30:23.293000",
          "content": "<p>Hi,\nI was thinking about that not with with the approach of this paper but rather with the conventional methods. However, I could not figure it out how to choose a proper reference image. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 493168,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-18T11:12:36.723000",
          "content": "<p>That was also my biggest problem. I couldn't find any good approaches anywhere. I tried to select the reference image randomly. However, I also came across the other problem. The images in the train set and test set with too many white pixels could not be transformed.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 493338,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2019-03-18T14:56:25.243000",
          "content": "<p>That's true. I had the same problem with the whitish images.\nAnother solution could be using the WSI stain normalization which was used by the Camelyon 2016 challenge winner. However, as we do not have the whole slide information for this dataset, it is not possible to use that approach. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 493613,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-03-18T21:42:29.733000",
          "content": "<p>Great ideas! And thanks for bringing this to my attention. </p>\n\n<p>Your kernel on the subject is very good, it’s good to show explicitly how the transforms work. I have started to experiment with stain normalization methods but it’s too early to report anything convincing. </p>\n\n<p>For anyone who wants to try this, aside from the kernel mentioned there is a nice python package called staintools that is quite easy to use (unless you are on windows). </p>\n\n<p>I agree that deciding on a reference image to normalize the rest to is not trivial or obvious. Trial and error here should be quite a bit faster than trying to train an infogan though, but that I have no idea how to do. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 493987,
          "author_name": "Aarya Patel",
          "author_url": "",
          "post_date": "2019-03-19T10:37:11.750000",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> Which kernel?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 493992,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2019-03-19T10:41:20.487000",
          "content": "<p>this one:\n<a href=\"https://www.kaggle.com/robotdreams/stain-normalization\">https://www.kaggle.com/robotdreams/stain-normalization</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 494197,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-03-19T15:11:27.353000",
          "content": "<p>Yep, that’s the one. Dima’s Kernel. Thanks masih. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 489956,
      "author_name": "SM",
      "author_url": "",
      "post_date": "2019-03-14T07:37:15.167000",
      "content": "<p><a href=\"/keremt\">@keremt</a> It is just matching of patches in the competition with original patches by mean pixel value. If there is only one patch in original dataset with the mean value, I used its WSI id, if multiple - I skipped it. </p>\n\n<p><a href=\"/tanlikesmath\">@tanlikesmath</a> No I did not. Resnet, densenet and some models from here <code>https://github.com/Cadene/pretrained-models.pytorch</code></p>\n\n<p><a href=\"/qitvision\">@qitvision</a> No I did not. I did not do stratified split by labels as I thought there are too many patches and random split should be good enough. Was thinking about stratified split by rounded pixel means but did not do that either.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2845742,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-05-30T16:56:05.533000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 488663,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-03-12T20:52:24.033000",
      "content": "<p>Thank you! I think that I figured out that WSI file.  If I'm not mistaken, we have to replace \"camelyon16_train_tumor_066\"  with 0. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 488658,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2019-03-12T20:39:36.907000",
      "content": "<p><a href=\"/sermakarevich\">@sermakarevich</a> , thank you for sharing WSIs IDs, I'm curious how did u get them? Another question is if u tried to use all 600 GB dataset including both camelyon16 and camelyon17 for training? Though I think it isn't worth doing it for a playground competition, the result may be much better.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 488678,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2019-03-12T21:15:38.497000",
          "content": "<p>It`s just a playground. I was just testing a pipeline that we work on at <a href=\"https://www.lilystyle.ai/\">https://www.lilystyle.ai/</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 488946,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2019-03-13T09:03:33.547000",
          "content": "<p>Very cool! I am also curious about how to get the WSI IDs :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2845761,
      "author_name": "ruigangge",
      "author_url": "",
      "post_date": "2024-05-30T17:03:55.170000",
      "content": "<p>I hope this message finds you well. I recently came across your forum post where you shared a link for the extracted slide IDs for almost every patch. Unfortunately, the link seems to have expired. Could you please resend the information or update the link directly to my email at ruigangge@gmail.com? It would be greatly appreciated.</p>\n<p>Additionally, I would like to inquire about the slide IDs. Are all slide IDs provided officially unique, and could you please share some insights into how these IDs are obtained?</p>\n<p>Thank you for your assistance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 496363,
      "author_name": "demonic toaster",
      "author_url": "",
      "post_date": "2019-03-22T04:35:42.223000",
      "content": "<p>Hi SM, many thanks for the awesome tips!!\nCould you share the code or explanation on how you extracted the slide IDs? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 499057,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2019-03-24T09:11:42.393000",
          "content": "<p>unfortunately I cant as thats a lib we write for prod at <code>www.lilystyle.ai</code>\nhowever some sketches of it you can find <a href=\"https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9\">here</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 494938,
      "author_name": "Taksh Kamlesh",
      "author_url": "",
      "post_date": "2019-03-20T12:59:25.813000",
      "content": "<p>What is a WSI??</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 492342,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2019-03-17T05:30:10.887000",
      "content": "<p>dataset isn't extremely imbalanced, but focal loss still provided a big advantage?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 492359,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-03-17T06:07:12.533000",
          "content": "<p>focal loss has not been any better than bce for me yet. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 492421,
          "author_name": "Joni Juvonen",
          "author_url": "",
          "post_date": "2019-03-17T09:06:03.030000",
          "content": "<p>Same for me. I didn't see any improvement to BCE and I tried focal loss with <code>alpha=0.25, gamma=2</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 491298,
      "author_name": "lvguofeng",
      "author_url": "",
      "post_date": "2019-03-15T12:48:09.380000",
      "content": "<p>Hi SM,\nThank you! I want to know you how to divide dataset through wsi_ids? If the validation dataset accounts for 20%，the train dataset contains all wsi_ids and takes 80% of them, or 80% of wsi_ids and takes all of them?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 491263,
      "author_name": "quinwu",
      "author_url": "",
      "post_date": "2019-03-15T12:01:56.367000",
      "content": "<p>Hi SM,\nMy single model is only 0.9711,\nI want to know which network model you are using.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 492007,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-03-16T14:04:01.400000",
          "content": "<p>Do you use TTA? TTA could help my models boost from around 0.97 to 0.974.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 490513,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-03-14T16:20:35.853000",
      "content": "<p>Is anyone willing to share their pipeline for WSI in a kernel? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 490565,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-14T17:08:42.393000",
          "content": "<p>I'm working on it. Maybe tomorrow or the day after that, when I finish. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490608,
          "author_name": "Gunther",
          "author_url": "",
          "post_date": "2019-03-14T17:49:28.490000",
          "content": "<p>Have a look at mine:\n<a href=\"https://www.kaggle.com/guntherthepenguin/fastai-v1-densenet169-with-wsi\">https://www.kaggle.com/guntherthepenguin/fastai-v1-densenet169-with-wsi</a></p>\n\n<p>Its still pretty rough on the edges but a good start I guess</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490630,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2019-03-14T18:10:42.060000",
          "content": "<p>It's still running. I'll check it out once the job finish running. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490658,
          "author_name": "Gunther",
          "author_url": "",
          "post_date": "2019-03-14T18:37:00.137000",
          "content": "<p>Yep will take some time.\nI also added Focal Loss for good measure</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 490722,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-14T19:46:39.067000",
          "content": "<p>Finished my version. Take a look: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a></p>\n\n<p>It generates train/cv split based on MSI. At the end of this script you will have train and cv ids as well as train and cv labels. </p>\n\n<p>It also works quite fast. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 554506,
          "author_name": "FeiZhuNiU",
          "author_url": "",
          "post_date": "2019-06-17T14:38:49.417000",
          "content": "<p><a href=\"/ivanpan\">@ivanpan</a>  What does cv/cv_label means here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 490175,
      "author_name": "Abiodun PK",
      "author_url": "",
      "post_date": "2019-03-14T10:49:03.267000",
      "content": "<p>Am I the only one who understands the meaning of Whole Slide Image, but still does understand what is going on here?\nHow does getting WSI id's help improve the model? I'm so absolutely lost here, and I would be glad if someone could give a detailed explanation for newbies like us. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 490188,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-14T11:01:29.117000",
          "content": "<p>Correct me if I'm wrong, but I understand it as follows: first of all, WSI is like a full image and images in our dataset are only patches (crops) of those original 'big' images. But why would getting these original big images improve the model? Because of data leakage. We should train on some WSIs, but validate on others. Otherwise, our estimation of performance of the model (that we get with CV) might (and in this case will be) too optimistic. </p>\n\n<p>So, instead of splitting to train/cv with random patches, we split using WSIs. Like, all patches from some particular WSI should either be in train or cv set. But if you split with random patches, some images from WSI will end up in train, and others (from the same WSI) in CV. Hence, data leakage, hence, our estimations are too optimistic. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 490194,
          "author_name": "Dimitrij Shulkin",
          "author_url": "",
          "post_date": "2019-03-14T11:14:31.940000",
          "content": "<p>Look at <a href=\"https://github.com/alexander-rakhlin/ICIAR2018\">https://github.com/alexander-rakhlin/ICIAR2018</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490204,
          "author_name": "Abiodun PK",
          "author_url": "",
          "post_date": "2019-03-14T11:38:03.403000",
          "content": "<p>Oh I totally understand now. So the main point of extracting the id's for each obtained WSI is to detect correlations between images belonging to the same WSI, and to have a solid validation during training.</p>\n\n<p>I have a slight reservation though. This means a particular WSI could contain patches that are both tumor positive and tumor negative; but the labels are quite definite: camelyon16train-tumor-078, camelyon16train-normal-078. This means a negative-labelled patch could still belong to one camelyon16train-tumor-106. \nIt would be great if <a href=\"/sermakarevich\">@sermakarevich</a> could explain the process behind his extraction of slide id's.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490206,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-14T11:42:17.620000",
          "content": "<p>Hm. Let me just get this straight: \"could contain\" or do contain? In other words, did you specifically check for that? </p>\n\n<p>P.S. If they do, then nice catch. I didn't think of that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490320,
          "author_name": "learner",
          "author_url": "",
          "post_date": "2019-03-14T13:31:11.130000",
          "content": "<p>Thanks for the nice explanation. One thing that I still do not understand is how it is beneficial to get better score at the end for the test data/competition validation data. With this proper data separation for the internal validation sets, the AUC score for the internal validation set in each fold may drop from 99+ to lets say 96-97. But How it will give any boost at the end for test data/competition validation data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 490330,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2019-03-14T13:37:09.463000",
          "content": "<p>Well, I'm not exactly sure, but I think it works like this: by having a more realistic AUC score for the cv, you can build a better model. In particular, the one that generalises better. I mean a had a model that easily achieved 99.7+ AUROC on CV, but it got worse LB score than my another model that had only 99.4+ AUROC on CV. The fact that there is no correlation between CV and LB score makes it really difficult to build a model that generalises well. So, we use WSIs. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 489543,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2019-03-14T00:45:57.830000",
      "content": "<p>Did you actually use a ResNet9 (I highly doubt)? I am not seeing any details regarding any cross validation... how was that performed?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 488711,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-03-12T22:43:07.707000",
      "content": "<p>Nice work! I thought about assembling the original images but I have not yet done so, but I am curious, it seems like it will improve the score, but I'm not sure its the sort of solution for pcam they are looking for. OTOH, I don't know the utility of the patches for medical imaging beyond giving tractable-sized images for working on, I don't know for sure but I would not expect pathologists to be working with such patches for actual diagnoses. </p>\n\n<p>Also, when you say your best single model, do you mean with or without fold averaging? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 499049,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-24T08:54:42.433000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 490233,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-14T12:20:41.900000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "488582": "main issues:\n- validation, needs to be done by isolating WSIs, not random patches\n- overfitting to WSIs in train\n\nsolutions:\n- I extracted slide ids for almost every patch, here they [are](https://drive.google.com/open?id=1NgE2Uuwhr3yDPVwVNpmSwJwQG1RIIJsC)\n- intense augmentations, dropout 3 x 90%, more Linear layers\n\ntraining:\n- image size - rescaled to model required input size \n- no cleaning\n- transformers: \n```\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n```\n- custom tail: avg and max pooling concat, multiple Linear layers with dropout and  batchnorm\n-  training:\n   - freeze all except custom tail\n   - lr * 100\n   - unfreeze all\n   - lr / 100\n   - cyclic lr (triangle)\n   - save model when score improved\n   - drop lr /2 if AUC on valid did not improve for 2-5 cycles (1 cycle is xxx iterations based on arch and bach size) and load best model\n   - lr restart *(10-100) after 4 lr drops\n- TTa predictions on all flips+rotations transforms mentioned above\n- single best model 0.975\n\nUPDT: forgot to mention, I use FocalLoss\n\n[starter kernel](https://www.kaggle.com/sermakarevich/complete-handcrafted-pipeline-in-pytorch-resnet9) ",
    "1762477": "I am late to the party, but this is very useful discussion. Thanks everyone!",
    "492571": "Why resize training image to model required input size, rather than using average pooling to replace pooling layer? Will there be any difference in score? Why?",
    "488888": "@ivanpan \nTTA stands for test-time-augmentation. I used one unique transformer at time to extract slightly different predictions for the same image in test set. It was already described on the forum. \n\n@Roshan Santhosh\nI use them to create proper train/valid split without leakage. \n\n@Dima Shulkin \n```\n  (1): Sequential(\n    (0): AdaptiveConcatPool2d(\n      (average_pool): AdaptiveAvgPool2d(output_size=1)\n      (max_pool): AdaptiveMaxPool2d(output_size=1)\n    )\n    (1): Flatten()\n    (2): BatchNorm1d(3072, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (3): Dropout(p=0.8)\n    (4): Linear(in_features=3072, out_features=512, bias=True)\n    (5): ReLU(inplace)\n    (6): BatchNorm1d(512, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (7): Dropout(p=0.8)\n    (8): Linear(in_features=512, out_features=256, bias=True)\n    (9): ReLU(inplace)\n    (10): BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (11): Dropout(p=0.8)\n    (12): Linear(in_features=256, out_features=1, bias=True)\n  )\n)\n```\n",
    "488694": "How exactly are you using the WSI IDs?",
    "490289": "Weird thing: started to check WSI ids and found something strange. For instance, there is a WSI called camelyon16_train_tumor_017. And it contains 544 patches. The thing is: non of them are labeled as \"tumor\" in the competition dataset. I mean not a single one. ",
    "489142": "Thanks for sharing WHOLE SLIDE IMAGE (WSI): A digitized histopathology glass slide that has been created on a slide scanner.  The digitized glass slide represents a high-resolution replica of the original glass that can then be manipulated through software to mimic microscope review and diagnosis. ",
    "489001": "Thank you for sharing, Can you explain the mean of WSI.",
    "488899": "Thank you for the slide IDs!\n\nSo in order to remove inter-WSI bias, we should partition train/validation and CV folds by WSIs. Splitting just randomly leads to too optimistic validation set estimates.\n\nWSI stats:\n- 27,273 training samples with unknown WSI\n- Minimum samples per WSI is 2 `wsi010`\n- Maximum samples per WSI is 3,363 `wsi094`\n\nDid you also stratify by labels when splitting? I'm not sure how to manage that...",
    "488877": "Hello SM,\nthank you for your contribution! \nWhat do you mean saying \"multiple linear layers\"? Linear activation functions? \nBest\nDima",
    "488596": "What do you mean by TTA predictions on all flips+rotations mentioned above? You mean we'll have 3 images (original, and 2 for each RandomChoice) or you mean 16 (original + 15 for every mentioned transformation)?",
    "493049": "Hello everyone,\ndid anybody perform stain normalization (https://arxiv.org/abs/1811.03815) ?",
    "489956": "@keremt It is just matching of patches in the competition with original patches by mean pixel value. If there is only one patch in original dataset with the mean value, I used its WSI id, if multiple - I skipped it. \n\n@tanlikesmath No I did not. Resnet, densenet and some models from here `https://github.com/Cadene/pretrained-models.pytorch`\n\n@qitvision No I did not. I did not do stratified split by labels as I thought there are too many patches and random split should be good enough. Was thinking about stratified split by rounded pixel means but did not do that either.  ",
    "488663": "Thank you! I think that I figured out that WSI file.  If I'm not mistaken, we have to replace \"camelyon16_train_tumor_066\"  with 0. ",
    "488658": "@sermakarevich , thank you for sharing WSIs IDs, I'm curious how did u get them? Another question is if u tried to use all 600 GB dataset including both camelyon16 and camelyon17 for training? Though I think it isn't worth doing it for a playground competition, the result may be much better.",
    "2845761": "I hope this message finds you well. I recently came across your forum post where you shared a link for the extracted slide IDs for almost every patch. Unfortunately, the link seems to have expired. Could you please resend the information or update the link directly to my email at ruigangge@gmail.com? It would be greatly appreciated.\n\nAdditionally, I would like to inquire about the slide IDs. Are all slide IDs provided officially unique, and could you please share some insights into how these IDs are obtained?\n\nThank you for your assistance.",
    "496363": "Hi SM, many thanks for the awesome tips!!\nCould you share the code or explanation on how you extracted the slide IDs? ",
    "494938": "What is a WSI??",
    "492342": " dataset isn't extremely imbalanced, but focal loss still provided a big advantage?",
    "491298": "Hi SM,\nThank you! I want to know you how to divide dataset through wsi_ids? If the validation dataset accounts for 20%，the train dataset contains all wsi_ids and takes 80% of them, or 80% of wsi_ids and takes all of them?",
    "491263": "Hi SM,\nMy single model is only 0.9711,\nI want to know which network model you are using.",
    "490513": "Is anyone willing to share their pipeline for WSI in a kernel? ",
    "490175": "Am I the only one who understands the meaning of Whole Slide Image, but still does understand what is going on here?\nHow does getting WSI id's help improve the model? I'm so absolutely lost here, and I would be glad if someone could give a detailed explanation for newbies like us. ",
    "489543": "Did you actually use a ResNet9 (I highly doubt)? I am not seeing any details regarding any cross validation... how was that performed?",
    "488711": "Nice work! I thought about assembling the original images but I have not yet done so, but I am curious, it seems like it will improve the score, but I'm not sure its the sort of solution for pcam they are looking for. OTOH, I don't know the utility of the patches for medical imaging beyond giving tractable-sized images for working on, I don't know for sure but I would not expect pathologists to be working with such patches for actual diagnoses. \n\nAlso, when you say your best single model, do you mean with or without fold averaging? ",
    "499049": "",
    "490233": ""
  }
}