{
  "id": 84790,
  "title": "Another way to get 0.98+ on LB",
  "url": "/competitions/histopathologic-cancer-detection/discussion/84790",
  "author_name": "",
  "post_date": "2019-03-19T18:49:48.275306100Z",
  "votes": 10,
  "comment_count": 20,
  "views": 0,
  "content": "<p>1) You should split data according to WSIs. This idea belongs to SM from <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760</a></p>\n\n<p>You can get a WSI split from my another discussion: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a></p>\n\n<p>It will allow you to have a more reliable metric. In other words, you won't be getting 0.99+ on validation, but barely 0.96 on LB.</p>\n\n<p>2) Use TTA. I tested augmentations from SM's discussion and they really work well. Here they are: </p>\n\n<p><code>\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n</code></p>\n\n<p>In order to make it work you can check my discussion regarding TTA in PyTorch (<a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056</a> ). In this case, you will have about 15 different versions of original image, so the training will become much slower. But it's worth it. </p>\n\n<p>Obviously, you should apply the same transformations to your CV set. </p>\n\n<p>3) I tested different architectures. For me, the best one so far is a pretrained se_resnet50 with dropout of 0.2 in conv layers, concatenation of average and max pooling (both global of course), several fully connected layers with intense dropout. The last idea also belongs to SM. </p>\n\n<p>2 poolings:\n<code>\nx1 = self.avg_pool(x)\nx2 = self.max_pool(x)\nx = torch.cat([x1,x2], 1)\n</code></p>\n\n<p>Linear layers:</p>\n\n<p><code>\nself.net.last_linear = nn.Sequential(nn.BatchNorm1d(4096, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(4096, 768, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(768, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(768, 256, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(256, 2))\n</code></p>\n\n<p>4) I resize images to 196x196 and use Adam with LR of 0.0007. Batch size is 150 for training and 10 for cv and test. If I change it to something bigger, I run out of cuda memory (Tesla K80)</p>\n\n<p>5) Save model every time AUROC is increased. By the end of the training save the best model. </p>\n\n<p>6) Apply ReduceLROnPlateau with the patience of 1-2 epocs.</p>\n\n<p>7) Train the added layers for 1-2 epoch and then the whole network for another 3-4. </p>\n\n<p>8) If you're using kaggle kernels, then train with one kernel and test with another. Otherwise, you will run out of time. The training takes about 6-7 hours, the testing — another 50-60 minutes.</p>\n\n<p>The single best model performs 0.9786 on LB. I haven't tested ensembles thoroughly, but I had no trouble breaking the 0.98 landmark with just 2 models. Will see how well it performs with 5 or 10 models.  </p>\n\n<p>If you have any questions, don't hesitate to ask. </p>",
  "messages": [
    {
      "id": "494356",
      "postDate": "03/19/2019 18:49:48",
      "content": "<p>1) You should split data according to WSIs. This idea belongs to SM from <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760</a></p>\n\n<p>You can get a WSI split from my another discussion: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a></p>\n\n<p>It will allow you to have a more reliable metric. In other words, you won't be getting 0.99+ on validation, but barely 0.96 on LB.</p>\n\n<p>2) Use TTA. I tested augmentations from SM's discussion and they really work well. Here they are: </p>\n\n<p><code>\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n</code></p>\n\n<p>In order to make it work you can check my discussion regarding TTA in PyTorch (<a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056</a> ). In this case, you will have about 15 different versions of original image, so the training will become much slower. But it's worth it. </p>\n\n<p>Obviously, you should apply the same transformations to your CV set. </p>\n\n<p>3) I tested different architectures. For me, the best one so far is a pretrained se_resnet50 with dropout of 0.2 in conv layers, concatenation of average and max pooling (both global of course), several fully connected layers with intense dropout. The last idea also belongs to SM. </p>\n\n<p>2 poolings:\n<code>\nx1 = self.avg_pool(x)\nx2 = self.max_pool(x)\nx = torch.cat([x1,x2], 1)\n</code></p>\n\n<p>Linear layers:</p>\n\n<p><code>\nself.net.last_linear = nn.Sequential(nn.BatchNorm1d(4096, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(4096, 768, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(768, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(768, 256, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(256, 2))\n</code></p>\n\n<p>4) I resize images to 196x196 and use Adam with LR of 0.0007. Batch size is 150 for training and 10 for cv and test. If I change it to something bigger, I run out of cuda memory (Tesla K80)</p>\n\n<p>5) Save model every time AUROC is increased. By the end of the training save the best model. </p>\n\n<p>6) Apply ReduceLROnPlateau with the patience of 1-2 epocs.</p>\n\n<p>7) Train the added layers for 1-2 epoch and then the whole network for another 3-4. </p>\n\n<p>8) If you're using kaggle kernels, then train with one kernel and test with another. Otherwise, you will run out of time. The training takes about 6-7 hours, the testing — another 50-60 minutes.</p>\n\n<p>The single best model performs 0.9786 on LB. I haven't tested ensembles thoroughly, but I had no trouble breaking the 0.98 landmark with just 2 models. Will see how well it performs with 5 or 10 models.  </p>\n\n<p>If you have any questions, don't hesitate to ask. </p>",
      "rawMarkdown": "1) You should split data according to WSIs. This idea belongs to SM from https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\n\nYou can get a WSI split from my another discussion: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\n\nIt will allow you to have a more reliable metric. In other words, you won't be getting 0.99+ on validation, but barely 0.96 on LB.\n\n2) Use TTA. I tested augmentations from SM's discussion and they really work well. Here they are: \n\n```\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n```\n\nIn order to make it work you can check my discussion regarding TTA in PyTorch (https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056 ). In this case, you will have about 15 different versions of original image, so the training will become much slower. But it's worth it. \n\nObviously, you should apply the same transformations to your CV set. \n\n3) I tested different architectures. For me, the best one so far is a pretrained se_resnet50 with dropout of 0.2 in conv layers, concatenation of average and max pooling (both global of course), several fully connected layers with intense dropout. The last idea also belongs to SM. \n\n2 poolings:\n```\nx1 = self.avg_pool(x)\nx2 = self.max_pool(x)\nx = torch.cat([x1,x2], 1)\n```\n\nLinear layers:\n\n```\nself.net.last_linear = nn.Sequential(nn.BatchNorm1d(4096, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(4096, 768, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(768, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(768, 256, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(256, 2))\n```\n\n4) I resize images to 196x196 and use Adam with LR of 0.0007. Batch size is 150 for training and 10 for cv and test. If I change it to something bigger, I run out of cuda memory (Tesla K80)\n\n5) Save model every time AUROC is increased. By the end of the training save the best model. \n\n6) Apply ReduceLROnPlateau with the patience of 1-2 epocs.\n\n7) Train the added layers for 1-2 epoch and then the whole network for another 3-4. \n\n8) If you're using kaggle kernels, then train with one kernel and test with another. Otherwise, you will run out of time. The training takes about 6-7 hours, the testing — another 50-60 minutes.\n\nThe single best model performs 0.9786 on LB. I haven't tested ensembles thoroughly, but I had no trouble breaking the 0.98 landmark with just 2 models. Will see how well it performs with 5 or 10 models.  \n\nIf you have any questions, don't hesitate to ask.",
      "votes": null
    },
    {
      "id": "494387",
      "postDate": "03/19/2019 19:31:05",
      "content": "<p>Add on to that, I tested serval models: SE nets are generally better than densenets for this dataset.</p>\n\n<p>Also, when you say: <code>dropout of 0.2</code>, do you mean you added dropout in the CNN layers? (because clearly, your linear layers' dropout is 0.8, not 0.2)</p>",
      "rawMarkdown": "Add on to that, I tested serval models: SE nets are generally better than densenets for this dataset.\n\nAlso, when you say: `dropout of 0.2`, do you mean you added dropout in the CNN layers? (because clearly, your linear layers' dropout is 0.8, not 0.2)",
      "votes": null
    },
    {
      "id": "494390",
      "postDate": "03/19/2019 19:32:49",
      "content": "<p>I second that \" SE nets are generally better than densenets for this dataset\".</p>\n\n<p>\"Also, when you say: dropout of 0.2, do you mean you added dropout in the CNN layers?\" yes, exactly</p>",
      "rawMarkdown": "I second that \" SE nets are generally better than densenets for this dataset\".\n\n\"Also, when you say: dropout of 0.2, do you mean you added dropout in the CNN layers?\" yes, exactly",
      "votes": null
    },
    {
      "id": "494395",
      "postDate": "03/19/2019 19:34:59",
      "content": "<p>Changed the post so that it's more clear.</p>",
      "rawMarkdown": "Changed the post so that it's more clear.",
      "votes": null
    },
    {
      "id": "494548",
      "postDate": "03/20/2019 01:36:12",
      "content": "<p>Hi, Ivan!! Really good results with only 2 models, may I ask why don't you use any activation function in the last linear layers?</p>",
      "rawMarkdown": "Hi, Ivan!! Really good results with only 2 models, may I ask why don't you use any activation function in the last linear layers?",
      "votes": null
    },
    {
      "id": "494556",
      "postDate": "03/20/2019 01:55:42",
      "content": "<p>It's kind of a typo. I do use activation functions of course. I use Relu + log softmax at the end. Updated the post. Thanks for noticing! </p>",
      "rawMarkdown": "It's kind of a typo. I do use activation functions of course. I use Relu + log softmax at the end. Updated the post. Thanks for noticing!",
      "votes": null
    },
    {
      "id": "494616",
      "postDate": "03/20/2019 03:41:15",
      "content": "<p>Why not try removing BN in the full connection layer? BN may not good in theory. Using SELU rather than BN+RELU in fc layers may have better result. <a href=\"https://arxiv.org/pdf/1706.02515.pdf\">SELU</a></p>",
      "rawMarkdown": "Why not try removing BN in the full connection layer? BN may not good in theory. Using SELU rather than BN+RELU in fc layers may have better result. [SELU](https://arxiv.org/pdf/1706.02515.pdf)",
      "votes": null
    },
    {
      "id": "494621",
      "postDate": "03/20/2019 03:45:39",
      "content": "<p>Hm, interesting. Never read that paper. Definitely will take a look. Thanks.</p>\n\n<p>But as far as a know, using BN with Relu usually gives better results that just Relu. </p>",
      "rawMarkdown": "Hm, interesting. Never read that paper. Definitely will take a look. Thanks.\n\nBut as far as a know, using BN with Relu usually gives better results that just Relu.",
      "votes": null
    },
    {
      "id": "494690",
      "postDate": "03/20/2019 06:22:57",
      "content": "<p>Hello Ivan,\nthank you for contribution! Do you use 15 random TTA according to your code?</p>",
      "rawMarkdown": "Hello Ivan,\nthank you for contribution! Do you use 15 random TTA according to your code?",
      "votes": null
    },
    {
      "id": "494693",
      "postDate": "03/20/2019 06:29:40",
      "content": "<p>Hi! You're welcome. </p>\n\n<p>Yes, of course. It's really slow (~15 times slower, than it could be :) ), but my experiments show that it helps getting a better performance, so it's worth it. </p>",
      "rawMarkdown": "Hi! You're welcome. \n\nYes, of course. It's really slow (~15 times slower, than it could be :) ), but my experiments show that it helps getting a better performance, so it's worth it.",
      "votes": null
    },
    {
      "id": "494822",
      "postDate": "03/20/2019 09:39:51",
      "content": "<p>Hi, </p>\n\n<p>do you take then simple mean average of 15 TTA?\nDo you also apply augmentation to the validtation set?</p>\n\n<p>Dima</p>",
      "rawMarkdown": "Hi, \n\ndo you take then simple mean average of 15 TTA?\nDo you also apply augmentation to the validtation set?\n\nDima",
      "votes": null
    },
    {
      "id": "494826",
      "postDate": "03/20/2019 09:44:31",
      "content": "<p>Yes and yes </p>",
      "rawMarkdown": "Yes and yes",
      "votes": null
    },
    {
      "id": "495014",
      "postDate": "03/20/2019 14:42:10",
      "content": "<p>Incorrect order? Actually, in most practice,  batch normalization should be performed right after FC/Conv instead of after ReLU. Although, some may place BN before FC/Conv. check this paper: Identity Mappings in Deep Residual Networks</p>",
      "rawMarkdown": "Incorrect order? Actually, in most practice,  batch normalization should be performed right after FC/Conv instead of after ReLU. Although, some may place BN before FC/Conv. check this paper: Identity Mappings in Deep Residual Networks",
      "votes": null
    },
    {
      "id": "496817",
      "postDate": "03/22/2019 16:08:15",
      "content": "<p>Hello Ivan,\nthank you for the very nice explanation. When you say that you break the 0.98 landmark with 2 models, do you mean that you average the predictions obtained with two different models? If so, how are the models differing from each other (different architectures, different hyperparams, etc.)?</p>",
      "rawMarkdown": "Hello Ivan,\nthank you for the very nice explanation. When you say that you break the 0.98 landmark with 2 models, do you mean that you average the predictions obtained with two different models? If so, how are the models differing from each other (different architectures, different hyperparams, etc.)?",
      "votes": null
    },
    {
      "id": "496835",
      "postDate": "03/22/2019 16:33:29",
      "content": "<p>Hi! You're welcome. </p>\n\n<p>Yes, I mean that I average the predictions. And those models don't really differ that much. Basically, there are 2 differences: they are trained on different subsets of WSIs (check my other discussion: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a> ) and they differ due to the random initialization of weigths. In other words, in terms of achitectures, hyperparameters, etc they are the same. </p>\n\n<p>P.S. Obviously, I will build an ensemble with other models. Something like densenet or se_densenet. I will try to update the post when I get the data. </p>",
      "rawMarkdown": "Hi! You're welcome. \n\nYes, I mean that I average the predictions. And those models don't really differ that much. Basically, there are 2 differences: they are trained on different subsets of WSIs (check my other discussion: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132 ) and they differ due to the random initialization of weigths. In other words, in terms of achitectures, hyperparameters, etc they are the same. \n\nP.S. Obviously, I will build an ensemble with other models. Something like densenet or se_densenet. I will try to update the post when I get the data.",
      "votes": null
    },
    {
      "id": "497709",
      "postDate": "03/23/2019 20:24:40",
      "content": "<p>Hello, Ivan. Thank you for contribution!\nDo you use l2 regularization?</p>",
      "rawMarkdown": "Hello, Ivan. Thank you for contribution!\nDo you use l2 regularization?",
      "votes": null
    },
    {
      "id": "498989",
      "postDate": "03/24/2019 07:01:49",
      "content": "<p>Hi! You're welcome. </p>\n\n<p>No, I don't. Mostly because (as far as I know) it's not a wise idea to use L2 and dropout together. </p>",
      "rawMarkdown": "Hi! You're welcome. \n\nNo, I don't. Mostly because (as far as I know) it's not a wise idea to use L2 and dropout together.",
      "votes": null
    },
    {
      "id": "500687",
      "postDate": "03/26/2019 12:23:39",
      "content": "<p><code>transforms.Normalize( mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])</code> \nthe mean - [0.485, 0.456, 0.406] and std - [0.229, 0.224, 0.225] is right ?</p>",
      "rawMarkdown": "`transforms.Normalize( mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])` \nthe mean - [0.485, 0.456, 0.406] and std - [0.229, 0.224, 0.225] is right ?",
      "votes": null
    },
    {
      "id": "500847",
      "postDate": "03/26/2019 16:07:02",
      "content": "<p>Yes, of course. Those are the mean and std values for the ImageNet dataset </p>",
      "rawMarkdown": "Yes, of course. Those are the mean and std values for the ImageNet dataset",
      "votes": null
    },
    {
      "id": "503776",
      "postDate": "03/30/2019 14:41:03",
      "content": "<p>\"The training takes about 6-7 hours, the testing — another 50-60 minutes.\"\nAll test data can be stored as single numpy array about 12 Gb. With simple generator for TTA this array can be used for prediction with relatively large batchsize (i use 640). In this case single test prediction takes about 80sec on 1060 6GB (I use DenseNet169). </p>",
      "rawMarkdown": "\"The training takes about 6-7 hours, the testing — another 50-60 minutes.\"\nAll test data can be stored as single numpy array about 12 Gb. With simple generator for TTA this array can be used for prediction with relatively large batchsize (i use 640). In this case single test prediction takes about 80sec on 1060 6GB (I use DenseNet169).",
      "votes": null
    },
    {
      "id": "503843",
      "postDate": "03/30/2019 16:19:26",
      "content": "<p>No, I do not think they are absolutely precise. I took them from my previous project of rectum/ileum inflammation detection. Should be similar :) </p>",
      "rawMarkdown": "No, I do not think they are absolutely precise. I took them from my previous project of rectum/ileum inflammation detection. Should be similar :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 494387,
      "author_name": "kokecacao",
      "author_url": "",
      "post_date": "03/19/2019 19:31:05",
      "content": "<p>Add on to that, I tested serval models: SE nets are generally better than densenets for this dataset.</p>\n\n<p>Also, when you say: <code>dropout of 0.2</code>, do you mean you added dropout in the CNN layers? (because clearly, your linear layers' dropout is 0.8, not 0.2)</p>",
      "votes": null,
      "replies": [
        {
          "id": 494390,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/19/2019 19:32:49",
          "content": "<p>I second that \" SE nets are generally better than densenets for this dataset\".</p>\n\n<p>\"Also, when you say: dropout of 0.2, do you mean you added dropout in the CNN layers?\" yes, exactly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494395,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/19/2019 19:34:59",
          "content": "<p>Changed the post so that it's more clear.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 494548,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "03/20/2019 01:36:12",
      "content": "<p>Hi, Ivan!! Really good results with only 2 models, may I ask why don't you use any activation function in the last linear layers?</p>",
      "votes": null,
      "replies": [
        {
          "id": 494556,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 01:55:42",
          "content": "<p>It's kind of a typo. I do use activation functions of course. I use Relu + log softmax at the end. Updated the post. Thanks for noticing! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 494616,
      "author_name": "seefun",
      "author_url": "",
      "post_date": "03/20/2019 03:41:15",
      "content": "<p>Why not try removing BN in the full connection layer? BN may not good in theory. Using SELU rather than BN+RELU in fc layers may have better result. <a href=\"https://arxiv.org/pdf/1706.02515.pdf\">SELU</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 494621,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 03:45:39",
          "content": "<p>Hm, interesting. Never read that paper. Definitely will take a look. Thanks.</p>\n\n<p>But as far as a know, using BN with Relu usually gives better results that just Relu. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 494690,
      "author_name": "robotdreams",
      "author_url": "",
      "post_date": "03/20/2019 06:22:57",
      "content": "<p>Hello Ivan,\nthank you for contribution! Do you use 15 random TTA according to your code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 494693,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 06:29:40",
          "content": "<p>Hi! You're welcome. </p>\n\n<p>Yes, of course. It's really slow (~15 times slower, than it could be :) ), but my experiments show that it helps getting a better performance, so it's worth it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494822,
          "author_name": "robotdreams",
          "author_url": "",
          "post_date": "03/20/2019 09:39:51",
          "content": "<p>Hi, </p>\n\n<p>do you take then simple mean average of 15 TTA?\nDo you also apply augmentation to the validtation set?</p>\n\n<p>Dima</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494826,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/20/2019 09:44:31",
          "content": "<p>Yes and yes </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 495014,
      "author_name": "samithuang",
      "author_url": "",
      "post_date": "03/20/2019 14:42:10",
      "content": "<p>Incorrect order? Actually, in most practice,  batch normalization should be performed right after FC/Conv instead of after ReLU. Although, some may place BN before FC/Conv. check this paper: Identity Mappings in Deep Residual Networks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 496817,
      "author_name": "mmmmarco",
      "author_url": "",
      "post_date": "03/22/2019 16:08:15",
      "content": "<p>Hello Ivan,\nthank you for the very nice explanation. When you say that you break the 0.98 landmark with 2 models, do you mean that you average the predictions obtained with two different models? If so, how are the models differing from each other (different architectures, different hyperparams, etc.)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496835,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/22/2019 16:33:29",
          "content": "<p>Hi! You're welcome. </p>\n\n<p>Yes, I mean that I average the predictions. And those models don't really differ that much. Basically, there are 2 differences: they are trained on different subsets of WSIs (check my other discussion: <a href=\"https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\">https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132</a> ) and they differ due to the random initialization of weigths. In other words, in terms of achitectures, hyperparameters, etc they are the same. </p>\n\n<p>P.S. Obviously, I will build an ensemble with other models. Something like densenet or se_densenet. I will try to update the post when I get the data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 497709,
      "author_name": "dolotovevgeniy",
      "author_url": "",
      "post_date": "03/23/2019 20:24:40",
      "content": "<p>Hello, Ivan. Thank you for contribution!\nDo you use l2 regularization?</p>",
      "votes": null,
      "replies": [
        {
          "id": 498989,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/24/2019 07:01:49",
          "content": "<p>Hi! You're welcome. </p>\n\n<p>No, I don't. Mostly because (as far as I know) it's not a wise idea to use L2 and dropout together. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 500687,
      "author_name": "chenzijian999",
      "author_url": "",
      "post_date": "03/26/2019 12:23:39",
      "content": "<p><code>transforms.Normalize( mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])</code> \nthe mean - [0.485, 0.456, 0.406] and std - [0.229, 0.224, 0.225] is right ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 500847,
          "author_name": "ivanpan",
          "author_url": "",
          "post_date": "03/26/2019 16:07:02",
          "content": "<p>Yes, of course. Those are the mean and std values for the ImageNet dataset </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 503843,
          "author_name": "sermakarevich",
          "author_url": "",
          "post_date": "03/30/2019 16:19:26",
          "content": "<p>No, I do not think they are absolutely precise. I took them from my previous project of rectum/ileum inflammation detection. Should be similar :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 503776,
      "author_name": "ajalnine",
      "author_url": "",
      "post_date": "03/30/2019 14:41:03",
      "content": "<p>\"The training takes about 6-7 hours, the testing — another 50-60 minutes.\"\nAll test data can be stored as single numpy array about 12 Gb. With simple generator for TTA this array can be used for prediction with relatively large batchsize (i use 640). In this case single test prediction takes about 80sec on 1060 6GB (I use DenseNet169). </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "494356": "1) You should split data according to WSIs. This idea belongs to SM from https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/83760\n\nYou can get a WSI split from my another discussion: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132\n\nIt will allow you to have a more reliable metric. In other words, you won't be getting 0.99+ on validation, but barely 0.96 on LB.\n\n2) Use TTA. I tested augmentations from SM's discussion and they really work well. Here they are: \n\n```\ntransforms.Compose([\n    transforms.Resize((size, size)),\n    transforms.RandomChoice([\n        transforms.ColorJitter(brightness=0.5),\n        transforms.ColorJitter(contrast=0.5), \n        transforms.ColorJitter(saturation=0.5),\n        transforms.ColorJitter(hue=0.5),\n        transforms.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1, hue=0.1), \n        transforms.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.3), \n        transforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5), \n    ]),\n    transforms.RandomChoice([\n        transforms.RandomRotation((0,0)),\n        transforms.RandomHorizontalFlip(p=1),\n        transforms.RandomVerticalFlip(p=1),\n        transforms.RandomRotation((90,90)),\n        transforms.RandomRotation((180,180)),\n        transforms.RandomRotation((270,270)),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((90,90)),\n        ]),\n        transforms.Compose([\n            transforms.RandomHorizontalFlip(p=1),\n            transforms.RandomRotation((270,270)),\n        ]) \n    ]),\n    transforms.ToTensor(),\n    transforms.Normalize(\n        mean=[0.485, 0.456, 0.406],\n        std=[0.229, 0.224, 0.225]\n    )\n])\n```\n\nIn order to make it work you can check my discussion regarding TTA in PyTorch (https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84056 ). In this case, you will have about 15 different versions of original image, so the training will become much slower. But it's worth it. \n\nObviously, you should apply the same transformations to your CV set. \n\n3) I tested different architectures. For me, the best one so far is a pretrained se_resnet50 with dropout of 0.2 in conv layers, concatenation of average and max pooling (both global of course), several fully connected layers with intense dropout. The last idea also belongs to SM. \n\n2 poolings:\n```\nx1 = self.avg_pool(x)\nx2 = self.max_pool(x)\nx = torch.cat([x1,x2], 1)\n```\n\nLinear layers:\n\n```\nself.net.last_linear = nn.Sequential(nn.BatchNorm1d(4096, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(4096, 768, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(768, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(768, 256, bias=True),\n                      nn.ReLU(inplace=True),\n                      nn.BatchNorm1d(256, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True),\n                      nn.Dropout(p=0.8),\n                      nn.Linear(256, 2))\n```\n\n4) I resize images to 196x196 and use Adam with LR of 0.0007. Batch size is 150 for training and 10 for cv and test. If I change it to something bigger, I run out of cuda memory (Tesla K80)\n\n5) Save model every time AUROC is increased. By the end of the training save the best model. \n\n6) Apply ReduceLROnPlateau with the patience of 1-2 epocs.\n\n7) Train the added layers for 1-2 epoch and then the whole network for another 3-4. \n\n8) If you're using kaggle kernels, then train with one kernel and test with another. Otherwise, you will run out of time. The training takes about 6-7 hours, the testing — another 50-60 minutes.\n\nThe single best model performs 0.9786 on LB. I haven't tested ensembles thoroughly, but I had no trouble breaking the 0.98 landmark with just 2 models. Will see how well it performs with 5 or 10 models.  \n\nIf you have any questions, don't hesitate to ask.",
    "494387": "Add on to that, I tested serval models: SE nets are generally better than densenets for this dataset.\n\nAlso, when you say: `dropout of 0.2`, do you mean you added dropout in the CNN layers? (because clearly, your linear layers' dropout is 0.8, not 0.2)",
    "494390": "I second that \" SE nets are generally better than densenets for this dataset\".\n\n\"Also, when you say: dropout of 0.2, do you mean you added dropout in the CNN layers?\" yes, exactly",
    "494395": "Changed the post so that it's more clear.",
    "494548": "Hi, Ivan!! Really good results with only 2 models, may I ask why don't you use any activation function in the last linear layers?",
    "494556": "It's kind of a typo. I do use activation functions of course. I use Relu + log softmax at the end. Updated the post. Thanks for noticing!",
    "494616": "Why not try removing BN in the full connection layer? BN may not good in theory. Using SELU rather than BN+RELU in fc layers may have better result. [SELU](https://arxiv.org/pdf/1706.02515.pdf)",
    "494621": "Hm, interesting. Never read that paper. Definitely will take a look. Thanks.\n\nBut as far as a know, using BN with Relu usually gives better results that just Relu.",
    "494690": "Hello Ivan,\nthank you for contribution! Do you use 15 random TTA according to your code?",
    "494693": "Hi! You're welcome. \n\nYes, of course. It's really slow (~15 times slower, than it could be :) ), but my experiments show that it helps getting a better performance, so it's worth it.",
    "494822": "Hi, \n\ndo you take then simple mean average of 15 TTA?\nDo you also apply augmentation to the validtation set?\n\nDima",
    "494826": "Yes and yes",
    "495014": "Incorrect order? Actually, in most practice,  batch normalization should be performed right after FC/Conv instead of after ReLU. Although, some may place BN before FC/Conv. check this paper: Identity Mappings in Deep Residual Networks",
    "496817": "Hello Ivan,\nthank you for the very nice explanation. When you say that you break the 0.98 landmark with 2 models, do you mean that you average the predictions obtained with two different models? If so, how are the models differing from each other (different architectures, different hyperparams, etc.)?",
    "496835": "Hi! You're welcome. \n\nYes, I mean that I average the predictions. And those models don't really differ that much. Basically, there are 2 differences: they are trained on different subsets of WSIs (check my other discussion: https://www.kaggle.com/c/histopathologic-cancer-detection/discussion/84132 ) and they differ due to the random initialization of weigths. In other words, in terms of achitectures, hyperparameters, etc they are the same. \n\nP.S. Obviously, I will build an ensemble with other models. Something like densenet or se_densenet. I will try to update the post when I get the data.",
    "497709": "Hello, Ivan. Thank you for contribution!\nDo you use l2 regularization?",
    "498989": "Hi! You're welcome. \n\nNo, I don't. Mostly because (as far as I know) it's not a wise idea to use L2 and dropout together.",
    "500687": "`transforms.Normalize( mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])` \nthe mean - [0.485, 0.456, 0.406] and std - [0.229, 0.224, 0.225] is right ?",
    "500847": "Yes, of course. Those are the mean and std values for the ImageNet dataset",
    "503776": "\"The training takes about 6-7 hours, the testing — another 50-60 minutes.\"\nAll test data can be stored as single numpy array about 12 Gb. With simple generator for TTA this array can be used for prediction with relatively large batchsize (i use 640). In this case single test prediction takes about 80sec on 1060 6GB (I use DenseNet169).",
    "503843": "No, I do not think they are absolutely precise. I took them from my previous project of rectum/ileum inflammation detection. Should be similar :)"
  },
  "source": "meta"
}