{
  "id": 81747,
  "title": "How to get 0.9788",
  "url": "/competitions/histopathologic-cancer-detection/discussion/81747",
  "author_name": "Alex Donchuk",
  "post_date": "2019-02-24T16:37:52.233000",
  "votes": 40,
  "comment_count": 70,
  "views": 0,
  "content": "<p>Hi! Everybody. Steps to get the score:</p>\n\n<ol>\n<li>Use se_resnext50 without maxpool. Imagenet </li>\n<li>Create ~20 folds</li>\n<li>User 4d TTA</li>\n<li>After 7 epoch start to validate often than 1 per epoch, after 10-12 epoch start to validate more often.(3-5 times) Find optimum between train, val, and roc. Don't overfit.</li>\n<li>Use lr = 0.00007</li>\n<li>Use batch 256(din't check if it important)</li>\n<li>ReduceOnPlateu - 2 epoch</li>\n<li>Simple np.mean() ensemble 5+ folds will give you 0.9780+</li>\n</ol>\n\n<p>p.s. Augs using albumentions\n        return Compose([\n            RandomRotate90(p=0.5),\n            Transpose(p=0.5),\n            Flip(p=0.5),\n            OneOf(\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomBrightnessContrast(),\n                JpegCompression(),\n                Blur(),\n                GaussNoise()\n            , p=0.5),\n            HueSaturationValue(p=0.5),\n            ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n            Normalize(p=1)</p>",
  "messages": [
    {
      "id": 477464,
      "postDate": "2019-02-24T16:37:52.233Z",
      "content": "<p>Hi! Everybody. Steps to get the score:</p>\n\n<ol>\n<li>Use se_resnext50 without maxpool. Imagenet </li>\n<li>Create ~20 folds</li>\n<li>User 4d TTA</li>\n<li>After 7 epoch start to validate often than 1 per epoch, after 10-12 epoch start to validate more often.(3-5 times) Find optimum between train, val, and roc. Don't overfit.</li>\n<li>Use lr = 0.00007</li>\n<li>Use batch 256(din't check if it important)</li>\n<li>ReduceOnPlateu - 2 epoch</li>\n<li>Simple np.mean() ensemble 5+ folds will give you 0.9780+</li>\n</ol>\n\n<p>p.s. Augs using albumentions\n        return Compose([\n            RandomRotate90(p=0.5),\n            Transpose(p=0.5),\n            Flip(p=0.5),\n            OneOf(\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomBrightnessContrast(),\n                JpegCompression(),\n                Blur(),\n                GaussNoise()\n            , p=0.5),\n            HueSaturationValue(p=0.5),\n            ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n            Normalize(p=1)</p>",
      "rawMarkdown": "Hi! Everybody. Steps to get the score:\n\n1. Use se_resnext50 without maxpool. Imagenet \n2. Create ~20 folds\n3. User 4d TTA\n4. After 7 epoch start to validate often than 1 per epoch, after 10-12 epoch start to validate more often.(3-5 times) Find optimum between train, val, and roc. Don't overfit.\n5. Use lr = 0.00007\n6. Use batch 256(din't check if it important)\n7. ReduceOnPlateu - 2 epoch\n8. Simple np.mean() ensemble 5+ folds will give you 0.9780+\n\np.s. Augs using albumentions\n        return Compose([\n            RandomRotate90(p=0.5),\n            Transpose(p=0.5),\n            Flip(p=0.5),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomBrightnessContrast(),\n                JpegCompression(),\n                Blur(),\n                GaussNoise()\n            ], p=0.5),\n            HueSaturationValue(p=0.5),\n            ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n            Normalize(p=1)\n",
      "votes": 39
    },
    {
      "id": 1715411,
      "postDate": "2022-03-08T01:16:35.510Z",
      "content": "<p>Hi friend!, there is only one thing I don't understand, you are using pre-trained ResNet150, however the imagenet dataset wasn't trained using these images/dataset, I mean you are using loading the pre-trained images with imagenet weights but training again for this data and reaching such results? Please, I hope not to bother, but I can't understand how you work with it. <br>\nThanks.</p>",
      "rawMarkdown": "Hi friend!, there is only one thing I don't understand, you are using pre-trained ResNet150, however the imagenet dataset wasn't trained using these images/dataset, I mean you are using loading the pre-trained images with imagenet weights but training again for this data and reaching such results? Please, I hope not to bother, but I can't understand how you work with it. \nThanks.",
      "votes": 2
    },
    {
      "id": 477767,
      "postDate": "2019-02-25T08:26:42.747Z",
      "content": "<p>Almost same for me, but I have only 5 folds, 12 random tta images, pretrain net on imagenet, starting with lr=3e-5 for last layer and 3e-5*0.001 for pretrained ones, batch of 64 images rescaled to 224x224 + albumentions  + mean of folds predictions </p>",
      "rawMarkdown": "Almost same for me, but I have only 5 folds, 12 random tta images, pretrain net on imagenet, starting with lr=3e-5 for last layer and 3e-5*0.001 for pretrained ones, batch of 64 images rescaled to 224x224 + albumentions  + mean of folds predictions ",
      "votes": 5,
      "replies": [
        {
          "id": 478040,
          "postDate": "2019-02-25T16:33:31.123Z",
          "content": "<p>Good job!</p>",
          "rawMarkdown": "Good job!"
        },
        {
          "id": 478044,
          "postDate": "2019-02-25T16:41:40.223Z",
          "content": "<p>@Kvigly are you using pytorch as well? @Alex are you using fast.ai library? </p>",
          "rawMarkdown": "@Kvigly are you using pytorch as well? @Alex are you using fast.ai library? ",
          "votes": 1
        },
        {
          "id": 478046,
          "postDate": "2019-02-25T16:45:23.733Z",
          "content": "<p>No, pure pytorch 1.0.1. </p>",
          "rawMarkdown": "No, pure pytorch 1.0.1. "
        },
        {
          "id": 478246,
          "postDate": "2019-02-25T23:35:54.667Z",
          "content": "<p>I am using pytorch as well </p>",
          "rawMarkdown": "I am using pytorch as well "
        },
        {
          "id": 478933,
          "postDate": "2019-02-26T20:12:34.627Z",
          "content": "<p>@kvigly: are you training the pre-trained network with a starting lr of 3e-5*0.001 = 3e-8? I'm just a little confused about the notation you used</p>",
          "rawMarkdown": "@kvigly: are you training the pre-trained network with a starting lr of 3e-5*0.001 = 3e-8? I'm just a little confused about the notation you used"
        },
        {
          "id": 478942,
          "postDate": "2019-02-26T20:29:18.060Z",
          "content": "<p>@Franchini\nI take pretrained seresnet50 \\ resnet34 \\ resnet50, replace last layer (because we have only one class). So we have \"random\" initialization only for last layer, so that, last layer, we train with lr=3e-5. Previous layers are pretrained. it has been shown that a good technique here is to freeze them for 1-2 epoch before training - to start training only the last layer. But in my case I still train them, just with lr waaaay smaller, than for last layer. I gradually increase that lr until lr for \"pretrained\" and last layer matches. \nRoughly like this \nepoch | lr last layer | lr pretrained layers\n1 | 3e-5 | 3e-8\n2 | 3e-5 | 1e-7\n3 | 3e-5 | 5e-7\n4 | 3e-5 | 1e-6\netc. Kind of inertia of pretrained layers\n(it is kind of more complicated, since resnet has several blocks, so each block has its own lr, starting from smaller is the very first layers (since they probably learned very common and basic patterns, which are useful perhaps even here) to bigger in latter ones (because imagenet pictures are quite different, higher layers learned probably useless patterns, so we want to \"forget them\"), I will maybe write a summary after this competition to check how much that added and for which models)</p>",
          "rawMarkdown": "@Franchini\nI take pretrained seresnet50 \\ resnet34 \\ resnet50, replace last layer (because we have only one class). So we have \"random\" initialization only for last layer, so that, last layer, we train with lr=3e-5. Previous layers are pretrained. it has been shown that a good technique here is to freeze them for 1-2 epoch before training - to start training only the last layer. But in my case I still train them, just with lr waaaay smaller, than for last layer. I gradually increase that lr until lr for \"pretrained\" and last layer matches. \nRoughly like this \nepoch | lr last layer | lr pretrained layers\n1 | 3e-5 | 3e-8\n2 | 3e-5 | 1e-7\n3 | 3e-5 | 5e-7\n4 | 3e-5 | 1e-6\netc. Kind of inertia of pretrained layers\n(it is kind of more complicated, since resnet has several blocks, so each block has its own lr, starting from smaller is the very first layers (since they probably learned very common and basic patterns, which are useful perhaps even here) to bigger in latter ones (because imagenet pictures are quite different, higher layers learned probably useless patterns, so we want to \"forget them\"), I will maybe write a summary after this competition to check how much that added and for which models)\n",
          "votes": 4
        },
        {
          "id": 478957,
          "postDate": "2019-02-26T20:57:50.997Z",
          "content": "<p>Thanks so much for taking the time to explain this!! This reminds me of a diffusive process (of information in this case). It would be very interesting to look at the type of curve for the learning rate along the network as a hyperparameter. Could be linear, exponential...\nReally looks like I have to make the move from Keras to Pytorch to play with this a bit.</p>",
          "rawMarkdown": "Thanks so much for taking the time to explain this!! This reminds me of a diffusive process (of information in this case). It would be very interesting to look at the type of curve for the learning rate along the network as a hyperparameter. Could be linear, exponential...\nReally looks like I have to make the move from Keras to Pytorch to play with this a bit."
        },
        {
          "id": 488723,
          "postDate": "2019-03-12T23:28:28.457Z",
          "content": "<p>Does rescaling the images improve the accuracy? If the images are 96x96. Why does making them larger work?</p>",
          "rawMarkdown": "Does rescaling the images improve the accuracy? If the images are 96x96. Why does making them larger work?"
        },
        {
          "id": 488727,
          "postDate": "2019-03-12T23:45:27.490Z",
          "content": "<p>I think the rescaling is done so the images fit into the model, which was pre-trained on Imagenet</p>",
          "rawMarkdown": "I think the rescaling is done so the images fit into the model, which was pre-trained on Imagenet"
        },
        {
          "id": 489022,
          "postDate": "2019-03-13T11:37:50.843Z",
          "content": "<p>Thanks Franchini</p>",
          "rawMarkdown": "Thanks Franchini"
        }
      ]
    },
    {
      "id": 506985,
      "postDate": "2019-04-04T05:54:08.860Z",
      "content": "<p>Thanks for your advice, Alex.  It really helped me a lot.</p>",
      "rawMarkdown": "Thanks for your advice, Alex.  It really helped me a lot.",
      "votes": 1
    },
    {
      "id": 485937,
      "postDate": "2019-03-08T05:40:12.750Z",
      "content": "<p>can this be achieved without a pre-trained network?</p>",
      "rawMarkdown": "can this be achieved without a pre-trained network?",
      "votes": 1,
      "replies": [
        {
          "id": 485985,
          "postDate": "2019-03-08T06:57:31.943Z",
          "content": "<p>I don't know, but I think you can.</p>",
          "rawMarkdown": "I don't know, but I think you can."
        },
        {
          "id": 486118,
          "postDate": "2019-03-08T10:40:14.320Z",
          "content": "<p>I tried this a long time ago but did not get a very good score on lb. it was ~094ish I think.</p>",
          "rawMarkdown": "I tried this a long time ago but did not get a very good score on lb. it was ~094ish I think.",
          "votes": 1
        }
      ]
    },
    {
      "id": 484229,
      "postDate": "2019-03-05T18:05:15.317Z",
      "content": "<p>Sorry for the dumb question, but how should one remove the maxpool? I tried just deleting the layer and it’s forward but that did not seem to work. I also tried replacing just the maxpool with identity but that seemed not to work also. </p>\n\n<p>Do you mean just the one maxpool after the first convolution in seresnext50? I’ll keep googling until I figure it out. Can you advise on the proper way to remove or change a layer in a pytorch pretrained model, If you don’t mind? Thanks for posting this!</p>",
      "rawMarkdown": "Sorry for the dumb question, but how should one remove the maxpool? I tried just deleting the layer and it’s forward but that did not seem to work. I also tried replacing just the maxpool with identity but that seemed not to work also. \n\nDo you mean just the one maxpool after the first convolution in seresnext50? I’ll keep googling until I figure it out. Can you advise on the proper way to remove or change a layer in a pytorch pretrained model, If you don’t mind? Thanks for posting this!",
      "votes": 1,
      "replies": [
        {
          "id": 484999,
          "postDate": "2019-03-06T18:56:29.570Z",
          "content": "<p>As there is only one maxpool layer in se_resnet50, I think the OP is talking about that only and Yeah, Directly removing the maxpool layer when the model was pretrained on imagenet (with the maxpool layer) <em>may</em> impact performance. If we were to remove the maxpool layer (say), we'll I either have to feed in 3x112x112 sized image or make some modifications in the FC layers for input shape 3x224x224</p>",
          "rawMarkdown": "As there is only one maxpool layer in se_resnet50, I think the OP is talking about that only and Yeah, Directly removing the maxpool layer when the model was pretrained on imagenet (with the maxpool layer) *may* impact performance. If we were to remove the maxpool layer (say), we'll I either have to feed in 3x112x112 sized image or make some modifications in the FC layers for input shape 3x224x224",
          "votes": 1
        },
        {
          "id": 485156,
          "postDate": "2019-03-07T02:18:50.343Z",
          "content": "<p>I'm using se_resNext50 and didn't resize image after deleting max pool. </p>",
          "rawMarkdown": "I'm using se_resNext50 and didn't resize image after deleting max pool. ",
          "votes": 1
        },
        {
          "id": 485187,
          "postDate": "2019-03-07T03:31:02.423Z",
          "content": "<p>Hey Alex, </p>\n\n<p>The last two layers of se_resNext50 are <code>AvgPool2d-269</code> and <code>Linear-270</code>, now if you use the standard network (with the maxpool layer) with input image of 3x224x224, the output of AvgPool2d layer is <code>[batch_size, 2048, 1, 1]</code> and the Linear layer is defined to take in 2048 features so it works. </p>\n\n<p>Now, if you remove the maxpool layer and feed in 3x224x224 image the output of AvgPool2d is <code>[batch_size, 2048, 8, 8]</code> which will throw a shape mismatch error until we modify the Linear layer to take in that many feature size. </p>\n\n<p>What input image size are you using?</p>",
          "rawMarkdown": "Hey Alex, \n\nThe last two layers of se_resNext50 are `AvgPool2d-269` and `Linear-270`, now if you use the standard network (with the maxpool layer) with input image of 3x224x224, the output of AvgPool2d layer is `[batch_size, 2048, 1, 1] ` and the Linear layer is defined to take in 2048 features so it works. \n\nNow, if you remove the maxpool layer and feed in 3x224x224 image the output of AvgPool2d is `[batch_size, 2048, 8, 8]` which will throw a shape mismatch error until we modify the Linear layer to take in that many feature size. \n\nWhat input image size are you using?",
          "votes": 1
        },
        {
          "id": 485982,
          "postDate": "2019-03-08T06:53:14.770Z",
          "content": "<p>96х96</p>",
          "rawMarkdown": "96х96",
          "votes": 1
        },
        {
          "id": 486593,
          "postDate": "2019-03-09T03:02:59.377Z",
          "content": "<p>Change avg pooling to adaptive avg</p>",
          "rawMarkdown": "Change avg pooling to adaptive avg",
          "votes": 1
        }
      ]
    },
    {
      "id": 484118,
      "postDate": "2019-03-05T15:40:26.423Z",
      "content": "<p>thanks for sharing your thoughts :) ! keep up the good job!</p>",
      "rawMarkdown": "thanks for sharing your thoughts :) ! keep up the good job!",
      "votes": 1
    },
    {
      "id": 478688,
      "postDate": "2019-02-26T13:42:53.490Z",
      "content": "<p>Thanks for sharing, Alex! Great job! Can I ask you which optimizer you used?</p>",
      "rawMarkdown": "Thanks for sharing, Alex! Great job! Can I ask you which optimizer you used?",
      "votes": 1,
      "replies": [
        {
          "id": 478749,
          "postDate": "2019-02-26T15:20:14.717Z",
          "content": "<p>Hi, Franchini! rmsprop but SGD, ADAM works the same here.</p>",
          "rawMarkdown": "Hi, Franchini! rmsprop but SGD, ADAM works the same here.",
          "votes": 1
        }
      ]
    },
    {
      "id": 485849,
      "postDate": "2019-03-08T01:43:56.830Z",
      "content": "<p>Thanks for sharing! I don't know if I understand it wrong.\nif you use ~20 folds, it means you will trained 20 model?\nand you use 4d TTA,  so each test images you will predict 20*4 times  and simple np.mean() it?</p>",
      "rawMarkdown": "Thanks for sharing! I don't know if I understand it wrong.\nif you use ~20 folds, it means you will trained 20 model?\nand you use 4d TTA,  so each test images you will predict 20*4 times  and simple np.mean() it?",
      "votes": 2,
      "replies": [
        {
          "id": 485984,
          "postDate": "2019-03-08T06:56:35.603Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes"
        },
        {
          "id": 486193,
          "postDate": "2019-03-08T12:13:18.583Z",
          "content": "<p>thank you！ i will try it</p>",
          "rawMarkdown": "thank you！ i will try it"
        }
      ]
    },
    {
      "id": 482612,
      "postDate": "2019-03-03T11:16:58.807Z",
      "content": "<p>Nice, thanks for sharing. What is your validation score ? \nI can easily achieve 0.996+ score on every cv fold (10)  vs public lb score ~0.967. Changes on CV score does not correlate with lb score. </p>",
      "rawMarkdown": "Nice, thanks for sharing. What is your validation score ? \nI can easily achieve 0.996+ score on every cv fold (10)  vs public lb score ~0.967. Changes on CV score does not correlate with lb score. ",
      "votes": 2,
      "replies": [
        {
          "id": 482965,
          "postDate": "2019-03-03T23:31:28.503Z",
          "content": "<p>The same for me. =) </p>",
          "rawMarkdown": "The same for me. =) "
        }
      ]
    },
    {
      "id": 2710532,
      "postDate": "2024-03-22T10:42:19.037Z",
      "content": "<p>need label test data csv file</p>",
      "rawMarkdown": "need label test data csv file"
    },
    {
      "id": 1105418,
      "postDate": "2020-12-07T21:12:56.010Z",
      "content": "<p><a href=\"https://www.kaggle.com/donchuk\" target=\"_blank\">@donchuk</a>  Alex, how much increase did image augmentation increase the accuracy?</p>\n<p>Thanks for the help and good job!</p>",
      "rawMarkdown": "@donchuk  Alex, how much increase did image augmentation increase the accuracy?\n\nThanks for the help and good job!"
    },
    {
      "id": 1095280,
      "postDate": "2020-11-29T12:42:11.170Z",
      "content": "<p>bro?<br>\nhow to calculate clahe manually?</p>",
      "rawMarkdown": "bro?\nhow to calculate clahe manually?"
    },
    {
      "id": 483853,
      "postDate": "2019-03-05T08:50:49.923Z",
      "content": "<p>what means User 4d TTA?</p>",
      "rawMarkdown": "what means User 4d TTA?",
      "replies": [
        {
          "id": 483877,
          "postDate": "2019-03-05T09:28:37.543Z",
          "content": "<p>I assume this is a typo... should be <code>use 4d TTA</code> ?\n<code>TTA</code> stands for <code>Test Time Augmentation</code>\nHowever i don't understand either the meaning of <code>4d</code> ;-)</p>",
          "rawMarkdown": "I assume this is a typo... should be `use 4d TTA` ?\n`TTA` stands for `Test Time Augmentation`\nHowever i don't understand either the meaning of `4d` ;-)"
        },
        {
          "id": 484010,
          "postDate": "2019-03-05T12:58:00.560Z",
          "content": "<p>Flips, Transpose, Rotations and its combination. </p>",
          "rawMarkdown": "Flips, Transpose, Rotations and its combination. ",
          "votes": 2
        },
        {
          "id": 484461,
          "postDate": "2019-03-06T03:33:19.263Z",
          "content": "<p>oh so you mean applying four augmentations in test time and vote to get the final class?</p>",
          "rawMarkdown": "oh so you mean applying four augmentations in test time and vote to get the final class?",
          "votes": 1
        },
        {
          "id": 485154,
          "postDate": "2019-03-07T02:16:58.093Z",
          "content": "<p>just simple mean</p>",
          "rawMarkdown": "just simple mean"
        },
        {
          "id": 485690,
          "postDate": "2019-03-07T19:42:38.973Z",
          "content": "<p>Hey Alex, do we apply 4d during test time? Because the model has been trained with a rigorous augmentation pipeline?</p>",
          "rawMarkdown": "Hey Alex, do we apply 4d during test time? Because the model has been trained with a rigorous augmentation pipeline?"
        },
        {
          "id": 485983,
          "postDate": "2019-03-08T06:56:12.117Z",
          "content": "<p>4d during test time - yes! </p>",
          "rawMarkdown": "4d during test time - yes! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 478331,
      "postDate": "2019-02-26T02:59:30.933Z",
      "content": "<p>Thanks for sharing. I am using kinda the same however have not yet reached .9780+. Have you cropped or resized the image? </p>",
      "rawMarkdown": "Thanks for sharing. I am using kinda the same however have not yet reached .9780+. Have you cropped or resized the image? ",
      "replies": [
        {
          "id": 478538,
          "postDate": "2019-02-26T09:55:40.977Z",
          "content": "<p>I use origin size. And I add some folds today and get 0.9792. </p>",
          "rawMarkdown": "I use origin size. And I add some folds today and get 0.9792. "
        },
        {
          "id": 478542,
          "postDate": "2019-02-26T10:00:35.847Z",
          "content": "<p>Additionally you can add pseudo labels.</p>",
          "rawMarkdown": "Additionally you can add pseudo labels."
        },
        {
          "id": 478978,
          "postDate": "2019-02-26T21:41:55.163Z",
          "content": "<p>Thanks for sharing your ideas! Good job of you! </p>",
          "rawMarkdown": "Thanks for sharing your ideas! Good job of you! ",
          "votes": 1
        },
        {
          "id": 483882,
          "postDate": "2019-03-05T09:33:53.307Z",
          "content": "<p>Can you elaborate more on <code>pseudo labels</code> in this context?\nDo you mean <code>Semi-Supervised Learning</code> ? \nBut this would require more unlabeled data?</p>",
          "rawMarkdown": "Can you elaborate more on `pseudo labels` in this context?\nDo you mean `Semi-Supervised Learning` ? \nBut this would require more unlabeled data?"
        },
        {
          "id": 484008,
          "postDate": "2019-03-05T12:56:52.900Z",
          "content": "<p>No! Add sample from test to train with very big score. </p>",
          "rawMarkdown": "No! Add sample from test to train with very big score. ",
          "votes": 1
        },
        {
          "id": 484187,
          "postDate": "2019-03-05T17:01:41.570Z",
          "content": "<p>Sorry, but i don't understand ;-(\nWhat do you mean by \"very big score\" ?</p>",
          "rawMarkdown": "Sorry, but i don't understand ;-(\nWhat do you mean by \"very big score\" ?"
        },
        {
          "id": 484192,
          "postDate": "2019-03-05T17:08:30.020Z",
          "content": "<p><a href=\"/donchuk\">@donchuk</a>  \"very big score\" means train a classifier, predict probabilities for the test set and add the images with the score very close to 0 or 1 to the training set and train again????</p>",
          "rawMarkdown": "@donchuk  \"very big score\" means train a classifier, predict probabilities for the test set and add the images with the score very close to 0 or 1 to the training set and train again????"
        },
        {
          "id": 485153,
          "postDate": "2019-03-07T02:16:30.123Z",
          "content": "<p>Yes, try &gt; 0.9999 and 0.0001 &lt;</p>",
          "rawMarkdown": "Yes, try &gt; 0.9999 and 0.0001 &lt;"
        },
        {
          "id": 485509,
          "postDate": "2019-03-07T13:32:22.207Z",
          "content": "<p><a href=\"/donchuk\">@donchuk</a> could you elaborate the motivation behind this strategy? Intuitively, I would have thought that the algorithm is not able to learn a lot from samples which it already perfectly predicts. </p>",
          "rawMarkdown": "@donchuk could you elaborate the motivation behind this strategy? Intuitively, I would have thought that the algorithm is not able to learn a lot from samples which it already perfectly predicts. "
        },
        {
          "id": 485691,
          "postDate": "2019-03-07T19:44:19.393Z",
          "content": "<p>@Franchini we do that so as to increase the training data size, as the model is pretty confident on these images we assume them to be their ground truth labels.</p>",
          "rawMarkdown": "@Franchini we do that so as to increase the training data size, as the model is pretty confident on these images we assume them to be their ground truth labels."
        },
        {
          "id": 485704,
          "postDate": "2019-03-07T20:02:46.700Z",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> \"as the model is pretty confident on these images we assume them to be their ground truth labels\" If the model is already perfectly able to predict those samples, what can it learn from them? Wouldn't we rather want the model to see more samples that it finds harder to predict?</p>",
          "rawMarkdown": "@rishabhiitbhu \"as the model is pretty confident on these images we assume them to be their ground truth labels\" If the model is already perfectly able to predict those samples, what can it learn from them? Wouldn't we rather want the model to see more samples that it finds harder to predict?"
        },
        {
          "id": 485735,
          "postDate": "2019-03-07T20:57:54.327Z",
          "content": "<p>Here is rationale <a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a> </p>",
          "rawMarkdown": "Here is rationale http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf ",
          "votes": 1
        },
        {
          "id": 485757,
          "postDate": "2019-03-07T21:51:28.417Z",
          "content": "<p>Pseudo-labeling is often used and most effective in cases where the number of training examples for a category is small. This technique is a part of several of the top scoring solutions to the whale competition that recently ended, for that one there were 5005 categories many of which only had one example in the training set. Such an imbalance is not the case with the binary classification here, but just because a model can already accurately (we assume) classify some test image does not mean there is not other useful information in those images that could be beneficial in training. To be safe I would only pseudo-label with images that different architectures independently gave high confidence to.</p>",
          "rawMarkdown": "Pseudo-labeling is often used and most effective in cases where the number of training examples for a category is small. This technique is a part of several of the top scoring solutions to the whale competition that recently ended, for that one there were 5005 categories many of which only had one example in the training set. Such an imbalance is not the case with the binary classification here, but just because a model can already accurately (we assume) classify some test image does not mean there is not other useful information in those images that could be beneficial in training. To be safe I would only pseudo-label with images that different architectures independently gave high confidence to.",
          "votes": 1
        },
        {
          "id": 485962,
          "postDate": "2019-03-08T06:29:56.277Z",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> and <a href=\"/sermakarevich\">@sermakarevich</a>, thanks for the background! I was confused because I thought it meant to take images from the test set of the kfolds. Now that I realised that you take images from the unlabeled set and add it to the training set it makes more sense!\nThanks guys, I'm learning a lot in this competition.</p>",
          "rawMarkdown": "@interneuron and @sermakarevich, thanks for the background! I was confused because I thought it meant to take images from the test set of the kfolds. Now that I realised that you take images from the unlabeled set and add it to the training set it makes more sense!\nThanks guys, I'm learning a lot in this competition.",
          "votes": 1
        }
      ]
    },
    {
      "id": 477947,
      "postDate": "2019-02-25T14:42:49.507Z",
      "content": "<p>Can you explain #4 a bit more?</p>",
      "rawMarkdown": "Can you explain #4 a bit more?",
      "replies": [
        {
          "id": 478034,
          "postDate": "2019-02-25T16:28:25.093Z",
          "content": "<p>You need to validate often than once per epoch! Do validation 3-5 times during epoch. Don't overfit. </p>",
          "rawMarkdown": "You need to validate often than once per epoch! Do validation 3-5 times during epoch. Don't overfit. "
        }
      ]
    },
    {
      "id": 477671,
      "postDate": "2019-02-25T04:19:17.423Z",
      "content": "<p>Can you explain #4? Are you using pytorch?</p>",
      "rawMarkdown": "Can you explain #4? Are you using pytorch?",
      "replies": [
        {
          "id": 477697,
          "postDate": "2019-02-25T05:33:49.177Z",
          "content": "<p>Yes, pure pytorch. \nValidate 3-5 times per epoch. Still not clear? </p>",
          "rawMarkdown": "Yes, pure pytorch. \nValidate 3-5 times per epoch. Still not clear? "
        },
        {
          "id": 485687,
          "postDate": "2019-03-07T19:39:31.240Z",
          "content": "<p>Hey Alex, validation for me is to test the model performance on the validation set, What does validate 3-5 times per epoch mean?  </p>",
          "rawMarkdown": "Hey Alex, validation for me is to test the model performance on the validation set, What does validate 3-5 times per epoch mean?  "
        },
        {
          "id": 488875,
          "postDate": "2019-03-13T06:04:31.997Z",
          "content": "<p>Hey Alex, still waiting for your reply :)</p>",
          "rawMarkdown": "Hey Alex, still waiting for your reply :)"
        }
      ]
    },
    {
      "id": 488400,
      "postDate": "2019-03-12T12:04:21.670Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 507023,
          "postDate": "2019-04-04T07:15:22.373Z",
          "content": "<p>+15 places on private</p>",
          "rawMarkdown": "+15 places on private"
        }
      ]
    },
    {
      "id": 478913,
      "postDate": "2019-02-26T19:50:21.493Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 478923,
          "postDate": "2019-02-26T20:05:50.937Z",
          "content": "<p>I asked exactly the same question and he answered it 6h ago ;-)</p>",
          "rawMarkdown": "I asked exactly the same question and he answered it 6h ago ;-)"
        },
        {
          "id": 478950,
          "postDate": "2019-02-26T20:46:57.167Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 482262,
      "postDate": "2019-03-02T15:48:11.363Z",
      "content": "<p>Great guide thanks a lot!</p>",
      "rawMarkdown": "Great guide thanks a lot!",
      "votes": 1
    },
    {
      "id": 480855,
      "postDate": "2019-02-28T19:15:20.527Z",
      "content": "<p>Thank you sooooo much!</p>",
      "rawMarkdown": "Thank you sooooo much!",
      "votes": 1
    },
    {
      "id": 478213,
      "postDate": "2019-02-25T22:17:38.300Z",
      "content": "<p>Wow! thanks a lot!!</p>",
      "rawMarkdown": "Wow! thanks a lot!!",
      "votes": 1
    },
    {
      "id": 477581,
      "postDate": "2019-02-24T22:39:04.287Z",
      "content": "<p>Awesome! Thank you</p>",
      "rawMarkdown": "Awesome! Thank you",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1715411,
      "author_name": "george saavedra",
      "author_url": "",
      "post_date": "2022-03-08T01:16:35.510000",
      "content": "<p>Hi friend!, there is only one thing I don't understand, you are using pre-trained ResNet150, however the imagenet dataset wasn't trained using these images/dataset, I mean you are using loading the pre-trained images with imagenet weights but training again for this data and reaching such results? Please, I hope not to bother, but I can't understand how you work with it. <br>\nThanks.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 477767,
      "author_name": "kvigly",
      "author_url": "",
      "post_date": "2019-02-25T08:26:42.747000",
      "content": "<p>Almost same for me, but I have only 5 folds, 12 random tta images, pretrain net on imagenet, starting with lr=3e-5 for last layer and 3e-5*0.001 for pretrained ones, batch of 64 images rescaled to 224x224 + albumentions  + mean of folds predictions </p>",
      "votes": 5,
      "replies": [
        {
          "id": 478040,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-25T16:33:31.123000",
          "content": "<p>Good job!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478044,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2019-02-25T16:41:40.223000",
          "content": "<p>@Kvigly are you using pytorch as well? @Alex are you using fast.ai library? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 478046,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-25T16:45:23.733000",
          "content": "<p>No, pure pytorch 1.0.1. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478246,
          "author_name": "kvigly",
          "author_url": "",
          "post_date": "2019-02-25T23:35:54.667000",
          "content": "<p>I am using pytorch as well </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478933,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-02-26T20:12:34.627000",
          "content": "<p>@kvigly: are you training the pre-trained network with a starting lr of 3e-5*0.001 = 3e-8? I'm just a little confused about the notation you used</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478942,
          "author_name": "kvigly",
          "author_url": "",
          "post_date": "2019-02-26T20:29:18.060000",
          "content": "<p>@Franchini\nI take pretrained seresnet50 \\ resnet34 \\ resnet50, replace last layer (because we have only one class). So we have \"random\" initialization only for last layer, so that, last layer, we train with lr=3e-5. Previous layers are pretrained. it has been shown that a good technique here is to freeze them for 1-2 epoch before training - to start training only the last layer. But in my case I still train them, just with lr waaaay smaller, than for last layer. I gradually increase that lr until lr for \"pretrained\" and last layer matches. \nRoughly like this \nepoch | lr last layer | lr pretrained layers\n1 | 3e-5 | 3e-8\n2 | 3e-5 | 1e-7\n3 | 3e-5 | 5e-7\n4 | 3e-5 | 1e-6\netc. Kind of inertia of pretrained layers\n(it is kind of more complicated, since resnet has several blocks, so each block has its own lr, starting from smaller is the very first layers (since they probably learned very common and basic patterns, which are useful perhaps even here) to bigger in latter ones (because imagenet pictures are quite different, higher layers learned probably useless patterns, so we want to \"forget them\"), I will maybe write a summary after this competition to check how much that added and for which models)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 478957,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-02-26T20:57:50.997000",
          "content": "<p>Thanks so much for taking the time to explain this!! This reminds me of a diffusive process (of information in this case). It would be very interesting to look at the type of curve for the learning rate along the network as a hyperparameter. Could be linear, exponential...\nReally looks like I have to make the move from Keras to Pytorch to play with this a bit.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488723,
          "author_name": "Mauro",
          "author_url": "",
          "post_date": "2019-03-12T23:28:28.457000",
          "content": "<p>Does rescaling the images improve the accuracy? If the images are 96x96. Why does making them larger work?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488727,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-03-12T23:45:27.490000",
          "content": "<p>I think the rescaling is done so the images fit into the model, which was pre-trained on Imagenet</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 489022,
          "author_name": "Mauro",
          "author_url": "",
          "post_date": "2019-03-13T11:37:50.843000",
          "content": "<p>Thanks Franchini</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 506985,
      "author_name": "RobinDong",
      "author_url": "",
      "post_date": "2019-04-04T05:54:08.860000",
      "content": "<p>Thanks for your advice, Alex.  It really helped me a lot.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 485937,
      "author_name": "Bogdan Barabanshchikov",
      "author_url": "",
      "post_date": "2019-03-08T05:40:12.750000",
      "content": "<p>can this be achieved without a pre-trained network?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 485985,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-08T06:57:31.943000",
          "content": "<p>I don't know, but I think you can.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 486118,
          "author_name": "Amirreza Mahbod",
          "author_url": "",
          "post_date": "2019-03-08T10:40:14.320000",
          "content": "<p>I tried this a long time ago but did not get a very good score on lb. it was ~094ish I think.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 484229,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-03-05T18:05:15.317000",
      "content": "<p>Sorry for the dumb question, but how should one remove the maxpool? I tried just deleting the layer and it’s forward but that did not seem to work. I also tried replacing just the maxpool with identity but that seemed not to work also. </p>\n\n<p>Do you mean just the one maxpool after the first convolution in seresnext50? I’ll keep googling until I figure it out. Can you advise on the proper way to remove or change a layer in a pytorch pretrained model, If you don’t mind? Thanks for posting this!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 484999,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-06T18:56:29.570000",
          "content": "<p>As there is only one maxpool layer in se_resnet50, I think the OP is talking about that only and Yeah, Directly removing the maxpool layer when the model was pretrained on imagenet (with the maxpool layer) <em>may</em> impact performance. If we were to remove the maxpool layer (say), we'll I either have to feed in 3x112x112 sized image or make some modifications in the FC layers for input shape 3x224x224</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485156,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-07T02:18:50.343000",
          "content": "<p>I'm using se_resNext50 and didn't resize image after deleting max pool. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485187,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-07T03:31:02.423000",
          "content": "<p>Hey Alex, </p>\n\n<p>The last two layers of se_resNext50 are <code>AvgPool2d-269</code> and <code>Linear-270</code>, now if you use the standard network (with the maxpool layer) with input image of 3x224x224, the output of AvgPool2d layer is <code>[batch_size, 2048, 1, 1]</code> and the Linear layer is defined to take in 2048 features so it works. </p>\n\n<p>Now, if you remove the maxpool layer and feed in 3x224x224 image the output of AvgPool2d is <code>[batch_size, 2048, 8, 8]</code> which will throw a shape mismatch error until we modify the Linear layer to take in that many feature size. </p>\n\n<p>What input image size are you using?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485982,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-08T06:53:14.770000",
          "content": "<p>96х96</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 486593,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-09T03:02:59.377000",
          "content": "<p>Change avg pooling to adaptive avg</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 484118,
      "author_name": "ratan rohith",
      "author_url": "",
      "post_date": "2019-03-05T15:40:26.423000",
      "content": "<p>thanks for sharing your thoughts :) ! keep up the good job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 478688,
      "author_name": "Franchini",
      "author_url": "",
      "post_date": "2019-02-26T13:42:53.490000",
      "content": "<p>Thanks for sharing, Alex! Great job! Can I ask you which optimizer you used?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 478749,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-26T15:20:14.717000",
          "content": "<p>Hi, Franchini! rmsprop but SGD, ADAM works the same here.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 485849,
      "author_name": "lvguofeng",
      "author_url": "",
      "post_date": "2019-03-08T01:43:56.830000",
      "content": "<p>Thanks for sharing! I don't know if I understand it wrong.\nif you use ~20 folds, it means you will trained 20 model?\nand you use 4d TTA,  so each test images you will predict 20*4 times  and simple np.mean() it?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 485984,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-08T06:56:35.603000",
          "content": "<p>Yes</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 486193,
          "author_name": "lvguofeng",
          "author_url": "",
          "post_date": "2019-03-08T12:13:18.583000",
          "content": "<p>thank you！ i will try it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 482612,
      "author_name": "SM",
      "author_url": "",
      "post_date": "2019-03-03T11:16:58.807000",
      "content": "<p>Nice, thanks for sharing. What is your validation score ? \nI can easily achieve 0.996+ score on every cv fold (10)  vs public lb score ~0.967. Changes on CV score does not correlate with lb score. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 482965,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-03T23:31:28.503000",
          "content": "<p>The same for me. =) </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2710532,
      "author_name": "MUHAMMAD DAWOOD RIZWAN",
      "author_url": "",
      "post_date": "2024-03-22T10:42:19.037000",
      "content": "<p>need label test data csv file</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1105418,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-12-07T21:12:56.010000",
      "content": "<p><a href=\"https://www.kaggle.com/donchuk\" target=\"_blank\">@donchuk</a>  Alex, how much increase did image augmentation increase the accuracy?</p>\n<p>Thanks for the help and good job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1095280,
      "author_name": "felix_indra",
      "author_url": "",
      "post_date": "2020-11-29T12:42:11.170000",
      "content": "<p>bro?<br>\nhow to calculate clahe manually?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 483853,
      "author_name": "quinwu",
      "author_url": "",
      "post_date": "2019-03-05T08:50:49.923000",
      "content": "<p>what means User 4d TTA?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 483877,
          "author_name": "Arvid Teichtmann",
          "author_url": "",
          "post_date": "2019-03-05T09:28:37.543000",
          "content": "<p>I assume this is a typo... should be <code>use 4d TTA</code> ?\n<code>TTA</code> stands for <code>Test Time Augmentation</code>\nHowever i don't understand either the meaning of <code>4d</code> ;-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 484010,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-05T12:58:00.560000",
          "content": "<p>Flips, Transpose, Rotations and its combination. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 484461,
          "author_name": "Shaohua Li",
          "author_url": "",
          "post_date": "2019-03-06T03:33:19.263000",
          "content": "<p>oh so you mean applying four augmentations in test time and vote to get the final class?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485154,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-07T02:16:58.093000",
          "content": "<p>just simple mean</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485690,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-07T19:42:38.973000",
          "content": "<p>Hey Alex, do we apply 4d during test time? Because the model has been trained with a rigorous augmentation pipeline?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485983,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-08T06:56:12.117000",
          "content": "<p>4d during test time - yes! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 478331,
      "author_name": "David Zhao",
      "author_url": "",
      "post_date": "2019-02-26T02:59:30.933000",
      "content": "<p>Thanks for sharing. I am using kinda the same however have not yet reached .9780+. Have you cropped or resized the image? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 478538,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-26T09:55:40.977000",
          "content": "<p>I use origin size. And I add some folds today and get 0.9792. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478542,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-26T10:00:35.847000",
          "content": "<p>Additionally you can add pseudo labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478978,
          "author_name": "David Zhao",
          "author_url": "",
          "post_date": "2019-02-26T21:41:55.163000",
          "content": "<p>Thanks for sharing your ideas! Good job of you! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 483882,
          "author_name": "Arvid Teichtmann",
          "author_url": "",
          "post_date": "2019-03-05T09:33:53.307000",
          "content": "<p>Can you elaborate more on <code>pseudo labels</code> in this context?\nDo you mean <code>Semi-Supervised Learning</code> ? \nBut this would require more unlabeled data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 484008,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-05T12:56:52.900000",
          "content": "<p>No! Add sample from test to train with very big score. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 484187,
          "author_name": "Arvid Teichtmann",
          "author_url": "",
          "post_date": "2019-03-05T17:01:41.570000",
          "content": "<p>Sorry, but i don't understand ;-(\nWhat do you mean by \"very big score\" ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 484192,
          "author_name": "Rohit Gupta",
          "author_url": "",
          "post_date": "2019-03-05T17:08:30.020000",
          "content": "<p><a href=\"/donchuk\">@donchuk</a>  \"very big score\" means train a classifier, predict probabilities for the test set and add the images with the score very close to 0 or 1 to the training set and train again????</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485153,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-03-07T02:16:30.123000",
          "content": "<p>Yes, try &gt; 0.9999 and 0.0001 &lt;</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485509,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-03-07T13:32:22.207000",
          "content": "<p><a href=\"/donchuk\">@donchuk</a> could you elaborate the motivation behind this strategy? Intuitively, I would have thought that the algorithm is not able to learn a lot from samples which it already perfectly predicts. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485691,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-07T19:44:19.393000",
          "content": "<p>@Franchini we do that so as to increase the training data size, as the model is pretty confident on these images we assume them to be their ground truth labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485704,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-03-07T20:02:46.700000",
          "content": "<p><a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> \"as the model is pretty confident on these images we assume them to be their ground truth labels\" If the model is already perfectly able to predict those samples, what can it learn from them? Wouldn't we rather want the model to see more samples that it finds harder to predict?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485735,
          "author_name": "SM",
          "author_url": "",
          "post_date": "2019-03-07T20:57:54.327000",
          "content": "<p>Here is rationale <a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485757,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-03-07T21:51:28.417000",
          "content": "<p>Pseudo-labeling is often used and most effective in cases where the number of training examples for a category is small. This technique is a part of several of the top scoring solutions to the whale competition that recently ended, for that one there were 5005 categories many of which only had one example in the training set. Such an imbalance is not the case with the binary classification here, but just because a model can already accurately (we assume) classify some test image does not mean there is not other useful information in those images that could be beneficial in training. To be safe I would only pseudo-label with images that different architectures independently gave high confidence to.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485962,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-03-08T06:29:56.277000",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> and <a href=\"/sermakarevich\">@sermakarevich</a>, thanks for the background! I was confused because I thought it meant to take images from the test set of the kfolds. Now that I realised that you take images from the unlabeled set and add it to the training set it makes more sense!\nThanks guys, I'm learning a lot in this competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 477947,
      "author_name": "Rohit Gupta",
      "author_url": "",
      "post_date": "2019-02-25T14:42:49.507000",
      "content": "<p>Can you explain #4 a bit more?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 478034,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-25T16:28:25.093000",
          "content": "<p>You need to validate often than once per epoch! Do validation 3-5 times during epoch. Don't overfit. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 477671,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-02-25T04:19:17.423000",
      "content": "<p>Can you explain #4? Are you using pytorch?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 477697,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-02-25T05:33:49.177000",
          "content": "<p>Yes, pure pytorch. \nValidate 3-5 times per epoch. Still not clear? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485687,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-07T19:39:31.240000",
          "content": "<p>Hey Alex, validation for me is to test the model performance on the validation set, What does validate 3-5 times per epoch mean?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488875,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-03-13T06:04:31.997000",
          "content": "<p>Hey Alex, still waiting for your reply :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 488400,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-12T12:04:21.670000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 507023,
          "author_name": "Alex Donchuk",
          "author_url": "",
          "post_date": "2019-04-04T07:15:22.373000",
          "content": "<p>+15 places on private</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 478913,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-26T19:50:21.493000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 478923,
          "author_name": "Franchini",
          "author_url": "",
          "post_date": "2019-02-26T20:05:50.937000",
          "content": "<p>I asked exactly the same question and he answered it 6h ago ;-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 478950,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-26T20:46:57.167000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 482262,
      "author_name": "Paul Spende",
      "author_url": "",
      "post_date": "2019-03-02T15:48:11.363000",
      "content": "<p>Great guide thanks a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 480855,
      "author_name": "Yangting",
      "author_url": "",
      "post_date": "2019-02-28T19:15:20.527000",
      "content": "<p>Thank you sooooo much!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 478213,
      "author_name": "Paco Diaz",
      "author_url": "",
      "post_date": "2019-02-25T22:17:38.300000",
      "content": "<p>Wow! thanks a lot!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 477581,
      "author_name": "William Green",
      "author_url": "",
      "post_date": "2019-02-24T22:39:04.287000",
      "content": "<p>Awesome! Thank you</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "477464": "Hi! Everybody. Steps to get the score:\n\n1. Use se_resnext50 without maxpool. Imagenet \n2. Create ~20 folds\n3. User 4d TTA\n4. After 7 epoch start to validate often than 1 per epoch, after 10-12 epoch start to validate more often.(3-5 times) Find optimum between train, val, and roc. Don't overfit.\n5. Use lr = 0.00007\n6. Use batch 256(din't check if it important)\n7. ReduceOnPlateu - 2 epoch\n8. Simple np.mean() ensemble 5+ folds will give you 0.9780+\n\np.s. Augs using albumentions\n        return Compose([\n            RandomRotate90(p=0.5),\n            Transpose(p=0.5),\n            Flip(p=0.5),\n            OneOf([\n                CLAHE(clip_limit=2),\n                IAASharpen(),\n                IAAEmboss(),\n                RandomBrightnessContrast(),\n                JpegCompression(),\n                Blur(),\n                GaussNoise()\n            ], p=0.5),\n            HueSaturationValue(p=0.5),\n            ShiftScaleRotate(shift_limit=0.15, scale_limit=0.15, rotate_limit=45, p=0.5),\n            Normalize(p=1)\n",
    "1715411": "Hi friend!, there is only one thing I don't understand, you are using pre-trained ResNet150, however the imagenet dataset wasn't trained using these images/dataset, I mean you are using loading the pre-trained images with imagenet weights but training again for this data and reaching such results? Please, I hope not to bother, but I can't understand how you work with it. \nThanks.",
    "477767": "Almost same for me, but I have only 5 folds, 12 random tta images, pretrain net on imagenet, starting with lr=3e-5 for last layer and 3e-5*0.001 for pretrained ones, batch of 64 images rescaled to 224x224 + albumentions  + mean of folds predictions ",
    "506985": "Thanks for your advice, Alex.  It really helped me a lot.",
    "485937": "can this be achieved without a pre-trained network?",
    "484229": "Sorry for the dumb question, but how should one remove the maxpool? I tried just deleting the layer and it’s forward but that did not seem to work. I also tried replacing just the maxpool with identity but that seemed not to work also. \n\nDo you mean just the one maxpool after the first convolution in seresnext50? I’ll keep googling until I figure it out. Can you advise on the proper way to remove or change a layer in a pytorch pretrained model, If you don’t mind? Thanks for posting this!",
    "484118": "thanks for sharing your thoughts :) ! keep up the good job!",
    "478688": "Thanks for sharing, Alex! Great job! Can I ask you which optimizer you used?",
    "485849": "Thanks for sharing! I don't know if I understand it wrong.\nif you use ~20 folds, it means you will trained 20 model?\nand you use 4d TTA,  so each test images you will predict 20*4 times  and simple np.mean() it?",
    "482612": "Nice, thanks for sharing. What is your validation score ? \nI can easily achieve 0.996+ score on every cv fold (10)  vs public lb score ~0.967. Changes on CV score does not correlate with lb score. ",
    "2710532": "need label test data csv file",
    "1105418": "@donchuk  Alex, how much increase did image augmentation increase the accuracy?\n\nThanks for the help and good job!",
    "1095280": "bro?\nhow to calculate clahe manually?",
    "483853": "what means User 4d TTA?",
    "478331": "Thanks for sharing. I am using kinda the same however have not yet reached .9780+. Have you cropped or resized the image? ",
    "477947": "Can you explain #4 a bit more?",
    "477671": "Can you explain #4? Are you using pytorch?",
    "488400": "",
    "478913": "",
    "482262": "Great guide thanks a lot!",
    "480855": "Thank you sooooo much!",
    "478213": "Wow! thanks a lot!!",
    "477581": "Awesome! Thank you"
  }
}