{
  "id": 105563,
  "title": "Struggling to improve (ResNets and EfficientNets)",
  "url": "/competitions/aptos2019-blindness-detection/discussion/105563",
  "author_name": "",
  "post_date": "2019-08-24T06:04:26.437186900Z",
  "votes": 28,
  "comment_count": 64,
  "views": 0,
  "content": "<p>I have been struggling to improve my score from my current position. My current score (0.793) is ResNet50 pure classification 5fold CV with crop_image_from_gray preprocessing pretrained on previous competition dataset, with significant augmentation (fastai defaults mainly) and no weight decay (but seems to overfit the training set).</p>\n\n<p>Additional things I have tried and have not increased my score:\n1. Add weight decay (does not overfit but also decreases performance on LB; kept in further iterations however)\n2. Add circle_crop \n3. Change to ResNet152\n4. Training longer\n5. Changing to regression</p>\n\n<p>At this point I decided to switch to EfficientNets to see what the hype was about. I mainly followed what @DrHB suggested. I pretrained an EfficientNetB3 on previous dataset, the whole network unfrozen (the Luke Melas PyTorch version does not seem to support freezing of the network) , no cropping. I then trained a 5-fold regression CV on the competition dataset with the cropping function, again all unfrozen. I used the correct size of 300x300 and again significant augmentation. It achieved nice results on the CV, but only achieved 0.778, even after playing a little bit with the seed.</p>\n\n<p>What am I doing wrong? Any tips either for the ResNet or the EfficientNet architectures?</p>",
  "messages": [
    {
      "id": "606817",
      "postDate": "08/24/2019 06:04:26",
      "content": "<p>I have been struggling to improve my score from my current position. My current score (0.793) is ResNet50 pure classification 5fold CV with crop_image_from_gray preprocessing pretrained on previous competition dataset, with significant augmentation (fastai defaults mainly) and no weight decay (but seems to overfit the training set).</p>\n\n<p>Additional things I have tried and have not increased my score:\n1. Add weight decay (does not overfit but also decreases performance on LB; kept in further iterations however)\n2. Add circle_crop \n3. Change to ResNet152\n4. Training longer\n5. Changing to regression</p>\n\n<p>At this point I decided to switch to EfficientNets to see what the hype was about. I mainly followed what @DrHB suggested. I pretrained an EfficientNetB3 on previous dataset, the whole network unfrozen (the Luke Melas PyTorch version does not seem to support freezing of the network) , no cropping. I then trained a 5-fold regression CV on the competition dataset with the cropping function, again all unfrozen. I used the correct size of 300x300 and again significant augmentation. It achieved nice results on the CV, but only achieved 0.778, even after playing a little bit with the seed.</p>\n\n<p>What am I doing wrong? Any tips either for the ResNet or the EfficientNet architectures?</p>",
      "rawMarkdown": "I have been struggling to improve my score from my current position. My current score (0.793) is ResNet50 pure classification 5fold CV with crop_image_from_gray preprocessing pretrained on previous competition dataset, with significant augmentation (fastai defaults mainly) and no weight decay (but seems to overfit the training set).\n\nAdditional things I have tried and have not increased my score:\n1. Add weight decay (does not overfit but also decreases performance on LB; kept in further iterations however)\n2. Add circle_crop \n3. Change to ResNet152\n4. Training longer\n5. Changing to regression\n\n\nAt this point I decided to switch to EfficientNets to see what the hype was about. I mainly followed what @DrHB suggested. I pretrained an EfficientNetB3 on previous dataset, the whole network unfrozen (the Luke Melas PyTorch version does not seem to support freezing of the network) , no cropping. I then trained a 5-fold regression CV on the competition dataset with the cropping function, again all unfrozen. I used the correct size of 300x300 and again significant augmentation. It achieved nice results on the CV, but only achieved 0.778, even after playing a little bit with the seed.\n\nWhat am I doing wrong? Any tips either for the ResNet or the EfficientNet architectures?",
      "votes": null
    },
    {
      "id": "606845",
      "postDate": "08/24/2019 07:16:24",
      "content": "<p>I have similar problems with B4 &amp; B5. I am sure there are others who did well with these networks. I guess we will get to learn more on the efficientnets after the competition ends.</p>\n\n<p>For B3, I too followed DrHB and did things similar to what you tried. I got LB ranging from 0.78 to 0.819. But, I cannot conclusively say that 0.819 LB is better than 0.8 LB</p>",
      "rawMarkdown": "I have similar problems with B4 &amp; B5. I am sure there are others who did well with these networks. I guess we will get to learn more on the efficientnets after the competition ends.\n\nFor B3, I too followed DrHB and did things similar to what you tried. I got LB ranging from 0.78 to 0.819. But, I cannot conclusively say that 0.819 LB is better than 0.8 LB",
      "votes": null
    },
    {
      "id": "606847",
      "postDate": "08/24/2019 07:21:44",
      "content": "<p>Did you try Res-Net 50? Moreover, try to balance the old data.</p>",
      "rawMarkdown": "Did you try Res-Net 50? Moreover, try to balance the old data.",
      "votes": null
    },
    {
      "id": "606903",
      "postDate": "08/24/2019 09:26:46",
      "content": "<p>Remember that public \"test\" is only 1,928 images (private one is 13,000). LB score represents only 12.9% of test data.\nSo your experiments does not necessarily will lead to improvements in LB score.</p>\n\n<p>From \"About this competition\":\n\"When Kaggle runs your Kernel privately, it substitutes the private test set and sample submission in place of the public ones. You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately)\". </p>",
      "rawMarkdown": "Remember that public \"test\" is only 1,928 images (private one is 13,000). LB score represents only 12.9% of test data.\nSo your experiments does not necessarily will lead to improvements in LB score.\n\nFrom \"About this competition\":\n\"When Kaggle runs your Kernel privately, it substitutes the private test set and sample submission in place of the public ones. You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately)\".",
      "votes": null
    },
    {
      "id": "606947",
      "postDate": "08/24/2019 11:03:49",
      "content": "<p>How does your training/valid curves look like?</p>\n\n<p>You also use Grad-CAM, did your models actually capture those scabs/wools correctly in both the training / validation data? My case is to try to make many examples of those missed scabs/wools as much as possible.</p>",
      "rawMarkdown": "How does your training/valid curves look like?\n\nYou also use Grad-CAM, did your models actually capture those scabs/wools correctly in both the training / validation data? My case is to try to make many examples of those missed scabs/wools as much as possible.",
      "votes": null
    },
    {
      "id": "606977",
      "postDate": "08/24/2019 11:49:56",
      "content": "<p>Me too.</p>",
      "rawMarkdown": "Me too.",
      "votes": null
    },
    {
      "id": "607011",
      "postDate": "08/24/2019 12:47:29",
      "content": "<p>For me, zooming in and TTA(horizontal+ vertical flips) seems to help performance </p>",
      "rawMarkdown": "For me, zooming in and TTA(horizontal+ vertical flips) seems to help performance",
      "votes": null
    },
    {
      "id": "607028",
      "postDate": "08/24/2019 13:28:36",
      "content": "<p>What optimizers/schedulers are you using for EfficientNets? My best results so far have been with AdamW+One cycle policy on B0, but I haven't been able to improve that with B2. I've been trying RMSprop+StepLR on B2, as that is the strategy used in the EfficientNet paper, but my results are worse so far.</p>",
      "rawMarkdown": "What optimizers/schedulers are you using for EfficientNets? My best results so far have been with AdamW+One cycle policy on B0, but I haven't been able to improve that with B2. I've been trying RMSprop+StepLR on B2, as that is the strategy used in the EfficientNet paper, but my results are worse so far.",
      "votes": null
    },
    {
      "id": "607085",
      "postDate": "08/24/2019 15:38:46",
      "content": "<p>For my case, Densenet perform better than Resnet. You may try that.</p>",
      "rawMarkdown": "For my case, Densenet perform better than Resnet. You may try that.",
      "votes": null
    },
    {
      "id": "607169",
      "postDate": "08/24/2019 17:45:17",
      "content": "<p>Following @DrHB's suggestion is sufficient to get <code>0.8+</code> LB score. But I find it extremely difficult to improve after <code>0.81</code>. My current LB score is an ensemble of 795, 794, 804 and 807 models. My LB score increased by only <code>0.001</code> after adding the 807 model. Any tips about how to boost my score will be appreciated. :)  </p>",
      "rawMarkdown": "Following @DrHB's suggestion is sufficient to get `0.8+` LB score. But I find it extremely difficult to improve after `0.81`. My current LB score is an ensemble of 795, 794, 804 and 807 models. My LB score increased by only `0.001` after adding the 807 model. Any tips about how to boost my score will be appreciated. :)",
      "votes": null
    },
    {
      "id": "607212",
      "postDate": "08/24/2019 20:01:27",
      "content": "<p>I did B4, with image size 256, cropping, pre-trained on 2015 data. As my single fold model now sitting comfortably on (0.795-0.8) and 5 fold model gives (0.801-0.804) but hard to push for 0.81... I noticed why people are merging... because you can ensemble easily to push high scores... </p>\n\n<p>So I am now training effb2 for non-cropping but heavy argumentation one, it is very interesting now single fold is around 0.790-0.795 (I alos clean the dataset but removing all same images with different file names). Just at the moment to do more h-para search, kaggle shut all of my kernels off and leave me only 1 to run...</p>\n\n<p>Another interesting experiment I did is remove the nn.Linear layer from efficient net, stick fastai gold head (which inspired by another on going kaggle steel competition), now the network is 2 layer_groups, all the sudden the freezing and unfreezing is start to work. The model single fold can reach 0.78-0.79 range, but since you have two layer_groups, it is very hard to train the network, the lr just kills the body (which is the efficient net backbone). Again, more experiments are needed, but no kernels available. </p>\n\n<p>I guess I am done with this competition, simply just I don't have enough resources to try ideas... all my kernel (just 1) now is doing ensemble of the previous models and is taking forever....</p>\n\n<p>sad :(</p>",
      "rawMarkdown": "I did B4, with image size 256, cropping, pre-trained on 2015 data. As my single fold model now sitting comfortably on (0.795-0.8) and 5 fold model gives (0.801-0.804) but hard to push for 0.81... I noticed why people are merging... because you can ensemble easily to push high scores... \n\nSo I am now training effb2 for non-cropping but heavy argumentation one, it is very interesting now single fold is around 0.790-0.795 (I alos clean the dataset but removing all same images with different file names). Just at the moment to do more h-para search, kaggle shut all of my kernels off and leave me only 1 to run...\n\nAnother interesting experiment I did is remove the nn.Linear layer from efficient net, stick fastai gold head (which inspired by another on going kaggle steel competition), now the network is 2 layer_groups, all the sudden the freezing and unfreezing is start to work. The model single fold can reach 0.78-0.79 range, but since you have two layer_groups, it is very hard to train the network, the lr just kills the body (which is the efficient net backbone). Again, more experiments are needed, but no kernels available. \n\nI guess I am done with this competition, simply just I don't have enough resources to try ideas... all my kernel (just 1) now is doing ensemble of the previous models and is taking forever....\n\nsad :(",
      "votes": null
    },
    {
      "id": "607363",
      "postDate": "08/25/2019 04:55:16",
      "content": "<p><a href=\"/heye0507\">@heye0507</a> I feel you. :/ </p>",
      "rawMarkdown": "heye0507 I feel you. :/",
      "votes": null
    },
    {
      "id": "607384",
      "postDate": "08/25/2019 06:11:51",
      "content": "<p>Hi, do you simply average them or did you tune a weighted average? Another weird thing I noticed in my approach is that applying slighty larger weight to a worse-performing model improved the LB... either its the model diversity or I'm simply overfitting to the LB :/ </p>",
      "rawMarkdown": "Hi, do you simply average them or did you tune a weighted average? Another weird thing I noticed in my approach is that applying slighty larger weight to a worse-performing model improved the LB... either its the model diversity or I'm simply overfitting to the LB :/",
      "votes": null
    },
    {
      "id": "607403",
      "postDate": "08/25/2019 07:09:19",
      "content": "<p>I am taking their weighted average. I agree with you. I found out that ensembling models with lower LB scores boost score comparatively higher than the high scoring models.  Weird. :/</p>",
      "rawMarkdown": "I am taking their weighted average. I agree with you. I found out that ensembling models with lower LB scores boost score comparatively higher than the high scoring models.  Weird. :/",
      "votes": null
    },
    {
      "id": "607413",
      "postDate": "08/25/2019 07:31:43",
      "content": "<p>If you don't mind answering, are all models in your ensemble from the same approach and how they differ from each other? I am currently experimenting model diversity in ensembling and I don't have a clear indication if diverse = better LB atm</p>",
      "rawMarkdown": "If you don't mind answering, are all models in your ensemble from the same approach and how they differ from each other? I am currently experimenting model diversity in ensembling and I don't have a clear indication if diverse = better LB atm",
      "votes": null
    },
    {
      "id": "607418",
      "postDate": "08/25/2019 08:00:44",
      "content": "<p>Yes. I am using Efficientnet models and using the same validation set for finetuning them. Only using circle cropping as preprocessing. Please let us know about the outcomes of your experiments. :) </p>",
      "rawMarkdown": "Yes. I am using Efficientnet models and using the same validation set for finetuning them. Only using circle cropping as preprocessing. Please let us know about the outcomes of your experiments. :)",
      "votes": null
    },
    {
      "id": "607454",
      "postDate": "08/25/2019 09:25:30",
      "content": "<p>I fit a LinearRegression model on the top of the outputs of my EfficientNet models. It gave me 1% boost. From 80.8 to 81.8</p>",
      "rawMarkdown": "I fit a LinearRegression model on the top of the outputs of my EfficientNet models. It gave me 1% boost. From 80.8 to 81.8",
      "votes": null
    },
    {
      "id": "607491",
      "postDate": "08/25/2019 11:56:58",
      "content": "<p>Can you please explain more <a href=\"/nemethpeti\">@nemethpeti</a> </p>",
      "rawMarkdown": "Can you please explain more @nemethpeti",
      "votes": null
    },
    {
      "id": "607502",
      "postDate": "08/25/2019 12:11:32",
      "content": "<ol>\n<li>Train a number of models on the same train/validation set</li>\n<li>Each will have a raw (float) output (for classifiers, you have to calculate an expected value)</li>\n<li>Predict the raw outputs on the VALIDATION set</li>\n<li>Predict the raw outputs on the TEST set</li>\n<li>Split the validation set into TRAIN2/VALIDATION2 sets</li>\n<li>Fit a Linear/RandomForest/XGBoost on TRAIN2 and optimize to have good result on VALIDATION2</li>\n<li>Take the results of step 4 as input and use the model of step 6 to predict the outputs of the TEST set</li>\n</ol>\n\n<p>You get a significant boost!\nI got over 1% from 3 models.</p>",
      "rawMarkdown": "1. Train a number of models on the same train/validation set\n2. Each will have a raw (float) output (for classifiers, you have to calculate an expected value)\n3. Predict the raw outputs on the VALIDATION set\n4. Predict the raw outputs on the TEST set\n5. Split the validation set into TRAIN2/VALIDATION2 sets\n6. Fit a Linear/RandomForest/XGBoost on TRAIN2 and optimize to have good result on VALIDATION2\n7. Take the results of step 4 as input and use the model of step 6 to predict the outputs of the TEST set\n\nYou get a significant boost!\nI got over 1% from 3 models.",
      "votes": null
    },
    {
      "id": "607508",
      "postDate": "08/25/2019 12:16:14",
      "content": "<p>Thanks for the detailed reply. I shall surely try this. </p>",
      "rawMarkdown": "Thanks for the detailed reply. I shall surely try this.",
      "votes": null
    },
    {
      "id": "607509",
      "postDate": "08/25/2019 12:17:30",
      "content": "<p>I used a Linear model with success. At the end it will give a weighted avearge, but the weights are calculated and not manually tuned. Manual tuning is hard if you have more than 2 models.</p>",
      "rawMarkdown": "I used a Linear model with success. At the end it will give a weighted avearge, but the weights are calculated and not manually tuned. Manual tuning is hard if you have more than 2 models.",
      "votes": null
    },
    {
      "id": "607666",
      "postDate": "08/25/2019 17:30:56",
      "content": "<p><a href=\"/nemethpeti\">@nemethpeti</a> Thanks for sharing. Will try your method. :) </p>",
      "rawMarkdown": "nemethpeti Thanks for sharing. Will try your method. :)",
      "votes": null
    },
    {
      "id": "607678",
      "postDate": "08/25/2019 17:49:09",
      "content": "<p>Good luck! Augmentations so far do not seem to work well for me</p>",
      "rawMarkdown": "Good luck! Augmentations so far do not seem to work well for me",
      "votes": null
    },
    {
      "id": "607697",
      "postDate": "08/25/2019 18:35:22",
      "content": "<p><a href=\"/dreamdragon\">@dreamdragon</a> quite interesting..\nHow you trained using old data, did you use entire old data which is quite huge..?</p>",
      "rawMarkdown": "dreamdragon quite interesting..\nHow you trained using old data, did you use entire old data which is quite huge..?",
      "votes": null
    },
    {
      "id": "607725",
      "postDate": "08/25/2019 19:13:16",
      "content": "<p><a href=\"/tahsin\">@tahsin</a> I will certainly update if I get any success, but with the recent downsizing in commits I hope I have time :(</p>",
      "rawMarkdown": "tahsin I will certainly update if I get any success, but with the recent downsizing in commits I hope I have time :(",
      "votes": null
    },
    {
      "id": "607882",
      "postDate": "08/26/2019 03:35:04",
      "content": "<p>you can get 0.815 from just ensembling efficientnets(regression) that all have lower than 0.8.</p>\n\n<p>Also to \"freeze\" layers you can iterate thru layers andn make certain ones not require grad and then pass all the params with requires grad to a new optimizer.</p>",
      "rawMarkdown": "you can get 0.815 from just ensembling efficientnets(regression) that all have lower than 0.8.\n\nAlso to \"freeze\" layers you can iterate thru layers andn make certain ones not require grad and then pass all the params with requires grad to a new optimizer.",
      "votes": null
    },
    {
      "id": "607884",
      "postDate": "08/26/2019 03:38:09",
      "content": "<p>thanks for your information! very useful</p>",
      "rawMarkdown": "thanks for your information! very useful",
      "votes": null
    },
    {
      "id": "607888",
      "postDate": "08/26/2019 03:45:27",
      "content": "<p>@peter thanks for input..\nwhat is raw output size  :  bs* Output layer nerurons * (1)  ?</p>",
      "rawMarkdown": "peter thanks for input..\nwhat is raw output size  :  bs* Output layer nerurons * (1)  ?",
      "votes": null
    },
    {
      "id": "607929",
      "postDate": "08/26/2019 05:07:35",
      "content": "<p>1 output per eye / model.\nBut the technic should work with 5 outputs as well.</p>",
      "rawMarkdown": "1 output per eye / model.\nBut the technic should work with 5 outputs as well.",
      "votes": null
    },
    {
      "id": "607946",
      "postDate": "08/26/2019 05:43:19",
      "content": "<p>Does freezing and then subsequently unfreezing provide any benefit compared to training on the unfrozen model?</p>",
      "rawMarkdown": "Does freezing and then subsequently unfreezing provide any benefit compared to training on the unfrozen model?",
      "votes": null
    },
    {
      "id": "607947",
      "postDate": "08/26/2019 05:43:45",
      "content": "<p>AdamW+One cycle policy is the default in fastai and what I am using...</p>",
      "rawMarkdown": "AdamW+One cycle policy is the default in fastai and what I am using...",
      "votes": null
    },
    {
      "id": "607958",
      "postDate": "08/26/2019 06:34:57",
      "content": "<p>In my case, linear regression on top of my models for ensemble degraded LB score. :/ Must have done something wrong. Here is my code:</p>\n\n<p>&gt; \nfrom sklearn.linear_model import LinearRegression\n&gt;\nimport numpy as np\n&gt;\nval_preds = np.hstack((val_pred1, val_pred2, val_pred3, val_pred4))\n&gt;\nreg = LinearRegression().fit(val_preds, val_label)\n&gt;\nW = reg.coef_\n&gt;\nb  = reg.intercept_\n&gt;\ntest_pred = test_pred1*W[0] + test_pred2*W[1] + test_pred3*W[2] + test_pred4*W[3] + b</p>",
      "rawMarkdown": "In my case, linear regression on top of my models for ensemble degraded LB score. :/ Must have done something wrong. Here is my code:\n\n&gt; \nfrom sklearn.linear_model import LinearRegression\n&gt;\nimport numpy as np\n&gt;\nval_preds = np.hstack((val_pred1, val_pred2, val_pred3, val_pred4))\n&gt;\nreg = LinearRegression().fit(val_preds, val_label)\n&gt;\nW = reg.coef_\n&gt;\nb  = reg.intercept_\n&gt;\ntest_pred = test_pred1*W[0] + test_pred2*W[1] + test_pred3*W[2] + test_pred4*W[3] + b",
      "votes": null
    },
    {
      "id": "607966",
      "postDate": "08/26/2019 06:51:29",
      "content": "<p>how many models, different backbones?</p>",
      "rawMarkdown": "how many models, different backbones?",
      "votes": null
    },
    {
      "id": "607978",
      "postDate": "08/26/2019 07:11:45",
      "content": "<p>opposite to you, I can't get a better score from resnet(34/50) or se-resnext50.Is it easy to overfitting than efficientnet?</p>",
      "rawMarkdown": "opposite to you, I can't get a better score from resnet(34/50) or se-resnext50.Is it easy to overfitting than efficientnet?",
      "votes": null
    },
    {
      "id": "608013",
      "postDate": "08/26/2019 07:57:36",
      "content": "<p>I assume valpredX values are floats!</p>",
      "rawMarkdown": "I assume valpredX values are floats!",
      "votes": null
    },
    {
      "id": "608021",
      "postDate": "08/26/2019 08:12:45",
      "content": "<p>yes. <a href=\"/nemethpeti\">@nemethpeti</a> </p>",
      "rawMarkdown": "yes. @nemethpeti",
      "votes": null
    },
    {
      "id": "608026",
      "postDate": "08/26/2019 08:24:48",
      "content": "<p>sorry ... trying to understand.. output layer of all eff or any model is  say like  nn.Linear(1024,1) \nso are you giving linear module output  to Xgboost <br>\nX= 4000* 1 ,Y=4000*1   ?</p>",
      "rawMarkdown": "sorry ... trying to understand.. output layer of all eff or any model is  say like  nn.Linear(1024,1) \nso are you giving linear module output  to Xgboost  \nX= 4000* 1 ,Y=4000*1   ?",
      "votes": null
    },
    {
      "id": "608030",
      "postDate": "08/26/2019 08:29:32",
      "content": "<p>How to balance the old data? I find it hard to train no matter which model i use</p>",
      "rawMarkdown": "How to balance the old data? I find it hard to train no matter which model i use",
      "votes": null
    },
    {
      "id": "608042",
      "postDate": "08/26/2019 08:43:45",
      "content": "<p>Playing with seed costs lots of time, any advice?</p>",
      "rawMarkdown": "Playing with seed costs lots of time, any advice?",
      "votes": null
    },
    {
      "id": "608089",
      "postDate": "08/26/2019 10:18:27",
      "content": "<p>hi neuron\ngood to see you second :)\nwould u like to tell the value_counts for output...this will be benchmark to see how far people below you are deviating from the top score :)\nCurrently by looking at value counts m able to figure out how much my submission score be around :)</p>",
      "rawMarkdown": "hi neuron\ngood to see you second :)\nwould u like to tell the value_counts for output...this will be benchmark to see how far people below you are deviating from the top score :)\nCurrently by looking at value counts m able to figure out how much my submission score be around :)",
      "votes": null
    },
    {
      "id": "608092",
      "postDate": "08/26/2019 10:22:32",
      "content": "<p>Don't do it. This will likely lead to overfitting. Just keep it fixed across experiments and try to improve</p>",
      "rawMarkdown": "Don't do it. This will likely lead to overfitting. Just keep it fixed across experiments and try to improve",
      "votes": null
    },
    {
      "id": "608322",
      "postDate": "08/26/2019 16:14:14",
      "content": "<p>agreed，ensemle models can easily boost the lb score</p>",
      "rawMarkdown": "agreed，ensemle models can easily boost the lb score",
      "votes": null
    },
    {
      "id": "608380",
      "postDate": "08/26/2019 17:17:09",
      "content": "<p>but i cast a doubt on ensembles how well  it can give results on bigger chunk of images in private set. \nPeople may have tried to fit to public ds their thresholds eg. one thing is clear that there are more than 1100 + images of category 2 ,any thing less than that will give you a score of less than 77 on public ds. </p>",
      "rawMarkdown": "but i cast a doubt on ensembles how well  it can give results on bigger chunk of images in private set. \nPeople may have tried to fit to public ds their thresholds eg. one thing is clear that there are more than 1100 + images of category 2 ,any thing less than that will give you a score of less than 77 on public ds.",
      "votes": null
    },
    {
      "id": "608394",
      "postDate": "08/26/2019 17:39:37",
      "content": "<p>thats also my concern，we are trying our best to fit the public ds while the private ds may have a quite different distribution</p>",
      "rawMarkdown": "thats also my concern，we are trying our best to fit the public ds while the private ds may have a quite different distribution",
      "votes": null
    },
    {
      "id": "608580",
      "postDate": "08/27/2019 01:09:58",
      "content": "<p>I agree with you. And there are a lot of images in public test set had been preprocessed. \nnot sure private test set will be same.</p>",
      "rawMarkdown": "I agree with you. And there are a lot of images in public test set had been preprocessed. \nnot sure private test set will be same.",
      "votes": null
    },
    {
      "id": "608588",
      "postDate": "08/27/2019 01:44:00",
      "content": "<p><a href=\"/leixiang\">@leixiang</a> <a href=\"/chanhu\">@chanhu</a> See my post over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-608543\">here</a></p>",
      "rawMarkdown": "leixiang @chanhu See my post over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-608543)",
      "votes": null
    },
    {
      "id": "608671",
      "postDate": "08/27/2019 04:04:37",
      "content": "<p>Thank you for you nice work. </p>",
      "rawMarkdown": "Thank you for you nice work.",
      "votes": null
    },
    {
      "id": "608704",
      "postDate": "08/27/2019 04:59:16",
      "content": "<p>thanks for your information</p>",
      "rawMarkdown": "thanks for your information",
      "votes": null
    },
    {
      "id": "609623",
      "postDate": "08/27/2019 23:46:15",
      "content": "<p>Couple things I want to update, if anyone is interested. </p>\n\n<p>All the things I talked about are pre-trained with 2015 data, used 2019 data as validation. Pick best valid loss after 10 epochs.</p>\n\n<ol>\n<li>Like DrHB's post, argumentation do help, but you have to make sure padding mode you are using if you use fastai, default padding mode is reflection pad, if you also rotate and pre-process the image, you will see the unwanted eye shape in the corner when you rotate. Also, if you pad zero, unwanted black edge will come back. I have very hard time to use nearest (boarder in fastai). Think first if you want to rotate the image (on my experiments, it worked)</li>\n</ol>\n\n<p>What I noticed is that pre-processing / train using original image gives similar LB scores.\nArgumentation used in fastai [max_zoom=1.3, max_rotate=360., flip] for both cases.</p>\n\n<p>depends on your pre-trained situation, effb4 with 8 epochs and lr=1e-3/2 with size=(256,256) gives around LB 0.79+ single fold with either valid = 0.1 or 0.2. But if you just want to have solid 0.79+ LB, set valid = 0.1 (assuming you are using above hyper-parameters)</p>\n\n<p>5 fold cv should give you around 0.801 - 0.804 range</p>\n\n<ol>\n<li><p>About adding fastai standard head, if you look at efficient-net, they stick conv_head then bn then linear, I tried following:</p></li>\n<li><p>remove linear layer, stick fastai-head, which is stacked (max and avg pool), bn, linear, dropout...etc</p></li>\n<li>remove conv_head, stick fastai-head, again, stacked (max and avg pool) ...  </li>\n</ol>\n\n<p>Split layers at head, so you will have efficient net b4 as body, fastai head as head.</p>\n\n<p>Here is what becomes interesting: </p>\n\n<p>In fastai, we said single avg pool gives some information (which is what most of models do, adaptive / global avg pool follow by linear), but if you gives another maxpool, and stack the results together, now you have twice as much as information, it should perform at least as good as previous model. </p>\n\n<p>For both 1/2 situation, the LB score is around 0.79 (single fold, 20% validation without rotate)</p>\n\n<p>Training steps are:</p>\n\n<ol>\n<li>freeze the body except the bn layers, train the head for 5-8 epochs, you will see the kappa score is around 0.9x</li>\n<li>unfreeze the model, \n2a. give head lr, body lr / 10 (standard fastai apporach)\n2b. give head lr, body lr / 100</li>\n</ol>\n\n<p>for both cases, you will see that the model becomes unstable. It barely improves anything. </p>\n\n<p>That makes me wonder, pre-trained on 2015 data is good enough? I am only fine tuning the head?</p>\n\n<p>In comparison,  I split the original effb4 model at model._fc, freeze the body, I found that after 1 epoch, the model is over-fitting,  which is as expected, we only fine tuning the bn and linear layers, which only has few parameters. </p>\n\n<p>I also tried split the model at _conv_head, repeat freeze() and unfreeze(). I found that it has similar result as you dont freeze, just train with consistent lr .</p>\n\n<p>In all cases, discriminative lr doesn't work. </p>\n\n<p>My naive conclusion is, if you just split the efficient model at head / fc layer, freeze / unfreeze can be void, or at least it doesn't seem help.</p>\n\n<p>I already learnt so much from the competition. Thank you all for the help.</p>\n\n<p>I guess next step is to learn ensemble tricks, which I hope I have enough time to try. </p>",
      "rawMarkdown": "Couple things I want to update, if anyone is interested. \n\nAll the things I talked about are pre-trained with 2015 data, used 2019 data as validation. Pick best valid loss after 10 epochs.\n\n1. Like DrHB's post, argumentation do help, but you have to make sure padding mode you are using if you use fastai, default padding mode is reflection pad, if you also rotate and pre-process the image, you will see the unwanted eye shape in the corner when you rotate. Also, if you pad zero, unwanted black edge will come back. I have very hard time to use nearest (boarder in fastai). Think first if you want to rotate the image (on my experiments, it worked)\n\nWhat I noticed is that pre-processing / train using original image gives similar LB scores.\nArgumentation used in fastai [max_zoom=1.3, max_rotate=360., flip] for both cases.\n\ndepends on your pre-trained situation, effb4 with 8 epochs and lr=1e-3/2 with size=(256,256) gives around LB 0.79+ single fold with either valid = 0.1 or 0.2. But if you just want to have solid 0.79+ LB, set valid = 0.1 (assuming you are using above hyper-parameters)\n\n5 fold cv should give you around 0.801 - 0.804 range\n\n2. About adding fastai standard head, if you look at efficient-net, they stick conv_head then bn then linear, I tried following:\n\n1. remove linear layer, stick fastai-head, which is stacked (max and avg pool), bn, linear, dropout...etc\n2. remove conv_head, stick fastai-head, again, stacked (max and avg pool) ...  \n\nSplit layers at head, so you will have efficient net b4 as body, fastai head as head.\n\nHere is what becomes interesting: \n\nIn fastai, we said single avg pool gives some information (which is what most of models do, adaptive / global avg pool follow by linear), but if you gives another maxpool, and stack the results together, now you have twice as much as information, it should perform at least as good as previous model. \n\nFor both 1/2 situation, the LB score is around 0.79 (single fold, 20% validation without rotate)\n\nTraining steps are:\n\n1. freeze the body except the bn layers, train the head for 5-8 epochs, you will see the kappa score is around 0.9x\n2. unfreeze the model, \n2a. give head lr, body lr / 10 (standard fastai apporach)\n2b. give head lr, body lr / 100\n\nfor both cases, you will see that the model becomes unstable. It barely improves anything. \n\nThat makes me wonder, pre-trained on 2015 data is good enough? I am only fine tuning the head?\n\nIn comparison,  I split the original effb4 model at model._fc, freeze the body, I found that after 1 epoch, the model is over-fitting,  which is as expected, we only fine tuning the bn and linear layers, which only has few parameters. \n\nI also tried split the model at _conv_head, repeat freeze() and unfreeze(). I found that it has similar result as you dont freeze, just train with consistent lr .\n\nIn all cases, discriminative lr doesn't work. \n\nMy naive conclusion is, if you just split the efficient model at head / fc layer, freeze / unfreeze can be void, or at least it doesn't seem help.\n\nI already learnt so much from the competition. Thank you all for the help.\n\nI guess next step is to learn ensemble tricks, which I hope I have enough time to try.",
      "votes": null
    },
    {
      "id": "609750",
      "postDate": "08/28/2019 04:25:17",
      "content": "<p>So just use fastai standard head(AdaptiveConcatPool2d, Flatten, blocks of [nn.BatchNorm1d, nn.Dropout, nn.Linear, nn.ReLU] )? And use consistent lr?</p>",
      "rawMarkdown": "So just use fastai standard head(AdaptiveConcatPool2d, Flatten, blocks of [nn.BatchNorm1d, nn.Dropout, nn.Linear, nn.ReLU] )? And use consistent lr?",
      "votes": null
    },
    {
      "id": "609850",
      "postDate": "08/28/2019 06:58:01",
      "content": "<p>Very informative. Thanks a lot.</p>",
      "rawMarkdown": "Very informative. Thanks a lot.",
      "votes": null
    },
    {
      "id": "610652",
      "postDate": "08/28/2019 23:45:05",
      "content": "<p><a href=\"/heye0507\">@heye0507</a> , thanks for the report, I have one question, on my experiments, my results change a lot depending on the batch size I use.</p>\n\n<p>When you use effnetb4 and img size 256, what's your batch size and do you use something like gradient accumulator?</p>",
      "rawMarkdown": "heye0507 , thanks for the report, I have one question, on my experiments, my results change a lot depending on the batch size I use.\n\nWhen you use effnetb4 and img size 256, what's your batch size and do you use something like gradient accumulator?",
      "votes": null
    },
    {
      "id": "610667",
      "postDate": "08/29/2019 00:02:35",
      "content": "<p><a href=\"/tahsin\">@tahsin</a> FYI, I tried adding ordinal regression model to my ensemble, which did not help, but adding a model with different pre-processing improved the score by small margin (~0.01). I am also trying to implement stacking so I can add lot more models without having to manually tune the weighted average.</p>",
      "rawMarkdown": "tahsin FYI, I tried adding ordinal regression model to my ensemble, which did not help, but adding a model with different pre-processing improved the score by small margin (~0.01). I am also trying to implement stacking so I can add lot more models without having to manually tune the weighted average.",
      "votes": null
    },
    {
      "id": "610930",
      "postDate": "08/29/2019 03:05:03",
      "content": "<p>I agree with you, batch size is important. I tired bs=64 and bs=32, and my best results are coming from bs=32. </p>",
      "rawMarkdown": "I agree with you, batch size is important. I tired bs=64 and bs=32, and my best results are coming from bs=32.",
      "votes": null
    },
    {
      "id": "611231",
      "postDate": "08/29/2019 07:15:40",
      "content": "<p><a href=\"/heye0507\">@heye0507</a>   did you treat the problem as regression or classification ?</p>",
      "rawMarkdown": "heye0507   did you treat the problem as regression or classification ?",
      "votes": null
    },
    {
      "id": "611255",
      "postDate": "08/29/2019 07:29:03",
      "content": "<p>Regression</p>",
      "rawMarkdown": "Regression",
      "votes": null
    },
    {
      "id": "612263",
      "postDate": "08/29/2019 17:26:12",
      "content": "<p>Just to follow up, I can confirm that ensemble works. \nHere is what I did, select the following models</p>\n\n<p>All model pretrained on 2015 data</p>\n\n<ol>\n<li>Effb4, non pre-processing, removed duplicates (local cv 0.912, LB 0.793)</li>\n<li>Effb4, pre-processing, using RAdam opt (local cv 0.9389, LB 0.800)</li>\n<li>Effb4, pre-processing, argumentation[follow DrHB discussion] , fastai apporach (local cv 0.9207, LB 0.800)</li>\n<li>Effb2, pre-processing, argumentation, fastai apporach, img size 260 [local cv 0.9266, LB 0.798]</li>\n</ol>\n\n<p>simply average above all models, boost my LB from 0.804 to 0.814</p>\n\n<p>I have 3 more models (different Effb4] that have LB scored range from 0.8 - 0.804 with 1 resnxt101 model to ensemble.</p>\n\n<p>But this makes me wonder... now it is just ensemble models, and how reliable are this scores? </p>\n\n<p>Also, ensemble is not fun... I have also tried fitting a Linear Regression model on top of them, but that's like weight Kappa to me, you will need to have a close validation = test set, otherwise, I just don't see the theory behind it (random split cv, and calculate weighting on that to approach test distribution? )</p>\n\n<p>Maybe I am wrong?</p>",
      "rawMarkdown": "Just to follow up, I can confirm that ensemble works. \nHere is what I did, select the following models\n\nAll model pretrained on 2015 data\n\n1. Effb4, non pre-processing, removed duplicates (local cv 0.912, LB 0.793)\n2. Effb4, pre-processing, using RAdam opt (local cv 0.9389, LB 0.800)\n3. Effb4, pre-processing, argumentation[follow DrHB discussion] , fastai apporach (local cv 0.9207, LB 0.800)\n4. Effb2, pre-processing, argumentation, fastai apporach, img size 260 [local cv 0.9266, LB 0.798]\n\nsimply average above all models, boost my LB from 0.804 to 0.814\n\nI have 3 more models (different Effb4] that have LB scored range from 0.8 - 0.804 with 1 resnxt101 model to ensemble.\n\nBut this makes me wonder... now it is just ensemble models, and how reliable are this scores? \n\nAlso, ensemble is not fun... I have also tried fitting a Linear Regression model on top of them, but that's like weight Kappa to me, you will need to have a close validation = test set, otherwise, I just don't see the theory behind it (random split cv, and calculate weighting on that to approach test distribution? )\n\nMaybe I am wrong?",
      "votes": null
    },
    {
      "id": "612293",
      "postDate": "08/29/2019 17:52:51",
      "content": "<p>in your ensemble all your models are regression?</p>",
      "rawMarkdown": "in your ensemble all your models are regression?",
      "votes": null
    },
    {
      "id": "612298",
      "postDate": "08/29/2019 17:54:39",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "615669",
      "postDate": "09/02/2019 08:19:44",
      "content": "<p>any idea on how to use discriminative learning here? as there is only one layer group in efficientnet.\nI am having a tough time finetuning. <a href=\"/tanlikesmath\">@tanlikesmath</a> <a href=\"/heye0507\">@heye0507</a> <a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> </p>",
      "rawMarkdown": "any idea on how to use discriminative learning here? as there is only one layer group in efficientnet.\nI am having a tough time finetuning. @tanlikesmath @heye0507 @sidhanthholalkere",
      "votes": null
    },
    {
      "id": "616079",
      "postDate": "09/02/2019 17:01:01",
      "content": "<p><a href=\"https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507\">https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507</a></p>\n\n<p>A novice cut for discriminative lr, sorry to give you a link outside of kaggle</p>\n\n<p>This is a summary of efficient-nets I have tested, I also commented on one of the DrHB's post, but don't want to search for that :(</p>\n\n<p>To be honest, it didn't work for me.</p>",
      "rawMarkdown": "https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507\n\nA novice cut for discriminative lr, sorry to give you a link outside of kaggle\n\nThis is a summary of efficient-nets I have tested, I also commented on one of the DrHB's post, but don't want to search for that :(\n\nTo be honest, it didn't work for me.",
      "votes": null
    },
    {
      "id": "617462",
      "postDate": "09/04/2019 06:44:05",
      "content": "<p>hi dream which is drhb discussion are you talking about..</p>",
      "rawMarkdown": "hi dream which is drhb discussion are you talking about..",
      "votes": null
    },
    {
      "id": "617476",
      "postDate": "09/04/2019 07:08:58",
      "content": "<p>An other great way of improving score is pseudo labelling:\n - train a classifier and take the high confidence test set predictions only (above a certain treshold)\n - or take the predictions of the test set, where many of your models agree on the prediction</p>\n\n<p>Add these test set images to the training set and retrain the model.\nIt gave me and my teammate a 1% boost, even though we had very different models.</p>\n\n<p>You may iterate these steps several times.</p>",
      "rawMarkdown": "An other great way of improving score is pseudo labelling:\n - train a classifier and take the high confidence test set predictions only (above a certain treshold)\n - or take the predictions of the test set, where many of your models agree on the prediction\n\nAdd these test set images to the training set and retrain the model.\nIt gave me and my teammate a 1% boost, even though we had very different models.\n\nYou may iterate these steps several times.",
      "votes": null
    },
    {
      "id": "617585",
      "postDate": "09/04/2019 09:21:53",
      "content": "<p>There might be a risk of overfitting to the public LB</p>",
      "rawMarkdown": "There might be a risk of overfitting to the public LB",
      "votes": null
    },
    {
      "id": "617654",
      "postDate": "09/04/2019 11:01:02",
      "content": "<p>True, but overtfitting to the public LB is better than overfitting to the train set 😃 </p>",
      "rawMarkdown": "True, but overtfitting to the public LB is better than overfitting to the train set 😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 606845,
      "author_name": "ravivadapalli",
      "author_url": "",
      "post_date": "08/24/2019 07:16:24",
      "content": "<p>I have similar problems with B4 &amp; B5. I am sure there are others who did well with these networks. I guess we will get to learn more on the efficientnets after the competition ends.</p>\n\n<p>For B3, I too followed DrHB and did things similar to what you tried. I got LB ranging from 0.78 to 0.819. But, I cannot conclusively say that 0.819 LB is better than 0.8 LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 606977,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/24/2019 11:49:56",
          "content": "<p>Me too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 606847,
      "author_name": "kir486680",
      "author_url": "",
      "post_date": "08/24/2019 07:21:44",
      "content": "<p>Did you try Res-Net 50? Moreover, try to balance the old data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608030,
          "author_name": "zhan2019",
          "author_url": "",
          "post_date": "08/26/2019 08:29:32",
          "content": "<p>How to balance the old data? I find it hard to train no matter which model i use</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 606903,
      "author_name": "agscin",
      "author_url": "",
      "post_date": "08/24/2019 09:26:46",
      "content": "<p>Remember that public \"test\" is only 1,928 images (private one is 13,000). LB score represents only 12.9% of test data.\nSo your experiments does not necessarily will lead to improvements in LB score.</p>\n\n<p>From \"About this competition\":\n\"When Kaggle runs your Kernel privately, it substitutes the private test set and sample submission in place of the public ones. You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately)\". </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 606947,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "08/24/2019 11:03:49",
      "content": "<p>How does your training/valid curves look like?</p>\n\n<p>You also use Grad-CAM, did your models actually capture those scabs/wools correctly in both the training / validation data? My case is to try to make many examples of those missed scabs/wools as much as possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 608089,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/26/2019 10:18:27",
          "content": "<p>hi neuron\ngood to see you second :)\nwould u like to tell the value_counts for output...this will be benchmark to see how far people below you are deviating from the top score :)\nCurrently by looking at value counts m able to figure out how much my submission score be around :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607011,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "08/24/2019 12:47:29",
      "content": "<p>For me, zooming in and TTA(horizontal+ vertical flips) seems to help performance </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 607028,
      "author_name": "larlia",
      "author_url": "",
      "post_date": "08/24/2019 13:28:36",
      "content": "<p>What optimizers/schedulers are you using for EfficientNets? My best results so far have been with AdamW+One cycle policy on B0, but I haven't been able to improve that with B2. I've been trying RMSprop+StepLR on B2, as that is the strategy used in the EfficientNet paper, but my results are worse so far.</p>",
      "votes": null,
      "replies": [
        {
          "id": 607947,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/26/2019 05:43:45",
          "content": "<p>AdamW+One cycle policy is the default in fastai and what I am using...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607085,
      "author_name": "ccjoshua",
      "author_url": "",
      "post_date": "08/24/2019 15:38:46",
      "content": "<p>For my case, Densenet perform better than Resnet. You may try that.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 607169,
      "author_name": "tahsin",
      "author_url": "",
      "post_date": "08/24/2019 17:45:17",
      "content": "<p>Following @DrHB's suggestion is sufficient to get <code>0.8+</code> LB score. But I find it extremely difficult to improve after <code>0.81</code>. My current LB score is an ensemble of 795, 794, 804 and 807 models. My LB score increased by only <code>0.001</code> after adding the 807 model. Any tips about how to boost my score will be appreciated. :)  </p>",
      "votes": null,
      "replies": [
        {
          "id": 607384,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/25/2019 06:11:51",
          "content": "<p>Hi, do you simply average them or did you tune a weighted average? Another weird thing I noticed in my approach is that applying slighty larger weight to a worse-performing model improved the LB... either its the model diversity or I'm simply overfitting to the LB :/ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607403,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/25/2019 07:09:19",
          "content": "<p>I am taking their weighted average. I agree with you. I found out that ensembling models with lower LB scores boost score comparatively higher than the high scoring models.  Weird. :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607413,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/25/2019 07:31:43",
          "content": "<p>If you don't mind answering, are all models in your ensemble from the same approach and how they differ from each other? I am currently experimenting model diversity in ensembling and I don't have a clear indication if diverse = better LB atm</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607418,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/25/2019 08:00:44",
          "content": "<p>Yes. I am using Efficientnet models and using the same validation set for finetuning them. Only using circle cropping as preprocessing. Please let us know about the outcomes of your experiments. :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607454,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/25/2019 09:25:30",
          "content": "<p>I fit a LinearRegression model on the top of the outputs of my EfficientNet models. It gave me 1% boost. From 80.8 to 81.8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607491,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/25/2019 11:56:58",
          "content": "<p>Can you please explain more <a href=\"/nemethpeti\">@nemethpeti</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607502,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/25/2019 12:11:32",
          "content": "<ol>\n<li>Train a number of models on the same train/validation set</li>\n<li>Each will have a raw (float) output (for classifiers, you have to calculate an expected value)</li>\n<li>Predict the raw outputs on the VALIDATION set</li>\n<li>Predict the raw outputs on the TEST set</li>\n<li>Split the validation set into TRAIN2/VALIDATION2 sets</li>\n<li>Fit a Linear/RandomForest/XGBoost on TRAIN2 and optimize to have good result on VALIDATION2</li>\n<li>Take the results of step 4 as input and use the model of step 6 to predict the outputs of the TEST set</li>\n</ol>\n\n<p>You get a significant boost!\nI got over 1% from 3 models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607508,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/25/2019 12:16:14",
          "content": "<p>Thanks for the detailed reply. I shall surely try this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607509,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/25/2019 12:17:30",
          "content": "<p>I used a Linear model with success. At the end it will give a weighted avearge, but the weights are calculated and not manually tuned. Manual tuning is hard if you have more than 2 models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607666,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/25/2019 17:30:56",
          "content": "<p><a href=\"/nemethpeti\">@nemethpeti</a> Thanks for sharing. Will try your method. :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607725,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/25/2019 19:13:16",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> I will certainly update if I get any success, but with the recent downsizing in commits I hope I have time :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607888,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/26/2019 03:45:27",
          "content": "<p>@peter thanks for input..\nwhat is raw output size  :  bs* Output layer nerurons * (1)  ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607929,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/26/2019 05:07:35",
          "content": "<p>1 output per eye / model.\nBut the technic should work with 5 outputs as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607958,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/26/2019 06:34:57",
          "content": "<p>In my case, linear regression on top of my models for ensemble degraded LB score. :/ Must have done something wrong. Here is my code:</p>\n\n<p>&gt; \nfrom sklearn.linear_model import LinearRegression\n&gt;\nimport numpy as np\n&gt;\nval_preds = np.hstack((val_pred1, val_pred2, val_pred3, val_pred4))\n&gt;\nreg = LinearRegression().fit(val_preds, val_label)\n&gt;\nW = reg.coef_\n&gt;\nb  = reg.intercept_\n&gt;\ntest_pred = test_pred1*W[0] + test_pred2*W[1] + test_pred3*W[2] + test_pred4*W[3] + b</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608013,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "08/26/2019 07:57:36",
          "content": "<p>I assume valpredX values are floats!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608021,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/26/2019 08:12:45",
          "content": "<p>yes. <a href=\"/nemethpeti\">@nemethpeti</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608026,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/26/2019 08:24:48",
          "content": "<p>sorry ... trying to understand.. output layer of all eff or any model is  say like  nn.Linear(1024,1) \nso are you giving linear module output  to Xgboost <br>\nX= 4000* 1 ,Y=4000*1   ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610667,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "08/29/2019 00:02:35",
          "content": "<p><a href=\"/tahsin\">@tahsin</a> FYI, I tried adding ordinal regression model to my ensemble, which did not help, but adding a model with different pre-processing improved the score by small margin (~0.01). I am also trying to implement stacking so I can add lot more models without having to manually tune the weighted average.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617476,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "09/04/2019 07:08:58",
          "content": "<p>An other great way of improving score is pseudo labelling:\n - train a classifier and take the high confidence test set predictions only (above a certain treshold)\n - or take the predictions of the test set, where many of your models agree on the prediction</p>\n\n<p>Add these test set images to the training set and retrain the model.\nIt gave me and my teammate a 1% boost, even though we had very different models.</p>\n\n<p>You may iterate these steps several times.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617585,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "09/04/2019 09:21:53",
          "content": "<p>There might be a risk of overfitting to the public LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617654,
          "author_name": "nemethpeti",
          "author_url": "",
          "post_date": "09/04/2019 11:01:02",
          "content": "<p>True, but overtfitting to the public LB is better than overfitting to the train set 😃 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607212,
      "author_name": "heye0507",
      "author_url": "",
      "post_date": "08/24/2019 20:01:27",
      "content": "<p>I did B4, with image size 256, cropping, pre-trained on 2015 data. As my single fold model now sitting comfortably on (0.795-0.8) and 5 fold model gives (0.801-0.804) but hard to push for 0.81... I noticed why people are merging... because you can ensemble easily to push high scores... </p>\n\n<p>So I am now training effb2 for non-cropping but heavy argumentation one, it is very interesting now single fold is around 0.790-0.795 (I alos clean the dataset but removing all same images with different file names). Just at the moment to do more h-para search, kaggle shut all of my kernels off and leave me only 1 to run...</p>\n\n<p>Another interesting experiment I did is remove the nn.Linear layer from efficient net, stick fastai gold head (which inspired by another on going kaggle steel competition), now the network is 2 layer_groups, all the sudden the freezing and unfreezing is start to work. The model single fold can reach 0.78-0.79 range, but since you have two layer_groups, it is very hard to train the network, the lr just kills the body (which is the efficient net backbone). Again, more experiments are needed, but no kernels available. </p>\n\n<p>I guess I am done with this competition, simply just I don't have enough resources to try ideas... all my kernel (just 1) now is doing ensemble of the previous models and is taking forever....</p>\n\n<p>sad :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 607363,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/25/2019 04:55:16",
          "content": "<p><a href=\"/heye0507\">@heye0507</a> I feel you. :/ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607678,
          "author_name": "chinhuic",
          "author_url": "",
          "post_date": "08/25/2019 17:49:09",
          "content": "<p>Good luck! Augmentations so far do not seem to work well for me</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607697,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/25/2019 18:35:22",
          "content": "<p><a href=\"/dreamdragon\">@dreamdragon</a> quite interesting..\nHow you trained using old data, did you use entire old data which is quite huge..?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607884,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/26/2019 03:38:09",
          "content": "<p>thanks for your information! very useful</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609623,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/27/2019 23:46:15",
          "content": "<p>Couple things I want to update, if anyone is interested. </p>\n\n<p>All the things I talked about are pre-trained with 2015 data, used 2019 data as validation. Pick best valid loss after 10 epochs.</p>\n\n<ol>\n<li>Like DrHB's post, argumentation do help, but you have to make sure padding mode you are using if you use fastai, default padding mode is reflection pad, if you also rotate and pre-process the image, you will see the unwanted eye shape in the corner when you rotate. Also, if you pad zero, unwanted black edge will come back. I have very hard time to use nearest (boarder in fastai). Think first if you want to rotate the image (on my experiments, it worked)</li>\n</ol>\n\n<p>What I noticed is that pre-processing / train using original image gives similar LB scores.\nArgumentation used in fastai [max_zoom=1.3, max_rotate=360., flip] for both cases.</p>\n\n<p>depends on your pre-trained situation, effb4 with 8 epochs and lr=1e-3/2 with size=(256,256) gives around LB 0.79+ single fold with either valid = 0.1 or 0.2. But if you just want to have solid 0.79+ LB, set valid = 0.1 (assuming you are using above hyper-parameters)</p>\n\n<p>5 fold cv should give you around 0.801 - 0.804 range</p>\n\n<ol>\n<li><p>About adding fastai standard head, if you look at efficient-net, they stick conv_head then bn then linear, I tried following:</p></li>\n<li><p>remove linear layer, stick fastai-head, which is stacked (max and avg pool), bn, linear, dropout...etc</p></li>\n<li>remove conv_head, stick fastai-head, again, stacked (max and avg pool) ...  </li>\n</ol>\n\n<p>Split layers at head, so you will have efficient net b4 as body, fastai head as head.</p>\n\n<p>Here is what becomes interesting: </p>\n\n<p>In fastai, we said single avg pool gives some information (which is what most of models do, adaptive / global avg pool follow by linear), but if you gives another maxpool, and stack the results together, now you have twice as much as information, it should perform at least as good as previous model. </p>\n\n<p>For both 1/2 situation, the LB score is around 0.79 (single fold, 20% validation without rotate)</p>\n\n<p>Training steps are:</p>\n\n<ol>\n<li>freeze the body except the bn layers, train the head for 5-8 epochs, you will see the kappa score is around 0.9x</li>\n<li>unfreeze the model, \n2a. give head lr, body lr / 10 (standard fastai apporach)\n2b. give head lr, body lr / 100</li>\n</ol>\n\n<p>for both cases, you will see that the model becomes unstable. It barely improves anything. </p>\n\n<p>That makes me wonder, pre-trained on 2015 data is good enough? I am only fine tuning the head?</p>\n\n<p>In comparison,  I split the original effb4 model at model._fc, freeze the body, I found that after 1 epoch, the model is over-fitting,  which is as expected, we only fine tuning the bn and linear layers, which only has few parameters. </p>\n\n<p>I also tried split the model at _conv_head, repeat freeze() and unfreeze(). I found that it has similar result as you dont freeze, just train with consistent lr .</p>\n\n<p>In all cases, discriminative lr doesn't work. </p>\n\n<p>My naive conclusion is, if you just split the efficient model at head / fc layer, freeze / unfreeze can be void, or at least it doesn't seem help.</p>\n\n<p>I already learnt so much from the competition. Thank you all for the help.</p>\n\n<p>I guess next step is to learn ensemble tricks, which I hope I have enough time to try. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609750,
          "author_name": "zhan2019",
          "author_url": "",
          "post_date": "08/28/2019 04:25:17",
          "content": "<p>So just use fastai standard head(AdaptiveConcatPool2d, Flatten, blocks of [nn.BatchNorm1d, nn.Dropout, nn.Linear, nn.ReLU] )? And use consistent lr?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 609850,
          "author_name": "tahsin",
          "author_url": "",
          "post_date": "08/28/2019 06:58:01",
          "content": "<p>Very informative. Thanks a lot.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610652,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "08/28/2019 23:45:05",
          "content": "<p><a href=\"/heye0507\">@heye0507</a> , thanks for the report, I have one question, on my experiments, my results change a lot depending on the batch size I use.</p>\n\n<p>When you use effnetb4 and img size 256, what's your batch size and do you use something like gradient accumulator?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 610930,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/29/2019 03:05:03",
          "content": "<p>I agree with you, batch size is important. I tired bs=64 and bs=32, and my best results are coming from bs=32. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611231,
          "author_name": "virajbagal",
          "author_url": "",
          "post_date": "08/29/2019 07:15:40",
          "content": "<p><a href=\"/heye0507\">@heye0507</a>   did you treat the problem as regression or classification ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 611255,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/29/2019 07:29:03",
          "content": "<p>Regression</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 612263,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/29/2019 17:26:12",
          "content": "<p>Just to follow up, I can confirm that ensemble works. \nHere is what I did, select the following models</p>\n\n<p>All model pretrained on 2015 data</p>\n\n<ol>\n<li>Effb4, non pre-processing, removed duplicates (local cv 0.912, LB 0.793)</li>\n<li>Effb4, pre-processing, using RAdam opt (local cv 0.9389, LB 0.800)</li>\n<li>Effb4, pre-processing, argumentation[follow DrHB discussion] , fastai apporach (local cv 0.9207, LB 0.800)</li>\n<li>Effb2, pre-processing, argumentation, fastai apporach, img size 260 [local cv 0.9266, LB 0.798]</li>\n</ol>\n\n<p>simply average above all models, boost my LB from 0.804 to 0.814</p>\n\n<p>I have 3 more models (different Effb4] that have LB scored range from 0.8 - 0.804 with 1 resnxt101 model to ensemble.</p>\n\n<p>But this makes me wonder... now it is just ensemble models, and how reliable are this scores? </p>\n\n<p>Also, ensemble is not fun... I have also tried fitting a Linear Regression model on top of them, but that's like weight Kappa to me, you will need to have a close validation = test set, otherwise, I just don't see the theory behind it (random split cv, and calculate weighting on that to approach test distribution? )</p>\n\n<p>Maybe I am wrong?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 612293,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/29/2019 17:52:51",
          "content": "<p>in your ensemble all your models are regression?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 612298,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "08/29/2019 17:54:39",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 617462,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/04/2019 06:44:05",
          "content": "<p>hi dream which is drhb discussion are you talking about..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607882,
      "author_name": "sidhanthholalkere",
      "author_url": "",
      "post_date": "08/26/2019 03:35:04",
      "content": "<p>you can get 0.815 from just ensembling efficientnets(regression) that all have lower than 0.8.</p>\n\n<p>Also to \"freeze\" layers you can iterate thru layers andn make certain ones not require grad and then pass all the params with requires grad to a new optimizer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 607946,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/26/2019 05:43:19",
          "content": "<p>Does freezing and then subsequently unfreezing provide any benefit compared to training on the unfrozen model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607966,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/26/2019 06:51:29",
          "content": "<p>how many models, different backbones?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608322,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/26/2019 16:14:14",
          "content": "<p>agreed，ensemle models can easily boost the lb score</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608380,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "08/26/2019 17:17:09",
          "content": "<p>but i cast a doubt on ensembles how well  it can give results on bigger chunk of images in private set. \nPeople may have tried to fit to public ds their thresholds eg. one thing is clear that there are more than 1100 + images of category 2 ,any thing less than that will give you a score of less than 77 on public ds. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608394,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/26/2019 17:39:37",
          "content": "<p>thats also my concern，we are trying our best to fit the public ds while the private ds may have a quite different distribution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608580,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/27/2019 01:09:58",
          "content": "<p>I agree with you. And there are a lot of images in public test set had been preprocessed. \nnot sure private test set will be same.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608588,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "08/27/2019 01:44:00",
          "content": "<p><a href=\"/leixiang\">@leixiang</a> <a href=\"/chanhu\">@chanhu</a> See my post over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-608543\">here</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608671,
          "author_name": "chanhu",
          "author_url": "",
          "post_date": "08/27/2019 04:04:37",
          "content": "<p>Thank you for you nice work. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608704,
          "author_name": "leixiang",
          "author_url": "",
          "post_date": "08/27/2019 04:59:16",
          "content": "<p>thanks for your information</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607978,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "08/26/2019 07:11:45",
      "content": "<p>opposite to you, I can't get a better score from resnet(34/50) or se-resnext50.Is it easy to overfitting than efficientnet?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 608042,
      "author_name": "zhan2019",
      "author_url": "",
      "post_date": "08/26/2019 08:43:45",
      "content": "<p>Playing with seed costs lots of time, any advice?</p>",
      "votes": null,
      "replies": [
        {
          "id": 608092,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "08/26/2019 10:22:32",
          "content": "<p>Don't do it. This will likely lead to overfitting. Just keep it fixed across experiments and try to improve</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 615669,
      "author_name": "nitishsingh41",
      "author_url": "",
      "post_date": "09/02/2019 08:19:44",
      "content": "<p>any idea on how to use discriminative learning here? as there is only one layer group in efficientnet.\nI am having a tough time finetuning. <a href=\"/tanlikesmath\">@tanlikesmath</a> <a href=\"/heye0507\">@heye0507</a> <a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 616079,
          "author_name": "heye0507",
          "author_url": "",
          "post_date": "09/02/2019 17:01:01",
          "content": "<p><a href=\"https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507\">https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507</a></p>\n\n<p>A novice cut for discriminative lr, sorry to give you a link outside of kaggle</p>\n\n<p>This is a summary of efficient-nets I have tested, I also commented on one of the DrHB's post, but don't want to search for that :(</p>\n\n<p>To be honest, it didn't work for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "606817": "I have been struggling to improve my score from my current position. My current score (0.793) is ResNet50 pure classification 5fold CV with crop_image_from_gray preprocessing pretrained on previous competition dataset, with significant augmentation (fastai defaults mainly) and no weight decay (but seems to overfit the training set).\n\nAdditional things I have tried and have not increased my score:\n1. Add weight decay (does not overfit but also decreases performance on LB; kept in further iterations however)\n2. Add circle_crop \n3. Change to ResNet152\n4. Training longer\n5. Changing to regression\n\n\nAt this point I decided to switch to EfficientNets to see what the hype was about. I mainly followed what @DrHB suggested. I pretrained an EfficientNetB3 on previous dataset, the whole network unfrozen (the Luke Melas PyTorch version does not seem to support freezing of the network) , no cropping. I then trained a 5-fold regression CV on the competition dataset with the cropping function, again all unfrozen. I used the correct size of 300x300 and again significant augmentation. It achieved nice results on the CV, but only achieved 0.778, even after playing a little bit with the seed.\n\nWhat am I doing wrong? Any tips either for the ResNet or the EfficientNet architectures?",
    "606845": "I have similar problems with B4 &amp; B5. I am sure there are others who did well with these networks. I guess we will get to learn more on the efficientnets after the competition ends.\n\nFor B3, I too followed DrHB and did things similar to what you tried. I got LB ranging from 0.78 to 0.819. But, I cannot conclusively say that 0.819 LB is better than 0.8 LB",
    "606847": "Did you try Res-Net 50? Moreover, try to balance the old data.",
    "606903": "Remember that public \"test\" is only 1,928 images (private one is 13,000). LB score represents only 12.9% of test data.\nSo your experiments does not necessarily will lead to improvements in LB score.\n\nFrom \"About this competition\":\n\"When Kaggle runs your Kernel privately, it substitutes the private test set and sample submission in place of the public ones. You can plan on the private test set consisting of 20GB of data across 13,000 images (approximately)\".",
    "606947": "How does your training/valid curves look like?\n\nYou also use Grad-CAM, did your models actually capture those scabs/wools correctly in both the training / validation data? My case is to try to make many examples of those missed scabs/wools as much as possible.",
    "606977": "Me too.",
    "607011": "For me, zooming in and TTA(horizontal+ vertical flips) seems to help performance",
    "607028": "What optimizers/schedulers are you using for EfficientNets? My best results so far have been with AdamW+One cycle policy on B0, but I haven't been able to improve that with B2. I've been trying RMSprop+StepLR on B2, as that is the strategy used in the EfficientNet paper, but my results are worse so far.",
    "607085": "For my case, Densenet perform better than Resnet. You may try that.",
    "607169": "Following @DrHB's suggestion is sufficient to get `0.8+` LB score. But I find it extremely difficult to improve after `0.81`. My current LB score is an ensemble of 795, 794, 804 and 807 models. My LB score increased by only `0.001` after adding the 807 model. Any tips about how to boost my score will be appreciated. :)",
    "607212": "I did B4, with image size 256, cropping, pre-trained on 2015 data. As my single fold model now sitting comfortably on (0.795-0.8) and 5 fold model gives (0.801-0.804) but hard to push for 0.81... I noticed why people are merging... because you can ensemble easily to push high scores... \n\nSo I am now training effb2 for non-cropping but heavy argumentation one, it is very interesting now single fold is around 0.790-0.795 (I alos clean the dataset but removing all same images with different file names). Just at the moment to do more h-para search, kaggle shut all of my kernels off and leave me only 1 to run...\n\nAnother interesting experiment I did is remove the nn.Linear layer from efficient net, stick fastai gold head (which inspired by another on going kaggle steel competition), now the network is 2 layer_groups, all the sudden the freezing and unfreezing is start to work. The model single fold can reach 0.78-0.79 range, but since you have two layer_groups, it is very hard to train the network, the lr just kills the body (which is the efficient net backbone). Again, more experiments are needed, but no kernels available. \n\nI guess I am done with this competition, simply just I don't have enough resources to try ideas... all my kernel (just 1) now is doing ensemble of the previous models and is taking forever....\n\nsad :(",
    "607363": "heye0507 I feel you. :/",
    "607384": "Hi, do you simply average them or did you tune a weighted average? Another weird thing I noticed in my approach is that applying slighty larger weight to a worse-performing model improved the LB... either its the model diversity or I'm simply overfitting to the LB :/",
    "607403": "I am taking their weighted average. I agree with you. I found out that ensembling models with lower LB scores boost score comparatively higher than the high scoring models.  Weird. :/",
    "607413": "If you don't mind answering, are all models in your ensemble from the same approach and how they differ from each other? I am currently experimenting model diversity in ensembling and I don't have a clear indication if diverse = better LB atm",
    "607418": "Yes. I am using Efficientnet models and using the same validation set for finetuning them. Only using circle cropping as preprocessing. Please let us know about the outcomes of your experiments. :)",
    "607454": "I fit a LinearRegression model on the top of the outputs of my EfficientNet models. It gave me 1% boost. From 80.8 to 81.8",
    "607491": "Can you please explain more @nemethpeti",
    "607502": "1. Train a number of models on the same train/validation set\n2. Each will have a raw (float) output (for classifiers, you have to calculate an expected value)\n3. Predict the raw outputs on the VALIDATION set\n4. Predict the raw outputs on the TEST set\n5. Split the validation set into TRAIN2/VALIDATION2 sets\n6. Fit a Linear/RandomForest/XGBoost on TRAIN2 and optimize to have good result on VALIDATION2\n7. Take the results of step 4 as input and use the model of step 6 to predict the outputs of the TEST set\n\nYou get a significant boost!\nI got over 1% from 3 models.",
    "607508": "Thanks for the detailed reply. I shall surely try this.",
    "607509": "I used a Linear model with success. At the end it will give a weighted avearge, but the weights are calculated and not manually tuned. Manual tuning is hard if you have more than 2 models.",
    "607666": "nemethpeti Thanks for sharing. Will try your method. :)",
    "607678": "Good luck! Augmentations so far do not seem to work well for me",
    "607697": "dreamdragon quite interesting..\nHow you trained using old data, did you use entire old data which is quite huge..?",
    "607725": "tahsin I will certainly update if I get any success, but with the recent downsizing in commits I hope I have time :(",
    "607882": "you can get 0.815 from just ensembling efficientnets(regression) that all have lower than 0.8.\n\nAlso to \"freeze\" layers you can iterate thru layers andn make certain ones not require grad and then pass all the params with requires grad to a new optimizer.",
    "607884": "thanks for your information! very useful",
    "607888": "peter thanks for input..\nwhat is raw output size  :  bs* Output layer nerurons * (1)  ?",
    "607929": "1 output per eye / model.\nBut the technic should work with 5 outputs as well.",
    "607946": "Does freezing and then subsequently unfreezing provide any benefit compared to training on the unfrozen model?",
    "607947": "AdamW+One cycle policy is the default in fastai and what I am using...",
    "607958": "In my case, linear regression on top of my models for ensemble degraded LB score. :/ Must have done something wrong. Here is my code:\n\n&gt; \nfrom sklearn.linear_model import LinearRegression\n&gt;\nimport numpy as np\n&gt;\nval_preds = np.hstack((val_pred1, val_pred2, val_pred3, val_pred4))\n&gt;\nreg = LinearRegression().fit(val_preds, val_label)\n&gt;\nW = reg.coef_\n&gt;\nb  = reg.intercept_\n&gt;\ntest_pred = test_pred1*W[0] + test_pred2*W[1] + test_pred3*W[2] + test_pred4*W[3] + b",
    "607966": "how many models, different backbones?",
    "607978": "opposite to you, I can't get a better score from resnet(34/50) or se-resnext50.Is it easy to overfitting than efficientnet?",
    "608013": "I assume valpredX values are floats!",
    "608021": "yes. @nemethpeti",
    "608026": "sorry ... trying to understand.. output layer of all eff or any model is  say like  nn.Linear(1024,1) \nso are you giving linear module output  to Xgboost  \nX= 4000* 1 ,Y=4000*1   ?",
    "608030": "How to balance the old data? I find it hard to train no matter which model i use",
    "608042": "Playing with seed costs lots of time, any advice?",
    "608089": "hi neuron\ngood to see you second :)\nwould u like to tell the value_counts for output...this will be benchmark to see how far people below you are deviating from the top score :)\nCurrently by looking at value counts m able to figure out how much my submission score be around :)",
    "608092": "Don't do it. This will likely lead to overfitting. Just keep it fixed across experiments and try to improve",
    "608322": "agreed，ensemle models can easily boost the lb score",
    "608380": "but i cast a doubt on ensembles how well  it can give results on bigger chunk of images in private set. \nPeople may have tried to fit to public ds their thresholds eg. one thing is clear that there are more than 1100 + images of category 2 ,any thing less than that will give you a score of less than 77 on public ds.",
    "608394": "thats also my concern，we are trying our best to fit the public ds while the private ds may have a quite different distribution",
    "608580": "I agree with you. And there are a lot of images in public test set had been preprocessed. \nnot sure private test set will be same.",
    "608588": "leixiang @chanhu See my post over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105763#latest-608543)",
    "608671": "Thank you for you nice work.",
    "608704": "thanks for your information",
    "609623": "Couple things I want to update, if anyone is interested. \n\nAll the things I talked about are pre-trained with 2015 data, used 2019 data as validation. Pick best valid loss after 10 epochs.\n\n1. Like DrHB's post, argumentation do help, but you have to make sure padding mode you are using if you use fastai, default padding mode is reflection pad, if you also rotate and pre-process the image, you will see the unwanted eye shape in the corner when you rotate. Also, if you pad zero, unwanted black edge will come back. I have very hard time to use nearest (boarder in fastai). Think first if you want to rotate the image (on my experiments, it worked)\n\nWhat I noticed is that pre-processing / train using original image gives similar LB scores.\nArgumentation used in fastai [max_zoom=1.3, max_rotate=360., flip] for both cases.\n\ndepends on your pre-trained situation, effb4 with 8 epochs and lr=1e-3/2 with size=(256,256) gives around LB 0.79+ single fold with either valid = 0.1 or 0.2. But if you just want to have solid 0.79+ LB, set valid = 0.1 (assuming you are using above hyper-parameters)\n\n5 fold cv should give you around 0.801 - 0.804 range\n\n2. About adding fastai standard head, if you look at efficient-net, they stick conv_head then bn then linear, I tried following:\n\n1. remove linear layer, stick fastai-head, which is stacked (max and avg pool), bn, linear, dropout...etc\n2. remove conv_head, stick fastai-head, again, stacked (max and avg pool) ...  \n\nSplit layers at head, so you will have efficient net b4 as body, fastai head as head.\n\nHere is what becomes interesting: \n\nIn fastai, we said single avg pool gives some information (which is what most of models do, adaptive / global avg pool follow by linear), but if you gives another maxpool, and stack the results together, now you have twice as much as information, it should perform at least as good as previous model. \n\nFor both 1/2 situation, the LB score is around 0.79 (single fold, 20% validation without rotate)\n\nTraining steps are:\n\n1. freeze the body except the bn layers, train the head for 5-8 epochs, you will see the kappa score is around 0.9x\n2. unfreeze the model, \n2a. give head lr, body lr / 10 (standard fastai apporach)\n2b. give head lr, body lr / 100\n\nfor both cases, you will see that the model becomes unstable. It barely improves anything. \n\nThat makes me wonder, pre-trained on 2015 data is good enough? I am only fine tuning the head?\n\nIn comparison,  I split the original effb4 model at model._fc, freeze the body, I found that after 1 epoch, the model is over-fitting,  which is as expected, we only fine tuning the bn and linear layers, which only has few parameters. \n\nI also tried split the model at _conv_head, repeat freeze() and unfreeze(). I found that it has similar result as you dont freeze, just train with consistent lr .\n\nIn all cases, discriminative lr doesn't work. \n\nMy naive conclusion is, if you just split the efficient model at head / fc layer, freeze / unfreeze can be void, or at least it doesn't seem help.\n\nI already learnt so much from the competition. Thank you all for the help.\n\nI guess next step is to learn ensemble tricks, which I hope I have enough time to try.",
    "609750": "So just use fastai standard head(AdaptiveConcatPool2d, Flatten, blocks of [nn.BatchNorm1d, nn.Dropout, nn.Linear, nn.ReLU] )? And use consistent lr?",
    "609850": "Very informative. Thanks a lot.",
    "610652": "heye0507 , thanks for the report, I have one question, on my experiments, my results change a lot depending on the batch size I use.\n\nWhen you use effnetb4 and img size 256, what's your batch size and do you use something like gradient accumulator?",
    "610667": "tahsin FYI, I tried adding ordinal regression model to my ensemble, which did not help, but adding a model with different pre-processing improved the score by small margin (~0.01). I am also trying to implement stacking so I can add lot more models without having to manually tune the weighted average.",
    "610930": "I agree with you, batch size is important. I tired bs=64 and bs=32, and my best results are coming from bs=32.",
    "611231": "heye0507   did you treat the problem as regression or classification ?",
    "611255": "Regression",
    "612263": "Just to follow up, I can confirm that ensemble works. \nHere is what I did, select the following models\n\nAll model pretrained on 2015 data\n\n1. Effb4, non pre-processing, removed duplicates (local cv 0.912, LB 0.793)\n2. Effb4, pre-processing, using RAdam opt (local cv 0.9389, LB 0.800)\n3. Effb4, pre-processing, argumentation[follow DrHB discussion] , fastai apporach (local cv 0.9207, LB 0.800)\n4. Effb2, pre-processing, argumentation, fastai apporach, img size 260 [local cv 0.9266, LB 0.798]\n\nsimply average above all models, boost my LB from 0.804 to 0.814\n\nI have 3 more models (different Effb4] that have LB scored range from 0.8 - 0.804 with 1 resnxt101 model to ensemble.\n\nBut this makes me wonder... now it is just ensemble models, and how reliable are this scores? \n\nAlso, ensemble is not fun... I have also tried fitting a Linear Regression model on top of them, but that's like weight Kappa to me, you will need to have a close validation = test set, otherwise, I just don't see the theory behind it (random split cv, and calculate weighting on that to approach test distribution? )\n\nMaybe I am wrong?",
    "612293": "in your ensemble all your models are regression?",
    "612298": "yes",
    "615669": "any idea on how to use discriminative learning here? as there is only one layer group in efficientnet.\nI am having a tough time finetuning. @tanlikesmath @heye0507 @sidhanthholalkere",
    "616079": "https://forums.fast.ai/t/efficientnet/46978/95?u=heye0507\n\nA novice cut for discriminative lr, sorry to give you a link outside of kaggle\n\nThis is a summary of efficient-nets I have tested, I also commented on one of the DrHB's post, but don't want to search for that :(\n\nTo be honest, it didn't work for me.",
    "617462": "hi dream which is drhb discussion are you talking about..",
    "617476": "An other great way of improving score is pseudo labelling:\n - train a classifier and take the high confidence test set predictions only (above a certain treshold)\n - or take the predictions of the test set, where many of your models agree on the prediction\n\nAdd these test set images to the training set and retrain the model.\nIt gave me and my teammate a 1% boost, even though we had very different models.\n\nYou may iterate these steps several times.",
    "617585": "There might be a risk of overfitting to the public LB",
    "617654": "True, but overtfitting to the public LB is better than overfitting to the train set 😃"
  },
  "source": "meta"
}