{
  "id": 106155,
  "title": "Rethink image size",
  "url": "/competitions/aptos2019-blindness-detection/discussion/106155",
  "author_name": "hirune924",
  "post_date": "2019-08-28T13:06:03.875000",
  "votes": 27,
  "comment_count": 110,
  "views": 0,
  "content": "<p>This is the last chance to rethink the image size because we have little time.\nI am currently using only 256. I've already tried a larger size but it only took a long time to train and didn't give good score. What is your best image size？</p>",
  "messages": [
    {
      "id": 610122,
      "postDate": "2019-08-28T13:06:03.877Z",
      "content": "<p>This is the last chance to rethink the image size because we have little time.\nI am currently using only 256. I've already tried a larger size but it only took a long time to train and didn't give good score. What is your best image size？</p>",
      "rawMarkdown": "This is the last chance to rethink the image size because we have little time.\nI am currently using only 256. I've already tried a larger size but it only took a long time to train and didn't give good score. What is your best image size？",
      "votes": 27
    },
    {
      "id": 612806,
      "postDate": "2019-08-30T02:39:47.713Z",
      "content": "<p>My single models are, <br>\n1. 300 Regression EfficientNet B3, LB 0.830\n2. 260 Regression EfficientNet B2, LB 0.820</p>\n\n<p>Our current place is from Model 1.  </p>",
      "rawMarkdown": "My single models are,  \n1. 300 Regression EfficientNet B3, LB 0.830\n2. 260 Regression EfficientNet B2, LB 0.820\n\nOur current place is from Model 1.  ",
      "votes": 9,
      "replies": [
        {
          "id": 614241,
          "postDate": "2019-08-31T08:46:05.973Z",
          "content": "<p>hi max well did you try any preprocessing like bens cropping or any other augs..and how much is you train loss for each at the time of making prediction</p>",
          "rawMarkdown": "hi max well did you try any preprocessing like bens cropping or any other augs..and how much is you train loss for each at the time of making prediction"
        },
        {
          "id": 614447,
          "postDate": "2019-08-31T13:55:30.117Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 614458,
          "postDate": "2019-08-31T14:14:21.790Z",
          "content": "<p><a href=\"/pranavpulijala\">@pranavpulijala</a> \nMy CV on <strong>APTOS 2019 train dataset</strong> is about 0.93.\nIf I add higher grade examples of DRD2015 for evaluation, that will go down.\nI think it would be meaningless to compare CV without using same datasets.  </p>",
          "rawMarkdown": "@pranavpulijala \nMy CV on **APTOS 2019 train dataset** is about 0.93.\nIf I add higher grade examples of DRD2015 for evaluation, that will go down.\nI think it would be meaningless to compare CV without using same datasets.  "
        },
        {
          "id": 614514,
          "postDate": "2019-08-31T15:50:24.747Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 611126,
      "postDate": "2019-08-29T05:50:34.947Z",
      "content": "<p>I use 256 preprocessed by mean subtraction.\nMy best single model (EfficientNet-B4) got 0.829.\nI've tried 320 and 512, but it was a waste of time.</p>",
      "rawMarkdown": "I use 256 preprocessed by mean subtraction.\nMy best single model (EfficientNet-B4) got 0.829.\nI've tried 320 and 512, but it was a waste of time.",
      "votes": 9,
      "replies": [
        {
          "id": 611224,
          "postDate": "2019-08-29T07:10:17.477Z",
          "content": "<p>Please give some tips.  What data you used, any other preprocessing than mean subtraction, data augmentations and whether you treated the problem as regression or classification ? Thanks in advance</p>",
          "rawMarkdown": "Please give some tips.  What data you used, any other preprocessing than mean subtraction, data augmentations and whether you treated the problem as regression or classification ? Thanks in advance"
        },
        {
          "id": 611268,
          "postDate": "2019-08-29T07:37:13.153Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> You are still use old and current data together？</p>",
          "rawMarkdown": "@octpath0302 You are still use old and current data together？"
        },
        {
          "id": 611327,
          "postDate": "2019-08-29T08:02:59.667Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> \nI use only rotation and horizontal flip in training phase. no tta. regression.\nSince there are many class 2 data in public test set, training data are sampled with adjusted distribution. My LB score so far improved best when the distribution was adjusted.\n<a href=\"/garybios\">@garybios</a> Yes. I use all the datasets listed in 'external data thread' at the same time.</p>",
          "rawMarkdown": "@virajbagal \nI use only rotation and horizontal flip in training phase. no tta. regression.\nSince there are many class 2 data in public test set, training data are sampled with adjusted distribution. My LB score so far improved best when the distribution was adjusted.\n@garybios Yes. I use all the datasets listed in 'external data thread' at the same time.",
          "votes": 3
        },
        {
          "id": 611434,
          "postDate": "2019-08-29T08:58:39.057Z",
          "content": "<p>nice, thanks a lot. Could you tell me how do you adjust the data distribution?</p>",
          "rawMarkdown": "nice, thanks a lot. Could you tell me how do you adjust the data distribution?"
        },
        {
          "id": 611504,
          "postDate": "2019-08-29T09:48:49.343Z",
          "content": "<p>A trained model was used to predict the class of public test sets and estimate the distribution.\nThere are various ways of sampling. Whether to do it only first, or every epoch, or to make a K-Fold.</p>",
          "rawMarkdown": "A trained model was used to predict the class of public test sets and estimate the distribution.\nThere are various ways of sampling. Whether to do it only first, or every epoch, or to make a K-Fold.",
          "votes": 1
        },
        {
          "id": 611534,
          "postDate": "2019-08-29T10:06:30.907Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  that this mean you do not crop your images? (you said you mean subtract only)</p>",
          "rawMarkdown": "@octpath0302  that this mean you do not crop your images? (you said you mean subtract only)"
        },
        {
          "id": 611548,
          "postDate": "2019-08-29T10:13:33.163Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  I thought of sampling based on the class distribution of public test dataset. But I thought that way would overfit public test dataset. I think the distribution of private test dataset is more similar to train dataset than private test dataset. What do you think about this？</p>",
          "rawMarkdown": "@octpath0302  I thought of sampling based on the class distribution of public test dataset. But I thought that way would overfit public test dataset. I think the distribution of private test dataset is more similar to train dataset than private test dataset. What do you think about this？",
          "votes": 1
        },
        {
          "id": 611563,
          "postDate": "2019-08-29T10:20:47.900Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> you can probe it and see. someone from the top has done it for sure.</p>",
          "rawMarkdown": "@hirune924 you can probe it and see. someone from the top has done it for sure."
        },
        {
          "id": 611566,
          "postDate": "2019-08-29T10:21:53.580Z",
          "content": "<p><a href=\"/valanm\">@valanm</a> Sorry for misleading! In the preprocessing stage, I use auto cropping which is similar to the 1st solution on the previous competition, and padding.\nI make preprocessing to get robustness against black area at corner of each image.</p>",
          "rawMarkdown": "@valanm Sorry for misleading! In the preprocessing stage, I use auto cropping which is similar to the 1st solution on the previous competition, and padding.\nI make preprocessing to get robustness against black area at corner of each image.",
          "votes": 2
        },
        {
          "id": 611568,
          "postDate": "2019-08-29T10:22:22.813Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> I also thought that way is easy to overfit public lb, but we do not know the distribution of private test dataset.</p>",
          "rawMarkdown": "@hirune924 I also thought that way is easy to overfit public lb, but we do not know the distribution of private test dataset."
        },
        {
          "id": 611577,
          "postDate": "2019-08-29T10:27:00.067Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> I also think there is a danger of overfit. However, it is generally possible that there is a lot of data in class 0 (normal) and class 2 (middle). Therefore, I assume that the public test set and the private test set are randomly divided, so far.</p>",
          "rawMarkdown": "@hirune924 I also think there is a danger of overfit. However, it is generally possible that there is a lot of data in class 0 (normal) and class 2 (middle). Therefore, I assume that the public test set and the private test set are randomly divided, so far.",
          "votes": 2
        },
        {
          "id": 611765,
          "postDate": "2019-08-29T12:24:41.440Z",
          "content": "<p><a href=\"/garybios\">@garybios</a> <a href=\"/valanm\">@valanm</a> <a href=\"/virajbagal\">@virajbagal</a>  How about you? Which size/preprocessing do you use?</p>",
          "rawMarkdown": "@garybios @valanm @virajbagal  How about you? Which size/preprocessing do you use?"
        },
        {
          "id": 611792,
          "postDate": "2019-08-29T12:46:45.253Z",
          "content": "<p>300px b3 classification gives me .822, for preprocessing just cropping the object.</p>",
          "rawMarkdown": "300px b3 classification gives me .822, for preprocessing just cropping the object.",
          "votes": 2
        },
        {
          "id": 611811,
          "postDate": "2019-08-29T12:59:18.987Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> 224, no preprocessing.\n<a href=\"/valanm\">@valanm</a> nice job, you made a great classification model.</p>",
          "rawMarkdown": "@octpath0302 224, no preprocessing.\n@valanm nice job, you made a great classification model.",
          "votes": 1
        },
        {
          "id": 611880,
          "postDate": "2019-08-29T13:37:50.370Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  256, cropping image and training only on new data gave me 0.796. Despite adding augmentations, my model trained on OLD Data with New Data as validation is overfitting and always predicting 0. I tried to balance the classes by downsampling and upsampling with each class around 5000 images but the score is really poor.  I am desperately trying again and again to train the model on old data. </p>",
          "rawMarkdown": "@octpath0302  256, cropping image and training only on new data gave me 0.796. Despite adding augmentations, my model trained on OLD Data with New Data as validation is overfitting and always predicting 0. I tried to balance the classes by downsampling and upsampling with each class around 5000 images but the score is really poor.  I am desperately trying again and again to train the model on old data. ",
          "votes": 2
        },
        {
          "id": 611929,
          "postDate": "2019-08-29T14:05:24.153Z",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> Hmm. . that's strange. But in my case too, augmentation is hardly working. And I use only downsampling (I have not tried upsampling).\nAre CV and LB (or losses) somewhat correlated? Do your CV get &gt;0.85?</p>",
          "rawMarkdown": "@virajbagal Hmm. . that's strange. But in my case too, augmentation is hardly working. And I use only downsampling (I have not tried upsampling).\nAre CV and LB (or losses) somewhat correlated? Do your CV get &gt;0.85?",
          "votes": 1
        },
        {
          "id": 611945,
          "postDate": "2019-08-29T14:11:41.657Z",
          "content": "<p>My CV goes upto 92. But predicts all zeros. </p>",
          "rawMarkdown": "My CV goes upto 92. But predicts all zeros. "
        },
        {
          "id": 611950,
          "postDate": "2019-08-29T14:18:49.520Z",
          "content": "<p>My CV has never reached to 0.9. So, I think that overfit or something is happening in your case..\n(too small validation set or .. I don't know.)\nIn my case, CV is around 0.87 to 0.88 and LB is around 0.81 to 0.83.</p>",
          "rawMarkdown": "My CV has never reached to 0.9. So, I think that overfit or something is happening in your case..\n(too small validation set or .. I don't know.)\nIn my case, CV is around 0.87 to 0.88 and LB is around 0.81 to 0.83.",
          "votes": 1
        },
        {
          "id": 612069,
          "postDate": "2019-08-29T15:30:18.047Z",
          "content": "<p>Me too.. </p>",
          "rawMarkdown": "Me too.. "
        },
        {
          "id": 612070,
          "postDate": "2019-08-29T15:31:28.803Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> , thanks for sharing the details. A quick question on the external data usage. External datasets have different labels. Some of them have DR levels (0 to 4 or 0 to 3), some have segmented masks, some have counts of microaneurisms, exudates, etc. How are you managing with such desparate labels. Are you somehow creating uniform labels for all the external datasets ?</p>",
          "rawMarkdown": "@octpath0302 , thanks for sharing the details. A quick question on the external data usage. External datasets have different labels. Some of them have DR levels (0 to 4 or 0 to 3), some have segmented masks, some have counts of microaneurisms, exudates, etc. How are you managing with such desparate labels. Are you somehow creating uniform labels for all the external datasets ?"
        },
        {
          "id": 612090,
          "postDate": "2019-08-29T15:46:40.437Z",
          "content": "<p><a href=\"/ravivadapalli\">@ravivadapalli</a> I use datasets which DR levels are 0 to 4 as labeled dataset, and the others are as unlabeled (pseudo-labeled or something) dataset. \nI don't try to use any other labels and masks because there is little time to Kaggle..</p>",
          "rawMarkdown": "@ravivadapalli I use datasets which DR levels are 0 to 4 as labeled dataset, and the others are as unlabeled (pseudo-labeled or something) dataset. \nI don't try to use any other labels and masks because there is little time to Kaggle..",
          "votes": 3
        },
        {
          "id": 612215,
          "postDate": "2019-08-29T16:58:51.740Z",
          "content": "<p>I see. thank you <a href=\"/octpath0302\">@octpath0302</a> for sharing the details. </p>",
          "rawMarkdown": "I see. thank you @octpath0302 for sharing the details. "
        },
        {
          "id": 612309,
          "postDate": "2019-08-29T18:02:23.947Z",
          "content": "<p>Hi, <a href=\"/octpath0302\">@octpath0302</a>. Thank you for sharing so much of your approach. Could you explain more how you apply mean subtraction to your images? </p>",
          "rawMarkdown": "Hi, @octpath0302. Thank you for sharing so much of your approach. Could you explain more how you apply mean subtraction to your images? "
        },
        {
          "id": 612602,
          "postDate": "2019-08-29T22:19:20.673Z",
          "content": "<p>I refer to the 1st place solution of the previous competition.\n<a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/discussion/15801#latest-370950\">https://www.kaggle.com/c/diabetic-retinopathy-detection/discussion/15801#latest-370950</a></p>",
          "rawMarkdown": "I refer to the 1st place solution of the previous competition.\nhttps://www.kaggle.com/c/diabetic-retinopathy-detection/discussion/15801#latest-370950",
          "votes": 1
        },
        {
          "id": 612609,
          "postDate": "2019-08-29T22:40:32.370Z",
          "content": "<p>Got it, thanks! </p>",
          "rawMarkdown": "Got it, thanks! "
        },
        {
          "id": 613421,
          "postDate": "2019-08-30T13:47:13.103Z",
          "content": "<p>May I ask what is your train setting?</p>",
          "rawMarkdown": "May I ask what is your train setting?"
        },
        {
          "id": 613587,
          "postDate": "2019-08-30T16:31:48.790Z",
          "content": "<p>Hi <a href=\"/octpath0302\">@octpath0302</a> , did you observe big drop in CV after applying downsampling but LB may be the same or even better? </p>",
          "rawMarkdown": "Hi @octpath0302 , did you observe big drop in CV after applying downsampling but LB may be the same or even better? "
        },
        {
          "id": 613917,
          "postDate": "2019-08-31T00:15:57.010Z",
          "content": "<p><a href=\"/andalpha\">@andalpha</a> I have provided a lot of information, so I can't do any more..\n<a href=\"/naivelamb\">@naivelamb</a>  I'm not sure because I've never trained by applying downsampling only. After applying downsampling and data addition, I observed that CV was almost unchanged and LB was getting better.</p>",
          "rawMarkdown": "@andalpha I have provided a lot of information, so I can't do any more..\n@naivelamb  I'm not sure because I've never trained by applying downsampling only. After applying downsampling and data addition, I observed that CV was almost unchanged and LB was getting better."
        },
        {
          "id": 613932,
          "postDate": "2019-08-31T00:37:24.717Z",
          "content": "<p>By data addition you mean mixing external data (2015 for example) with current data together to train the model, right?</p>",
          "rawMarkdown": "By data addition you mean mixing external data (2015 for example) with current data together to train the model, right?"
        },
        {
          "id": 613943,
          "postDate": "2019-08-31T00:43:31.403Z",
          "content": "<p>yes.</p>",
          "rawMarkdown": "yes."
        },
        {
          "id": 614227,
          "postDate": "2019-08-31T08:28:11.460Z",
          "content": "<p>hi oct when you training distribution sampling wrt to public test set.. what you mean\nis it taking more class 2 data or how is it ...</p>",
          "rawMarkdown": "hi oct when you training distribution sampling wrt to public test set.. what you mean\nis it taking more class 2 data or how is it ...\n"
        },
        {
          "id": 614234,
          "postDate": "2019-08-31T08:40:13.860Z",
          "content": "<p>@valann by b3 classification you mean not using regression?</p>",
          "rawMarkdown": "@valann by b3 classification you mean not using regression?"
        },
        {
          "id": 614247,
          "postDate": "2019-08-31T08:52:22.140Z",
          "content": "<p>yes I only use classification</p>",
          "rawMarkdown": "yes I only use classification"
        },
        {
          "id": 614256,
          "postDate": "2019-08-31T09:23:07.180Z",
          "content": "<p><a href=\"/octpath\">@octpath</a> which external data set you using.. \nThanks in advance</p>",
          "rawMarkdown": "@octpath which external data set you using.. \nThanks in advance"
        },
        {
          "id": 614660,
          "postDate": "2019-08-31T21:13:24.103Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 614702,
          "postDate": "2019-08-31T23:52:08.133Z",
          "content": "<p>do you use optimized thresholds?</p>",
          "rawMarkdown": "do you use optimized thresholds?"
        },
        {
          "id": 614756,
          "postDate": "2019-09-01T02:53:29.483Z",
          "content": "<p>@valan  how did you work with wierd crop patches produced by fastai ...I found some blacks borders with zero padding mode,or circle arcs at border with reflection ...</p>",
          "rawMarkdown": "@valan  how did you work with wierd crop patches produced by fastai ...I found some blacks borders with zero padding mode,or circle arcs at border with reflection ..."
        },
        {
          "id": 615136,
          "postDate": "2019-09-01T14:01:40.210Z",
          "content": "<p><a href=\"/ggoldszmidt\">@ggoldszmidt</a> I mean local mean subtraction (high-pass filtering), so not standardized. Almost the same as 1st place solution of the previous competition.\n<a href=\"/tanlikesmath\">@tanlikesmath</a> Yes. Optimized.\n<a href=\"/jaideepvalani\">@jaideepvalani</a> \n&gt; hi oct when you training distribution sampling wrt to public test set.. what you mean\n&gt; is it taking more class 2 data or how is it …</p>\n\n<p>Sorry, I can't catch your meaning. You can calculate the public test class distribution with a model (e. g. from public kernel). \nExternal data I'm using are listed up on the official thread.</p>",
          "rawMarkdown": "@ggoldszmidt I mean local mean subtraction (high-pass filtering), so not standardized. Almost the same as 1st place solution of the previous competition.\n@tanlikesmath Yes. Optimized.\n@jaideepvalani \n&gt; hi oct when you training distribution sampling wrt to public test set.. what you mean\n&gt; is it taking more class 2 data or how is it …\n\nSorry, I can't catch your meaning. You can calculate the public test class distribution with a model (e. g. from public kernel). \nExternal data I'm using are listed up on the official thread."
        },
        {
          "id": 615160,
          "postDate": "2019-09-01T14:50:28.137Z",
          "content": "<p>hi oct thanks for reply..\n1) in training set we have 3000 eg . how we decide on the ratio of distribution while sampling say we have batch size of 32..\n2) I tried to follow the external thread none of them have labels ,if i missed any link could u past the one that you used in ..\n3) which public kernel are talking about</p>",
          "rawMarkdown": "hi oct thanks for reply..\n1) in training set we have 3000 eg . how we decide on the ratio of distribution while sampling say we have batch size of 32..\n2) I tried to follow the external thread none of them have labels ,if i missed any link could u past the one that you used in ..\n3) which public kernel are talking about\n"
        },
        {
          "id": 615176,
          "postDate": "2019-09-01T15:05:11.730Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \n1) I mean not sampling mini-batch, but how to re-construct labeled dataset.\n2) 3) Most of the external datasets are unlabeled datasets. But you can pseudo-label any images by using high-performance model from public kernels. If you already have good models, you can use them. </p>",
          "rawMarkdown": "@jaideepvalani \n1) I mean not sampling mini-batch, but how to re-construct labeled dataset.\n2) 3) Most of the external datasets are unlabeled datasets. But you can pseudo-label any images by using high-performance model from public kernels. If you already have good models, you can use them. "
        },
        {
          "id": 615200,
          "postDate": "2019-09-01T15:46:54.470Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> i read the documentation and adopted the code. I default would be good enough then data science wouldn't be hot on job market. </p>",
          "rawMarkdown": "@jaideepvalani i read the documentation and adopted the code. I default would be good enough then data science wouldn't be hot on job market. "
        }
      ]
    },
    {
      "id": 615912,
      "postDate": "2019-09-02T13:47:37.107Z",
      "content": "<p>Here are some of my results.  I did on image size.  All 5x CV </p>\n\n<p><code>EfficientNet B5,  224 LB: 0.807</code>\n<code>EfficientNet B5,  300 LB: 0.812</code>\n<code>EfficientNet B5,  380 LB: 0.816</code>\n<code>EfficientNet B5,  456 LB: 0.823</code></p>\n\n<p><code>EfficientNet B4,  224 LB: 0.800</code>\n<code>EfficientNet B4,  300 LB: 0.811</code>\n<code>EfficientNet B4,  380 LB: 0.816</code></p>",
      "rawMarkdown": "Here are some of my results.  I did on image size.  All 5x CV \n\n`EfficientNet B5,  224 LB: 0.807`\n`EfficientNet B5,  300 LB: 0.812`\n`EfficientNet B5,  380 LB: 0.816`\n`EfficientNet B5,  456 LB: 0.823`\n\n`EfficientNet B4,  224 LB: 0.800`\n`EfficientNet B4,  300 LB: 0.811`\n`EfficientNet B4,  380 LB: 0.816`\n\n",
      "votes": 10,
      "replies": [
        {
          "id": 615960,
          "postDate": "2019-09-02T14:46:31.283Z",
          "content": "<p>@dir HB  good results,your kernel public appreaciated by many kagglers and including me :)\nwould you like to tell the class distribution..  </p>",
          "rawMarkdown": "@dir HB  good results,your kernel public appreaciated by many kagglers and including me :)\nwould you like to tell the class distribution..  ",
          "votes": 2
        },
        {
          "id": 615973,
          "postDate": "2019-09-02T14:55:46.567Z",
          "content": "<p>Glad to hear that kernel helped =) My final distributions are very similar to some that found in this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#614198\">thread</a> . =) </p>",
          "rawMarkdown": "Glad to hear that kernel helped =) My final distributions are very similar to some that found in this [thread](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#614198) . =) "
        },
        {
          "id": 615975,
          "postDate": "2019-09-02T14:59:08.097Z",
          "content": "<p>I tried doing 5x CV with EfficientNet B5 but I only get .79-.792. I dont know what else to do.</p>",
          "rawMarkdown": "I tried doing 5x CV with EfficientNet B5 but I only get .79-.792. I dont know what else to do.",
          "votes": 1
        },
        {
          "id": 615985,
          "postDate": "2019-09-02T15:07:34.797Z",
          "content": "<p>They are diffrent ways how you can improve score.  Have a look some threads like [this] (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105100#latest-615864\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105100#latest-615864</a>) . All of them have good tips. </p>\n\n<p>But in genral. \n1) Try diffrent preprocessing \n2) Try diffrent augmentations \n3) Try combining of diffrent dataset. </p>\n\n<p>Forum is full of tips =) \nGood luck to you =) </p>",
          "rawMarkdown": "They are diffrent ways how you can improve score.  Have a look some threads like [this] (https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105100#latest-615864) . All of them have good tips. \n\nBut in genral. \n1) Try diffrent preprocessing \n2) Try diffrent augmentations \n3) Try combining of diffrent dataset. \n\nForum is full of tips =) \nGood luck to you =) \n\n",
          "votes": 2
        },
        {
          "id": 615987,
          "postDate": "2019-09-02T15:08:00.310Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> when you increase img size, do you also increase your parameters, like learning rate or epochs? I'm asking this because for me it's harder to keep models with large img size learning.</p>",
          "rawMarkdown": "@drhabib when you increase img size, do you also increase your parameters, like learning rate or epochs? I'm asking this because for me it's harder to keep models with large img size learning.",
          "votes": 1
        },
        {
          "id": 615996,
          "postDate": "2019-09-02T15:17:51.187Z",
          "content": "<p>Good question. </p>\n\n<p>I ususally run lr_finder and select learning rate based on this, but in most casses its very similar across diffrent sizes. Also regarding number of Epochs if you are intrested,  for 224 I run for 20 epochs and for rest image sizes 5 epochs is enogh. Again it depends also on which learning rate policy you are using in my case its one_cycle_policy.</p>\n\n<p>Hope it helps =)</p>",
          "rawMarkdown": "Good question. \n\nI ususally run lr_finder and select learning rate based on this, but in most casses its very similar across diffrent sizes. Also regarding number of Epochs if you are intrested,  for 224 I run for 20 epochs and for rest image sizes 5 epochs is enogh. Again it depends also on which learning rate policy you are using in my case its one_cycle_policy.\n\nHope it helps =)\n",
          "votes": 5
        },
        {
          "id": 616011,
          "postDate": "2019-09-02T15:44:01.563Z",
          "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a> , I have tried many policies, also I use Keras, but the bigger models don't train longer than 10 epochs reliably.</p>",
          "rawMarkdown": "Thanks @drhabib , I have tried many policies, also I use Keras, but the bigger models don't train longer than 10 epochs reliably.",
          "votes": 1
        },
        {
          "id": 616016,
          "postDate": "2019-09-02T15:47:02.880Z",
          "content": "<p>which optimizer are you using ?</p>",
          "rawMarkdown": "which optimizer are you using ?"
        },
        {
          "id": 616022,
          "postDate": "2019-09-02T15:58:23.130Z",
          "content": "<p>For me what seems to work is either Adam or RAdam, but also with LR policies, like Cosine or cycle.</p>\n\n<p>So far my best single model result comes from B5 224, using Adam and cosine LR (almost 0.8)</p>",
          "rawMarkdown": "For me what seems to work is either Adam or RAdam, but also with LR policies, like Cosine or cycle.\n\nSo far my best single model result comes from B5 224, using Adam and cosine LR (almost 0.8)",
          "votes": 1
        },
        {
          "id": 616051,
          "postDate": "2019-09-02T16:30:58.663Z",
          "content": "<p>Do you use weights from lower resolution or u always start from imagenet?</p>",
          "rawMarkdown": "Do you use weights from lower resolution or u always start from imagenet?",
          "votes": 1
        },
        {
          "id": 616062,
          "postDate": "2019-09-02T16:41:17.920Z",
          "content": "<p><a href=\"/valanm\">@valanm</a> , I'm starting from imagenet. By lower resolution, do you mean pretraining from the older dataset?</p>",
          "rawMarkdown": "@valanm , I'm starting from imagenet. By lower resolution, do you mean pretraining from the older dataset?",
          "votes": 1
        },
        {
          "id": 616068,
          "postDate": "2019-09-02T16:49:07.713Z",
          "content": "<p>For instance, train 20 epochs on 224 and then 5 more with increased resolution to 456</p>",
          "rawMarkdown": "For instance, train 20 epochs on 224 and then 5 more with increased resolution to 456",
          "votes": 4
        },
        {
          "id": 616072,
          "postDate": "2019-09-02T16:54:05.180Z",
          "content": "<p>I think the issue might be that when you go to higher size you have to use small batch size. \n1) Have you tried to do gradient accumulation ?\n2) Try Mish activation function. Basicity its very simple to implement here is the <a href=\"https://arxiv.org/pdf/1908.08681v1.pdf\">paper</a>. (<a href=\"https://arxiv.org/pdf/1908.08681v1.pdf\">https://arxiv.org/pdf/1908.08681v1.pdf</a>) . You can see in on figure 7 that decreasing batch size has little effect in accuracy in comparison to ReLu. In my case it helped a bit. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F3fa4d3a563388f552a6302bc050c53e9%2FScreen%20Shot%202019-09-02%20at%2012.52.41%20PM.png?generation=1567443170774352&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"/valanm\">@valanm</a> I reuse weight for lower resolution to train higher \nHope this helps </p>",
          "rawMarkdown": "I think the issue might be that when you go to higher size you have to use small batch size. \n1) Have you tried to do gradient accumulation ?\n2) Try Mish activation function. Basicity its very simple to implement here is the [paper](https://arxiv.org/pdf/1908.08681v1.pdf). (https://arxiv.org/pdf/1908.08681v1.pdf) . You can see in on figure 7 that decreasing batch size has little effect in accuracy in comparison to ReLu. In my case it helped a bit. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F3fa4d3a563388f552a6302bc050c53e9%2FScreen%20Shot%202019-09-02%20at%2012.52.41%20PM.png?generation=1567443170774352&amp;alt=media)\n\n\n@valanm I reuse weight for lower resolution to train higher \nHope this helps ",
          "votes": 8
        },
        {
          "id": 616094,
          "postDate": "2019-09-02T17:17:21.433Z",
          "content": "<p>Thank <a href=\"/valanm\">@valanm</a> and <a href=\"/drhabib\">@drhabib</a> , both suggestions are very interesting, I'll give a try.</p>",
          "rawMarkdown": "Thank @valanm and @drhabib , both suggestions are very interesting, I'll give a try.",
          "votes": 1
        },
        {
          "id": 617975,
          "postDate": "2019-09-04T17:18:55.830Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> i was just wondering , if increasing the image size should take more training as we have more to learn. Can you explain something about this?</p>",
          "rawMarkdown": "@drhabib i was just wondering , if increasing the image size should take more training as we have more to learn. Can you explain something about this?",
          "votes": 1
        },
        {
          "id": 618051,
          "postDate": "2019-09-04T19:04:55.393Z",
          "content": "<p>Hi.. I will refer you to this two blog posts. \n<a href=\"https://miguel-data-sc.github.io/2017-11-23-second/\">https://miguel-data-sc.github.io/2017-11-23-second/</a>\n<a href=\"https://towardsdatascience.com/boost-your-cnn-image-classifier-performance-with-progressive-resizing-in-keras-a7d96da06e20\">https://towardsdatascience.com/boost-your-cnn-image-classifier-performance-with-progressive-resizing-in-keras-a7d96da06e20</a></p>",
          "rawMarkdown": "Hi.. I will refer you to this two blog posts. \nhttps://miguel-data-sc.github.io/2017-11-23-second/\nhttps://towardsdatascience.com/boost-your-cnn-image-classifier-performance-with-progressive-resizing-in-keras-a7d96da06e20",
          "votes": 1
        },
        {
          "id": 619846,
          "postDate": "2019-09-06T16:50:48.257Z",
          "content": "<p>@drhb i was wondering were u able to run wit b5 456 sz image  on kaggles kernel,how much was bs..</p>",
          "rawMarkdown": "@drhb i was wondering were u able to run wit b5 456 sz image  on kaggles kernel,how much was bs..\n",
          "votes": 1
        },
        {
          "id": 619848,
          "postDate": "2019-09-06T16:52:08.863Z",
          "content": "<p>for interference only =) \nbs = 64 did a job, but i think going higher is possible but i havent tried =) </p>",
          "rawMarkdown": "for interference only =) \nbs = 64 did a job, but i think going higher is possible but i havent tried =) "
        },
        {
          "id": 619853,
          "postDate": "2019-09-06T16:54:57.507Z",
          "content": "<p>ok so training on smaller image size and but inference on a higher one could work well ? this is new for me i was thinking that we should always give same size image for inference as training..\nPlease confirm.. i will add this in my learning book :)</p>",
          "rawMarkdown": "ok so training on smaller image size and but inference on a higher one could work well ? this is new for me i was thinking that we should always give same size image for inference as training..\nPlease confirm.. i will add this in my learning book :)",
          "votes": 1
        },
        {
          "id": 619871,
          "postDate": "2019-09-06T17:19:27.677Z",
          "content": "<p>When you are doing interference you only do forward pass and you dont have to compute gradients as a result more memory is available for gpu and therefore you can use big batch size =) </p>",
          "rawMarkdown": "When you are doing interference you only do forward pass and you dont have to compute gradients as a result more memory is available for gpu and therefore you can use big batch size =) "
        },
        {
          "id": 620725,
          "postDate": "2019-09-07T22:51:17.453Z",
          "content": "<p>in my last submission i get some thing similar to what is posted without doing any mechanical adjustments much.. \nbut dont know will i get also similar score as tilll now for all submissions only the ones having 2 above 1200 have seen score above 0.80.. ,1 and 3 between  130 and 160, the posted ones are very different to these..\nWas that bar graph post  a competition gimmick hahah  :) ...</p>",
          "rawMarkdown": "in my last submission i get some thing similar to what is posted without doing any mechanical adjustments much.. \nbut dont know will i get also similar score as tilll now for all submissions only the ones having 2 above 1200 have seen score above 0.80.. ,1 and 3 between  130 and 160, the posted ones are very different to these..\nWas that bar graph post  a competition gimmick hahah  :) ..."
        },
        {
          "id": 621395,
          "postDate": "2019-09-08T13:33:19.477Z",
          "content": "<p>hi hb good to see you gained the rank in private lb..\ncongrats\ncould you tell me how you replaced the default act functions for the model.. </p>",
          "rawMarkdown": "hi hb good to see you gained the rank in private lb..\ncongrats\ncould you tell me how you replaced the default act functions for the model.. \n"
        }
      ]
    },
    {
      "id": 610215,
      "postDate": "2019-08-28T14:49:24.983Z",
      "content": "<p>456, efficientnetb5  lb:0,824 single model</p>",
      "rawMarkdown": "456, efficientnetb5  lb:0,824 single model",
      "votes": 6,
      "replies": [
        {
          "id": 610243,
          "postDate": "2019-08-28T15:20:50.643Z",
          "content": "<p>a nice lb score.</p>",
          "rawMarkdown": "a nice lb score."
        },
        {
          "id": 610699,
          "postDate": "2019-08-29T00:19:29.260Z",
          "content": "<p>nice work! my best lb score of 256 single model is 0.821.</p>",
          "rawMarkdown": "nice work! my best lb score of 256 single model is 0.821."
        },
        {
          "id": 610786,
          "postDate": "2019-08-29T01:15:25.363Z",
          "content": "<p>whats kind of network?</p>",
          "rawMarkdown": "whats kind of network?"
        },
        {
          "id": 610896,
          "postDate": "2019-08-29T02:44:11.283Z",
          "content": "<p>efficientnetb5</p>",
          "rawMarkdown": "efficientnetb5"
        },
        {
          "id": 611029,
          "postDate": "2019-08-29T04:21:47.390Z",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> <a href=\"/leixiang\">@leixiang</a> Nice scores !!\nWhat kind of preprocessing you guys used?\nDid you first pretrain on old train and finetune on new </p>",
          "rawMarkdown": "@hirune924 @leixiang Nice scores !!\nWhat kind of preprocessing you guys used?\nDid you first pretrain on old train and finetune on new "
        },
        {
          "id": 611757,
          "postDate": "2019-08-29T12:19:18.590Z",
          "content": "<p><a href=\"/leixiang\">@leixiang</a> <a href=\"/hirune924\">@hirune924</a> nice to see single model reach the score. May I ask for If you have any change of the output header of the efficientnet like adding BN or concat [avgpool, maxpool]? Thank you all.</p>",
          "rawMarkdown": "@leixiang @hirune924 nice to see single model reach the score. May I ask for If you have any change of the output header of the efficientnet like adding BN or concat [avgpool, maxpool]? Thank you all."
        },
        {
          "id": 612317,
          "postDate": "2019-08-29T18:06:14.153Z",
          "content": "<p>Hi, </p>\n\n<p>You are asking others, but I would like to offer my experience: Using MaxPool with EfficientNet seemed to hurt the model's performance. I tried several times with different experiments and no good. I thought ConcatPool should improve performance, but it seems that the information from the MaxPool only brings the performance of the AvgPool down. Of course, this is only data from my limited experiments. </p>",
          "rawMarkdown": "Hi, \n\nYou are asking others, but I would like to offer my experience: Using MaxPool with EfficientNet seemed to hurt the model's performance. I tried several times with different experiments and no good. I thought ConcatPool should improve performance, but it seems that the information from the MaxPool only brings the performance of the AvgPool down. Of course, this is only data from my limited experiments. "
        },
        {
          "id": 612656,
          "postDate": "2019-08-29T23:28:22.537Z",
          "content": "<p>This idea is borrowed from the bi-lstm model. It seems no one who has high rank in the discussion has told this feature recently... you are right and I make a mistake... I just search my log and find that using concat just decrease my CV from 0.8091 to 0.7993 because I have too many trials and i get confused... Maybe I need to figure out the difference between them.</p>",
          "rawMarkdown": "This idea is borrowed from the bi-lstm model. It seems no one who has high rank in the discussion has told this feature recently... you are right and I make a mistake... I just search my log and find that using concat just decrease my CV from 0.8091 to 0.7993 because I have too many trials and i get confused... Maybe I need to figure out the difference between them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 612000,
      "postDate": "2019-08-29T14:52:56.757Z",
      "content": "<p>328 B5 Classification, reach 0.825. <br>\n328 B5 Regression, reach 0.823.\n256 B4 Regression, reach 0.821</p>\n\n<p>[Uptate 2019/08/30 13:18]\n256 B4 Classification, reach 0.820</p>",
      "rawMarkdown": "328 B5 Classification, reach 0.825.   \n328 B5 Regression, reach 0.823.\n256 B4 Regression, reach 0.821\n\n[Uptate 2019/08/30 13:18]\n256 B4 Classification, reach 0.820",
      "votes": 1,
      "replies": [
        {
          "id": 612012,
          "postDate": "2019-08-29T15:04:45.767Z",
          "content": "<p>and ensemble gives you your current score 0.833 !?</p>",
          "rawMarkdown": "and ensemble gives you your current score 0.833 !?"
        },
        {
          "id": 612058,
          "postDate": "2019-08-29T15:27:09.777Z",
          "content": "<p>Yes.... Only ensemble B5 regression and B4 regression</p>",
          "rawMarkdown": "Yes.... Only ensemble B5 regression and B4 regression",
          "votes": 1
        },
        {
          "id": 612124,
          "postDate": "2019-08-29T15:58:43.767Z",
          "content": "<p>0.823 &amp; 0.812 -&gt; 0.833 ? This is a great synergy effect!!\nDo you think that the difference in size has an effect?</p>",
          "rawMarkdown": "0.823 &amp; 0.812 -&gt; 0.833 ? This is a great synergy effect!!\nDo you think that the difference in size has an effect?"
        },
        {
          "id": 612148,
          "postDate": "2019-08-29T16:11:55.400Z",
          "content": "<p>i got similar result with him just ensemble a lb 0.824 eftnet b5 456 and lb 0.814 eftnet b4      380，maybe different image size boost the score</p>",
          "rawMarkdown": "i got similar result with him just ensemble a lb 0.824 eftnet b5 456 and lb 0.814 eftnet b4      380，maybe different image size boost the score",
          "votes": 1
        },
        {
          "id": 612182,
          "postDate": "2019-08-29T16:34:44.243Z",
          "content": "<p>I tried voting between 4 classification models \nb3 300 0.808\nb3 416  0.812\nb5 384 0.812\nb5 456 0.817\nbut got 0.785, couldn't find any bug so I concluded its better to use averaging, but i haven't tried yet.  </p>",
          "rawMarkdown": "I tried voting between 4 classification models \nb3 300 0.808\nb3 416  0.812\nb5 384 0.812\nb5 456 0.817\nbut got 0.785, couldn't find any bug so I concluded its better to use averaging, but i haven't tried yet.  ",
          "votes": 1
        },
        {
          "id": 612660,
          "postDate": "2019-08-29T23:33:59.357Z",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> I am not sure maybe not.\nIn my experience, combining different structure model output together may boost score.\n I am still trying to improve my single model score right now. (struggling)</p>",
          "rawMarkdown": "@octpath0302 I am not sure maybe not.\nIn my experience, combining different structure model output together may boost score.\n I am still trying to improve my single model score right now. (struggling)",
          "votes": 2
        },
        {
          "id": 612667,
          "postDate": "2019-08-29T23:44:59.013Z",
          "content": "<p><a href=\"/valanm\">@valanm</a>, I think you use majority voting, it did not work for my.</p>\n\n<p>I use this way to ensemble two classification models. \nit slightly improve my score(b4:0.815, b5:0.825 -&gt; 0.829),</p>\n\n<ol>\n<li><p>for one model prediction\nfor _ in range(TTA):\nwith torch.no_grad():\n    for i, data in tqdm(enumerate(test_loader)):\n        images = data\n        images = images.cuda()\n        pred   = model(images)\n        pred   = torch.softmax(pred, dim=-1)\n        test_pred[i * 64:(i + 1) * 64] += pred</p></li>\n<li><p>add two model prediction together\noutput = np.argmax(((test_pred1 / TTA) +  (test_pred2 / TTA)) / 2, axis=1)</p></li>\n</ol>",
          "rawMarkdown": "@valanm, I think you use majority voting, it did not work for my.\n\nI use this way to ensemble two classification models. \nit slightly improve my score(b4:0.815, b5:0.825 -&gt; 0.829),\n\n1. for one model prediction\nfor _ in range(TTA):\n    with torch.no_grad():\n        for i, data in tqdm(enumerate(test_loader)):\n            images = data\n            images = images.cuda()\n            pred   = model(images)\n            pred   = torch.softmax(pred, dim=-1)\n            test_pred[i * 64:(i + 1) * 64] += pred\n\n\n2. add two model prediction together\noutput = np.argmax(((test_pred1 / TTA) +  (test_pred2 / TTA)) / 2, axis=1)",
          "votes": 4
        },
        {
          "id": 613248,
          "postDate": "2019-08-30T10:57:58.107Z",
          "content": "<p><a href=\"/chanhu\">@chanhu</a> This also happened to our case. Ensembling (averaging) B5 (0. 819) and B0 (0.812) made our current leaderboard score (0.837). Although, the training procedure for each model was completely different, which might have caused such synergy!</p>",
          "rawMarkdown": "@chanhu This also happened to our case. Ensembling (averaging) B5 (0. 819) and B0 (0.812) made our current leaderboard score (0.837). Although, the training procedure for each model was completely different, which might have caused such synergy!",
          "votes": 3
        },
        {
          "id": 613251,
          "postDate": "2019-08-30T11:01:34.387Z",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a> that's just awesome synergy</p>",
          "rawMarkdown": "@udaykamal that's just awesome synergy",
          "votes": 1
        },
        {
          "id": 614383,
          "postDate": "2019-08-31T12:21:09.477Z",
          "content": "<p><a href=\"/chanhu\">@chanhu</a> <a href=\"/udaykamal\">@udaykamal</a>  how many tta transform do you use? flip,transpose  and multi-scale?</p>",
          "rawMarkdown": "@chanhu @udaykamal  how many tta transform do you use? flip,transpose  and multi-scale?"
        },
        {
          "id": 614397,
          "postDate": "2019-08-31T12:41:10.700Z",
          "content": "<p>tta 5 or 10. only flip, and rotation.</p>",
          "rawMarkdown": "tta 5 or 10. only flip, and rotation."
        },
        {
          "id": 614420,
          "postDate": "2019-08-31T13:10:01.410Z",
          "content": "<p><a href=\"/chanhu\">@chanhu</a>  have you checked how much boost this gives you. i tried horizontal flip but the score got worse 0.001</p>",
          "rawMarkdown": "@chanhu  have you checked how much boost this gives you. i tried horizontal flip but the score got worse 0.001"
        },
        {
          "id": 614483,
          "postDate": "2019-08-31T14:49:45.930Z",
          "content": "<p><a href=\"/hustkevin1037\">@hustkevin1037</a> we have not tried with TTA yet. Out current leaderboard score is simple average of two predictions from two different model. No TTA, no K-fold.</p>",
          "rawMarkdown": "@hustkevin1037 we have not tried with TTA yet. Out current leaderboard score is simple average of two predictions from two different model. No TTA, no K-fold."
        },
        {
          "id": 614515,
          "postDate": "2019-08-31T15:54:09.760Z",
          "content": "<p>Hey uday did you try regression only for both ?</p>",
          "rawMarkdown": "Hey uday did you try regression only for both ?"
        },
        {
          "id": 614745,
          "postDate": "2019-09-01T01:55:32.897Z",
          "content": "<p><a href=\"/valanm\">@valanm</a> do you use torch or keras? </p>",
          "rawMarkdown": "@valanm do you use torch or keras? "
        },
        {
          "id": 614748,
          "postDate": "2019-09-01T02:07:41.613Z",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a>  did you use optimized threshold or just [0.5,1,5,2,5,3,5] ?</p>",
          "rawMarkdown": "@udaykamal  did you use optimized threshold or just [0.5,1,5,2,5,3,5] ?"
        },
        {
          "id": 614903,
          "postDate": "2019-09-01T08:03:41.890Z",
          "content": "<p><a href=\"/valanm\">@valanm</a>  about +0.004.</p>",
          "rawMarkdown": "@valanm  about +0.004."
        },
        {
          "id": 615011,
          "postDate": "2019-09-01T11:03:34.163Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> yes\n<a href=\"/hustkevin1037\">@hustkevin1037</a> optimized threshold resulted in lower score in our case. We are using just [0.5,1,5,2,5,3,5]</p>",
          "rawMarkdown": "@jaideepvalani yes\n@hustkevin1037 optimized threshold resulted in lower score in our case. We are using just [0.5,1,5,2,5,3,5]"
        },
        {
          "id": 615207,
          "postDate": "2019-09-01T15:56:06.917Z",
          "content": "<p><a href=\"/custodiogabriel\">@custodiogabriel</a> i am trying fastai for this competition\nthis might be a dumb question but are you guys ensembling logits or softmax outputs, does this matter at all?</p>",
          "rawMarkdown": "@custodiogabriel i am trying fastai for this competition\nthis might be a dumb question but are you guys ensembling logits or softmax outputs, does this matter at all?"
        },
        {
          "id": 615508,
          "postDate": "2019-09-02T03:15:49.100Z",
          "content": "<p><a href=\"/valanm\">@valanm</a> If you are doing classification, softmax outputs makes more sense. </p>",
          "rawMarkdown": "@valanm If you are doing classification, softmax outputs makes more sense. ",
          "votes": 1
        },
        {
          "id": 615547,
          "postDate": "2019-09-02T04:54:49.937Z",
          "content": "<p>@ uday thanks...\ncould you please tell    what is distribution of each class  in submission score..</p>",
          "rawMarkdown": "@ uday thanks...\ncould you please tell    what is distribution of each class  in submission score.."
        },
        {
          "id": 615574,
          "postDate": "2019-09-02T05:44:04.953Z",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> thanks. Is the reason scaling? Would ligits make sense for stacking?</p>",
          "rawMarkdown": "@naivelamb thanks. Is the reason scaling? Would ligits make sense for stacking?"
        },
        {
          "id": 615596,
          "postDate": "2019-09-02T06:23:07.523Z",
          "content": "<p>It depends on what stacking methods you are using. Some are obviously sensitive to scaling. </p>",
          "rawMarkdown": "It depends on what stacking methods you are using. Some are obviously sensitive to scaling. "
        },
        {
          "id": 615733,
          "postDate": "2019-09-02T09:58:29.877Z",
          "content": "<p>may I know what batch size u are using? I am training with 16 only due to limited gpu ram. Considering to rent a machine if batch size matters</p>",
          "rawMarkdown": "may I know what batch size u are using? I am training with 16 only due to limited gpu ram. Considering to rent a machine if batch size matters"
        },
        {
          "id": 616046,
          "postDate": "2019-09-02T16:25:06.247Z",
          "content": "<p>Averaging softmax outputs instead of logits gave the same score. However, 12 images (0.6%) got different label.</p>",
          "rawMarkdown": "Averaging softmax outputs instead of logits gave the same score. However, 12 images (0.6%) got different label."
        },
        {
          "id": 620049,
          "postDate": "2019-09-07T00:08:12.497Z",
          "content": "<p><a href=\"/valanm\">@valanm</a> I hope you can get good score in private test set</p>",
          "rawMarkdown": "@valanm I hope you can get good score in private test set"
        }
      ]
    },
    {
      "id": 616493,
      "postDate": "2019-09-03T07:21:07.007Z",
      "content": "<p>EfficientNet b5, no CV or LB improvement after 256</p>",
      "rawMarkdown": "EfficientNet b5, no CV or LB improvement after 256"
    },
    {
      "id": 612028,
      "postDate": "2019-08-29T15:11:56.193Z",
      "content": "<p>Is there a reason to use img size 256x256 with efficient nets? I ask this because it's not the default size for any of its models, but people seem to get good results anyway.</p>",
      "rawMarkdown": "Is there a reason to use img size 256x256 with efficient nets? I ask this because it's not the default size for any of its models, but people seem to get good results anyway.",
      "replies": [
        {
          "id": 612076,
          "postDate": "2019-08-29T15:37:22.917Z",
          "content": "<p>In my understanding, the default sizes have been searched on a specific dataset. And the dataset are different from this competition dataset. So, I think it is not always necessary to resize this dataset according to the default size.\nMany people have reported that the accuracy is good even if it is different from the default size, so I think it supports this idea.</p>",
          "rawMarkdown": "In my understanding, the default sizes have been searched on a specific dataset. And the dataset are different from this competition dataset. So, I think it is not always necessary to resize this dataset according to the default size.\nMany people have reported that the accuracy is good even if it is different from the default size, so I think it supports this idea.",
          "votes": 3
        },
        {
          "id": 612098,
          "postDate": "2019-08-29T15:50:34.097Z",
          "content": "<p>I understand <a href=\"/octpath0302\">@octpath0302</a> , this makes sense.</p>",
          "rawMarkdown": "I understand @octpath0302 , this makes sense."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 612806,
      "author_name": "Maxwell",
      "author_url": "",
      "post_date": "2019-08-30T02:39:47.713000",
      "content": "<p>My single models are, <br>\n1. 300 Regression EfficientNet B3, LB 0.830\n2. 260 Regression EfficientNet B2, LB 0.820</p>\n\n<p>Our current place is from Model 1.  </p>",
      "votes": 9,
      "replies": [
        {
          "id": 614241,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-31T08:46:05.973000",
          "content": "<p>hi max well did you try any preprocessing like bens cropping or any other augs..and how much is you train loss for each at the time of making prediction</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614447,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-31T13:55:30.117000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 614458,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2019-08-31T14:14:21.790000",
          "content": "<p><a href=\"/pranavpulijala\">@pranavpulijala</a> \nMy CV on <strong>APTOS 2019 train dataset</strong> is about 0.93.\nIf I add higher grade examples of DRD2015 for evaluation, that will go down.\nI think it would be meaningless to compare CV without using same datasets.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614514,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-31T15:50:24.747000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 611126,
      "author_name": "oct_path",
      "author_url": "",
      "post_date": "2019-08-29T05:50:34.947000",
      "content": "<p>I use 256 preprocessed by mean subtraction.\nMy best single model (EfficientNet-B4) got 0.829.\nI've tried 320 and 512, but it was a waste of time.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 611224,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2019-08-29T07:10:17.477000",
          "content": "<p>Please give some tips.  What data you used, any other preprocessing than mean subtraction, data augmentations and whether you treated the problem as regression or classification ? Thanks in advance</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611268,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-29T07:37:13.153000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> You are still use old and current data together？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611327,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T08:02:59.667000",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> \nI use only rotation and horizontal flip in training phase. no tta. regression.\nSince there are many class 2 data in public test set, training data are sampled with adjusted distribution. My LB score so far improved best when the distribution was adjusted.\n<a href=\"/garybios\">@garybios</a> Yes. I use all the datasets listed in 'external data thread' at the same time.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 611434,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-29T08:58:39.057000",
          "content": "<p>nice, thanks a lot. Could you tell me how do you adjust the data distribution?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611504,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T09:48:49.343000",
          "content": "<p>A trained model was used to predict the class of public test sets and estimate the distribution.\nThere are various ways of sampling. Whether to do it only first, or every epoch, or to make a K-Fold.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 611534,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-29T10:06:30.907000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  that this mean you do not crop your images? (you said you mean subtract only)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611548,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2019-08-29T10:13:33.163000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  I thought of sampling based on the class distribution of public test dataset. But I thought that way would overfit public test dataset. I think the distribution of private test dataset is more similar to train dataset than private test dataset. What do you think about this？</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 611563,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-29T10:20:47.900000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> you can probe it and see. someone from the top has done it for sure.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611566,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T10:21:53.580000",
          "content": "<p><a href=\"/valanm\">@valanm</a> Sorry for misleading! In the preprocessing stage, I use auto cropping which is similar to the 1st solution on the previous competition, and padding.\nI make preprocessing to get robustness against black area at corner of each image.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 611568,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-29T10:22:22.813000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> I also thought that way is easy to overfit public lb, but we do not know the distribution of private test dataset.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611577,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T10:27:00.067000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> I also think there is a danger of overfit. However, it is generally possible that there is a lot of data in class 0 (normal) and class 2 (middle). Therefore, I assume that the public test set and the private test set are randomly divided, so far.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 611765,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T12:24:41.440000",
          "content": "<p><a href=\"/garybios\">@garybios</a> <a href=\"/valanm\">@valanm</a> <a href=\"/virajbagal\">@virajbagal</a>  How about you? Which size/preprocessing do you use?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611792,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-29T12:46:45.253000",
          "content": "<p>300px b3 classification gives me .822, for preprocessing just cropping the object.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 611811,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-29T12:59:18.987000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> 224, no preprocessing.\n<a href=\"/valanm\">@valanm</a> nice job, you made a great classification model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 611880,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2019-08-29T13:37:50.370000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a>  256, cropping image and training only on new data gave me 0.796. Despite adding augmentations, my model trained on OLD Data with New Data as validation is overfitting and always predicting 0. I tried to balance the classes by downsampling and upsampling with each class around 5000 images but the score is really poor.  I am desperately trying again and again to train the model on old data. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 611929,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T14:05:24.153000",
          "content": "<p><a href=\"/virajbagal\">@virajbagal</a> Hmm. . that's strange. But in my case too, augmentation is hardly working. And I use only downsampling (I have not tried upsampling).\nAre CV and LB (or losses) somewhat correlated? Do your CV get &gt;0.85?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 611945,
          "author_name": "Viraj Bagal",
          "author_url": "",
          "post_date": "2019-08-29T14:11:41.657000",
          "content": "<p>My CV goes upto 92. But predicts all zeros. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611950,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T14:18:49.520000",
          "content": "<p>My CV has never reached to 0.9. So, I think that overfit or something is happening in your case..\n(too small validation set or .. I don't know.)\nIn my case, CV is around 0.87 to 0.88 and LB is around 0.81 to 0.83.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 612069,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-08-29T15:30:18.047000",
          "content": "<p>Me too.. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612070,
          "author_name": "Ravi Kumar Vadapalli",
          "author_url": "",
          "post_date": "2019-08-29T15:31:28.803000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> , thanks for sharing the details. A quick question on the external data usage. External datasets have different labels. Some of them have DR levels (0 to 4 or 0 to 3), some have segmented masks, some have counts of microaneurisms, exudates, etc. How are you managing with such desparate labels. Are you somehow creating uniform labels for all the external datasets ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612090,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T15:46:40.437000",
          "content": "<p><a href=\"/ravivadapalli\">@ravivadapalli</a> I use datasets which DR levels are 0 to 4 as labeled dataset, and the others are as unlabeled (pseudo-labeled or something) dataset. \nI don't try to use any other labels and masks because there is little time to Kaggle..</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 612215,
          "author_name": "Ravi Kumar Vadapalli",
          "author_url": "",
          "post_date": "2019-08-29T16:58:51.740000",
          "content": "<p>I see. thank you <a href=\"/octpath0302\">@octpath0302</a> for sharing the details. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612309,
          "author_name": "Andrew Minh Nguyen",
          "author_url": "",
          "post_date": "2019-08-29T18:02:23.947000",
          "content": "<p>Hi, <a href=\"/octpath0302\">@octpath0302</a>. Thank you for sharing so much of your approach. Could you explain more how you apply mean subtraction to your images? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612602,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T22:19:20.673000",
          "content": "<p>I refer to the 1st place solution of the previous competition.\n<a href=\"https://www.kaggle.com/c/diabetic-retinopathy-detection/discussion/15801#latest-370950\">https://www.kaggle.com/c/diabetic-retinopathy-detection/discussion/15801#latest-370950</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 612609,
          "author_name": "Andrew Minh Nguyen",
          "author_url": "",
          "post_date": "2019-08-29T22:40:32.370000",
          "content": "<p>Got it, thanks! </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 613421,
          "author_name": "RAMPAGE",
          "author_url": "",
          "post_date": "2019-08-30T13:47:13.103000",
          "content": "<p>May I ask what is your train setting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 613587,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-08-30T16:31:48.790000",
          "content": "<p>Hi <a href=\"/octpath0302\">@octpath0302</a> , did you observe big drop in CV after applying downsampling but LB may be the same or even better? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 613917,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-31T00:15:57.010000",
          "content": "<p><a href=\"/andalpha\">@andalpha</a> I have provided a lot of information, so I can't do any more..\n<a href=\"/naivelamb\">@naivelamb</a>  I'm not sure because I've never trained by applying downsampling only. After applying downsampling and data addition, I observed that CV was almost unchanged and LB was getting better.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 613932,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-08-31T00:37:24.717000",
          "content": "<p>By data addition you mean mixing external data (2015 for example) with current data together to train the model, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 613943,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-31T00:43:31.403000",
          "content": "<p>yes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614227,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-31T08:28:11.460000",
          "content": "<p>hi oct when you training distribution sampling wrt to public test set.. what you mean\nis it taking more class 2 data or how is it ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614234,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-31T08:40:13.860000",
          "content": "<p>@valann by b3 classification you mean not using regression?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614247,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-31T08:52:22.140000",
          "content": "<p>yes I only use classification</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614256,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-31T09:23:07.180000",
          "content": "<p><a href=\"/octpath\">@octpath</a> which external data set you using.. \nThanks in advance</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614660,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-31T21:13:24.103000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614702,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-08-31T23:52:08.133000",
          "content": "<p>do you use optimized thresholds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614756,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-01T02:53:29.483000",
          "content": "<p>@valan  how did you work with wierd crop patches produced by fastai ...I found some blacks borders with zero padding mode,or circle arcs at border with reflection ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615136,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-09-01T14:01:40.210000",
          "content": "<p><a href=\"/ggoldszmidt\">@ggoldszmidt</a> I mean local mean subtraction (high-pass filtering), so not standardized. Almost the same as 1st place solution of the previous competition.\n<a href=\"/tanlikesmath\">@tanlikesmath</a> Yes. Optimized.\n<a href=\"/jaideepvalani\">@jaideepvalani</a> \n&gt; hi oct when you training distribution sampling wrt to public test set.. what you mean\n&gt; is it taking more class 2 data or how is it …</p>\n\n<p>Sorry, I can't catch your meaning. You can calculate the public test class distribution with a model (e. g. from public kernel). \nExternal data I'm using are listed up on the official thread.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615160,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-01T14:50:28.137000",
          "content": "<p>hi oct thanks for reply..\n1) in training set we have 3000 eg . how we decide on the ratio of distribution while sampling say we have batch size of 32..\n2) I tried to follow the external thread none of them have labels ,if i missed any link could u past the one that you used in ..\n3) which public kernel are talking about</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615176,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-09-01T15:05:11.730000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \n1) I mean not sampling mini-batch, but how to re-construct labeled dataset.\n2) 3) Most of the external datasets are unlabeled datasets. But you can pseudo-label any images by using high-performance model from public kernels. If you already have good models, you can use them. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615200,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-01T15:46:54.470000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> i read the documentation and adopted the code. I default would be good enough then data science wouldn't be hot on job market. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 615912,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-09-02T13:47:37.107000",
      "content": "<p>Here are some of my results.  I did on image size.  All 5x CV </p>\n\n<p><code>EfficientNet B5,  224 LB: 0.807</code>\n<code>EfficientNet B5,  300 LB: 0.812</code>\n<code>EfficientNet B5,  380 LB: 0.816</code>\n<code>EfficientNet B5,  456 LB: 0.823</code></p>\n\n<p><code>EfficientNet B4,  224 LB: 0.800</code>\n<code>EfficientNet B4,  300 LB: 0.811</code>\n<code>EfficientNet B4,  380 LB: 0.816</code></p>",
      "votes": 10,
      "replies": [
        {
          "id": 615960,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-02T14:46:31.283000",
          "content": "<p>@dir HB  good results,your kernel public appreaciated by many kagglers and including me :)\nwould you like to tell the class distribution..  </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 615973,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-02T14:55:46.567000",
          "content": "<p>Glad to hear that kernel helped =) My final distributions are very similar to some that found in this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/106844#614198\">thread</a> . =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615975,
          "author_name": "Benson Jin",
          "author_url": "",
          "post_date": "2019-09-02T14:59:08.097000",
          "content": "<p>I tried doing 5x CV with EfficientNet B5 but I only get .79-.792. I dont know what else to do.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 615985,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-02T15:07:34.797000",
          "content": "<p>They are diffrent ways how you can improve score.  Have a look some threads like [this] (<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105100#latest-615864\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/105100#latest-615864</a>) . All of them have good tips. </p>\n\n<p>But in genral. \n1) Try diffrent preprocessing \n2) Try diffrent augmentations \n3) Try combining of diffrent dataset. </p>\n\n<p>Forum is full of tips =) \nGood luck to you =) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 615987,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-09-02T15:08:00.310000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> when you increase img size, do you also increase your parameters, like learning rate or epochs? I'm asking this because for me it's harder to keep models with large img size learning.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 615996,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-02T15:17:51.187000",
          "content": "<p>Good question. </p>\n\n<p>I ususally run lr_finder and select learning rate based on this, but in most casses its very similar across diffrent sizes. Also regarding number of Epochs if you are intrested,  for 224 I run for 20 epochs and for rest image sizes 5 epochs is enogh. Again it depends also on which learning rate policy you are using in my case its one_cycle_policy.</p>\n\n<p>Hope it helps =)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 616011,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-09-02T15:44:01.563000",
          "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a> , I have tried many policies, also I use Keras, but the bigger models don't train longer than 10 epochs reliably.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 616016,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-02T15:47:02.880000",
          "content": "<p>which optimizer are you using ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 616022,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-09-02T15:58:23.130000",
          "content": "<p>For me what seems to work is either Adam or RAdam, but also with LR policies, like Cosine or cycle.</p>\n\n<p>So far my best single model result comes from B5 224, using Adam and cosine LR (almost 0.8)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 616051,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-02T16:30:58.663000",
          "content": "<p>Do you use weights from lower resolution or u always start from imagenet?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 616062,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-09-02T16:41:17.920000",
          "content": "<p><a href=\"/valanm\">@valanm</a> , I'm starting from imagenet. By lower resolution, do you mean pretraining from the older dataset?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 616068,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-02T16:49:07.713000",
          "content": "<p>For instance, train 20 epochs on 224 and then 5 more with increased resolution to 456</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 616072,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-02T16:54:05.180000",
          "content": "<p>I think the issue might be that when you go to higher size you have to use small batch size. \n1) Have you tried to do gradient accumulation ?\n2) Try Mish activation function. Basicity its very simple to implement here is the <a href=\"https://arxiv.org/pdf/1908.08681v1.pdf\">paper</a>. (<a href=\"https://arxiv.org/pdf/1908.08681v1.pdf\">https://arxiv.org/pdf/1908.08681v1.pdf</a>) . You can see in on figure 7 that decreasing batch size has little effect in accuracy in comparison to ReLu. In my case it helped a bit. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F3fa4d3a563388f552a6302bc050c53e9%2FScreen%20Shot%202019-09-02%20at%2012.52.41%20PM.png?generation=1567443170774352&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"/valanm\">@valanm</a> I reuse weight for lower resolution to train higher \nHope this helps </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 616094,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-09-02T17:17:21.433000",
          "content": "<p>Thank <a href=\"/valanm\">@valanm</a> and <a href=\"/drhabib\">@drhabib</a> , both suggestions are very interesting, I'll give a try.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 617975,
          "author_name": "Nitish Kumar",
          "author_url": "",
          "post_date": "2019-09-04T17:18:55.830000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> i was just wondering , if increasing the image size should take more training as we have more to learn. Can you explain something about this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 618051,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-04T19:04:55.393000",
          "content": "<p>Hi.. I will refer you to this two blog posts. \n<a href=\"https://miguel-data-sc.github.io/2017-11-23-second/\">https://miguel-data-sc.github.io/2017-11-23-second/</a>\n<a href=\"https://towardsdatascience.com/boost-your-cnn-image-classifier-performance-with-progressive-resizing-in-keras-a7d96da06e20\">https://towardsdatascience.com/boost-your-cnn-image-classifier-performance-with-progressive-resizing-in-keras-a7d96da06e20</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 619846,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-06T16:50:48.257000",
          "content": "<p>@drhb i was wondering were u able to run wit b5 456 sz image  on kaggles kernel,how much was bs..</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 619848,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-06T16:52:08.863000",
          "content": "<p>for interference only =) \nbs = 64 did a job, but i think going higher is possible but i havent tried =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 619853,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-06T16:54:57.507000",
          "content": "<p>ok so training on smaller image size and but inference on a higher one could work well ? this is new for me i was thinking that we should always give same size image for inference as training..\nPlease confirm.. i will add this in my learning book :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 619871,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-09-06T17:19:27.677000",
          "content": "<p>When you are doing interference you only do forward pass and you dont have to compute gradients as a result more memory is available for gpu and therefore you can use big batch size =) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 620725,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-07T22:51:17.453000",
          "content": "<p>in my last submission i get some thing similar to what is posted without doing any mechanical adjustments much.. \nbut dont know will i get also similar score as tilll now for all submissions only the ones having 2 above 1200 have seen score above 0.80.. ,1 and 3 between  130 and 160, the posted ones are very different to these..\nWas that bar graph post  a competition gimmick hahah  :) ...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621395,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-08T13:33:19.477000",
          "content": "<p>hi hb good to see you gained the rank in private lb..\ncongrats\ncould you tell me how you replaced the default act functions for the model.. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 610215,
      "author_name": "leixiang@AInnovation",
      "author_url": "",
      "post_date": "2019-08-28T14:49:24.983000",
      "content": "<p>456, efficientnetb5  lb:0,824 single model</p>",
      "votes": 6,
      "replies": [
        {
          "id": 610243,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-28T15:20:50.643000",
          "content": "<p>a nice lb score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 610699,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2019-08-29T00:19:29.260000",
          "content": "<p>nice work! my best lb score of 256 single model is 0.821.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 610786,
          "author_name": "leixiang@AInnovation",
          "author_url": "",
          "post_date": "2019-08-29T01:15:25.363000",
          "content": "<p>whats kind of network?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 610896,
          "author_name": "hirune924",
          "author_url": "",
          "post_date": "2019-08-29T02:44:11.283000",
          "content": "<p>efficientnetb5</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611029,
          "author_name": "Ankit",
          "author_url": "",
          "post_date": "2019-08-29T04:21:47.390000",
          "content": "<p><a href=\"/hirune924\">@hirune924</a> <a href=\"/leixiang\">@leixiang</a> Nice scores !!\nWhat kind of preprocessing you guys used?\nDid you first pretrain on old train and finetune on new </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 611757,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2019-08-29T12:19:18.590000",
          "content": "<p><a href=\"/leixiang\">@leixiang</a> <a href=\"/hirune924\">@hirune924</a> nice to see single model reach the score. May I ask for If you have any change of the output header of the efficientnet like adding BN or concat [avgpool, maxpool]? Thank you all.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612317,
          "author_name": "Andrew Minh Nguyen",
          "author_url": "",
          "post_date": "2019-08-29T18:06:14.153000",
          "content": "<p>Hi, </p>\n\n<p>You are asking others, but I would like to offer my experience: Using MaxPool with EfficientNet seemed to hurt the model's performance. I tried several times with different experiments and no good. I thought ConcatPool should improve performance, but it seems that the information from the MaxPool only brings the performance of the AvgPool down. Of course, this is only data from my limited experiments. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612656,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2019-08-29T23:28:22.537000",
          "content": "<p>This idea is borrowed from the bi-lstm model. It seems no one who has high rank in the discussion has told this feature recently... you are right and I make a mistake... I just search my log and find that using concat just decrease my CV from 0.8091 to 0.7993 because I have too many trials and i get confused... Maybe I need to figure out the difference between them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 612000,
      "author_name": "Chanhu",
      "author_url": "",
      "post_date": "2019-08-29T14:52:56.757000",
      "content": "<p>328 B5 Classification, reach 0.825. <br>\n328 B5 Regression, reach 0.823.\n256 B4 Regression, reach 0.821</p>\n\n<p>[Uptate 2019/08/30 13:18]\n256 B4 Classification, reach 0.820</p>",
      "votes": 1,
      "replies": [
        {
          "id": 612012,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-29T15:04:45.767000",
          "content": "<p>and ensemble gives you your current score 0.833 !?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612058,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-08-29T15:27:09.777000",
          "content": "<p>Yes.... Only ensemble B5 regression and B4 regression</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 612124,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T15:58:43.767000",
          "content": "<p>0.823 &amp; 0.812 -&gt; 0.833 ? This is a great synergy effect!!\nDo you think that the difference in size has an effect?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 612148,
          "author_name": "leixiang@AInnovation",
          "author_url": "",
          "post_date": "2019-08-29T16:11:55.400000",
          "content": "<p>i got similar result with him just ensemble a lb 0.824 eftnet b5 456 and lb 0.814 eftnet b4      380，maybe different image size boost the score</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 612182,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-29T16:34:44.243000",
          "content": "<p>I tried voting between 4 classification models \nb3 300 0.808\nb3 416  0.812\nb5 384 0.812\nb5 456 0.817\nbut got 0.785, couldn't find any bug so I concluded its better to use averaging, but i haven't tried yet.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 612660,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-08-29T23:33:59.357000",
          "content": "<p><a href=\"/octpath0302\">@octpath0302</a> I am not sure maybe not.\nIn my experience, combining different structure model output together may boost score.\n I am still trying to improve my single model score right now. (struggling)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 612667,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-08-29T23:44:59.013000",
          "content": "<p><a href=\"/valanm\">@valanm</a>, I think you use majority voting, it did not work for my.</p>\n\n<p>I use this way to ensemble two classification models. \nit slightly improve my score(b4:0.815, b5:0.825 -&gt; 0.829),</p>\n\n<ol>\n<li><p>for one model prediction\nfor _ in range(TTA):\nwith torch.no_grad():\n    for i, data in tqdm(enumerate(test_loader)):\n        images = data\n        images = images.cuda()\n        pred   = model(images)\n        pred   = torch.softmax(pred, dim=-1)\n        test_pred[i * 64:(i + 1) * 64] += pred</p></li>\n<li><p>add two model prediction together\noutput = np.argmax(((test_pred1 / TTA) +  (test_pred2 / TTA)) / 2, axis=1)</p></li>\n</ol>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 613248,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-08-30T10:57:58.107000",
          "content": "<p><a href=\"/chanhu\">@chanhu</a> This also happened to our case. Ensembling (averaging) B5 (0. 819) and B0 (0.812) made our current leaderboard score (0.837). Although, the training procedure for each model was completely different, which might have caused such synergy!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 613251,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-30T11:01:34.387000",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a> that's just awesome synergy</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 614383,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-08-31T12:21:09.477000",
          "content": "<p><a href=\"/chanhu\">@chanhu</a> <a href=\"/udaykamal\">@udaykamal</a>  how many tta transform do you use? flip,transpose  and multi-scale?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614397,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-08-31T12:41:10.700000",
          "content": "<p>tta 5 or 10. only flip, and rotation.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614420,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-31T13:10:01.410000",
          "content": "<p><a href=\"/chanhu\">@chanhu</a>  have you checked how much boost this gives you. i tried horizontal flip but the score got worse 0.001</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614483,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-08-31T14:49:45.930000",
          "content": "<p><a href=\"/hustkevin1037\">@hustkevin1037</a> we have not tried with TTA yet. Out current leaderboard score is simple average of two predictions from two different model. No TTA, no K-fold.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614515,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-08-31T15:54:09.760000",
          "content": "<p>Hey uday did you try regression only for both ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614745,
          "author_name": "Gabriel",
          "author_url": "",
          "post_date": "2019-09-01T01:55:32.897000",
          "content": "<p><a href=\"/valanm\">@valanm</a> do you use torch or keras? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614748,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-09-01T02:07:41.613000",
          "content": "<p><a href=\"/udaykamal\">@udaykamal</a>  did you use optimized threshold or just [0.5,1,5,2,5,3,5] ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 614903,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-09-01T08:03:41.890000",
          "content": "<p><a href=\"/valanm\">@valanm</a>  about +0.004.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615011,
          "author_name": "Uday Kamal",
          "author_url": "",
          "post_date": "2019-09-01T11:03:34.163000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> yes\n<a href=\"/hustkevin1037\">@hustkevin1037</a> optimized threshold resulted in lower score in our case. We are using just [0.5,1,5,2,5,3,5]</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615207,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-01T15:56:06.917000",
          "content": "<p><a href=\"/custodiogabriel\">@custodiogabriel</a> i am trying fastai for this competition\nthis might be a dumb question but are you guys ensembling logits or softmax outputs, does this matter at all?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615508,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-02T03:15:49.100000",
          "content": "<p><a href=\"/valanm\">@valanm</a> If you are doing classification, softmax outputs makes more sense. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 615547,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-09-02T04:54:49.937000",
          "content": "<p>@ uday thanks...\ncould you please tell    what is distribution of each class  in submission score..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615574,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-02T05:44:04.953000",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> thanks. Is the reason scaling? Would ligits make sense for stacking?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615596,
          "author_name": "Xuan Cao",
          "author_url": "",
          "post_date": "2019-09-02T06:23:07.523000",
          "content": "<p>It depends on what stacking methods you are using. Some are obviously sensitive to scaling. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 615733,
          "author_name": "yu",
          "author_url": "",
          "post_date": "2019-09-02T09:58:29.877000",
          "content": "<p>may I know what batch size u are using? I am training with 16 only due to limited gpu ram. Considering to rent a machine if batch size matters</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 616046,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-09-02T16:25:06.247000",
          "content": "<p>Averaging softmax outputs instead of logits gave the same score. However, 12 images (0.6%) got different label.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 620049,
          "author_name": "Chanhu",
          "author_url": "",
          "post_date": "2019-09-07T00:08:12.497000",
          "content": "<p><a href=\"/valanm\">@valanm</a> I hope you can get good score in private test set</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 616493,
      "author_name": "Atabay Ziyaden",
      "author_url": "",
      "post_date": "2019-09-03T07:21:07.007000",
      "content": "<p>EfficientNet b5, no CV or LB improvement after 256</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 612028,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-08-29T15:11:56.193000",
      "content": "<p>Is there a reason to use img size 256x256 with efficient nets? I ask this because it's not the default size for any of its models, but people seem to get good results anyway.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 612076,
          "author_name": "oct_path",
          "author_url": "",
          "post_date": "2019-08-29T15:37:22.917000",
          "content": "<p>In my understanding, the default sizes have been searched on a specific dataset. And the dataset are different from this competition dataset. So, I think it is not always necessary to resize this dataset according to the default size.\nMany people have reported that the accuracy is good even if it is different from the default size, so I think it supports this idea.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 612098,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-08-29T15:50:34.097000",
          "content": "<p>I understand <a href=\"/octpath0302\">@octpath0302</a> , this makes sense.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "610122": "This is the last chance to rethink the image size because we have little time.\nI am currently using only 256. I've already tried a larger size but it only took a long time to train and didn't give good score. What is your best image size？",
    "612806": "My single models are,  \n1. 300 Regression EfficientNet B3, LB 0.830\n2. 260 Regression EfficientNet B2, LB 0.820\n\nOur current place is from Model 1.  ",
    "611126": "I use 256 preprocessed by mean subtraction.\nMy best single model (EfficientNet-B4) got 0.829.\nI've tried 320 and 512, but it was a waste of time.",
    "615912": "Here are some of my results.  I did on image size.  All 5x CV \n\n`EfficientNet B5,  224 LB: 0.807`\n`EfficientNet B5,  300 LB: 0.812`\n`EfficientNet B5,  380 LB: 0.816`\n`EfficientNet B5,  456 LB: 0.823`\n\n`EfficientNet B4,  224 LB: 0.800`\n`EfficientNet B4,  300 LB: 0.811`\n`EfficientNet B4,  380 LB: 0.816`\n\n",
    "610215": "456, efficientnetb5  lb:0,824 single model",
    "612000": "328 B5 Classification, reach 0.825.   \n328 B5 Regression, reach 0.823.\n256 B4 Regression, reach 0.821\n\n[Uptate 2019/08/30 13:18]\n256 B4 Classification, reach 0.820",
    "616493": "EfficientNet b5, no CV or LB improvement after 256",
    "612028": "Is there a reason to use img size 256x256 with efficient nets? I ask this because it's not the default size for any of its models, but people seem to get good results anyway."
  }
}