{
  "id": 108119,
  "title": "46th place with classification and barely using external data.",
  "url": "/competitions/aptos2019-blindness-detection/writeups/0-0-eyeing-eyers-46th-place-with-classification-an",
  "author_name": "",
  "post_date": "2019-10-03T12:07:12.127Z",
  "votes": 24,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Congrats to all the winners and those who learned something new! I certainly did.</p>\n\n<p>I wanted to experiment with fastai and I entered this competition because I thought the dataset is small. Oh boy, how wrong I was. </p>\n\n<p><strong>Software:</strong> Fastai</p>\n\n<p><strong>Resources:</strong>  Kaggle Kernels only</p>\n\n<p><strong>Networks</strong> : EfficientNet 2xB3 , 2xB5 and 1x seresnext101 each with 5 folds</p>\n\n<p><strong>Preprocessing</strong>: some used cropping and other models no preprocessing.</p>\n\n<p><strong>Augmentation</strong>: train transform = brightness, contrast, hflip(), vflip() (some models), and two custom processes randomzoom with different ratios from 1vs1 to 4vs3, sometimes crop to ratio4vs3.</p>\n\n<p><strong>Optimizer</strong> : Adam</p>\n\n<p><strong>Loss</strong> : CE with label smoothing</p>\n\n<p><strong>Learning Rate Scheduler</strong>：one_cycle with max 6e-4</p>\n\n<p><strong>Validation</strong>: 5Fold stratified official data, and if used external added only to train not to hold out </p>\n\n<p><strong>Score</strong>. I did fast submission so only Public LB known:</p>\n\n<p>only official data\nB3  0.813 (448px ) \nB5  0.817 (456px )\nseresnext101 0.799 (416px) </p>\n\n<p>official + external train clipped at 2000 samples per class:\nB3  0.822 (456px)\nB5 no_clue but CV improved to 0.94+ so I added it (512px) here I trusted my gut feeling to trust your CV</p>\n\n<p><strong>Final Submission</strong>:\n0. Combining all EfficientNet outputs together no TTA. Public LB:0.831, Private 0.925)\n1. Combining all EfficientNet outputs together + h_flip. Public LB:0.828, Private 0.925)\n2. Combining all outputs together + h_flip. Public LB:0.829, Private 0.926) \n- excluded sub1 because I thought with h_flip will be more robust. Good decision\n- In some previous competition I would ensemble only models with best LB score but here only did  those with correlation &lt;0.9\n- I added seresnext last day although it didn't seem to work as good on Public LB, but it turned out it actually did pretty good on Private. A few of my early trials with seresnext scored ~0.79Public which transfered to ~0.92 Private </p>\n\n<p><strong>CONCLUSION</strong>: I am very satisfied given \n- I treated this as classification and not regression as it would be more appropriate\n- I barely used external data (GPU limit killed my primary plan and I lost the motivation first time)\n- I used fastai which is easy to get you going but I found it not so straightforward to implement your own ideas such as multiheaded outputs and combination of different losses. (so I lost the motivation for the second time) \n- I did not use pseudolabels\n- And I also started sharing solutions and discussing topics</p>\n\n<p><strong>Fun fact</strong>: I am a veterinarian on parental leave and I did most of my experimenting via phone</p>",
  "messages": [
    {
      "id": "622062",
      "postDate": "09/09/2019 08:25:44",
      "content": "<p>Congrats to all the winners and those who learned something new! I certainly did.</p>\n\n<p>I wanted to experiment with fastai and I entered this competition because I thought the dataset is small. Oh boy, how wrong I was. </p>\n\n<p><strong>Software:</strong> Fastai</p>\n\n<p><strong>Resources:</strong>  Kaggle Kernels only</p>\n\n<p><strong>Networks</strong> : EfficientNet 2xB3 , 2xB5 and 1x seresnext101 each with 5 folds</p>\n\n<p><strong>Preprocessing</strong>: some used cropping and other models no preprocessing.</p>\n\n<p><strong>Augmentation</strong>: train transform = brightness, contrast, hflip(), vflip() (some models), and two custom processes randomzoom with different ratios from 1vs1 to 4vs3, sometimes crop to ratio4vs3.</p>\n\n<p><strong>Optimizer</strong> : Adam</p>\n\n<p><strong>Loss</strong> : CE with label smoothing</p>\n\n<p><strong>Learning Rate Scheduler</strong>：one_cycle with max 6e-4</p>\n\n<p><strong>Validation</strong>: 5Fold stratified official data, and if used external added only to train not to hold out </p>\n\n<p><strong>Score</strong>. I did fast submission so only Public LB known:</p>\n\n<p>only official data\nB3  0.813 (448px ) \nB5  0.817 (456px )\nseresnext101 0.799 (416px) </p>\n\n<p>official + external train clipped at 2000 samples per class:\nB3  0.822 (456px)\nB5 no_clue but CV improved to 0.94+ so I added it (512px) here I trusted my gut feeling to trust your CV</p>\n\n<p><strong>Final Submission</strong>:\n0. Combining all EfficientNet outputs together no TTA. Public LB:0.831, Private 0.925)\n1. Combining all EfficientNet outputs together + h_flip. Public LB:0.828, Private 0.925)\n2. Combining all outputs together + h_flip. Public LB:0.829, Private 0.926) \n- excluded sub1 because I thought with h_flip will be more robust. Good decision\n- In some previous competition I would ensemble only models with best LB score but here only did  those with correlation &lt;0.9\n- I added seresnext last day although it didn't seem to work as good on Public LB, but it turned out it actually did pretty good on Private. A few of my early trials with seresnext scored ~0.79Public which transfered to ~0.92 Private </p>\n\n<p><strong>CONCLUSION</strong>: I am very satisfied given \n- I treated this as classification and not regression as it would be more appropriate\n- I barely used external data (GPU limit killed my primary plan and I lost the motivation first time)\n- I used fastai which is easy to get you going but I found it not so straightforward to implement your own ideas such as multiheaded outputs and combination of different losses. (so I lost the motivation for the second time) \n- I did not use pseudolabels\n- And I also started sharing solutions and discussing topics</p>\n\n<p><strong>Fun fact</strong>: I am a veterinarian on parental leave and I did most of my experimenting via phone</p>",
      "rawMarkdown": "Congrats to all the winners and those who learned something new! I certainly did.\n\nI wanted to experiment with fastai and I entered this competition because I thought the dataset is small. Oh boy, how wrong I was. \n\n**Software:** Fastai\n\n**Resources:**  Kaggle Kernels only\n\n**Networks** : EfficientNet 2xB3 , 2xB5 and 1x seresnext101 each with 5 folds\n\n**Preprocessing**: some used cropping and other models no preprocessing.\n\n**Augmentation**: train transform = brightness, contrast, hflip(), vflip() (some models), and two custom processes randomzoom with different ratios from 1vs1 to 4vs3, sometimes crop to ratio4vs3.\n\n**Optimizer** : Adam\n\n**Loss** : CE with label smoothing\n\n**Learning Rate Scheduler**：one_cycle with max 6e-4\n\n**Validation**: 5Fold stratified official data, and if used external added only to train not to hold out \n\n**Score**. I did fast submission so only Public LB known:\n\nonly official data\nB3  0.813 (448px ) \nB5  0.817 (456px )\nseresnext101 0.799 (416px) \n\nofficial + external train clipped at 2000 samples per class:\nB3  0.822 (456px)\nB5 no_clue but CV improved to 0.94+ so I added it (512px) here I trusted my gut feeling to trust your CV\n\n**Final Submission**:\n0. Combining all EfficientNet outputs together no TTA. Public LB:0.831, Private 0.925)\n1. Combining all EfficientNet outputs together + h_flip. Public LB:0.828, Private 0.925)\n2. Combining all outputs together + h_flip. Public LB:0.829, Private 0.926) \n- excluded sub1 because I thought with h_flip will be more robust. Good decision\n- In some previous competition I would ensemble only models with best LB score but here only did  those with correlation &lt;0.9\n- I added seresnext last day although it didn't seem to work as good on Public LB, but it turned out it actually did pretty good on Private. A few of my early trials with seresnext scored ~0.79Public which transfered to ~0.92 Private \n\n**CONCLUSION**: I am very satisfied given \n- I treated this as classification and not regression as it would be more appropriate\n- I barely used external data (GPU limit killed my primary plan and I lost the motivation first time)\n- I used fastai which is easy to get you going but I found it not so straightforward to implement your own ideas such as multiheaded outputs and combination of different losses. (so I lost the motivation for the second time) \n- I did not use pseudolabels\n- And I also started sharing solutions and discussing topics\n\n**Fun fact**: I am a veterinarian on parental leave and I did most of my experimenting via phone",
      "votes": null
    },
    {
      "id": "622164",
      "postDate": "09/09/2019 10:51:53",
      "content": "<p>Good jod, you did a great classification work😎 </p>",
      "rawMarkdown": "Good jod, you did a great classification work😎",
      "votes": null
    },
    {
      "id": "622173",
      "postDate": "09/09/2019 11:05:54",
      "content": "<p>Great work. Strangely, I could only get B3 to hit 0.765 on competition data. May be it's because I used the image size of 300x300 or you use better augmentations through fast.ai (like contrast etc helped you more. I use Keras ImageDataGenerator with random flips, zoom, brightness, shear, and rotate. I am using categorical cross entropy loss without smoothing. Any suggestions on what could be the reason of difference between 0.765 and 0.813 would be helpful. Thanks.</p>",
      "rawMarkdown": "Great work. Strangely, I could only get B3 to hit 0.765 on competition data. May be it's because I used the image size of 300x300 or you use better augmentations through fast.ai (like contrast etc helped you more. I use Keras ImageDataGenerator with random flips, zoom, brightness, shear, and rotate. I am using categorical cross entropy loss without smoothing. Any suggestions on what could be the reason of difference between 0.765 and 0.813 would be helpful. Thanks.",
      "votes": null
    },
    {
      "id": "622189",
      "postDate": "09/09/2019 11:19:24",
      "content": "<p>But you kicked my but with pseudolabels, again :) </p>",
      "rawMarkdown": "But you kicked my but with pseudolabels, again :)",
      "votes": null
    },
    {
      "id": "622196",
      "postDate": "09/09/2019 11:27:26",
      "content": "<p>It looks unreal I got here without regression and pseudolabels and with only 1/10 of external data </p>",
      "rawMarkdown": "It looks unreal I got here without regression and pseudolabels and with only 1/10 of external data",
      "votes": null
    },
    {
      "id": "622205",
      "postDate": "09/09/2019 11:32:00",
      "content": "<p>looking at the info you wrote, I could say that:\n1. augmentation helps when you do it right and you find the right one by experimenting ( it helped me ) for instance custom cropping ratio was helpful\n2. the dataset is pretty noisy and smoothing labels addresses this (it helped in my case)\n3. larger input helps (read 1st solution)\nPretty much all you said I would change to some extent :)</p>",
      "rawMarkdown": "looking at the info you wrote, I could say that:\n1. augmentation helps when you do it right and you find the right one by experimenting ( it helped me ) for instance custom cropping ratio was helpful\n2. the dataset is pretty noisy and smoothing labels addresses this (it helped in my case)\n3. larger input helps (read 1st solution)\nPretty much all you said I would change to some extent :)",
      "votes": null
    },
    {
      "id": "622420",
      "postDate": "09/09/2019 16:25:37",
      "content": "<p>Congratulations\nGreat Work...\nThanks for Sharing..... <a href=\"/valanm\">@valanm</a> </p>",
      "rawMarkdown": "Congratulations\nGreat Work...\nThanks for Sharing..... @valanm",
      "votes": null
    },
    {
      "id": "622525",
      "postDate": "09/09/2019 19:10:57",
      "content": "<p>Congratulations! Glad to see a fellow fastai user and Kaggle Kernels do so well in this competition, and also just using classification! </p>\n\n<p>I am curious, if Kaggle Kernels is your primary resource, what will you do for future competitions, as GPU usage is limited?</p>",
      "rawMarkdown": "Congratulations! Glad to see a fellow fastai user and Kaggle Kernels do so well in this competition, and also just using classification! \n\nI am curious, if Kaggle Kernels is your primary resource, what will you do for future competitions, as GPU usage is limited?",
      "votes": null
    },
    {
      "id": "622551",
      "postDate": "09/09/2019 20:00:11",
      "content": "<p>Kernels are my only resource for now and I don't plan to change it soon. I guess I will prioritize and reduce experiments</p>",
      "rawMarkdown": "Kernels are my only resource for now and I don't plan to change it soon. I guess I will prioritize and reduce experiments",
      "votes": null
    },
    {
      "id": "622577",
      "postDate": "09/09/2019 20:47:15",
      "content": "<p>Ok. Kernels and Google Colab are also my only resources as well, hence I was curious. I hope there wil be easier access for GPUs for competitions in the future, whether it is through more GCP credits, or increasing GPU limits, or through other means.</p>",
      "rawMarkdown": "Ok. Kernels and Google Colab are also my only resources as well, hence I was curious. I hope there wil be easier access for GPUs for competitions in the future, whether it is through more GCP credits, or increasing GPU limits, or through other means.",
      "votes": null
    },
    {
      "id": "622830",
      "postDate": "09/10/2019 06:18:46",
      "content": "<p>Thx</p>",
      "rawMarkdown": "Thx",
      "votes": null
    },
    {
      "id": "623752",
      "postDate": "09/11/2019 09:18:23",
      "content": "<p>Great result with classification, congratulations! I see that your single models succeeded in beating 0.8 LB. Personally, I tried to train EfficientNets (mainly B4 and B5) with different setups, but couldn't reach 0.8 LB. I also did 5-fold cv, and just averaged models on 5 folds to get one single prediction. What do you think made it possible to beat 0.8 LB? The way you trained the network (with one_cycle)? Smart augmentations? Also, how can you explain the way you smoothed labels, how strongly you did it?</p>",
      "rawMarkdown": "Great result with classification, congratulations! I see that your single models succeeded in beating 0.8 LB. Personally, I tried to train EfficientNets (mainly B4 and B5) with different setups, but couldn't reach 0.8 LB. I also did 5-fold cv, and just averaged models on 5 folds to get one single prediction. What do you think made it possible to beat 0.8 LB? The way you trained the network (with one_cycle)? Smart augmentations? Also, how can you explain the way you smoothed labels, how strongly you did it?",
      "votes": null
    },
    {
      "id": "623775",
      "postDate": "09/11/2019 09:42:21",
      "content": "<p>Augmentation gave significant boost while training on official data. I experimented with augmentation only with official data and that's why I got 0.817 and only 0.822 after including 3x more external images which have different properties so the augmentation was supposed to be different. </p>\n\n<p>To better understand smoothing read this <a href=\"https://arxiv.org/pdf/1906.02629.pdf\">paper</a> (and references therein) from a father of backpropagation and deep CNNs <strong>Geoffrey Hinton</strong>.</p>",
      "rawMarkdown": "Augmentation gave significant boost while training on official data. I experimented with augmentation only with official data and that's why I got 0.817 and only 0.822 after including 3x more external images which have different properties so the augmentation was supposed to be different. \n\nTo better understand smoothing read this [paper](https://arxiv.org/pdf/1906.02629.pdf) (and references therein) from a father of backpropagation and deep CNNs **Geoffrey Hinton**.",
      "votes": null
    },
    {
      "id": "623812",
      "postDate": "09/11/2019 10:20:00",
      "content": "<p>Thank you very much for the answer and the link to the paper! We also tried label smoothing and it worked the best for us. However, I'm afraid that we used too small value for this - we changed the correct class value from 1 to 0.96 and correspondingly the wrong classes from 0 to 0.01. Which values did you use?</p>",
      "rawMarkdown": "Thank you very much for the answer and the link to the paper! We also tried label smoothing and it worked the best for us. However, I'm afraid that we used too small value for this - we changed the correct class value from 1 to 0.96 and correspondingly the wrong classes from 0 to 0.01. Which values did you use?",
      "votes": null
    },
    {
      "id": "623855",
      "postDate": "09/11/2019 10:53:21",
      "content": "<p>I started small and I  ended up with .8 and .05\n:)</p>",
      "rawMarkdown": "I started small and I  ended up with .8 and .05\n:)",
      "votes": null
    },
    {
      "id": "623891",
      "postDate": "09/11/2019 11:38:42",
      "content": "<p>Thank you! We certainly should have tried them)</p>",
      "rawMarkdown": "Thank you! We certainly should have tried them)",
      "votes": null
    },
    {
      "id": "624194",
      "postDate": "09/11/2019 19:39:34",
      "content": "<p>Beautiful. Congratulations.</p>",
      "rawMarkdown": "Beautiful. Congratulations.",
      "votes": null
    },
    {
      "id": "624468",
      "postDate": "09/12/2019 05:44:17",
      "content": "<p>Good job <a href=\"/valanm\">@valanm</a>. Congrats. </p>",
      "rawMarkdown": "Good job @valanm. Congrats.",
      "votes": null
    },
    {
      "id": "625120",
      "postDate": "09/12/2019 18:11:42",
      "content": "<p>Haha, I also did some experimenting via phone! Congrats on the medal, and thanks for sharing.</p>",
      "rawMarkdown": "Haha, I also did some experimenting via phone! Congrats on the medal, and thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 622164,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "09/09/2019 10:51:53",
      "content": "<p>Good jod, you did a great classification work😎 </p>",
      "votes": null,
      "replies": [
        {
          "id": 622189,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/09/2019 11:19:24",
          "content": "<p>But you kicked my but with pseudolabels, again :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622196,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/09/2019 11:27:26",
          "content": "<p>It looks unreal I got here without regression and pseudolabels and with only 1/10 of external data </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 622173,
      "author_name": "basitsheikh",
      "author_url": "",
      "post_date": "09/09/2019 11:05:54",
      "content": "<p>Great work. Strangely, I could only get B3 to hit 0.765 on competition data. May be it's because I used the image size of 300x300 or you use better augmentations through fast.ai (like contrast etc helped you more. I use Keras ImageDataGenerator with random flips, zoom, brightness, shear, and rotate. I am using categorical cross entropy loss without smoothing. Any suggestions on what could be the reason of difference between 0.765 and 0.813 would be helpful. Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 622205,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/09/2019 11:32:00",
          "content": "<p>looking at the info you wrote, I could say that:\n1. augmentation helps when you do it right and you find the right one by experimenting ( it helped me ) for instance custom cropping ratio was helpful\n2. the dataset is pretty noisy and smoothing labels addresses this (it helped in my case)\n3. larger input helps (read 1st solution)\nPretty much all you said I would change to some extent :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 622420,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "09/09/2019 16:25:37",
      "content": "<p>Congratulations\nGreat Work...\nThanks for Sharing..... <a href=\"/valanm\">@valanm</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 622830,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/10/2019 06:18:46",
          "content": "<p>Thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 622525,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "09/09/2019 19:10:57",
      "content": "<p>Congratulations! Glad to see a fellow fastai user and Kaggle Kernels do so well in this competition, and also just using classification! </p>\n\n<p>I am curious, if Kaggle Kernels is your primary resource, what will you do for future competitions, as GPU usage is limited?</p>",
      "votes": null,
      "replies": [
        {
          "id": 622551,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/09/2019 20:00:11",
          "content": "<p>Kernels are my only resource for now and I don't plan to change it soon. I guess I will prioritize and reduce experiments</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 622577,
          "author_name": "tanlikesmath",
          "author_url": "",
          "post_date": "09/09/2019 20:47:15",
          "content": "<p>Ok. Kernels and Google Colab are also my only resources as well, hence I was curious. I hope there wil be easier access for GPUs for competitions in the future, whether it is through more GCP credits, or increasing GPU limits, or through other means.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 623752,
      "author_name": "blackitten13",
      "author_url": "",
      "post_date": "09/11/2019 09:18:23",
      "content": "<p>Great result with classification, congratulations! I see that your single models succeeded in beating 0.8 LB. Personally, I tried to train EfficientNets (mainly B4 and B5) with different setups, but couldn't reach 0.8 LB. I also did 5-fold cv, and just averaged models on 5 folds to get one single prediction. What do you think made it possible to beat 0.8 LB? The way you trained the network (with one_cycle)? Smart augmentations? Also, how can you explain the way you smoothed labels, how strongly you did it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 623775,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/11/2019 09:42:21",
          "content": "<p>Augmentation gave significant boost while training on official data. I experimented with augmentation only with official data and that's why I got 0.817 and only 0.822 after including 3x more external images which have different properties so the augmentation was supposed to be different. </p>\n\n<p>To better understand smoothing read this <a href=\"https://arxiv.org/pdf/1906.02629.pdf\">paper</a> (and references therein) from a father of backpropagation and deep CNNs <strong>Geoffrey Hinton</strong>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623812,
          "author_name": "blackitten13",
          "author_url": "",
          "post_date": "09/11/2019 10:20:00",
          "content": "<p>Thank you very much for the answer and the link to the paper! We also tried label smoothing and it worked the best for us. However, I'm afraid that we used too small value for this - we changed the correct class value from 1 to 0.96 and correspondingly the wrong classes from 0 to 0.01. Which values did you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623855,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "09/11/2019 10:53:21",
          "content": "<p>I started small and I  ended up with .8 and .05\n:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 623891,
          "author_name": "blackitten13",
          "author_url": "",
          "post_date": "09/11/2019 11:38:42",
          "content": "<p>Thank you! We certainly should have tried them)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 624194,
      "author_name": "abualabed",
      "author_url": "",
      "post_date": "09/11/2019 19:39:34",
      "content": "<p>Beautiful. Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 624468,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "09/12/2019 05:44:17",
      "content": "<p>Good job <a href=\"/valanm\">@valanm</a>. Congrats. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 625120,
      "author_name": "amqdnguyen",
      "author_url": "",
      "post_date": "09/12/2019 18:11:42",
      "content": "<p>Haha, I also did some experimenting via phone! Congrats on the medal, and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "622062": "Congrats to all the winners and those who learned something new! I certainly did.\n\nI wanted to experiment with fastai and I entered this competition because I thought the dataset is small. Oh boy, how wrong I was. \n\n**Software:** Fastai\n\n**Resources:**  Kaggle Kernels only\n\n**Networks** : EfficientNet 2xB3 , 2xB5 and 1x seresnext101 each with 5 folds\n\n**Preprocessing**: some used cropping and other models no preprocessing.\n\n**Augmentation**: train transform = brightness, contrast, hflip(), vflip() (some models), and two custom processes randomzoom with different ratios from 1vs1 to 4vs3, sometimes crop to ratio4vs3.\n\n**Optimizer** : Adam\n\n**Loss** : CE with label smoothing\n\n**Learning Rate Scheduler**：one_cycle with max 6e-4\n\n**Validation**: 5Fold stratified official data, and if used external added only to train not to hold out \n\n**Score**. I did fast submission so only Public LB known:\n\nonly official data\nB3  0.813 (448px ) \nB5  0.817 (456px )\nseresnext101 0.799 (416px) \n\nofficial + external train clipped at 2000 samples per class:\nB3  0.822 (456px)\nB5 no_clue but CV improved to 0.94+ so I added it (512px) here I trusted my gut feeling to trust your CV\n\n**Final Submission**:\n0. Combining all EfficientNet outputs together no TTA. Public LB:0.831, Private 0.925)\n1. Combining all EfficientNet outputs together + h_flip. Public LB:0.828, Private 0.925)\n2. Combining all outputs together + h_flip. Public LB:0.829, Private 0.926) \n- excluded sub1 because I thought with h_flip will be more robust. Good decision\n- In some previous competition I would ensemble only models with best LB score but here only did  those with correlation &lt;0.9\n- I added seresnext last day although it didn't seem to work as good on Public LB, but it turned out it actually did pretty good on Private. A few of my early trials with seresnext scored ~0.79Public which transfered to ~0.92 Private \n\n**CONCLUSION**: I am very satisfied given \n- I treated this as classification and not regression as it would be more appropriate\n- I barely used external data (GPU limit killed my primary plan and I lost the motivation first time)\n- I used fastai which is easy to get you going but I found it not so straightforward to implement your own ideas such as multiheaded outputs and combination of different losses. (so I lost the motivation for the second time) \n- I did not use pseudolabels\n- And I also started sharing solutions and discussing topics\n\n**Fun fact**: I am a veterinarian on parental leave and I did most of my experimenting via phone",
    "622164": "Good jod, you did a great classification work😎",
    "622173": "Great work. Strangely, I could only get B3 to hit 0.765 on competition data. May be it's because I used the image size of 300x300 or you use better augmentations through fast.ai (like contrast etc helped you more. I use Keras ImageDataGenerator with random flips, zoom, brightness, shear, and rotate. I am using categorical cross entropy loss without smoothing. Any suggestions on what could be the reason of difference between 0.765 and 0.813 would be helpful. Thanks.",
    "622189": "But you kicked my but with pseudolabels, again :)",
    "622196": "It looks unreal I got here without regression and pseudolabels and with only 1/10 of external data",
    "622205": "looking at the info you wrote, I could say that:\n1. augmentation helps when you do it right and you find the right one by experimenting ( it helped me ) for instance custom cropping ratio was helpful\n2. the dataset is pretty noisy and smoothing labels addresses this (it helped in my case)\n3. larger input helps (read 1st solution)\nPretty much all you said I would change to some extent :)",
    "622420": "Congratulations\nGreat Work...\nThanks for Sharing..... @valanm",
    "622525": "Congratulations! Glad to see a fellow fastai user and Kaggle Kernels do so well in this competition, and also just using classification! \n\nI am curious, if Kaggle Kernels is your primary resource, what will you do for future competitions, as GPU usage is limited?",
    "622551": "Kernels are my only resource for now and I don't plan to change it soon. I guess I will prioritize and reduce experiments",
    "622577": "Ok. Kernels and Google Colab are also my only resources as well, hence I was curious. I hope there wil be easier access for GPUs for competitions in the future, whether it is through more GCP credits, or increasing GPU limits, or through other means.",
    "622830": "Thx",
    "623752": "Great result with classification, congratulations! I see that your single models succeeded in beating 0.8 LB. Personally, I tried to train EfficientNets (mainly B4 and B5) with different setups, but couldn't reach 0.8 LB. I also did 5-fold cv, and just averaged models on 5 folds to get one single prediction. What do you think made it possible to beat 0.8 LB? The way you trained the network (with one_cycle)? Smart augmentations? Also, how can you explain the way you smoothed labels, how strongly you did it?",
    "623775": "Augmentation gave significant boost while training on official data. I experimented with augmentation only with official data and that's why I got 0.817 and only 0.822 after including 3x more external images which have different properties so the augmentation was supposed to be different. \n\nTo better understand smoothing read this [paper](https://arxiv.org/pdf/1906.02629.pdf) (and references therein) from a father of backpropagation and deep CNNs **Geoffrey Hinton**.",
    "623812": "Thank you very much for the answer and the link to the paper! We also tried label smoothing and it worked the best for us. However, I'm afraid that we used too small value for this - we changed the correct class value from 1 to 0.96 and correspondingly the wrong classes from 0 to 0.01. Which values did you use?",
    "623855": "I started small and I  ended up with .8 and .05\n:)",
    "623891": "Thank you! We certainly should have tried them)",
    "624194": "Beautiful. Congratulations.",
    "624468": "Good job @valanm. Congrats.",
    "625120": "Haha, I also did some experimenting via phone! Congrats on the medal, and thanks for sharing."
  },
  "source": "meta"
}