{
  "id": 99353,
  "title": "Some Tips",
  "url": "/competitions/aptos2019-blindness-detection/discussion/99353",
  "author_name": "DrHB",
  "post_date": "2019-07-10T15:39:47.621000",
  "votes": 190,
  "comment_count": 114,
  "views": 0,
  "content": "<p>This is a post about some experiments I performed which did not help me so much, but eventually leaded to my current position in LB. </p>\n\n<p>My project setup:\n<code>\nIMAGE SIZE = 256\nVALIDATION = 20%\nno TTA\nno x-fold CV\n</code></p>\n\n<p>EXPERIMENT 01:</p>\n\n<p>```\nepoch: 10\nmodel: resnet50 (pretrained = True)\nCV: 0.932291026\nLB: 0.755</p>\n\n<p>```\nEXPERIMENT 02:\nTrained as a classification problem, and then retrained as a regression problem. </p>\n\n<p><code>\nmodel: resnet50 (pretrained = True)\nepoch: 10\nCV: 0.930\nLB: 0.750\n</code></p>\n\n<p>EXPERIMENT 03:\nIdea came from the old competition where the winners could achieve good score with progressively resizing image size</p>\n\n<p>round 1: trained on 128 (epoch: 10)\nround 2: finetuned on 256 (epoch: 5)\nround 3: finetuned on 512 (epoch: 10)</p>\n\n<p><code>\nmodel: resnet50 (pretrained = True)\nCV: 0.91938504\nLB: 0.734\n</code></p>\n\n<p>EXPERIMENT 04:\nThanks to the <a href=\"/ilovescience\">@ilovescience</a> for the dataset, I tried to balance all classes and trained the network. </p>\n\n<p>```\nepoch: 15\nCV: 0.933592607\nLB:  0.751\nmodel: resnet50 (pretrained = True)</p>\n\n<p>```\nEXPERIMENT 05:\nSame as experiment 4 but I have added label Smoothing \n<a href=\"https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0\">https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0</a> (thanks to <a href=\"/dimitreoliveira\">@dimitreoliveira</a>).</p>\n\n<p>round 1: first trained as a classification problem (so I can use label smoothing)\nround 2: finetuned as a regression problem</p>\n\n<p>```\nmodel: resnet50 (pretrained = True)\nepoch: 29\nCV: 0.927\nLB:  0.0.732</p>\n\n<p>```</p>\n\n<p>Other experiments that failed:\n-Using Auto-encoders to generate images \n-Using Auto -encoders latent layes as input to NN\n-GANs to generate images of the classes that have low examples\n-Diffrent types of preprocessing (gray, edges, different filters)</p>\n\n<p>Tips: \n-please pay attentions to image transformations\n-if you have resized images on local computer make sure you do same resizing when you use Kaggle kernel for submission \n-in the beginning of competition don't waste your time doing x fold CVs, it’s better to focus on building one robust model with same validation set across different experimenting</p>\n\n<p>Concerns:\n-I am still concerned about the gap between CV and LB.\n-People have pointed out that there is a lot of noise in labels.. I haven’t tried but one way to solve this issue would be to use smartly label smoothening in regression.  </p>\n\n<p>EDIT 1:\ninteresting research from google AI talking about how to judge if your NN is generalizing <a href=\"https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html\">https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html</a></p>\n\n<p>EDIT 2:\nTable of all submissions same validation set across experiments (without classification). Hope it helps</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdf51224302a7656a2e41efd20d849821%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562960483994331&amp;alt=media\" alt=\"\"></p>\n\n<p>EDIT 3:\nRecently there was paper that claimed that CNN's tend to learn texture information, but if we push training toward shapes it can improve accuracy. Its pretty easy to implement and I had some success in terms of training. For people who are interested here is a link <a href=\"https://openreview.net/forum?id=Bygh9j09KX\">https://openreview.net/forum?id=Bygh9j09KX</a> and link to GitHub GitHub <a href=\"https://github.com/rgeirhos/texture-vs-shape\">https://github.com/rgeirhos/texture-vs-shape</a></p>",
  "messages": [
    {
      "id": 572186,
      "postDate": "2019-07-10T15:39:47.623Z",
      "content": "<p>This is a post about some experiments I performed which did not help me so much, but eventually leaded to my current position in LB. </p>\n\n<p>My project setup:\n<code>\nIMAGE SIZE = 256\nVALIDATION = 20%\nno TTA\nno x-fold CV\n</code></p>\n\n<p>EXPERIMENT 01:</p>\n\n<p>```\nepoch: 10\nmodel: resnet50 (pretrained = True)\nCV: 0.932291026\nLB: 0.755</p>\n\n<p>```\nEXPERIMENT 02:\nTrained as a classification problem, and then retrained as a regression problem. </p>\n\n<p><code>\nmodel: resnet50 (pretrained = True)\nepoch: 10\nCV: 0.930\nLB: 0.750\n</code></p>\n\n<p>EXPERIMENT 03:\nIdea came from the old competition where the winners could achieve good score with progressively resizing image size</p>\n\n<p>round 1: trained on 128 (epoch: 10)\nround 2: finetuned on 256 (epoch: 5)\nround 3: finetuned on 512 (epoch: 10)</p>\n\n<p><code>\nmodel: resnet50 (pretrained = True)\nCV: 0.91938504\nLB: 0.734\n</code></p>\n\n<p>EXPERIMENT 04:\nThanks to the <a href=\"/ilovescience\">@ilovescience</a> for the dataset, I tried to balance all classes and trained the network. </p>\n\n<p>```\nepoch: 15\nCV: 0.933592607\nLB:  0.751\nmodel: resnet50 (pretrained = True)</p>\n\n<p>```\nEXPERIMENT 05:\nSame as experiment 4 but I have added label Smoothing \n<a href=\"https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0\">https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0</a> (thanks to <a href=\"/dimitreoliveira\">@dimitreoliveira</a>).</p>\n\n<p>round 1: first trained as a classification problem (so I can use label smoothing)\nround 2: finetuned as a regression problem</p>\n\n<p>```\nmodel: resnet50 (pretrained = True)\nepoch: 29\nCV: 0.927\nLB:  0.0.732</p>\n\n<p>```</p>\n\n<p>Other experiments that failed:\n-Using Auto-encoders to generate images \n-Using Auto -encoders latent layes as input to NN\n-GANs to generate images of the classes that have low examples\n-Diffrent types of preprocessing (gray, edges, different filters)</p>\n\n<p>Tips: \n-please pay attentions to image transformations\n-if you have resized images on local computer make sure you do same resizing when you use Kaggle kernel for submission \n-in the beginning of competition don't waste your time doing x fold CVs, it’s better to focus on building one robust model with same validation set across different experimenting</p>\n\n<p>Concerns:\n-I am still concerned about the gap between CV and LB.\n-People have pointed out that there is a lot of noise in labels.. I haven’t tried but one way to solve this issue would be to use smartly label smoothening in regression.  </p>\n\n<p>EDIT 1:\ninteresting research from google AI talking about how to judge if your NN is generalizing <a href=\"https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html\">https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html</a></p>\n\n<p>EDIT 2:\nTable of all submissions same validation set across experiments (without classification). Hope it helps</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdf51224302a7656a2e41efd20d849821%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562960483994331&amp;alt=media\" alt=\"\"></p>\n\n<p>EDIT 3:\nRecently there was paper that claimed that CNN's tend to learn texture information, but if we push training toward shapes it can improve accuracy. Its pretty easy to implement and I had some success in terms of training. For people who are interested here is a link <a href=\"https://openreview.net/forum?id=Bygh9j09KX\">https://openreview.net/forum?id=Bygh9j09KX</a> and link to GitHub GitHub <a href=\"https://github.com/rgeirhos/texture-vs-shape\">https://github.com/rgeirhos/texture-vs-shape</a></p>",
      "rawMarkdown": "This is a post about some experiments I performed which did not help me so much, but eventually leaded to my current position in LB. \n\nMy project setup:\n```\nIMAGE SIZE = 256\nVALIDATION = 20%\nno TTA\nno x-fold CV\n```\n\nEXPERIMENT 01:\n\n```\nepoch: 10\nmodel: resnet50 (pretrained = True)\nCV: 0.932291026\nLB: 0.755\n\n```\nEXPERIMENT 02:\nTrained as a classification problem, and then retrained as a regression problem. \n\n```\nmodel: resnet50 (pretrained = True)\nepoch: 10\nCV: 0.930\nLB: 0.750\n```\n\nEXPERIMENT 03:\nIdea came from the old competition where the winners could achieve good score with progressively resizing image size\n\nround 1: trained on 128 (epoch: 10)\nround 2: finetuned on 256 (epoch: 5)\nround 3: finetuned on 512 (epoch: 10)\n\n```\nmodel: resnet50 (pretrained = True)\nCV: 0.91938504\nLB: 0.734\n```\n\n\nEXPERIMENT 04:\nThanks to the @ilovescience for the dataset, I tried to balance all classes and trained the network. \n\n```\nepoch: 15\nCV: 0.933592607\nLB:  0.751\nmodel: resnet50 (pretrained = True)\n\n```\nEXPERIMENT 05:\nSame as experiment 4 but I have added label Smoothing \nhttps://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0 (thanks to @dimitreoliveira).\n\nround 1: first trained as a classification problem (so I can use label smoothing)\nround 2: finetuned as a regression problem\n\n```\nmodel: resnet50 (pretrained = True)\nepoch: 29\nCV: 0.927\nLB:  0.0.732\n\n```\n\nOther experiments that failed:\n-Using Auto-encoders to generate images \n-Using Auto -encoders latent layes as input to NN\n-GANs to generate images of the classes that have low examples\n-Diffrent types of preprocessing (gray, edges, different filters)\n\n\nTips: \n-please pay attentions to image transformations\n-if you have resized images on local computer make sure you do same resizing when you use Kaggle kernel for submission \n-in the beginning of competition don't waste your time doing x fold CVs, it’s better to focus on building one robust model with same validation set across different experimenting\n\nConcerns:\n-I am still concerned about the gap between CV and LB.\n-People have pointed out that there is a lot of noise in labels.. I haven’t tried but one way to solve this issue would be to use smartly label smoothening in regression.  \n\nEDIT 1:\ninteresting research from google AI talking about how to judge if your NN is generalizing https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html\n\nEDIT 2:\nTable of all submissions same validation set across experiments (without classification). Hope it helps\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdf51224302a7656a2e41efd20d849821%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562960483994331&amp;alt=media)\n\n\nEDIT 3:\nRecently there was paper that claimed that CNN's tend to learn texture information, but if we push training toward shapes it can improve accuracy. Its pretty easy to implement and I had some success in terms of training. For people who are interested here is a link https://openreview.net/forum?id=Bygh9j09KX and link to GitHub GitHub https://github.com/rgeirhos/texture-vs-shape",
      "votes": 188
    },
    {
      "id": 572236,
      "postDate": "2019-07-10T16:54:56.123Z",
      "content": "<ol>\n<li>The difference between train set and public test set target distribution is large. 2. The target distribution of public test set is very different with 2015 competition dataset too. 3. The difference between target distribution of train set and 2015 competition data is not so large, one has (73% level 0, 15% level 2) and one has (50% level 0, 25% level 2). There are too many 2s for public test set based on my submission (50% level 2, 21% level 0).  I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.</li>\n</ol>",
      "rawMarkdown": "1. The difference between train set and public test set target distribution is large. 2. The target distribution of public test set is very different with 2015 competition dataset too. 3. The difference between target distribution of train set and 2015 competition data is not so large, one has (73% level 0, 15% level 2) and one has (50% level 0, 25% level 2). There are too many 2s for public test set based on my submission (50% level 2, 21% level 0).  I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.",
      "votes": 11,
      "replies": [
        {
          "id": 572283,
          "postDate": "2019-07-10T18:09:38.483Z",
          "content": "<p>Thanks for your insights. I agree one should be really careful using optimized thresholds with regressions. I haven't done experiments yet with x-folds CV but I assume simple averaging across the folds will be better (maybe safer) than optimized kappa over folds... </p>",
          "rawMarkdown": "Thanks for your insights. I agree one should be really careful using optimized thresholds with regressions. I haven't done experiments yet with x-folds CV but I assume simple averaging across the folds will be better (maybe safer) than optimized kappa over folds... ",
          "votes": 2
        },
        {
          "id": 572959,
          "postDate": "2019-07-11T15:44:56.517Z",
          "content": "<p>&gt; I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.</p>\n\n<p>Exactly, I'm using threshold optimised at the validation set and I have this feeling that my validation set doesn't really represent the test set out there that's the reason I'm getting bad results.</p>",
          "rawMarkdown": "&gt; I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.\n\nExactly, I'm using threshold optimised at the validation set and I have this feeling that my validation set doesn't really represent the test set out there that's the reason I'm getting bad results."
        }
      ]
    },
    {
      "id": 573231,
      "postDate": "2019-07-12T02:05:25.177Z",
      "content": "<p>Just a small fix, the label smoothing link is incorrect, you should remove the \")\" at the end of the hyperlink ; )</p>\n\n<p>like this: <a href=\"https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0\">https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0</a></p>",
      "rawMarkdown": "Just a small fix, the label smoothing link is incorrect, you should remove the \")\" at the end of the hyperlink ; )\n\nlike this: https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0",
      "votes": 7,
      "replies": [
        {
          "id": 573569,
          "postDate": "2019-07-12T12:27:57.160Z",
          "content": "<p>thank you very much!!</p>",
          "rawMarkdown": "thank you very much!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 579437,
      "postDate": "2019-07-18T20:38:51.017Z",
      "content": "<p>Thanks for sharing. My base experiment setup is similar. Although I am finding slight increase in both my local cv and lb with progressive resizing to 128-&gt;256-&gt;299 for resnet50.</p>",
      "rawMarkdown": "Thanks for sharing. My base experiment setup is similar. Although I am finding slight increase in both my local cv and lb with progressive resizing to 128-&gt;256-&gt;299 for resnet50.",
      "votes": 6,
      "replies": [
        {
          "id": 580720,
          "postDate": "2019-07-20T16:10:29.250Z",
          "content": "<p>I can confirm, after playing around a little bit image size increase leads to better lb score =)</p>",
          "rawMarkdown": "I can confirm, after playing around a little bit image size increase leads to better lb score =)",
          "votes": 3
        }
      ]
    },
    {
      "id": 573979,
      "postDate": "2019-07-13T04:23:24.010Z",
      "content": "<p>Thanks for the post. Do you think your LB scores are mainly due to careful image transformations in your training set? I've tried training resnet50 for a similar number of epochs, I get over 0.9 kappa in CV but LB scores are around 0.65, and I'm not sure what I should be doing differently. Would you suggest focusing on hyperparameters, image transformations, or expanding the training set?</p>",
      "rawMarkdown": "Thanks for the post. Do you think your LB scores are mainly due to careful image transformations in your training set? I've tried training resnet50 for a similar number of epochs, I get over 0.9 kappa in CV but LB scores are around 0.65, and I'm not sure what I should be doing differently. Would you suggest focusing on hyperparameters, image transformations, or expanding the training set?",
      "votes": 3,
      "replies": [
        {
          "id": 574054,
          "postDate": "2019-07-13T07:02:11.020Z",
          "content": "<p>Hey, Generalisation is a big issue in this competition. The LB is decided by a private test set which is ~10x the public test set. People are using optimized kappa but that would be useless if we don't have a good validation set (good: which is as close to test distribution as possible). As far as expanding the training set (using the previous year's data) is concerned I haven't had good results so far, I'm still experimenting with that, fingers crossed.</p>",
          "rawMarkdown": "Hey, Generalisation is a big issue in this competition. The LB is decided by a private test set which is ~10x the public test set. People are using optimized kappa but that would be useless if we don't have a good validation set (good: which is as close to test distribution as possible). As far as expanding the training set (using the previous year's data) is concerned I haven't had good results so far, I'm still experimenting with that, fingers crossed.",
          "votes": 2
        }
      ]
    },
    {
      "id": 594922,
      "postDate": "2019-08-08T16:26:32.007Z",
      "content": "<p>Very well documented. I like the research quoted. \nI tried the things you mentioned above (autoencoders, gan, various architectures and image sizes) and only now I found your report. \nI can confirm your experiments. I tried a few extra things such as image filtering including Ben's processing. Did not seem to help. </p>",
      "rawMarkdown": "Very well documented. I like the research quoted. \nI tried the things you mentioned above (autoencoders, gan, various architectures and image sizes) and only now I found your report. \nI can confirm your experiments. I tried a few extra things such as image filtering including Ben's processing. Did not seem to help. \n",
      "votes": 1,
      "replies": [
        {
          "id": 594938,
          "postDate": "2019-08-08T17:01:39.527Z",
          "content": "<p>it always good to hear when people can reproduce your results! Good luck in this competition =) </p>",
          "rawMarkdown": "it always good to hear when people can reproduce your results! Good luck in this competition =) ",
          "votes": 1
        },
        {
          "id": 594979,
          "postDate": "2019-08-08T17:41:29.587Z",
          "content": "<p>Thanks, to you, too! Looks like we all need some luck in this competition.</p>",
          "rawMarkdown": "Thanks, to you, too! Looks like we all need some luck in this competition.",
          "votes": 1
        },
        {
          "id": 608603,
          "postDate": "2019-08-27T02:37:41.787Z",
          "content": "<p>Hi, a little new to this but is using GANs to create images with lesser classes a common thing to do?\nFirst time I've ever heard of it, hence the question.</p>",
          "rawMarkdown": "Hi, a little new to this but is using GANs to create images with lesser classes a common thing to do?\nFirst time I've ever heard of it, hence the question."
        }
      ]
    },
    {
      "id": 584495,
      "postDate": "2019-07-26T04:42:30.303Z",
      "content": "<p>When using multiple models that have similar input shape, do the most important preprocessing (cropping and resize) beforehand and load it to disk. Now you are wondering how to fit all of this on the disk. </p>\n\n<ol>\n<li>First, you need to ready your model (do not perform training on the same kernel)</li>\n<li>Separate inferences to several batches. </li>\n</ol>\n\n<p>Basically, the workflow on each batch will be :\n1. Crop and resize (you can further speed this up with multiprocessing) and save to disk\n2. Perform inference, you can use TTA if you want. TTA should be quick enough since you are not performing resizing.\n3. After finish one batch clear these images and repeat with next batch.</p>\n\n<p>Since most of the images are big enough, it takes quite a lot of time doing preprocessing and with gpu on you will not have that much computing power in a quick fashion, hence I think my tricks should help.\nThis could work if your model is not too big and you can load all your ensemble model in one-go</p>",
      "rawMarkdown": "When using multiple models that have similar input shape, do the most important preprocessing (cropping and resize) beforehand and load it to disk. Now you are wondering how to fit all of this on the disk. \n\n1. First, you need to ready your model (do not perform training on the same kernel)\n2. Separate inferences to several batches. \n\nBasically, the workflow on each batch will be :\n1. Crop and resize (you can further speed this up with multiprocessing) and save to disk\n2. Perform inference, you can use TTA if you want. TTA should be quick enough since you are not performing resizing.\n3. After finish one batch clear these images and repeat with next batch.\n\nSince most of the images are big enough, it takes quite a lot of time doing preprocessing and with gpu on you will not have that much computing power in a quick fashion, hence I think my tricks should help.\nThis could work if your model is not too big and you can load all your ensemble model in one-go\n",
      "votes": 1
    },
    {
      "id": 579092,
      "postDate": "2019-07-18T13:33:59.723Z",
      "content": "<p>Great.. thanks for sharing .. so in general optimized_kappa gave better results in LB </p>",
      "rawMarkdown": "Great.. thanks for sharing .. so in general optimized_kappa gave better results in LB ",
      "votes": 1,
      "replies": [
        {
          "id": 580721,
          "postDate": "2019-07-20T16:11:23.443Z",
          "content": "<p>Its very risky to use optimized kappa, unless you are very confident that your valid data matches test distribution =) </p>",
          "rawMarkdown": "Its very risky to use optimized kappa, unless you are very confident that your valid data matches test distribution =) ",
          "votes": 1
        },
        {
          "id": 586368,
          "postDate": "2019-07-29T04:25:10.053Z",
          "content": "<p>I agree with DrHB. Since qwk is closely related to mean_squared_error, i think optimized mse is better alternative to qwk.</p>",
          "rawMarkdown": "I agree with DrHB. Since qwk is closely related to mean_squared_error, i think optimized mse is better alternative to qwk.",
          "votes": 1
        },
        {
          "id": 602367,
          "postDate": "2019-08-19T02:11:58.650Z",
          "content": "<p>Would optimised kappa help with smoothing out the effects of imbalanced classes? </p>",
          "rawMarkdown": "Would optimised kappa help with smoothing out the effects of imbalanced classes? "
        }
      ]
    },
    {
      "id": 578238,
      "postDate": "2019-07-17T13:59:59.870Z",
      "content": "<p>Thanks, DrHB for your kind sharing. I noticed that you did many experiments. May I know how long it would take you to approach one model on average? I am asking this because it takes me hours to run one experiment with my local machine(MacBook pro i5 8GB). I am thinking to rent an AWS server to do the experiments faster.</p>",
      "rawMarkdown": "Thanks, DrHB for your kind sharing. I noticed that you did many experiments. May I know how long it would take you to approach one model on average? I am asking this because it takes me hours to run one experiment with my local machine(MacBook pro i5 8GB). I am thinking to rent an AWS server to do the experiments faster.",
      "votes": 1,
      "replies": [
        {
          "id": 578257,
          "postDate": "2019-07-17T14:22:41.447Z",
          "content": "<p>AWS is costly. Use vast.ai </p>",
          "rawMarkdown": "AWS is costly. Use vast.ai ",
          "votes": 2
        },
        {
          "id": 580785,
          "postDate": "2019-07-20T18:48:30.170Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 580860,
          "postDate": "2019-07-20T23:07:11.317Z",
          "content": "<p>Gcloud is also a good option. If you register you get 300 dollars to use towards computing =) </p>",
          "rawMarkdown": "Gcloud is also a good option. If you register you get 300 dollars to use towards computing =) ",
          "votes": 1
        },
        {
          "id": 581130,
          "postDate": "2019-07-21T13:30:55.347Z",
          "content": "<p>Thanks for sharing. I have a concern with cloud though. If I use a very powerful machine in the cloud and it takes me 4 hrs to run. But it would take 10 hrs to run in kaggle kernel. The limitation of the kaggle kernel is only 9 hrs. Any thoughts on this situation?</p>",
          "rawMarkdown": "Thanks for sharing. I have a concern with cloud though. If I use a very powerful machine in the cloud and it takes me 4 hrs to run. But it would take 10 hrs to run in kaggle kernel. The limitation of the kaggle kernel is only 9 hrs. Any thoughts on this situation?"
        },
        {
          "id": 581134,
          "postDate": "2019-07-21T13:45:31.307Z",
          "content": "<p><a href=\"/bingwu0\">@bingwu0</a> You can train your models in cloud, download trained weights, write a inference kernel to use those weights and predict on the dataset. You don't have to train your models while you submit your predictions, just use the trained weights. I hope you get it.</p>",
          "rawMarkdown": "@bingwu0 You can train your models in cloud, download trained weights, write a inference kernel to use those weights and predict on the dataset. You don't have to train your models while you submit your predictions, just use the trained weights. I hope you get it.",
          "votes": 3
        },
        {
          "id": 581308,
          "postDate": "2019-07-21T18:32:10.213Z",
          "content": "<p>Yes,  I see your point. Thanks for your sharing!</p>",
          "rawMarkdown": "Yes,  I see your point. Thanks for your sharing!"
        }
      ]
    },
    {
      "id": 575076,
      "postDate": "2019-07-15T02:51:01.407Z",
      "content": "<p>Please try use ResNet-101 or ResNet-152 if your device accept this model</p>",
      "rawMarkdown": "Please try use ResNet-101 or ResNet-152 if your device accept this model",
      "votes": 1
    },
    {
      "id": 575066,
      "postDate": "2019-07-15T02:17:33.687Z",
      "content": "<p>Good</p>",
      "rawMarkdown": "Good",
      "votes": 1
    },
    {
      "id": 574842,
      "postDate": "2019-07-14T14:52:31.240Z",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting",
      "votes": 1
    },
    {
      "id": 574819,
      "postDate": "2019-07-14T14:22:17.677Z",
      "content": "<p>Thanks a lot for the information. I had the similar concern. My model is getting better in CV but got worse in LB. Made me confuse about the performance of my model.</p>",
      "rawMarkdown": "Thanks a lot for the information. I had the similar concern. My model is getting better in CV but got worse in LB. Made me confuse about the performance of my model.",
      "votes": 1
    },
    {
      "id": 574611,
      "postDate": "2019-07-14T07:09:56.373Z",
      "content": "<p>nice tips for me, thx.</p>",
      "rawMarkdown": "nice tips for me, thx.",
      "votes": 1
    },
    {
      "id": 574430,
      "postDate": "2019-07-13T19:55:53.533Z",
      "content": "<p>interesting! </p>",
      "rawMarkdown": "interesting! ",
      "votes": 1
    },
    {
      "id": 574338,
      "postDate": "2019-07-13T16:38:29.290Z",
      "content": "<p>Cool! People usually only show what works and ignore posting how they got there.</p>",
      "rawMarkdown": "Cool! People usually only show what works and ignore posting how they got there.",
      "votes": 1
    },
    {
      "id": 573445,
      "postDate": "2019-07-12T09:40:17.033Z",
      "content": "<p>Interesting</p>",
      "rawMarkdown": "Interesting\n",
      "votes": 1
    },
    {
      "id": 572522,
      "postDate": "2019-07-11T03:13:12.560Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> . Thanks for sharing. Have you tried x-fold cross-validation (no. of folds) yet? How about split train samples into Train-Valid-Test(0.1)?</p>",
      "rawMarkdown": "@drhabib . Thanks for sharing. Have you tried x-fold cross-validation (no. of folds) yet? How about split train samples into Train-Valid-Test(0.1)?",
      "votes": 1,
      "replies": [
        {
          "id": 572791,
          "postDate": "2019-07-11T12:08:19.200Z",
          "content": "<p>I haven't tried cross fold validation. Test-Valid (0.1) I feel it will lead to overfitting.. but I might be wrong </p>",
          "rawMarkdown": "I haven't tried cross fold validation. Test-Valid (0.1) I feel it will lead to overfitting.. but I might be wrong "
        },
        {
          "id": 599231,
          "postDate": "2019-08-14T17:32:14.263Z",
          "content": "<p>Wouldn't dropouts help with that?</p>",
          "rawMarkdown": "Wouldn't dropouts help with that?"
        }
      ]
    },
    {
      "id": 572490,
      "postDate": "2019-07-11T01:52:35.153Z",
      "content": "<p>Hi,man,thanks  for sharing.\nDo  you  train   with  an  external  data    or  you  just  train  in   the  official  dataset   to   get   13th?</p>",
      "rawMarkdown": "Hi,man,thanks  for sharing.\nDo  you  train   with  an  external  data    or  you  just  train  in   the  official  dataset   to   get   13th?",
      "votes": 1,
      "replies": [
        {
          "id": 572792,
          "postDate": "2019-07-11T12:08:41.153Z",
          "content": "<p>I have combination of both. </p>",
          "rawMarkdown": "I have combination of both. "
        }
      ]
    },
    {
      "id": 572482,
      "postDate": "2019-07-11T01:42:37.577Z",
      "content": "<p>So what's your image transformations, guys ?</p>",
      "rawMarkdown": "So what's your image transformations, guys ?",
      "votes": 1,
      "replies": [
        {
          "id": 572491,
          "postDate": "2019-07-11T02:02:04.417Z",
          "content": "<p>```\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]</p>\n\n<p>train_transform = transforms.Compose([\n    transforms.Resize((IMAGE_SIZE, IMAGE_SIZE)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ToTensor(),\n    transforms.Normalize(IMG_MEAN, IMG_STD)\n])\n```</p>",
          "rawMarkdown": "```\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]\n\ntrain_transform = transforms.Compose([\n    transforms.Resize((IMAGE_SIZE, IMAGE_SIZE)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ToTensor(),\n    transforms.Normalize(IMG_MEAN, IMG_STD)\n])\n```",
          "votes": 4
        },
        {
          "id": 572663,
          "postDate": "2019-07-11T08:34:29.323Z",
          "content": "<p><code>\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]\n</code>\nShould we be using Image-Net stats here as well?</p>",
          "rawMarkdown": "```\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]\n```\nShould we be using Image-Net stats here as well?",
          "votes": 1
        },
        {
          "id": 572789,
          "postDate": "2019-07-11T12:07:06.650Z",
          "content": "<p>yes, at least in my case. </p>",
          "rawMarkdown": "yes, at least in my case. "
        },
        {
          "id": 579345,
          "postDate": "2019-07-18T18:48:39.447Z",
          "content": "<p>Thanks <a href=\"/rinnqd\">@rinnqd</a> how did you calculate the image stats?</p>\n\n<p>I did the mean and stdev of each image.  And then the mean of all image means (per RGB channel), and the mean of all images stdev.\nBut my results were slightly different</p>",
          "rawMarkdown": "Thanks @rinnqd how did you calculate the image stats?\n\nI did the mean and stdev of each image.  And then the mean of all image means (per RGB channel), and the mean of all images stdev.\nBut my results were slightly different"
        },
        {
          "id": 580717,
          "postDate": "2019-07-20T16:08:18.127Z",
          "content": "<p>I think he is refraining to standard image net stats. If you are using pretreated model its  better to use image stats that were used to train the model = )</p>",
          "rawMarkdown": "I think he is refraining to standard image net stats. If you are using pretreated model its  better to use image stats that were used to train the model = )"
        }
      ]
    },
    {
      "id": 572469,
      "postDate": "2019-07-11T01:25:41.280Z",
      "content": "<p>Very good, thanks for sharing, this post gave me some ideas.</p>",
      "rawMarkdown": "Very good, thanks for sharing, this post gave me some ideas.",
      "votes": 1,
      "replies": [
        {
          "id": 572790,
          "postDate": "2019-07-11T12:07:21.217Z",
          "content": "<p>Hope it helps!</p>",
          "rawMarkdown": "Hope it helps!",
          "votes": 2
        }
      ]
    },
    {
      "id": 572386,
      "postDate": "2019-07-10T21:18:49.263Z",
      "content": "<p>Thanks for sharing ! \nAre you considering it as a classification problem?</p>",
      "rawMarkdown": "Thanks for sharing ! \nAre you considering it as a classification problem?",
      "votes": 1,
      "replies": [
        {
          "id": 572388,
          "postDate": "2019-07-10T21:22:56.910Z",
          "content": "<p>You are welcome! \nMy best results I got when I treated it as a regression problem! </p>",
          "rawMarkdown": "You are welcome! \nMy best results I got when I treated it as a regression problem! ",
          "votes": 2
        }
      ]
    },
    {
      "id": 572342,
      "postDate": "2019-07-10T19:39:38.863Z",
      "content": "<p>Regarding diffdifferences between training and test distribution, I summarized some of those concerns in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361\">this post</a></p>",
      "rawMarkdown": "Regarding diffdifferences between training and test distribution, I summarized some of those concerns in [this post](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361)",
      "votes": 1,
      "replies": [
        {
          "id": 572793,
          "postDate": "2019-07-11T12:08:57.753Z",
          "content": "<p>Thanks for doing great work!</p>",
          "rawMarkdown": "Thanks for doing great work!"
        }
      ]
    },
    {
      "id": 572210,
      "postDate": "2019-07-10T16:08:58.093Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1,
      "replies": [
        {
          "id": 572284,
          "postDate": "2019-07-10T18:11:44.223Z",
          "content": "<p>you are welcome! </p>",
          "rawMarkdown": "you are welcome! ",
          "votes": 1
        }
      ]
    },
    {
      "id": 572192,
      "postDate": "2019-07-10T15:44:51.560Z",
      "content": "<p>Thanks for sharing, you've tried a lot of things. I saw a gap between CV and LB also for all the models which I tried by now, not sure how to reduce (besides the things you've already mentioned).</p>",
      "rawMarkdown": "Thanks for sharing, you've tried a lot of things. I saw a gap between CV and LB also for all the models which I tried by now, not sure how to reduce (besides the things you've already mentioned).",
      "votes": 1,
      "replies": [
        {
          "id": 572232,
          "postDate": "2019-07-10T16:47:47.287Z",
          "content": "<p>Generalisation is a big issue. This might help to judge <a href=\"http://bit.ly/2YGRmK5\">http://bit.ly/2YGRmK5</a></p>",
          "rawMarkdown": "Generalisation is a big issue. This might help to judge http://bit.ly/2YGRmK5"
        }
      ]
    },
    {
      "id": 577813,
      "postDate": "2019-07-17T05:43:33.620Z",
      "content": "<p>Nice kernel</p>",
      "rawMarkdown": "Nice kernel",
      "votes": 2
    },
    {
      "id": 573580,
      "postDate": "2019-07-12T12:45:52.080Z",
      "content": "<p>try autoAI, you do not need to change hyperparms on your own now</p>",
      "rawMarkdown": "try autoAI, you do not need to change hyperparms on your own now",
      "votes": 2
    },
    {
      "id": 573346,
      "postDate": "2019-07-12T06:36:01.097Z",
      "content": "<p>Hello, thanks for sharing, are you using only data from this competition?</p>",
      "rawMarkdown": "Hello, thanks for sharing, are you using only data from this competition?",
      "votes": 2,
      "replies": [
        {
          "id": 573572,
          "postDate": "2019-07-12T12:30:31.543Z",
          "content": "<p>yes, and some data from old competition as well. </p>",
          "rawMarkdown": "yes, and some data from old competition as well. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 572732,
      "postDate": "2019-07-11T10:41:02.333Z",
      "content": "<p>Thanks!\nCan you share preprocessing method? Ben's or another method?? </p>",
      "rawMarkdown": "Thanks!\nCan you share preprocessing method? Ben's or another method?? ",
      "votes": 2,
      "replies": [
        {
          "id": 572795,
          "postDate": "2019-07-11T12:11:09.470Z",
          "content": "<p>I haven't tried yet Ben's Method. I don't do anything unusual just making sure that everything is in center and remove bad quality images. </p>",
          "rawMarkdown": "I haven't tried yet Ben's Method. I don't do anything unusual just making sure that everything is in center and remove bad quality images. "
        },
        {
          "id": 572952,
          "postDate": "2019-07-11T15:40:29.740Z",
          "content": "<p>Hey <a href=\"/drhabib\">@drhabib</a>, what's your definition for \"bad quality images\" as some images which may be bad for us may not be bad for the model, right?</p>",
          "rawMarkdown": "Hey @drhabib, what's your definition for \"bad quality images\" as some images which may be bad for us may not be bad for the model, right?",
          "votes": 1
        },
        {
          "id": 577249,
          "postDate": "2019-07-16T14:35:47.393Z",
          "content": "<p>Here is an example ...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fc1c817e30335cf545a64eca19bef85c1%2F12905_right.jpeg?generation=1563287697698109&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F002bf8af7d9afb9dc5ad1745b1d833f2%2F12905_left.jpeg?generation=1563287697637071&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Here is an example ...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fc1c817e30335cf545a64eca19bef85c1%2F12905_right.jpeg?generation=1563287697698109&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F002bf8af7d9afb9dc5ad1745b1d833f2%2F12905_left.jpeg?generation=1563287697637071&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 577492,
          "postDate": "2019-07-16T18:18:08.977Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> ,You remove them by manually inspection or do you use some kind of script?</p>",
          "rawMarkdown": "@drhabib ,You remove them by manually inspection or do you use some kind of script?"
        },
        {
          "id": 577498,
          "postDate": "2019-07-16T18:21:51.200Z",
          "content": "<p>good old manual inspection =)</p>",
          "rawMarkdown": "good old manual inspection =)",
          "votes": 4
        },
        {
          "id": 580333,
          "postDate": "2019-07-20T02:31:05.153Z",
          "content": "<p>wondering how you identify that they are bad</p>",
          "rawMarkdown": "wondering how you identify that they are bad"
        },
        {
          "id": 580719,
          "postDate": "2019-07-20T16:09:25.560Z",
          "content": "<p>Nothing special.. just visual inspection =)</p>",
          "rawMarkdown": "Nothing special.. just visual inspection =)\n"
        }
      ]
    },
    {
      "id": 572763,
      "postDate": "2019-07-11T11:31:27.357Z",
      "content": "<p>Why so many various libraries. . .</p>",
      "rawMarkdown": "Why so many various libraries. . .",
      "votes": -2
    },
    {
      "id": 591489,
      "postDate": "2019-08-03T19:41:45.700Z",
      "content": "<p><a href=\"/drhabib\">@drhabib</a>, can you share your data balancing approach.</p>",
      "rawMarkdown": "@drhabib, can you share your data balancing approach."
    },
    {
      "id": 589563,
      "postDate": "2019-08-01T05:16:47.740Z",
      "content": "<p>Hi , thank you for mentioning your various approaches. I want to ask whether you have used multilabels or just normal multiclass classification ? </p>",
      "rawMarkdown": "Hi , thank you for mentioning your various approaches. I want to ask whether you have used multilabels or just normal multiclass classification ? "
    },
    {
      "id": 589527,
      "postDate": "2019-08-01T03:24:34.533Z",
      "content": "<p>Thanks for your useful tips!\nIn experiment 5, you trained the model as a classification to apply label smoothing and then finetuned it as a regression problem. I thought the weights must be very different between the two.  Can you clarify the way you transfer a classification to regression?</p>\n\n<p>Do you have any idea to apply label smoothing directly to regression models for ordinary classification like this one?</p>",
      "rawMarkdown": "Thanks for your useful tips!\nIn experiment 5, you trained the model as a classification to apply label smoothing and then finetuned it as a regression problem. I thought the weights must be very different between the two.  Can you clarify the way you transfer a classification to regression?\n\nDo you have any idea to apply label smoothing directly to regression models for ordinary classification like this one?"
    },
    {
      "id": 584147,
      "postDate": "2019-07-25T13:38:56.880Z",
      "content": "<p>very informational post</p>",
      "rawMarkdown": "very informational post"
    },
    {
      "id": 582913,
      "postDate": "2019-07-23T19:10:51.197Z",
      "content": "<p>Thanks for sharing! Sharing knowledge is the best we can do!</p>",
      "rawMarkdown": "Thanks for sharing! Sharing knowledge is the best we can do!"
    },
    {
      "id": 579882,
      "postDate": "2019-07-19T11:10:32.193Z",
      "content": "<p>good to see many results of experiments. thanks! hope it saves time :)</p>",
      "rawMarkdown": "good to see many results of experiments. thanks! hope it saves time :)"
    },
    {
      "id": 579739,
      "postDate": "2019-07-19T06:50:13.317Z",
      "content": "<p>Thanks for sharing. 2 additions from my end, as per my experimentation: \n- he-uniform initialization didn't play as well with adam, as it does with SGD with Nesterov momentum\n- In my experimentations, using adam with concatenated globalmaxpool and globalAvgPool, preforms poorly as compared to using only globalmaxpool.</p>\n\n<p>Also wondering, if we should this thread as official thread to gather all tips and tricks at one place?</p>",
      "rawMarkdown": "Thanks for sharing. 2 additions from my end, as per my experimentation: \n- he-uniform initialization didn't play as well with adam, as it does with SGD with Nesterov momentum\n- In my experimentations, using adam with concatenated globalmaxpool and globalAvgPool, preforms poorly as compared to using only globalmaxpool.\n\nAlso wondering, if we should this thread as official thread to gather all tips and tricks at one place?"
    },
    {
      "id": 579603,
      "postDate": "2019-07-19T02:41:31.190Z",
      "content": "<p>Thank you for sharing. It was very helpful.</p>",
      "rawMarkdown": "Thank you for sharing. It was very helpful."
    },
    {
      "id": 579356,
      "postDate": "2019-07-18T19:04:47.913Z",
      "content": "<p>Thanks <a href=\"/drhabib\">@drhabib</a>!</p>\n\n<p>In the experiment-03, how did you increase the image size of a resnet50.</p>\n\n<p>In a Resnet, the head input (feature map) is proportional to the input image size.</p>\n\n<p>Did you retrain a new head on each change of image size?</p>\n\n<p>Instead of that, did you use a gapnet-resnet50?  Or just global average pooling for the last layer?</p>",
      "rawMarkdown": "Thanks @drhabib!\n\nIn the experiment-03, how did you increase the image size of a resnet50.\n\nIn a Resnet, the head input (feature map) is proportional to the input image size.\n\nDid you retrain a new head on each change of image size?\n\nInstead of that, did you use a gapnet-resnet50?  Or just global average pooling for the last layer?",
      "replies": [
        {
          "id": 580715,
          "postDate": "2019-07-20T16:07:17.343Z",
          "content": "<p>Hi=)\nI just change the image size without changing model head. Pytorch by default uses Adaptive layers in there models. </p>\n\n<p>-First I train on small image size \n-Freeze the model except last layer\n-train the model with large image size\n-Unfreeze the last layer and fine tune for few more epochs. </p>",
          "rawMarkdown": "Hi=)\nI just change the image size without changing model head. Pytorch by default uses Adaptive layers in there models. \n\n-First I train on small image size \n-Freeze the model except last layer\n-train the model with large image size\n-Unfreeze the last layer and fine tune for few more epochs. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 579305,
      "postDate": "2019-07-18T17:55:29.390Z",
      "content": "<p>Good job</p>",
      "rawMarkdown": "Good job"
    },
    {
      "id": 578893,
      "postDate": "2019-07-18T09:06:09.603Z",
      "content": "<p>Great work!\nIt's very helpful for me:)</p>",
      "rawMarkdown": "Great work!\nIt's very helpful for me:)"
    },
    {
      "id": 578250,
      "postDate": "2019-07-17T14:16:04.110Z",
      "content": "<p>Great!</p>",
      "rawMarkdown": "Great!"
    },
    {
      "id": 577970,
      "postDate": "2019-07-17T08:39:04.827Z",
      "content": "<p>very helpful</p>",
      "rawMarkdown": "very helpful"
    },
    {
      "id": 577893,
      "postDate": "2019-07-17T07:10:35.303Z",
      "content": "<p>Very helpful tips!! Thankyou</p>",
      "rawMarkdown": "Very helpful tips!! Thankyou"
    },
    {
      "id": 577772,
      "postDate": "2019-07-17T04:00:39.290Z",
      "content": "<p>GOOD WORK</p>",
      "rawMarkdown": "GOOD WORK"
    },
    {
      "id": 577573,
      "postDate": "2019-07-16T19:52:17.907Z",
      "content": "<p>Thank you for sharing. It is very helpful for us.</p>",
      "rawMarkdown": "Thank you for sharing. It is very helpful for us."
    },
    {
      "id": 577119,
      "postDate": "2019-07-16T11:40:17.630Z",
      "content": "<p>Thnaks!\nVery helpful</p>",
      "rawMarkdown": "Thnaks!\nVery helpful"
    },
    {
      "id": 575500,
      "postDate": "2019-07-15T13:58:01.667Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 573766,
      "postDate": "2019-07-12T18:03:01.030Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 574407,
      "postDate": "2019-07-13T19:23:36.830Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 3
    },
    {
      "id": 575522,
      "postDate": "2019-07-15T14:42:20.573Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 575331,
      "postDate": "2019-07-15T09:43:23.847Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1
    },
    {
      "id": 575251,
      "postDate": "2019-07-15T08:09:11.040Z",
      "content": "<p>Thank you. It is very helpful</p>",
      "rawMarkdown": "Thank you. It is very helpful",
      "votes": 1
    },
    {
      "id": 575224,
      "postDate": "2019-07-15T07:29:01.367Z",
      "content": "<p>thank you for sharing</p>",
      "rawMarkdown": "thank you for sharing",
      "votes": 1
    },
    {
      "id": 574841,
      "postDate": "2019-07-14T14:52:10.930Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": 1
    },
    {
      "id": 574840,
      "postDate": "2019-07-14T14:51:48.723Z",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": 1
    },
    {
      "id": 574035,
      "postDate": "2019-07-13T06:20:38.217Z",
      "content": "<p>Thanks! Good job.</p>",
      "rawMarkdown": "Thanks! Good job.",
      "votes": 1
    },
    {
      "id": 574617,
      "postDate": "2019-07-14T07:24:03.213Z",
      "content": "<p>Thank You!</p>",
      "rawMarkdown": "Thank You!"
    },
    {
      "id": 2053467,
      "postDate": "2022-12-03T09:59:31.483Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 588907,
      "postDate": "2019-07-31T08:05:46.947Z",
      "content": "<p>Thanks for your informative post!!!</p>",
      "rawMarkdown": "Thanks for your informative post!!!"
    },
    {
      "id": 585774,
      "postDate": "2019-07-28T03:55:03.097Z",
      "content": "<p>Thanks! Interesting read.</p>",
      "rawMarkdown": "Thanks! Interesting read."
    },
    {
      "id": 580951,
      "postDate": "2019-07-21T06:05:44.030Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 580237,
      "postDate": "2019-07-19T21:50:37.213Z",
      "content": "<p>Thanks </p>",
      "rawMarkdown": "Thanks "
    },
    {
      "id": 579525,
      "postDate": "2019-07-18T23:32:26.883Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!"
    },
    {
      "id": 579357,
      "postDate": "2019-07-18T19:06:21.703Z",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks"
    },
    {
      "id": 579165,
      "postDate": "2019-07-18T14:57:27.287Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 579103,
      "postDate": "2019-07-18T13:54:56.707Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    },
    {
      "id": 578971,
      "postDate": "2019-07-18T11:07:46.133Z",
      "content": "<p>Thanks! </p>",
      "rawMarkdown": "Thanks! "
    },
    {
      "id": 578191,
      "postDate": "2019-07-17T13:00:20.557Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!"
    },
    {
      "id": 578084,
      "postDate": "2019-07-17T10:35:34.263Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 577018,
      "postDate": "2019-07-16T08:44:57.497Z",
      "content": "<p>Thanks for sharing this.</p>",
      "rawMarkdown": "Thanks for sharing this."
    },
    {
      "id": 577008,
      "postDate": "2019-07-16T08:38:13.827Z",
      "content": "<p>Thank you.It is helpful</p>",
      "rawMarkdown": "Thank you.It is helpful"
    },
    {
      "id": 576972,
      "postDate": "2019-07-16T08:06:18.550Z",
      "content": "<p>Thanks!!</p>",
      "rawMarkdown": "Thanks!!"
    },
    {
      "id": 576810,
      "postDate": "2019-07-16T03:39:20.027Z",
      "content": "<p>Awesome! Thanks</p>",
      "rawMarkdown": "Awesome! Thanks"
    },
    {
      "id": 576809,
      "postDate": "2019-07-16T03:33:19.843Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 578182,
      "postDate": "2019-07-17T12:49:24.563Z",
      "content": "<p>Thank you for sharing !</p>",
      "rawMarkdown": "Thank you for sharing !",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 572236,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-07-10T16:54:56.123000",
      "content": "<ol>\n<li>The difference between train set and public test set target distribution is large. 2. The target distribution of public test set is very different with 2015 competition dataset too. 3. The difference between target distribution of train set and 2015 competition data is not so large, one has (73% level 0, 15% level 2) and one has (50% level 0, 25% level 2). There are too many 2s for public test set based on my submission (50% level 2, 21% level 0).  I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.</li>\n</ol>",
      "votes": 11,
      "replies": [
        {
          "id": 572283,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-10T18:09:38.483000",
          "content": "<p>Thanks for your insights. I agree one should be really careful using optimized thresholds with regressions. I haven't done experiments yet with x-folds CV but I assume simple averaging across the folds will be better (maybe safer) than optimized kappa over folds... </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 572959,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-11T15:44:56.517000",
          "content": "<p>&gt; I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.</p>\n\n<p>Exactly, I'm using threshold optimised at the validation set and I have this feeling that my validation set doesn't really represent the test set out there that's the reason I'm getting bad results.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 573231,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-07-12T02:05:25.177000",
      "content": "<p>Just a small fix, the label smoothing link is incorrect, you should remove the \")\" at the end of the hyperlink ; )</p>\n\n<p>like this: <a href=\"https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0\">https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 573569,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-12T12:27:57.160000",
          "content": "<p>thank you very much!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 579437,
      "author_name": "Kartik Godawat",
      "author_url": "",
      "post_date": "2019-07-18T20:38:51.017000",
      "content": "<p>Thanks for sharing. My base experiment setup is similar. Although I am finding slight increase in both my local cv and lb with progressive resizing to 128-&gt;256-&gt;299 for resnet50.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 580720,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:10:29.250000",
          "content": "<p>I can confirm, after playing around a little bit image size increase leads to better lb score =)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 573979,
      "author_name": "lsimoneau",
      "author_url": "",
      "post_date": "2019-07-13T04:23:24.010000",
      "content": "<p>Thanks for the post. Do you think your LB scores are mainly due to careful image transformations in your training set? I've tried training resnet50 for a similar number of epochs, I get over 0.9 kappa in CV but LB scores are around 0.65, and I'm not sure what I should be doing differently. Would you suggest focusing on hyperparameters, image transformations, or expanding the training set?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 574054,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-13T07:02:11.020000",
          "content": "<p>Hey, Generalisation is a big issue in this competition. The LB is decided by a private test set which is ~10x the public test set. People are using optimized kappa but that would be useless if we don't have a good validation set (good: which is as close to test distribution as possible). As far as expanding the training set (using the previous year's data) is concerned I haven't had good results so far, I'm still experimenting with that, fingers crossed.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 594922,
      "author_name": "Verne",
      "author_url": "",
      "post_date": "2019-08-08T16:26:32.007000",
      "content": "<p>Very well documented. I like the research quoted. \nI tried the things you mentioned above (autoencoders, gan, various architectures and image sizes) and only now I found your report. \nI can confirm your experiments. I tried a few extra things such as image filtering including Ben's processing. Did not seem to help. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 594938,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-08-08T17:01:39.527000",
          "content": "<p>it always good to hear when people can reproduce your results! Good luck in this competition =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 594979,
          "author_name": "Verne",
          "author_url": "",
          "post_date": "2019-08-08T17:41:29.587000",
          "content": "<p>Thanks, to you, too! Looks like we all need some luck in this competition.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 608603,
          "author_name": "Rohit Midha",
          "author_url": "",
          "post_date": "2019-08-27T02:37:41.787000",
          "content": "<p>Hi, a little new to this but is using GANs to create images with lesser classes a common thing to do?\nFirst time I've ever heard of it, hence the question.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 584495,
      "author_name": "Yohanes Alfredo",
      "author_url": "",
      "post_date": "2019-07-26T04:42:30.303000",
      "content": "<p>When using multiple models that have similar input shape, do the most important preprocessing (cropping and resize) beforehand and load it to disk. Now you are wondering how to fit all of this on the disk. </p>\n\n<ol>\n<li>First, you need to ready your model (do not perform training on the same kernel)</li>\n<li>Separate inferences to several batches. </li>\n</ol>\n\n<p>Basically, the workflow on each batch will be :\n1. Crop and resize (you can further speed this up with multiprocessing) and save to disk\n2. Perform inference, you can use TTA if you want. TTA should be quick enough since you are not performing resizing.\n3. After finish one batch clear these images and repeat with next batch.</p>\n\n<p>Since most of the images are big enough, it takes quite a lot of time doing preprocessing and with gpu on you will not have that much computing power in a quick fashion, hence I think my tricks should help.\nThis could work if your model is not too big and you can load all your ensemble model in one-go</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 579092,
      "author_name": "Nabeelma",
      "author_url": "",
      "post_date": "2019-07-18T13:33:59.723000",
      "content": "<p>Great.. thanks for sharing .. so in general optimized_kappa gave better results in LB </p>",
      "votes": 1,
      "replies": [
        {
          "id": 580721,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:11:23.443000",
          "content": "<p>Its very risky to use optimized kappa, unless you are very confident that your valid data matches test distribution =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 586368,
          "author_name": "Quan",
          "author_url": "",
          "post_date": "2019-07-29T04:25:10.053000",
          "content": "<p>I agree with DrHB. Since qwk is closely related to mean_squared_error, i think optimized mse is better alternative to qwk.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 602367,
          "author_name": "Antonio De Perio",
          "author_url": "",
          "post_date": "2019-08-19T02:11:58.650000",
          "content": "<p>Would optimised kappa help with smoothing out the effects of imbalanced classes? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 578238,
      "author_name": "Bing Wu",
      "author_url": "",
      "post_date": "2019-07-17T13:59:59.870000",
      "content": "<p>Thanks, DrHB for your kind sharing. I noticed that you did many experiments. May I know how long it would take you to approach one model on average? I am asking this because it takes me hours to run one experiment with my local machine(MacBook pro i5 8GB). I am thinking to rent an AWS server to do the experiments faster.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 578257,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-17T14:22:41.447000",
          "content": "<p>AWS is costly. Use vast.ai </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 580785,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-20T18:48:30.170000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 580860,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T23:07:11.317000",
          "content": "<p>Gcloud is also a good option. If you register you get 300 dollars to use towards computing =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581130,
          "author_name": "Bing Wu",
          "author_url": "",
          "post_date": "2019-07-21T13:30:55.347000",
          "content": "<p>Thanks for sharing. I have a concern with cloud though. If I use a very powerful machine in the cloud and it takes me 4 hrs to run. But it would take 10 hrs to run in kaggle kernel. The limitation of the kaggle kernel is only 9 hrs. Any thoughts on this situation?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581134,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-21T13:45:31.307000",
          "content": "<p><a href=\"/bingwu0\">@bingwu0</a> You can train your models in cloud, download trained weights, write a inference kernel to use those weights and predict on the dataset. You don't have to train your models while you submit your predictions, just use the trained weights. I hope you get it.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 581308,
          "author_name": "Bing Wu",
          "author_url": "",
          "post_date": "2019-07-21T18:32:10.213000",
          "content": "<p>Yes,  I see your point. Thanks for your sharing!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 575076,
      "author_name": "ForesightKing",
      "author_url": "",
      "post_date": "2019-07-15T02:51:01.407000",
      "content": "<p>Please try use ResNet-101 or ResNet-152 if your device accept this model</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 575066,
      "author_name": "Praveen Kumar Ulasa",
      "author_url": "",
      "post_date": "2019-07-15T02:17:33.687000",
      "content": "<p>Good</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574842,
      "author_name": "AG07",
      "author_url": "",
      "post_date": "2019-07-14T14:52:31.240000",
      "content": "<p>Interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574819,
      "author_name": "dexter6855",
      "author_url": "",
      "post_date": "2019-07-14T14:22:17.677000",
      "content": "<p>Thanks a lot for the information. I had the similar concern. My model is getting better in CV but got worse in LB. Made me confuse about the performance of my model.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574611,
      "author_name": "s1mple",
      "author_url": "",
      "post_date": "2019-07-14T07:09:56.373000",
      "content": "<p>nice tips for me, thx.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574430,
      "author_name": "Sergio Filatov",
      "author_url": "",
      "post_date": "2019-07-13T19:55:53.533000",
      "content": "<p>interesting! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574338,
      "author_name": "Jonathan",
      "author_url": "",
      "post_date": "2019-07-13T16:38:29.290000",
      "content": "<p>Cool! People usually only show what works and ignore posting how they got there.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 573445,
      "author_name": "Darshan Chheda",
      "author_url": "",
      "post_date": "2019-07-12T09:40:17.033000",
      "content": "<p>Interesting</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 572522,
      "author_name": "Himanshu Gamit",
      "author_url": "",
      "post_date": "2019-07-11T03:13:12.560000",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> . Thanks for sharing. Have you tried x-fold cross-validation (no. of folds) yet? How about split train samples into Train-Valid-Test(0.1)?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572791,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:08:19.200000",
          "content": "<p>I haven't tried cross fold validation. Test-Valid (0.1) I feel it will lead to overfitting.. but I might be wrong </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 599231,
          "author_name": "Dylan Schoenmakers",
          "author_url": "",
          "post_date": "2019-08-14T17:32:14.263000",
          "content": "<p>Wouldn't dropouts help with that?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 572490,
      "author_name": "哈尔的移动城堡",
      "author_url": "",
      "post_date": "2019-07-11T01:52:35.153000",
      "content": "<p>Hi,man,thanks  for sharing.\nDo  you  train   with  an  external  data    or  you  just  train  in   the  official  dataset   to   get   13th?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572792,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:08:41.153000",
          "content": "<p>I have combination of both. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 572482,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "2019-07-11T01:42:37.577000",
      "content": "<p>So what's your image transformations, guys ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572491,
          "author_name": "Firas Baba",
          "author_url": "",
          "post_date": "2019-07-11T02:02:04.417000",
          "content": "<p>```\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]</p>\n\n<p>train_transform = transforms.Compose([\n    transforms.Resize((IMAGE_SIZE, IMAGE_SIZE)),\n    transforms.RandomHorizontalFlip(),\n    transforms.RandomVerticalFlip(),\n    transforms.ToTensor(),\n    transforms.Normalize(IMG_MEAN, IMG_STD)\n])\n```</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 572663,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2019-07-11T08:34:29.323000",
          "content": "<p><code>\nIMG_MEAN = [0.485, 0.456, 0.406]\nIMG_STD = [0.229, 0.224, 0.225]\n</code>\nShould we be using Image-Net stats here as well?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 572789,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:07:06.650000",
          "content": "<p>yes, at least in my case. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 579345,
          "author_name": "Virilo Tejedor Aguilera",
          "author_url": "",
          "post_date": "2019-07-18T18:48:39.447000",
          "content": "<p>Thanks <a href=\"/rinnqd\">@rinnqd</a> how did you calculate the image stats?</p>\n\n<p>I did the mean and stdev of each image.  And then the mean of all image means (per RGB channel), and the mean of all images stdev.\nBut my results were slightly different</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 580717,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:08:18.127000",
          "content": "<p>I think he is refraining to standard image net stats. If you are using pretreated model its  better to use image stats that were used to train the model = )</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 572469,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-07-11T01:25:41.280000",
      "content": "<p>Very good, thanks for sharing, this post gave me some ideas.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572790,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:07:21.217000",
          "content": "<p>Hope it helps!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 572386,
      "author_name": "Firas Baba",
      "author_url": "",
      "post_date": "2019-07-10T21:18:49.263000",
      "content": "<p>Thanks for sharing ! \nAre you considering it as a classification problem?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572388,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-10T21:22:56.910000",
          "content": "<p>You are welcome! \nMy best results I got when I treated it as a regression problem! </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 572342,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2019-07-10T19:39:38.863000",
      "content": "<p>Regarding diffdifferences between training and test distribution, I summarized some of those concerns in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361\">this post</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 572793,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:08:57.753000",
          "content": "<p>Thanks for doing great work!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 572210,
      "author_name": "makogarei",
      "author_url": "",
      "post_date": "2019-07-10T16:08:58.093000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572284,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-10T18:11:44.223000",
          "content": "<p>you are welcome! </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 572192,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2019-07-10T15:44:51.560000",
      "content": "<p>Thanks for sharing, you've tried a lot of things. I saw a gap between CV and LB also for all the models which I tried by now, not sure how to reduce (besides the things you've already mentioned).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 572232,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-10T16:47:47.287000",
          "content": "<p>Generalisation is a big issue. This might help to judge <a href=\"http://bit.ly/2YGRmK5\">http://bit.ly/2YGRmK5</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 577813,
      "author_name": "Ekrem Bayar",
      "author_url": "",
      "post_date": "2019-07-17T05:43:33.620000",
      "content": "<p>Nice kernel</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 573580,
      "author_name": "Vivek Sharma",
      "author_url": "",
      "post_date": "2019-07-12T12:45:52.080000",
      "content": "<p>try autoAI, you do not need to change hyperparms on your own now</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 573346,
      "author_name": "wayfarer",
      "author_url": "",
      "post_date": "2019-07-12T06:36:01.097000",
      "content": "<p>Hello, thanks for sharing, are you using only data from this competition?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 573572,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-12T12:30:31.543000",
          "content": "<p>yes, and some data from old competition as well. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 572732,
      "author_name": "SeongwoongCho",
      "author_url": "",
      "post_date": "2019-07-11T10:41:02.333000",
      "content": "<p>Thanks!\nCan you share preprocessing method? Ben's or another method?? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 572795,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-11T12:11:09.470000",
          "content": "<p>I haven't tried yet Ben's Method. I don't do anything unusual just making sure that everything is in center and remove bad quality images. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 572952,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-11T15:40:29.740000",
          "content": "<p>Hey <a href=\"/drhabib\">@drhabib</a>, what's your definition for \"bad quality images\" as some images which may be bad for us may not be bad for the model, right?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 577249,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-16T14:35:47.393000",
          "content": "<p>Here is an example ...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fc1c817e30335cf545a64eca19bef85c1%2F12905_right.jpeg?generation=1563287697698109&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F002bf8af7d9afb9dc5ad1745b1d833f2%2F12905_left.jpeg?generation=1563287697637071&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 577492,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-07-16T18:18:08.977000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> ,You remove them by manually inspection or do you use some kind of script?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 577498,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-16T18:21:51.200000",
          "content": "<p>good old manual inspection =)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 580333,
          "author_name": "Vishnu Subramanian",
          "author_url": "",
          "post_date": "2019-07-20T02:31:05.153000",
          "content": "<p>wondering how you identify that they are bad</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 580719,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-20T16:09:25.560000",
          "content": "<p>Nothing special.. just visual inspection =)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 572763,
      "author_name": "Santhosh M Kunthe",
      "author_url": "",
      "post_date": "2019-07-11T11:31:27.357000",
      "content": "<p>Why so many various libraries. . .</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 591489,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-03T19:41:45.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 589563,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-01T05:16:47.740000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 589527,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-01T03:24:34.533000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 584147,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-25T13:38:56.880000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 582913,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-23T19:10:51.197000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579882,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-19T11:10:32.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579739,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-19T06:50:13.317000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579603,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-19T02:41:31.190000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579356,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T19:04:47.913000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 580715,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-07-20T16:07:17.343000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 579305,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T17:55:29.390000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578893,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T09:06:09.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578250,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T14:16:04.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577970,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T08:39:04.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577893,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T07:10:35.303000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577772,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T04:00:39.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577573,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T19:52:17.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577119,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T11:40:17.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 575500,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-15T13:58:01.667000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 573766,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-12T18:03:01.030000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574407,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-13T19:23:36.830000",
      "content": "",
      "votes": 3,
      "replies": []
    },
    {
      "id": 575522,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-15T14:42:20.573000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 575331,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-15T09:43:23.847000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 575251,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-15T08:09:11.040000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 575224,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-15T07:29:01.367000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574841,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-14T14:52:10.930000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574840,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-14T14:51:48.723000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574035,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-13T06:20:38.217000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 574617,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-14T07:24:03.213000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2053467,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-12-03T09:59:31.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 588907,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-31T08:05:46.947000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 585774,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-28T03:55:03.097000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 580951,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-21T06:05:44.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 580237,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-19T21:50:37.213000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579525,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T23:32:26.883000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579357,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T19:06:21.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579165,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T14:57:27.287000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 579103,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T13:54:56.707000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578971,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-18T11:07:46.133000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578191,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T13:00:20.557000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578084,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T10:35:34.263000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577018,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T08:44:57.497000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 577008,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T08:38:13.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 576972,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T08:06:18.550000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 576810,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T03:39:20.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 576809,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-16T03:33:19.843000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 578182,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-17T12:49:24.563000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "572186": "This is a post about some experiments I performed which did not help me so much, but eventually leaded to my current position in LB. \n\nMy project setup:\n```\nIMAGE SIZE = 256\nVALIDATION = 20%\nno TTA\nno x-fold CV\n```\n\nEXPERIMENT 01:\n\n```\nepoch: 10\nmodel: resnet50 (pretrained = True)\nCV: 0.932291026\nLB: 0.755\n\n```\nEXPERIMENT 02:\nTrained as a classification problem, and then retrained as a regression problem. \n\n```\nmodel: resnet50 (pretrained = True)\nepoch: 10\nCV: 0.930\nLB: 0.750\n```\n\nEXPERIMENT 03:\nIdea came from the old competition where the winners could achieve good score with progressively resizing image size\n\nround 1: trained on 128 (epoch: 10)\nround 2: finetuned on 256 (epoch: 5)\nround 3: finetuned on 512 (epoch: 10)\n\n```\nmodel: resnet50 (pretrained = True)\nCV: 0.91938504\nLB: 0.734\n```\n\n\nEXPERIMENT 04:\nThanks to the @ilovescience for the dataset, I tried to balance all classes and trained the network. \n\n```\nepoch: 15\nCV: 0.933592607\nLB:  0.751\nmodel: resnet50 (pretrained = True)\n\n```\nEXPERIMENT 05:\nSame as experiment 4 but I have added label Smoothing \nhttps://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0 (thanks to @dimitreoliveira).\n\nround 1: first trained as a classification problem (so I can use label smoothing)\nround 2: finetuned as a regression problem\n\n```\nmodel: resnet50 (pretrained = True)\nepoch: 29\nCV: 0.927\nLB:  0.0.732\n\n```\n\nOther experiments that failed:\n-Using Auto-encoders to generate images \n-Using Auto -encoders latent layes as input to NN\n-GANs to generate images of the classes that have low examples\n-Diffrent types of preprocessing (gray, edges, different filters)\n\n\nTips: \n-please pay attentions to image transformations\n-if you have resized images on local computer make sure you do same resizing when you use Kaggle kernel for submission \n-in the beginning of competition don't waste your time doing x fold CVs, it’s better to focus on building one robust model with same validation set across different experimenting\n\nConcerns:\n-I am still concerned about the gap between CV and LB.\n-People have pointed out that there is a lot of noise in labels.. I haven’t tried but one way to solve this issue would be to use smartly label smoothening in regression.  \n\nEDIT 1:\ninteresting research from google AI talking about how to judge if your NN is generalizing https://ai.googleblog.com/2019/07/predicting-generalization-gap-in-deep.html\n\nEDIT 2:\nTable of all submissions same validation set across experiments (without classification). Hope it helps\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fdf51224302a7656a2e41efd20d849821%2FScreen%20Shot%202019-07-12%20at%202.28.45%20PM.png?generation=1562960483994331&amp;alt=media)\n\n\nEDIT 3:\nRecently there was paper that claimed that CNN's tend to learn texture information, but if we push training toward shapes it can improve accuracy. Its pretty easy to implement and I had some success in terms of training. For people who are interested here is a link https://openreview.net/forum?id=Bygh9j09KX and link to GitHub GitHub https://github.com/rgeirhos/texture-vs-shape",
    "572236": "1. The difference between train set and public test set target distribution is large. 2. The target distribution of public test set is very different with 2015 competition dataset too. 3. The difference between target distribution of train set and 2015 competition data is not so large, one has (73% level 0, 15% level 2) and one has (50% level 0, 25% level 2). There are too many 2s for public test set based on my submission (50% level 2, 21% level 0).  I'm using optimized threshold with regression, this is very dangerous since it highly depends on the data distribution.",
    "573231": "Just a small fix, the label smoothing link is incorrect, you should remove the \")\" at the end of the hyperlink ; )\n\nlike this: https://towardsdatascience.com/label-smoothing-making-model-robust-to-incorrect-labels-2fae037ffbd0",
    "579437": "Thanks for sharing. My base experiment setup is similar. Although I am finding slight increase in both my local cv and lb with progressive resizing to 128-&gt;256-&gt;299 for resnet50.",
    "573979": "Thanks for the post. Do you think your LB scores are mainly due to careful image transformations in your training set? I've tried training resnet50 for a similar number of epochs, I get over 0.9 kappa in CV but LB scores are around 0.65, and I'm not sure what I should be doing differently. Would you suggest focusing on hyperparameters, image transformations, or expanding the training set?",
    "594922": "Very well documented. I like the research quoted. \nI tried the things you mentioned above (autoencoders, gan, various architectures and image sizes) and only now I found your report. \nI can confirm your experiments. I tried a few extra things such as image filtering including Ben's processing. Did not seem to help. \n",
    "584495": "When using multiple models that have similar input shape, do the most important preprocessing (cropping and resize) beforehand and load it to disk. Now you are wondering how to fit all of this on the disk. \n\n1. First, you need to ready your model (do not perform training on the same kernel)\n2. Separate inferences to several batches. \n\nBasically, the workflow on each batch will be :\n1. Crop and resize (you can further speed this up with multiprocessing) and save to disk\n2. Perform inference, you can use TTA if you want. TTA should be quick enough since you are not performing resizing.\n3. After finish one batch clear these images and repeat with next batch.\n\nSince most of the images are big enough, it takes quite a lot of time doing preprocessing and with gpu on you will not have that much computing power in a quick fashion, hence I think my tricks should help.\nThis could work if your model is not too big and you can load all your ensemble model in one-go\n",
    "579092": "Great.. thanks for sharing .. so in general optimized_kappa gave better results in LB ",
    "578238": "Thanks, DrHB for your kind sharing. I noticed that you did many experiments. May I know how long it would take you to approach one model on average? I am asking this because it takes me hours to run one experiment with my local machine(MacBook pro i5 8GB). I am thinking to rent an AWS server to do the experiments faster.",
    "575076": "Please try use ResNet-101 or ResNet-152 if your device accept this model",
    "575066": "Good",
    "574842": "Interesting",
    "574819": "Thanks a lot for the information. I had the similar concern. My model is getting better in CV but got worse in LB. Made me confuse about the performance of my model.",
    "574611": "nice tips for me, thx.",
    "574430": "interesting! ",
    "574338": "Cool! People usually only show what works and ignore posting how they got there.",
    "573445": "Interesting\n",
    "572522": "@drhabib . Thanks for sharing. Have you tried x-fold cross-validation (no. of folds) yet? How about split train samples into Train-Valid-Test(0.1)?",
    "572490": "Hi,man,thanks  for sharing.\nDo  you  train   with  an  external  data    or  you  just  train  in   the  official  dataset   to   get   13th?",
    "572482": "So what's your image transformations, guys ?",
    "572469": "Very good, thanks for sharing, this post gave me some ideas.",
    "572386": "Thanks for sharing ! \nAre you considering it as a classification problem?",
    "572342": "Regarding diffdifferences between training and test distribution, I summarized some of those concerns in [this post](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361)",
    "572210": "Thanks for sharing!",
    "572192": "Thanks for sharing, you've tried a lot of things. I saw a gap between CV and LB also for all the models which I tried by now, not sure how to reduce (besides the things you've already mentioned).",
    "577813": "Nice kernel",
    "573580": "try autoAI, you do not need to change hyperparms on your own now",
    "573346": "Hello, thanks for sharing, are you using only data from this competition?",
    "572732": "Thanks!\nCan you share preprocessing method? Ben's or another method?? ",
    "572763": "Why so many various libraries. . .",
    "591489": "@drhabib, can you share your data balancing approach.",
    "589563": "Hi , thank you for mentioning your various approaches. I want to ask whether you have used multilabels or just normal multiclass classification ? ",
    "589527": "Thanks for your useful tips!\nIn experiment 5, you trained the model as a classification to apply label smoothing and then finetuned it as a regression problem. I thought the weights must be very different between the two.  Can you clarify the way you transfer a classification to regression?\n\nDo you have any idea to apply label smoothing directly to regression models for ordinary classification like this one?",
    "584147": "very informational post",
    "582913": "Thanks for sharing! Sharing knowledge is the best we can do!",
    "579882": "good to see many results of experiments. thanks! hope it saves time :)",
    "579739": "Thanks for sharing. 2 additions from my end, as per my experimentation: \n- he-uniform initialization didn't play as well with adam, as it does with SGD with Nesterov momentum\n- In my experimentations, using adam with concatenated globalmaxpool and globalAvgPool, preforms poorly as compared to using only globalmaxpool.\n\nAlso wondering, if we should this thread as official thread to gather all tips and tricks at one place?",
    "579603": "Thank you for sharing. It was very helpful.",
    "579356": "Thanks @drhabib!\n\nIn the experiment-03, how did you increase the image size of a resnet50.\n\nIn a Resnet, the head input (feature map) is proportional to the input image size.\n\nDid you retrain a new head on each change of image size?\n\nInstead of that, did you use a gapnet-resnet50?  Or just global average pooling for the last layer?",
    "579305": "Good job",
    "578893": "Great work!\nIt's very helpful for me:)",
    "578250": "Great!",
    "577970": "very helpful",
    "577893": "Very helpful tips!! Thankyou",
    "577772": "GOOD WORK",
    "577573": "Thank you for sharing. It is very helpful for us.",
    "577119": "Thnaks!\nVery helpful",
    "575500": "",
    "573766": "",
    "574407": "Thank you!",
    "575522": "Thanks!",
    "575331": "Thank you!",
    "575251": "Thank you. It is very helpful",
    "575224": "thank you for sharing",
    "574841": "Thank you!",
    "574840": "Thank you",
    "574035": "Thanks! Good job.",
    "574617": "Thank You!",
    "2053467": "Thank you!",
    "588907": "Thanks for your informative post!!!",
    "585774": "Thanks! Interesting read.",
    "580951": "Thank you!",
    "580237": "Thanks ",
    "579525": "Thanks!",
    "579357": "thanks",
    "579165": "Thanks",
    "579103": "Thank you!",
    "578971": "Thanks! ",
    "578191": "Thanks!",
    "578084": "Thanks",
    "577018": "Thanks for sharing this.",
    "577008": "Thank you.It is helpful",
    "576972": "Thanks!!",
    "576810": "Awesome! Thanks",
    "576809": "Thanks",
    "578182": "Thank you for sharing !"
  }
}