{
  "id": 108065,
  "title": "1st place solution summary",
  "url": "/competitions/aptos2019-blindness-detection/discussion/108065",
  "author_name": "Guanshuo Xu",
  "post_date": "2019-09-08T20:21:22.508000",
  "votes": 303,
  "comment_count": 115,
  "views": 0,
  "content": "<p>Thanks to APTOS and Kaggle for hosting this interesting competition. \nI also would like to thank those who generously contributed in the kernels and discussions in this competition, and to the top teams that shared solutions and findings in the 2015 competition. Those findings and solutions have greatly impacted my strategy.</p>\n\n<p><strong>Validation Strategy</strong></p>\n\n<p>One of the most popular topics in every competition are proper validation strategies. During the early stage, I tried using 2015 data (both train and test) as train set, and 2019 train data as validation set. Unfortunately, the validation results and public LB were very different, I was not able to make their performance correlate well. In some discussions and kernels, other participants were also reporting inconsistent performance between CV and LB. Because of this, I did not know how to move forward, so I shifted to some other competitions for a few weeks. When I came back to this competition, I made the decision to combine the whole 2015 and 2019 data as train set, and solely relied on public LB for validation.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>I don't think it's necessary to preprocess images to help with the modelling, the image qualities are perfect as input for deep neural networks. So, no special preprocessing, just plain resizing.</p>\n\n<p><strong>Models and Input sizes</strong></p>\n\n<p>My final submission was a simple average of the following eight models. Inceptions and ResNets usually blend well. If I could have two more weeks I would definitely add some EfficientNets. </p>\n\n<p><code>\n2 x inception_resnet_v2, input size 512\n2 x inception_v4, input size 512\n2 x seresnext50, input size 512\n2 x seresnext101, input size 384\n</code></p>\n\n<p>The input size was mainly determined by observations in the 2015 competition that larger input size brought better performance. Even though I did not find a lot beneficial to go beyond 384 based on the public LB feedback, I still push to the extreme the input size because the private set might benefit.</p>\n\n<p><strong>Loss, Augmentations, Pooling</strong></p>\n\n<p>I used only <em>nn.SmoothL1Loss()</em> as the loss function. Other loss functions may work well too. I sticked to this single loss just to simplify the emsembling process.</p>\n\n<p>For augmentations, the following were helpful</p>\n\n<p><code>\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n</code></p>\n\n<p>For the last pooling layer, I found the generalized mean pooling (<a href=\"https://arxiv.org/pdf/1711.02512.pdf\">https://arxiv.org/pdf/1711.02512.pdf</a>) better than the original average pooling. Code copied from <a href=\"https://github.com/filipradenovic/cnnimageretrieval-pytorch\">https://github.com/filipradenovic/cnnimageretrieval-pytorch</a>.</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\nmodel = se_resnet50(num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = GeM()\n</code></p>\n\n<p><strong>Training and Testing</strong></p>\n\n<p>The training process can be divided into two stages. In the first stage, I routinely trained the eight models and validated each of them on the public LB. To get more stable results, models were evaluated in pairs (with different seeds), that's why I have 2x for each type of model. When probing LB, I tried to reduce the degree of freedom of hyperparemeters to alleviate overfitting, for example, to determine the best number of epochs for training I used a step size of five. The following are the optimized results after stage1 training:</p>\n\n<p>inception_resnet_v2   public: 0.831 private: 0.927\ninception_v4            public: 0.826 private: 0.924\nseresnext50            public: 0.826 private: 0.931\nseresnext101          public: 0.819 private: 0.923   (the 2nd best result, missing the best)\nensemble                public: 0.844 private: 0.934</p>\n\n<p>In the second stage of training, I added pseudo-labelled (soft version) public test data and two additional external data - the Idrid and the Messidor dataset, to the stage1 trainset. The labels of Idrid also have five levels, to mitigate the labeling bias (my guess), the Idrid labels were smoothed by averaging the provided labels with the predicted labels from stage1 models. For the Messidor dataset which we only have four levels, I grouped the stage1 predicted soft labels by the provided groundtruth labels, calculated the mean of each group, and bounded the outliers using the mean values. For example, if the mean value of a group is 2.2, and within the same group an image has stage1 prediction of 1.1, then the label is adjusted to 2.2-0.5=1.7. After the preparation of all the labels and data, each stage1 model were trained for 10 more epoch. Finally, the LB improved from public:0.850 and private:0.935. Hours before the deadline, I made the last shot by changing the qwk thresholds from [0.5, 1.5, 2.5, 3.5] to [0.7, 1.5, 2.5, 3.5] and private improved to 0.936 ...</p>",
  "messages": [
    {
      "id": 2657843,
      "postDate": "2024-02-18T19:19:29.957Z",
      "content": "<p>Could you please provide me with your answer submission file? I need it to train a model and test it on the test data.</p>",
      "rawMarkdown": "Could you please provide me with your answer submission file? I need it to train a model and test it on the test data.",
      "votes": 1,
      "replies": [
        {
          "id": 3092015,
          "postDate": "2025-01-09T04:50:15.417Z",
          "content": "<p>me too. If you found it, can you kindly share it w me?</p>",
          "rawMarkdown": "me too. If you found it, can you kindly share it w me?",
          "replies": [
            {
              "id": 3427977,
              "postDate": "2026-03-24T18:37:21.157Z",
              "content": "<p>did you find it?</p>",
              "rawMarkdown": "did you find it?\n"
            }
          ]
        }
      ]
    },
    {
      "id": 2402441,
      "postDate": "2023-08-22T06:42:46.523Z",
      "content": "<p>Is the source code available somewhere? I wanted to integrate the model with some hardware for an on the go product as a part of my project semester</p>",
      "rawMarkdown": "Is the source code available somewhere? I wanted to integrate the model with some hardware for an on the go product as a part of my project semester",
      "votes": 1
    },
    {
      "id": 621672,
      "postDate": "2019-09-08T20:21:22.507Z",
      "content": "<p>Thanks to APTOS and Kaggle for hosting this interesting competition. \nI also would like to thank those who generously contributed in the kernels and discussions in this competition, and to the top teams that shared solutions and findings in the 2015 competition. Those findings and solutions have greatly impacted my strategy.</p>\n\n<p><strong>Validation Strategy</strong></p>\n\n<p>One of the most popular topics in every competition are proper validation strategies. During the early stage, I tried using 2015 data (both train and test) as train set, and 2019 train data as validation set. Unfortunately, the validation results and public LB were very different, I was not able to make their performance correlate well. In some discussions and kernels, other participants were also reporting inconsistent performance between CV and LB. Because of this, I did not know how to move forward, so I shifted to some other competitions for a few weeks. When I came back to this competition, I made the decision to combine the whole 2015 and 2019 data as train set, and solely relied on public LB for validation.</p>\n\n<p><strong>Preprocessing</strong></p>\n\n<p>I don't think it's necessary to preprocess images to help with the modelling, the image qualities are perfect as input for deep neural networks. So, no special preprocessing, just plain resizing.</p>\n\n<p><strong>Models and Input sizes</strong></p>\n\n<p>My final submission was a simple average of the following eight models. Inceptions and ResNets usually blend well. If I could have two more weeks I would definitely add some EfficientNets. </p>\n\n<p><code>\n2 x inception_resnet_v2, input size 512\n2 x inception_v4, input size 512\n2 x seresnext50, input size 512\n2 x seresnext101, input size 384\n</code></p>\n\n<p>The input size was mainly determined by observations in the 2015 competition that larger input size brought better performance. Even though I did not find a lot beneficial to go beyond 384 based on the public LB feedback, I still push to the extreme the input size because the private set might benefit.</p>\n\n<p><strong>Loss, Augmentations, Pooling</strong></p>\n\n<p>I used only <em>nn.SmoothL1Loss()</em> as the loss function. Other loss functions may work well too. I sticked to this single loss just to simplify the emsembling process.</p>\n\n<p>For augmentations, the following were helpful</p>\n\n<p><code>\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n</code></p>\n\n<p>For the last pooling layer, I found the generalized mean pooling (<a href=\"https://arxiv.org/pdf/1711.02512.pdf\">https://arxiv.org/pdf/1711.02512.pdf</a>) better than the original average pooling. Code copied from <a href=\"https://github.com/filipradenovic/cnnimageretrieval-pytorch\">https://github.com/filipradenovic/cnnimageretrieval-pytorch</a>.</p>\n\n<p><code>\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps) <br>\n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\nmodel = se_resnet50(num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = GeM()\n</code></p>\n\n<p><strong>Training and Testing</strong></p>\n\n<p>The training process can be divided into two stages. In the first stage, I routinely trained the eight models and validated each of them on the public LB. To get more stable results, models were evaluated in pairs (with different seeds), that's why I have 2x for each type of model. When probing LB, I tried to reduce the degree of freedom of hyperparemeters to alleviate overfitting, for example, to determine the best number of epochs for training I used a step size of five. The following are the optimized results after stage1 training:</p>\n\n<p>inception_resnet_v2   public: 0.831 private: 0.927\ninception_v4            public: 0.826 private: 0.924\nseresnext50            public: 0.826 private: 0.931\nseresnext101          public: 0.819 private: 0.923   (the 2nd best result, missing the best)\nensemble                public: 0.844 private: 0.934</p>\n\n<p>In the second stage of training, I added pseudo-labelled (soft version) public test data and two additional external data - the Idrid and the Messidor dataset, to the stage1 trainset. The labels of Idrid also have five levels, to mitigate the labeling bias (my guess), the Idrid labels were smoothed by averaging the provided labels with the predicted labels from stage1 models. For the Messidor dataset which we only have four levels, I grouped the stage1 predicted soft labels by the provided groundtruth labels, calculated the mean of each group, and bounded the outliers using the mean values. For example, if the mean value of a group is 2.2, and within the same group an image has stage1 prediction of 1.1, then the label is adjusted to 2.2-0.5=1.7. After the preparation of all the labels and data, each stage1 model were trained for 10 more epoch. Finally, the LB improved from public:0.850 and private:0.935. Hours before the deadline, I made the last shot by changing the qwk thresholds from [0.5, 1.5, 2.5, 3.5] to [0.7, 1.5, 2.5, 3.5] and private improved to 0.936 ...</p>",
      "rawMarkdown": "Thanks to APTOS and Kaggle for hosting this interesting competition. \nI also would like to thank those who generously contributed in the kernels and discussions in this competition, and to the top teams that shared solutions and findings in the 2015 competition. Those findings and solutions have greatly impacted my strategy.\n\n**Validation Strategy**\n\nOne of the most popular topics in every competition are proper validation strategies. During the early stage, I tried using 2015 data (both train and test) as train set, and 2019 train data as validation set. Unfortunately, the validation results and public LB were very different, I was not able to make their performance correlate well. In some discussions and kernels, other participants were also reporting inconsistent performance between CV and LB. Because of this, I did not know how to move forward, so I shifted to some other competitions for a few weeks. When I came back to this competition, I made the decision to combine the whole 2015 and 2019 data as train set, and solely relied on public LB for validation.\n\n**Preprocessing**\n\nI don't think it's necessary to preprocess images to help with the modelling, the image qualities are perfect as input for deep neural networks. So, no special preprocessing, just plain resizing.\n\n**Models and Input sizes**\n\nMy final submission was a simple average of the following eight models. Inceptions and ResNets usually blend well. If I could have two more weeks I would definitely add some EfficientNets. \n\n```\n2 x inception_resnet_v2, input size 512\n2 x inception_v4, input size 512\n2 x seresnext50, input size 512\n2 x seresnext101, input size 384\n```\n\nThe input size was mainly determined by observations in the 2015 competition that larger input size brought better performance. Even though I did not find a lot beneficial to go beyond 384 based on the public LB feedback, I still push to the extreme the input size because the private set might benefit.\n\n**Loss, Augmentations, Pooling**\n\nI used only *nn.SmoothL1Loss()* as the loss function. Other loss functions may work well too. I sticked to this single loss just to simplify the emsembling process.\n\nFor augmentations, the following were helpful\n\n```\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n```\n\nFor the last pooling layer, I found the generalized mean pooling (https://arxiv.org/pdf/1711.02512.pdf) better than the original average pooling. Code copied from https://github.com/filipradenovic/cnnimageretrieval-pytorch.\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\nmodel = se_resnet50(num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = GeM()\n```\n\n**Training and Testing**\n\nThe training process can be divided into two stages. In the first stage, I routinely trained the eight models and validated each of them on the public LB. To get more stable results, models were evaluated in pairs (with different seeds), that's why I have 2x for each type of model. When probing LB, I tried to reduce the degree of freedom of hyperparemeters to alleviate overfitting, for example, to determine the best number of epochs for training I used a step size of five. The following are the optimized results after stage1 training:\n\ninception_resnet_v2   public: 0.831 private: 0.927\ninception_v4            public: 0.826 private: 0.924\nseresnext50            public: 0.826 private: 0.931\nseresnext101          public: 0.819 private: 0.923   (the 2nd best result, missing the best)\nensemble                public: 0.844 private: 0.934\n\nIn the second stage of training, I added pseudo-labelled (soft version) public test data and two additional external data - the Idrid and the Messidor dataset, to the stage1 trainset. The labels of Idrid also have five levels, to mitigate the labeling bias (my guess), the Idrid labels were smoothed by averaging the provided labels with the predicted labels from stage1 models. For the Messidor dataset which we only have four levels, I grouped the stage1 predicted soft labels by the provided groundtruth labels, calculated the mean of each group, and bounded the outliers using the mean values. For example, if the mean value of a group is 2.2, and within the same group an image has stage1 prediction of 1.1, then the label is adjusted to 2.2-0.5=1.7. After the preparation of all the labels and data, each stage1 model were trained for 10 more epoch. Finally, the LB improved from public:0.850 and private:0.935. Hours before the deadline, I made the last shot by changing the qwk thresholds from [0.5, 1.5, 2.5, 3.5] to [0.7, 1.5, 2.5, 3.5] and private improved to 0.936 ...",
      "votes": 303
    },
    {
      "id": 752006,
      "postDate": "2020-02-20T17:47:18.137Z",
      "content": "<p>Tq u for sharing Awesome </p>",
      "rawMarkdown": "Tq u for sharing Awesome ",
      "votes": 13
    },
    {
      "id": 625604,
      "postDate": "2019-09-13T08:10:07.407Z",
      "content": "<p>Congratulations! And it's really amazing that you became the solo winner while being short of time.... Just to summarize:\n1. A good loss function. I think smooth l1 loss is indeed better than mse, for it deals better with mis-labelled outliers. \n2. More data. You used all the train data, two external datasets and pseudo-labelling. Also the external data were carefully labelled. This should help quite a lot.\nOther techniques include heavy augmentations, generalized mean pooling, ensembling various of networks, validating with public lb and trying another threshold, I don't know which ones of these help more, but augmentations should help quite a lot(maybe this can be included as part of \"more data\").</p>\n\n<p>From your writeup, I believe 0.936 isn't the upper bound. This model might be further improved by:\n1. Adding some ensembles with preprocessing. This can at least add some variety to your models.\n2. The smooth l1 loss. Perhaps it can be furthered improved by e.g. setting the gradient with respect to outliers to be less than 1. Of course this hasn't been tried out yet.\n3. Adding efficientnets to the ensembles. For many of us, efficient is significantly better than all other conv nets, had you tried it out, the score could have been even higher...</p>",
      "rawMarkdown": "Congratulations! And it's really amazing that you became the solo winner while being short of time.... Just to summarize:\n1. A good loss function. I think smooth l1 loss is indeed better than mse, for it deals better with mis-labelled outliers. \n2. More data. You used all the train data, two external datasets and pseudo-labelling. Also the external data were carefully labelled. This should help quite a lot.\nOther techniques include heavy augmentations, generalized mean pooling, ensembling various of networks, validating with public lb and trying another threshold, I don't know which ones of these help more, but augmentations should help quite a lot(maybe this can be included as part of \"more data\").\n\nFrom your writeup, I believe 0.936 isn't the upper bound. This model might be further improved by:\n1. Adding some ensembles with preprocessing. This can at least add some variety to your models.\n2. The smooth l1 loss. Perhaps it can be furthered improved by e.g. setting the gradient with respect to outliers to be less than 1. Of course this hasn't been tried out yet.\n3. Adding efficientnets to the ensembles. For many of us, efficient is significantly better than all other conv nets, had you tried it out, the score could have been even higher...",
      "votes": 12,
      "replies": [
        {
          "id": 627484,
          "postDate": "2019-09-16T03:44:35.033Z",
          "content": "<p>Great summary</p>",
          "rawMarkdown": "Great summary"
        }
      ]
    },
    {
      "id": 623667,
      "postDate": "2019-09-11T07:33:08.063Z",
      "content": "<p>Thank you very much for your insights, this is very informative!\n1. Could you please elaborate on how did you train the models? What lr/batchsize did you use? Which optimizer? Did you use any lr scheduling?\n2. I was eagerly waiting for the end of this competition to learn about the validation strategy from the top solutions. And then you say that you didn't use any tricks - you even didn't use absolutely anything special, just use a leaderboard score for validation. So this is what blows my mind. How could you be so sure that you wouldn't be tragically affected by shakeup? You certainly can believe that your models generalize well, but how could you be so sure about their performance on private data?\n3. A similar question is about the thresholding. It seemed that one careless step in the thresholding could lead to a catastrophic overfitting. So this is why your success in switching 0.5 -&gt; 0.7 is interesting. I am wondering what made you do this? Why didn't you change other numbers, why did you change this one?\n4. I know many people who made a mistake by selecting not the best submission. So what did you choose for the final 2 submissions? What was your strategy of selecting submissions?</p>",
      "rawMarkdown": "Thank you very much for your insights, this is very informative!\n1. Could you please elaborate on how did you train the models? What lr/batchsize did you use? Which optimizer? Did you use any lr scheduling?\n2. I was eagerly waiting for the end of this competition to learn about the validation strategy from the top solutions. And then you say that you didn't use any tricks - you even didn't use absolutely anything special, just use a leaderboard score for validation. So this is what blows my mind. How could you be so sure that you wouldn't be tragically affected by shakeup? You certainly can believe that your models generalize well, but how could you be so sure about their performance on private data?\n3. A similar question is about the thresholding. It seemed that one careless step in the thresholding could lead to a catastrophic overfitting. So this is why your success in switching 0.5 -&gt; 0.7 is interesting. I am wondering what made you do this? Why didn't you change other numbers, why did you change this one?\n4. I know many people who made a mistake by selecting not the best submission. So what did you choose for the final 2 submissions? What was your strategy of selecting submissions?",
      "votes": 7,
      "replies": [
        {
          "id": 624017,
          "postDate": "2019-09-11T14:10:26.930Z",
          "content": "<ol>\n<li>I used Adem with lr=0.0002 and trained 30-50 epochs, depending on the LB performance, then lowered to 0.00002 and continued for 5-10 more epochs.</li>\n<li>Of course I could not be sure about shakeup and my model's performance on private data. But I did try to avoid unnecessary overfitting to the public LB, I have introduced my methodology of LB probing  in the main thread. Here, I would like to emphasize that there is no generic method of LB probing that will guarantee success because every competition is different. </li>\n<li>Because I noticed public test data has smaller ratio of normal cases compare with the train data, my model could be biased to predicting less normal cases due to the LB probing. In case the private test data has higher ratio of normal case, just like the train data, the obvious way is to adjust the threshold to let the model predict more 0s.</li>\n<li>I don't have other obviously better submissions to choose in this competition. Luck always play some part in final submission selection, I also missed the best private submissions in some previous competitions. </li>\n</ol>",
          "rawMarkdown": "1. I used Adem with lr=0.0002 and trained 30-50 epochs, depending on the LB performance, then lowered to 0.00002 and continued for 5-10 more epochs.\n2. Of course I could not be sure about shakeup and my model's performance on private data. But I did try to avoid unnecessary overfitting to the public LB, I have introduced my methodology of LB probing  in the main thread. Here, I would like to emphasize that there is no generic method of LB probing that will guarantee success because every competition is different. \n3. Because I noticed public test data has smaller ratio of normal cases compare with the train data, my model could be biased to predicting less normal cases due to the LB probing. In case the private test data has higher ratio of normal case, just like the train data, the obvious way is to adjust the threshold to let the model predict more 0s.\n4. I don't have other obviously better submissions to choose in this competition. Luck always play some part in final submission selection, I also missed the best private submissions in some previous competitions. ",
          "votes": 11
        },
        {
          "id": 625712,
          "postDate": "2019-09-13T11:08:25.177Z",
          "content": "<p>Thank you very much for your answers! Congratulations and good luck in your next competitions!</p>",
          "rawMarkdown": "Thank you very much for your answers! Congratulations and good luck in your next competitions!"
        },
        {
          "id": 627156,
          "postDate": "2019-09-15T14:06:31.693Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 627186,
          "postDate": "2019-09-15T14:41:49.050Z",
          "content": "<p>Sorry I don't get your question. What I did was, for example, submitting to LB with saved weights of epoch 30,35,40,45, and obtaining results, say, 0.80, 0.81, 0.82, 0.81, respectively. Then I would lower the learning rate and continue training from the weights of epoch 40. </p>",
          "rawMarkdown": "Sorry I don't get your question. What I did was, for example, submitting to LB with saved weights of epoch 30,35,40,45, and obtaining results, say, 0.80, 0.81, 0.82, 0.81, respectively. Then I would lower the learning rate and continue training from the weights of epoch 40. ",
          "votes": 5
        },
        {
          "id": 627339,
          "postDate": "2019-09-15T19:40:39.293Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 622000,
      "postDate": "2019-09-09T06:36:20.470Z",
      "content": "<p>I am amazed how abandoning common rules like \"never trust public LB alone\" can pay off, when done in proper circumstances. Congrats on your first win! Great job! And thanks for sharing your solution in details.. as you always do 😁 </p>",
      "rawMarkdown": "I am amazed how abandoning common rules like \"never trust public LB alone\" can pay off, when done in proper circumstances. Congrats on your first win! Great job! And thanks for sharing your solution in details.. as you always do 😁 ",
      "votes": 5,
      "replies": [
        {
          "id": 622544,
          "postDate": "2019-09-09T19:43:47.937Z",
          "content": "<p>Congratulations for upgrading to GM</p>",
          "rawMarkdown": "Congratulations for upgrading to GM",
          "votes": 1
        }
      ]
    },
    {
      "id": 621849,
      "postDate": "2019-09-09T02:54:35.210Z",
      "content": "<p>Wow, congratulation!\nYour solution is totally out of my imagination 😂 No EfficientNet, No image preprocessing!</p>",
      "rawMarkdown": "Wow, congratulation!\nYour solution is totally out of my imagination 😂 No EfficientNet, No image preprocessing!",
      "votes": 5
    },
    {
      "id": 628246,
      "postDate": "2019-09-17T02:22:18.023Z",
      "content": "<p>Congrats for the top 1, Guanshuo, my I ask a question about the idea of Generalized Mean Pooling? How could you tune the p parameter? I am adapting this idea, but it is quite unstable to train, i.e it produce gradient overflow after just a few epochs. My current approach is to set initial p=1 (so that GeM acts like average pooling) and use longer warmup to stabilized the training process, but this setting lower my model performance. </p>",
      "rawMarkdown": "Congrats for the top 1, Guanshuo, my I ask a question about the idea of Generalized Mean Pooling? How could you tune the p parameter? I am adapting this idea, but it is quite unstable to train, i.e it produce gradient overflow after just a few epochs. My current approach is to set initial p=1 (so that GeM acts like average pooling) and use longer warmup to stabilized the training process, but this setting lower my model performance. ",
      "votes": 3,
      "replies": [
        {
          "id": 628820,
          "postDate": "2019-09-18T00:49:55.190Z",
          "content": "<p>I used the default p=3 as initial value.</p>",
          "rawMarkdown": "I used the default p=3 as initial value.",
          "votes": 2
        },
        {
          "id": 1201789,
          "postDate": "2021-02-15T16:57:32.953Z",
          "content": "<p>exactly sir</p>",
          "rawMarkdown": "exactly sir\n"
        }
      ]
    },
    {
      "id": 622044,
      "postDate": "2019-09-09T07:57:32.113Z",
      "content": "<p>I'm glad the winner also found out that preprocessing strategies didn't improve model performance. All the cropping and equalization didn't really add much wrt plain resizing. What makes the difference is larger input size and ensembling. Congratulations on your first place!</p>",
      "rawMarkdown": "I'm glad the winner also found out that preprocessing strategies didn't improve model performance. All the cropping and equalization didn't really add much wrt plain resizing. What makes the difference is larger input size and ensembling. Congratulations on your first place!",
      "votes": 1,
      "replies": [
        {
          "id": 622171,
          "postDate": "2019-09-09T11:03:08.207Z",
          "content": "<p>In fact, my statement on preprocessing is totally subjective, I have not tried any preprocessing.</p>",
          "rawMarkdown": "In fact, my statement on preprocessing is totally subjective, I have not tried any preprocessing.",
          "votes": 2
        }
      ]
    },
    {
      "id": 622013,
      "postDate": "2019-09-09T07:01:59.367Z",
      "content": "<p>Congratulations and thanks for your summary. Your approach seems amazing to me, first because of validation strategy, while everyone is complaining how different was public and private test sets, you relied on public LB and won! Then I thought that the winner would use ensemble of EfficientNets and something else and your solution doesn't use EfficientNets at all. Real surprise!</p>",
      "rawMarkdown": "Congratulations and thanks for your summary. Your approach seems amazing to me, first because of validation strategy, while everyone is complaining how different was public and private test sets, you relied on public LB and won! Then I thought that the winner would use ensemble of EfficientNets and something else and your solution doesn't use EfficientNets at all. Real surprise!",
      "votes": 1
    },
    {
      "id": 621870,
      "postDate": "2019-09-09T03:53:55.457Z",
      "content": "<p>Congrats and appreciate that sharing the best strategy!!!</p>\n\n<p>I have some little detailed questions, the first one is that since I treat this competition as classification problem, I am not pretty sure how to set the boundary like [0.5, 1.5, 2.5, 3.5] for regression problem, is that come from stage1 model predictions of each class group?</p>\n\n<p>Second is that how to set final prediction using the boundary(I guess setting 0 for prediction smaller than 0.5; 1 for prediction in [0.5, 1.5); and so on...)?</p>\n\n<p>Last one is that i am confused when you give the example about bounding the outlier for the Messidor dataset,  where is 0.5 in \"2.2-0.5=1.7\" come from?</p>\n\n<p>Thanks!!</p>",
      "rawMarkdown": "Congrats and appreciate that sharing the best strategy!!!\n\nI have some little detailed questions, the first one is that since I treat this competition as classification problem, I am not pretty sure how to set the boundary like [0.5, 1.5, 2.5, 3.5] for regression problem, is that come from stage1 model predictions of each class group?\n\nSecond is that how to set final prediction using the boundary(I guess setting 0 for prediction smaller than 0.5; 1 for prediction in [0.5, 1.5); and so on...)?\n\nLast one is that i am confused when you give the example about bounding the outlier for the Messidor dataset,  where is 0.5 in \"2.2-0.5=1.7\" come from?\n\nThanks!!",
      "votes": 1,
      "replies": [
        {
          "id": 622166,
          "postDate": "2019-09-09T10:56:36.677Z",
          "content": "<ul>\n<li>I'm not sure how to adjust the predicted probabilities in a scientific way, given I know the target label distribution bias. There might be some way in the literature.</li>\n<li>Yes</li>\n<li>0.5 is a dispersion value I manually set. It may not be optimal.</li>\n</ul>",
          "rawMarkdown": "- I'm not sure how to adjust the predicted probabilities in a scientific way, given I know the target label distribution bias. There might be some way in the literature.\n- Yes\n- 0.5 is a dispersion value I manually set. It may not be optimal.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 630561,
      "postDate": "2019-09-20T11:58:20.540Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>  Congrats on your solution !! It's amazing. As a beginner to Kaggle, have some questions.</p>\n\n<ol>\n<li>You said you were using 2 x RTX titan instances. How do I get access to these instances? Should I go for AWS/Azure services, which I pay on demand?</li>\n<li>\"My final submission was a simple average of the following eight models.\" - Is it that you do ensemble, and take the average score of all the models?</li>\n<li>How do you change the learning rate after some epochs? Is it like you save the weights, and retrain in another session with a different learning rate.</li>\n<li>As a beginner, I don't have the patience to wait till my commit finishes and I get my score. Do you guys do it differently? Or do you do it offline in your own higher end GPUs?</li>\n</ol>",
      "rawMarkdown": "@wowfattie  Congrats on your solution !! It's amazing. As a beginner to Kaggle, have some questions.\n\n1. You said you were using 2 x RTX titan instances. How do I get access to these instances? Should I go for AWS/Azure services, which I pay on demand?\n2. \"My final submission was a simple average of the following eight models.\" - Is it that you do ensemble, and take the average score of all the models?\n3. How do you change the learning rate after some epochs? Is it like you save the weights, and retrain in another session with a different learning rate.\n4. As a beginner, I don't have the patience to wait till my commit finishes and I get my score. Do you guys do it differently? Or do you do it offline in your own higher end GPUs?",
      "votes": 2,
      "replies": [
        {
          "id": 630856,
          "postDate": "2019-09-20T21:27:10.263Z",
          "content": "<ol>\n<li>I bought the graphics cards myself.</li>\n<li>Yes</li>\n<li>Train until the validation results (the public LB for me, in this competition) are not increasing, then reduce lr. </li>\n<li>I did most of the hyperparameter tuning on kernels with a smaller CNN model, until Kaggle enforced the new kernel usage limit. The final ensemble was trained on my own machines. </li>\n</ol>",
          "rawMarkdown": "1. I bought the graphics cards myself.\n2. Yes\n3. Train until the validation results (the public LB for me, in this competition) are not increasing, then reduce lr. \n4. I did most of the hyperparameter tuning on kernels with a smaller CNN model, until Kaggle enforced the new kernel usage limit. The final ensemble was trained on my own machines. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 621767,
      "postDate": "2019-09-08T23:09:41.577Z",
      "content": "<p>Thanks for sharing! I have always seen you alone go straight to the top for every single competition we have. Not sure how can you do it, looking forward to learning from you again in the next comp.!</p>",
      "rawMarkdown": "Thanks for sharing! I have always seen you alone go straight to the top for every single competition we have. Not sure how can you do it, looking forward to learning from you again in the next comp.!",
      "votes": 2
    },
    {
      "id": 621697,
      "postDate": "2019-09-08T20:48:00.507Z",
      "content": "<p>Big congratz, you are always doing amazing work. Fitting public LB done right I guess :)</p>",
      "rawMarkdown": "Big congratz, you are always doing amazing work. Fitting public LB done right I guess :)",
      "votes": 2,
      "replies": [
        {
          "id": 621713,
          "postDate": "2019-09-08T21:10:29.093Z",
          "content": "<p>Thanks. \nYour 4 gold medals in 5 uncorrelated competitions is not legit😄 </p>",
          "rawMarkdown": "Thanks. \nYour 4 gold medals in 5 uncorrelated competitions is not legit😄 ",
          "votes": 7
        }
      ]
    },
    {
      "id": 621819,
      "postDate": "2019-09-09T01:34:43.737Z",
      "content": "<p>大佬大佬恭喜恭喜</p>",
      "rawMarkdown": "大佬大佬恭喜恭喜"
    },
    {
      "id": 3098129,
      "postDate": "2025-01-16T05:52:38.603Z",
      "content": "<p>i want code can anyone share</p>",
      "rawMarkdown": "i want code can anyone share\n"
    },
    {
      "id": 2508113,
      "postDate": "2023-11-01T13:37:47.923Z",
      "content": "<p>Great Work!</p>",
      "rawMarkdown": "Great Work!"
    },
    {
      "id": 1276158,
      "postDate": "2021-04-17T07:20:20.947Z",
      "content": "<p>Hello Guanshuo Xu. It may be too late already, but congratulations for this accomplishment. I am new to machine learning and kaggle, and I have been exploring around. I read some of the solutions you posted on the competitions where  you ranked 1st, and although I cannot understand most of the details (like the solutions themselves), I was deeply amazed on how machine learning can be used to solve real world problems. I learned on some of your posts that building a successful ML model requires critical thinking, resourcefulness (like using other people's kernel) and sometimes even luck. Can you please give me some advise on how I can be as successful as you are now? I hope I can be like you someday. :-)<br>\nThanks for reading and have a nice day.</p>",
      "rawMarkdown": "Hello Guanshuo Xu. It may be too late already, but congratulations for this accomplishment. I am new to machine learning and kaggle, and I have been exploring around. I read some of the solutions you posted on the competitions where  you ranked 1st, and although I cannot understand most of the details (like the solutions themselves), I was deeply amazed on how machine learning can be used to solve real world problems. I learned on some of your posts that building a successful ML model requires critical thinking, resourcefulness (like using other people's kernel) and sometimes even luck. Can you please give me some advise on how I can be as successful as you are now? I hope I can be like you someday. :-)\nThanks for reading and have a nice day."
    },
    {
      "id": 1104711,
      "postDate": "2020-12-07T07:24:53.507Z",
      "content": "<p>Thanks for sharing, this is awesome and will be helpful for a beginner like me</p>",
      "rawMarkdown": "Thanks for sharing, this is awesome and will be helpful for a beginner like me"
    },
    {
      "id": 821865,
      "postDate": "2020-04-26T13:26:48.027Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>  this was great summary i read after long time\n1) if you have git hub repo could u point me to ?\n2)which library did u  use for this augs\n<code>\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n</code>\n2) what do you think could u been a major contributor</p>",
      "rawMarkdown": "@wowfattie  this was great summary i read after long time\n1) if you have git hub repo could u point me to ?\n2)which library did u  use for this augs\n```\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n```\n2) what do you think could u been a major contributor\n"
    },
    {
      "id": 636669,
      "postDate": "2019-09-30T02:41:12.260Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> Congratulations on winning Gold. I have read through your comments to various questions on your strategy. It's just amazing. I'm very new to Kaggle and data science in general.  I have definitely picked up a lot of valuable knowledgefrom your explanations. Let me go back to read the comments. 👍 </p>",
      "rawMarkdown": "@wowfattie Congratulations on winning Gold. I have read through your comments to various questions on your strategy. It's just amazing. I'm very new to Kaggle and data science in general.  I have definitely picked up a lot of valuable knowledgefrom your explanations. Let me go back to read the comments. 👍 "
    },
    {
      "id": 634929,
      "postDate": "2019-09-27T00:49:46.963Z",
      "content": "<p>Congratulation!  BTW  may I ask you to share your kernel,  after all the competition is finished, so everyone can learn so much from your great kernel.</p>",
      "rawMarkdown": "Congratulation!  BTW  may I ask you to share your kernel,  after all the competition is finished, so everyone can learn so much from your great kernel."
    },
    {
      "id": 630509,
      "postDate": "2019-09-20T09:44:44.610Z",
      "content": "<p>Good job!))and my congratulations!</p>",
      "rawMarkdown": "Good job!))and my congratulations!"
    },
    {
      "id": 629734,
      "postDate": "2019-09-19T05:40:55.027Z",
      "content": "<p>谢分享！ 恭喜</p>",
      "rawMarkdown": "谢分享！ 恭喜"
    },
    {
      "id": 629283,
      "postDate": "2019-09-18T15:46:22.620Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 627873,
      "postDate": "2019-09-16T14:01:59.147Z",
      "content": "<p>Congratulations for your first place <a href=\"/wowfattie\">@wowfattie</a> and for becoming 2nd in the global rank! Well done! </p>\n\n<p>And thanks a lot for sharing.</p>\n\n<p>Did you use any crop?</p>\n\n<p>How did you deal with the unbalanced classes?</p>",
      "rawMarkdown": "Congratulations for your first place @wowfattie and for becoming 2nd in the global rank! Well done! \n\nAnd thanks a lot for sharing.\n\nDid you use any crop?\n\nHow did you deal with the unbalanced classes?",
      "replies": [
        {
          "id": 628823,
          "postDate": "2019-09-18T00:51:30.250Z",
          "content": "<p>No, I did not perform any cropping.\nI did nothing to deal with the class imbalance</p>",
          "rawMarkdown": "No, I did not perform any cropping.\nI did nothing to deal with the class imbalance"
        }
      ]
    },
    {
      "id": 626666,
      "postDate": "2019-09-14T16:39:34.587Z",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> a wonderful job，Guanshuo! Congrats for the 1st place winning as well as for achieving the 2nd overall kaggle user ranking!</p>",
      "rawMarkdown": "@wowfattie a wonderful job，Guanshuo! Congrats for the 1st place winning as well as for achieving the 2nd overall kaggle user ranking!",
      "replies": [
        {
          "id": 627179,
          "postDate": "2019-09-15T14:34:34.827Z",
          "content": "<p>Thank you, Shize</p>",
          "rawMarkdown": "Thank you, Shize"
        }
      ]
    },
    {
      "id": 626466,
      "postDate": "2019-09-14T11:26:30.920Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 624835,
      "postDate": "2019-09-12T12:19:40.217Z",
      "content": "<p>Congratulations on your win!</p>\n\n<p>Probably a silly question sorry, (new here!)  What does LB/ CV mean?</p>\n\n<p>S</p>",
      "rawMarkdown": "Congratulations on your win!\n\nProbably a silly question sorry, (new here!)  What does LB/ CV mean?\n\nS",
      "replies": [
        {
          "id": 624847,
          "postDate": "2019-09-12T12:39:17.013Z",
          "content": "<p>LB -&gt; Leaderboard\nCV -&gt; Cross Validation</p>",
          "rawMarkdown": "LB -&gt; Leaderboard\nCV -&gt; Cross Validation"
        }
      ]
    },
    {
      "id": 624712,
      "postDate": "2019-09-12T09:58:17.187Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 624621,
      "postDate": "2019-09-12T08:07:28.467Z",
      "content": "<p>Congratulations! Great Work!</p>",
      "rawMarkdown": "Congratulations! Great Work!"
    },
    {
      "id": 624579,
      "postDate": "2019-09-12T07:33:37.457Z",
      "content": "<p>Wow, congratulations!</p>",
      "rawMarkdown": "Wow, congratulations!"
    },
    {
      "id": 624245,
      "postDate": "2019-09-11T21:01:37.950Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!"
    },
    {
      "id": 624163,
      "postDate": "2019-09-11T18:51:26.790Z",
      "content": "<p>Congrats! What augmentation module or package do you use? Or can you post the augmentation code? Thanks!</p>",
      "rawMarkdown": "Congrats! What augmentation module or package do you use? Or can you post the augmentation code? Thanks!",
      "replies": [
        {
          "id": 624210,
          "postDate": "2019-09-11T20:03:45.457Z",
          "content": "<p>I have my own version of augmentation using opencv and numpy. Some of them are just copy/modified from imgaug or albumentation. </p>",
          "rawMarkdown": "I have my own version of augmentation using opencv and numpy. Some of them are just copy/modified from imgaug or albumentation. "
        },
        {
          "id": 624537,
          "postDate": "2019-09-12T07:08:29.263Z",
          "content": "<p>really curious what motivates you writing your own augmentation library?</p>",
          "rawMarkdown": "really curious what motivates you writing your own augmentation library?"
        }
      ]
    },
    {
      "id": 623746,
      "postDate": "2019-09-11T09:02:22.900Z",
      "content": "<p>Congratulations </p>",
      "rawMarkdown": "Congratulations "
    },
    {
      "id": 623735,
      "postDate": "2019-09-11T08:52:26.913Z",
      "content": "<p>Congratulations...and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations...and thanks for sharing."
    },
    {
      "id": 623693,
      "postDate": "2019-09-11T08:02:20.757Z",
      "content": "<p>Many congratulations.  Thank you for sharing important insights.  Ensemble is next i need to work on. </p>",
      "rawMarkdown": "Many congratulations.  Thank you for sharing important insights.  Ensemble is next i need to work on. "
    },
    {
      "id": 623246,
      "postDate": "2019-09-10T16:25:01.607Z",
      "content": "<p>Congratulation!!!!!</p>",
      "rawMarkdown": "Congratulation!!!!!"
    },
    {
      "id": 622948,
      "postDate": "2019-09-10T09:45:01.230Z",
      "content": "<p>Congrates. Impressive.\nI can feel your excitement to jump from rank 6 to rank 1.</p>\n\n<p>Learn a lot from this competition. Thanks for sharing.</p>",
      "rawMarkdown": "Congrates. Impressive.\nI can feel your excitement to jump from rank 6 to rank 1.\n\nLearn a lot from this competition. Thanks for sharing."
    },
    {
      "id": 622859,
      "postDate": "2019-09-10T07:06:48.690Z",
      "content": "<p>Congrats on your Win. Great job.</p>",
      "rawMarkdown": "Congrats on your Win. Great job."
    },
    {
      "id": 622772,
      "postDate": "2019-09-10T04:39:20.730Z",
      "content": "<p>Congratulations on your win!</p>\n\n<p>Do you plan on sharing your code?</p>",
      "rawMarkdown": "Congratulations on your win!\n\nDo you plan on sharing your code?"
    },
    {
      "id": 622769,
      "postDate": "2019-09-10T04:36:22.617Z",
      "content": "<p>I bet on Seresnext too, but didn`t do my best(\nThanks for sharing! Good job</p>",
      "rawMarkdown": "I bet on Seresnext too, but didn`t do my best(\nThanks for sharing! Good job"
    },
    {
      "id": 622753,
      "postDate": "2019-09-10T03:39:00.803Z",
      "content": "<p>Congratulations! Great Work!</p>",
      "rawMarkdown": "Congratulations! Great Work!"
    },
    {
      "id": 622709,
      "postDate": "2019-09-10T01:47:31.693Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 622697,
      "postDate": "2019-09-10T01:15:04.977Z",
      "content": "<p>Congratulations!!!</p>",
      "rawMarkdown": "Congratulations!!!"
    },
    {
      "id": 622547,
      "postDate": "2019-09-09T19:53:10.730Z",
      "content": "<p>Congratulations <a href=\"/wowfattie\">@wowfattie</a> for a well deserved win. Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations @wowfattie for a well deserved win. Thanks for sharing."
    },
    {
      "id": 622410,
      "postDate": "2019-09-09T16:05:10.630Z",
      "content": "<p>congrats on your Win.</p>",
      "rawMarkdown": "congrats on your Win."
    },
    {
      "id": 622092,
      "postDate": "2019-09-09T09:13:04.253Z",
      "content": "<p>Congratulations for your great results!</p>\n\n<p>I was taking a look at your strategy and I have a question: as the EyePACS dataset is so imbalanced, did you use some balancing or class weighting strategy on training time?</p>",
      "rawMarkdown": "Congratulations for your great results!\n\nI was taking a look at your strategy and I have a question: as the EyePACS dataset is so imbalanced, did you use some balancing or class weighting strategy on training time?",
      "replies": [
        {
          "id": 622175,
          "postDate": "2019-09-09T11:06:14.643Z",
          "content": "<p>What is EyePACS dataset?\nI did not use any balancing or class weighting</p>",
          "rawMarkdown": "What is EyePACS dataset?\nI did not use any balancing or class weighting"
        },
        {
          "id": 622284,
          "postDate": "2019-09-09T13:13:37.127Z",
          "content": "<p>Thanks for your reply. EyePACS dataset is the one used in the 2015 competition.</p>",
          "rawMarkdown": "Thanks for your reply. EyePACS dataset is the one used in the 2015 competition."
        }
      ]
    },
    {
      "id": 622063,
      "postDate": "2019-09-09T08:25:58.470Z",
      "content": "<p>Congrats! Did you apply TTA when ensemble?</p>",
      "rawMarkdown": "Congrats! Did you apply TTA when ensemble?",
      "replies": [
        {
          "id": 622172,
          "postDate": "2019-09-09T11:03:30.770Z",
          "content": "<p>Yes, only horizontal flip to reduce running time. I also tried some other rotate90 degree or vertical flip but results were almost same</p>",
          "rawMarkdown": "Yes, only horizontal flip to reduce running time. I also tried some other rotate90 degree or vertical flip but results were almost same"
        }
      ]
    },
    {
      "id": 621911,
      "postDate": "2019-09-09T04:56:42.177Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations"
    },
    {
      "id": 621879,
      "postDate": "2019-09-09T04:12:18.283Z",
      "content": "<p>Congratulations, Thank you for writing solution summary.</p>",
      "rawMarkdown": "Congratulations, Thank you for writing solution summary."
    },
    {
      "id": 621863,
      "postDate": "2019-09-09T03:35:48.423Z",
      "content": "<p>Thank you for sharing your superb solution !\nWithout any efficient-nets &amp; pre-processing, you achieved this result. I am wondering, what will be the score if you add them.. </p>",
      "rawMarkdown": "Thank you for sharing your superb solution !\nWithout any efficient-nets &amp; pre-processing, you achieved this result. I am wondering, what will be the score if you add them.. \n"
    },
    {
      "id": 621822,
      "postDate": "2019-09-09T01:45:30.013Z",
      "content": "<p>Thanks for sharing and congratz for 1st place!!!!</p>",
      "rawMarkdown": "Thanks for sharing and congratz for 1st place!!!!"
    },
    {
      "id": 621820,
      "postDate": "2019-09-09T01:40:20.027Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 621815,
      "postDate": "2019-09-09T01:30:01.827Z",
      "content": "<p>wow.. congrats.. just fitting the LB, amazing work😏 </p>",
      "rawMarkdown": "wow.. congrats.. just fitting the LB, amazing work😏 "
    },
    {
      "id": 621725,
      "postDate": "2019-09-08T21:20:19.653Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    },
    {
      "id": 621723,
      "postDate": "2019-09-08T21:17:40.517Z",
      "content": "<p>Congratulations! \nI like your brave validation strategy a lot, it reminds me of our validation approach 😄 </p>",
      "rawMarkdown": "Congratulations! \nI like your brave validation strategy a lot, it reminds me of our validation approach 😄 "
    },
    {
      "id": 621722,
      "postDate": "2019-09-08T21:16:38.237Z",
      "content": "<p>Big Congrats!!! 👍 </p>",
      "rawMarkdown": "Big Congrats!!! 👍 "
    },
    {
      "id": 621688,
      "postDate": "2019-09-08T20:38:59.930Z",
      "content": "<p>Congratulations! Thank you for sharing your solution.\nI would like to ask you one thing.\nWhy did you decide to make the last threshold change?</p>",
      "rawMarkdown": "Congratulations! Thank you for sharing your solution.\nI would like to ask you one thing.\nWhy did you decide to make the last threshold change?",
      "replies": [
        {
          "id": 621692,
          "postDate": "2019-09-08T20:43:25.153Z",
          "content": "<p>The private set may have more normal cases, so I changed the threshold to accomodate</p>",
          "rawMarkdown": "The private set may have more normal cases, so I changed the threshold to accomodate"
        },
        {
          "id": 621704,
          "postDate": "2019-09-08T20:53:36.457Z",
          "content": "<p>I understood your idea. Thank you!\nIn fact, the private set seems to have more class 0 than the public. </p>",
          "rawMarkdown": "I understood your idea. Thank you!\nIn fact, the private set seems to have more class 0 than the public. "
        }
      ]
    },
    {
      "id": 621680,
      "postDate": "2019-09-08T20:32:26.743Z",
      "content": "<p>Congratulations! Finally a top solution without EfficientNets :p </p>",
      "rawMarkdown": "Congratulations! Finally a top solution without EfficientNets :p ",
      "replies": [
        {
          "id": 625619,
          "postDate": "2019-09-13T08:27:21.710Z",
          "content": "<p>Doesn't  it  is  a  joke   ?  :)</p>",
          "rawMarkdown": "Doesn't  it  is  a  joke   ?  :)"
        }
      ]
    },
    {
      "id": 621677,
      "postDate": "2019-09-08T20:30:52.957Z",
      "content": "<p>Congrats and Thank you <a href=\"/wowfattie\">@wowfattie</a> for summarizing strategy. Few queries:</p>\n\n<ol>\n<li>What was the size of the training set from the 2015 dataset you used?</li>\n<li>What machine configuration you used to train models? Model, Epochs and Training time?</li>\n</ol>\n\n<p>Thanks,\nHimanshu</p>",
      "rawMarkdown": "Congrats and Thank you @wowfattie for summarizing strategy. Few queries:\n\n1. What was the size of the training set from the 2015 dataset you used?\n2. What machine configuration you used to train models? Model, Epochs and Training time?\n\nThanks,\nHimanshu",
      "replies": [
        {
          "id": 621693,
          "postDate": "2019-09-08T20:44:36.890Z",
          "content": "<p>You mean number of images? Should be around 90000.\n2 x RTX titan</p>",
          "rawMarkdown": "You mean number of images? Should be around 90000.\n2 x RTX titan",
          "votes": 1
        },
        {
          "id": 621700,
          "postDate": "2019-09-08T20:51:06.140Z",
          "content": "<p>Thanks <a href=\"/wowfattie\">@wowfattie</a> . There were many dark and dirty images. I was expecting you might have cleaned those up manually. Didn't you?</p>",
          "rawMarkdown": "Thanks @wowfattie . There were many dark and dirty images. I was expecting you might have cleaned those up manually. Didn't you?"
        },
        {
          "id": 621715,
          "postDate": "2019-09-08T21:11:25.223Z",
          "content": "<p>No, I did not clean any training data</p>",
          "rawMarkdown": "No, I did not clean any training data"
        }
      ]
    },
    {
      "id": 981235,
      "postDate": "2020-08-22T10:02:13.553Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 673088,
      "postDate": "2019-11-14T13:58:00.943Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 631361,
      "postDate": "2019-09-21T22:06:05.627Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 626746,
      "postDate": "2019-09-14T19:24:24.167Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 627182,
          "postDate": "2019-09-15T14:36:44.623Z",
          "content": "<p>It's hard to say how much time I spent, because even when I was doing something else I could be thinking about how to solve the problem</p>",
          "rawMarkdown": "It's hard to say how much time I spent, because even when I was doing something else I could be thinking about how to solve the problem"
        }
      ]
    },
    {
      "id": 626483,
      "postDate": "2019-09-14T12:12:34.640Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 626510,
          "postDate": "2019-09-14T12:50:45.900Z",
          "content": "<ol>\n<li>I was talking about optimizing the number of training epochs. For example, I tested LB score on the saved weights of epochs [25, 30, 35, 40] ...  And I selected the one with the best LB score.</li>\n<li>Yes.</li>\n<li>Please refer to this great kernel about setting thresholds.\n<a href=\"https://www.kaggle.com/abhishek/pytorch-inference-kernel-lazy-tta\">https://www.kaggle.com/abhishek/pytorch-inference-kernel-lazy-tta</a></li>\n</ol>",
          "rawMarkdown": "1. I was talking about optimizing the number of training epochs. For example, I tested LB score on the saved weights of epochs [25, 30, 35, 40] ...  And I selected the one with the best LB score.\n2. Yes.\n3. Please refer to this great kernel about setting thresholds.\nhttps://www.kaggle.com/abhishek/pytorch-inference-kernel-lazy-tta",
          "votes": 1
        },
        {
          "id": 626576,
          "postDate": "2019-09-14T14:08:00.607Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 626582,
          "postDate": "2019-09-14T14:11:33.920Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 626615,
          "postDate": "2019-09-14T15:01:14.103Z",
          "content": "<p>Compared with keras, pytorch is faster and has more complete pretrained model zoos, for example, <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>\nI suggest shift to pytorch ASAP</p>",
          "rawMarkdown": "Compared with keras, pytorch is faster and has more complete pretrained model zoos, for example, https://github.com/Cadene/pretrained-models.pytorch\nI suggest shift to pytorch ASAP",
          "votes": 3
        },
        {
          "id": 626800,
          "postDate": "2019-09-14T21:41:41.303Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 632447,
          "postDate": "2019-09-23T15:43:53.513Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 632689,
          "postDate": "2019-09-23T23:15:56.667Z",
          "content": "<p>If you train 100 epochs and find out that the best is epoch30, then the other 70 epochs of training are wasted.</p>",
          "rawMarkdown": "If you train 100 epochs and find out that the best is epoch30, then the other 70 epochs of training are wasted."
        },
        {
          "id": 633072,
          "postDate": "2019-09-24T12:13:52.157Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 622667,
      "postDate": "2019-09-10T00:26:46.040Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 622182,
      "postDate": "2019-09-09T11:15:26.850Z",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "isDeleted": true
    },
    {
      "id": 621896,
      "postDate": "2019-09-09T04:41:11.280Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 621818,
      "postDate": "2019-09-09T01:33:01.777Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 621825,
          "postDate": "2019-09-09T01:54:52.093Z",
          "content": "<p>Averaging</p>",
          "rawMarkdown": "Averaging"
        }
      ]
    },
    {
      "id": 621771,
      "postDate": "2019-09-08T23:33:05.467Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1201788,
      "postDate": "2021-02-15T16:57:07.410Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing\n"
    },
    {
      "id": 623691,
      "postDate": "2019-09-11T08:01:05.080Z",
      "content": "<p>Thanks for sharing and Congratulation!!</p>",
      "rawMarkdown": "Thanks for sharing and Congratulation!!"
    },
    {
      "id": 623119,
      "postDate": "2019-09-10T13:45:56.660Z",
      "content": "<p>Thanks for your contributions. 😁 </p>",
      "rawMarkdown": "Thanks for your contributions. 😁 "
    },
    {
      "id": 621795,
      "postDate": "2019-09-09T00:44:31.800Z",
      "content": "<p>Big congrats! Thanks for sharing! </p>",
      "rawMarkdown": "Big congrats! Thanks for sharing! "
    },
    {
      "id": 823037,
      "postDate": "2020-04-27T10:59:33.057Z",
      "content": "<p>Thanks for sharing!!!</p>",
      "rawMarkdown": "Thanks for sharing!!!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2657843,
      "author_name": "joo",
      "author_url": "",
      "post_date": "2024-02-18T19:19:29.957000",
      "content": "<p>Could you please provide me with your answer submission file? I need it to train a model and test it on the test data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3092015,
          "author_name": "Sreeja Pottabathula",
          "author_url": "",
          "post_date": "2025-01-09T04:50:15.417000",
          "content": "<p>me too. If you found it, can you kindly share it w me?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3427977,
              "author_name": "Sharan Magesh",
              "author_url": "",
              "post_date": "2026-03-24T18:37:21.157000",
              "content": "<p>did you find it?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2402441,
      "author_name": "Naman Chopra",
      "author_url": "",
      "post_date": "2023-08-22T06:42:46.523000",
      "content": "<p>Is the source code available somewhere? I wanted to integrate the model with some hardware for an on the go product as a part of my project semester</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 752006,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T17:47:18.137000",
      "content": "<p>Tq u for sharing Awesome </p>",
      "votes": 13,
      "replies": []
    },
    {
      "id": 625604,
      "author_name": "Homoalways",
      "author_url": "",
      "post_date": "2019-09-13T08:10:07.407000",
      "content": "<p>Congratulations! And it's really amazing that you became the solo winner while being short of time.... Just to summarize:\n1. A good loss function. I think smooth l1 loss is indeed better than mse, for it deals better with mis-labelled outliers. \n2. More data. You used all the train data, two external datasets and pseudo-labelling. Also the external data were carefully labelled. This should help quite a lot.\nOther techniques include heavy augmentations, generalized mean pooling, ensembling various of networks, validating with public lb and trying another threshold, I don't know which ones of these help more, but augmentations should help quite a lot(maybe this can be included as part of \"more data\").</p>\n\n<p>From your writeup, I believe 0.936 isn't the upper bound. This model might be further improved by:\n1. Adding some ensembles with preprocessing. This can at least add some variety to your models.\n2. The smooth l1 loss. Perhaps it can be furthered improved by e.g. setting the gradient with respect to outliers to be less than 1. Of course this hasn't been tried out yet.\n3. Adding efficientnets to the ensembles. For many of us, efficient is significantly better than all other conv nets, had you tried it out, the score could have been even higher...</p>",
      "votes": 12,
      "replies": [
        {
          "id": 627484,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2019-09-16T03:44:35.033000",
          "content": "<p>Great summary</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 623667,
      "author_name": "Evgeny Kovalev",
      "author_url": "",
      "post_date": "2019-09-11T07:33:08.063000",
      "content": "<p>Thank you very much for your insights, this is very informative!\n1. Could you please elaborate on how did you train the models? What lr/batchsize did you use? Which optimizer? Did you use any lr scheduling?\n2. I was eagerly waiting for the end of this competition to learn about the validation strategy from the top solutions. And then you say that you didn't use any tricks - you even didn't use absolutely anything special, just use a leaderboard score for validation. So this is what blows my mind. How could you be so sure that you wouldn't be tragically affected by shakeup? You certainly can believe that your models generalize well, but how could you be so sure about their performance on private data?\n3. A similar question is about the thresholding. It seemed that one careless step in the thresholding could lead to a catastrophic overfitting. So this is why your success in switching 0.5 -&gt; 0.7 is interesting. I am wondering what made you do this? Why didn't you change other numbers, why did you change this one?\n4. I know many people who made a mistake by selecting not the best submission. So what did you choose for the final 2 submissions? What was your strategy of selecting submissions?</p>",
      "votes": 7,
      "replies": [
        {
          "id": 624017,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-11T14:10:26.930000",
          "content": "<ol>\n<li>I used Adem with lr=0.0002 and trained 30-50 epochs, depending on the LB performance, then lowered to 0.00002 and continued for 5-10 more epochs.</li>\n<li>Of course I could not be sure about shakeup and my model's performance on private data. But I did try to avoid unnecessary overfitting to the public LB, I have introduced my methodology of LB probing  in the main thread. Here, I would like to emphasize that there is no generic method of LB probing that will guarantee success because every competition is different. </li>\n<li>Because I noticed public test data has smaller ratio of normal cases compare with the train data, my model could be biased to predicting less normal cases due to the LB probing. In case the private test data has higher ratio of normal case, just like the train data, the obvious way is to adjust the threshold to let the model predict more 0s.</li>\n<li>I don't have other obviously better submissions to choose in this competition. Luck always play some part in final submission selection, I also missed the best private submissions in some previous competitions. </li>\n</ol>",
          "votes": 11,
          "replies": []
        },
        {
          "id": 625712,
          "author_name": "Evgeny Kovalev",
          "author_url": "",
          "post_date": "2019-09-13T11:08:25.177000",
          "content": "<p>Thank you very much for your answers! Congratulations and good luck in your next competitions!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 627156,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-15T14:06:31.693000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 627186,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-15T14:41:49.050000",
          "content": "<p>Sorry I don't get your question. What I did was, for example, submitting to LB with saved weights of epoch 30,35,40,45, and obtaining results, say, 0.80, 0.81, 0.82, 0.81, respectively. Then I would lower the learning rate and continue training from the weights of epoch 40. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 627339,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-15T19:40:39.293000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 622000,
      "author_name": "dott",
      "author_url": "",
      "post_date": "2019-09-09T06:36:20.470000",
      "content": "<p>I am amazed how abandoning common rules like \"never trust public LB alone\" can pay off, when done in proper circumstances. Congrats on your first win! Great job! And thanks for sharing your solution in details.. as you always do 😁 </p>",
      "votes": 5,
      "replies": [
        {
          "id": 622544,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-09T19:43:47.937000",
          "content": "<p>Congratulations for upgrading to GM</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 621849,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2019-09-09T02:54:35.210000",
      "content": "<p>Wow, congratulation!\nYour solution is totally out of my imagination 😂 No EfficientNet, No image preprocessing!</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 628246,
      "author_name": "nan",
      "author_url": "",
      "post_date": "2019-09-17T02:22:18.023000",
      "content": "<p>Congrats for the top 1, Guanshuo, my I ask a question about the idea of Generalized Mean Pooling? How could you tune the p parameter? I am adapting this idea, but it is quite unstable to train, i.e it produce gradient overflow after just a few epochs. My current approach is to set initial p=1 (so that GeM acts like average pooling) and use longer warmup to stabilized the training process, but this setting lower my model performance. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 628820,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-18T00:49:55.190000",
          "content": "<p>I used the default p=3 as initial value.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1201789,
          "author_name": "Muhammad Ali",
          "author_url": "",
          "post_date": "2021-02-15T16:57:32.953000",
          "content": "<p>exactly sir</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 622044,
      "author_name": "Francesco Ramoni",
      "author_url": "",
      "post_date": "2019-09-09T07:57:32.113000",
      "content": "<p>I'm glad the winner also found out that preprocessing strategies didn't improve model performance. All the cropping and equalization didn't really add much wrt plain resizing. What makes the difference is larger input size and ensembling. Congratulations on your first place!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 622171,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-09T11:03:08.207000",
          "content": "<p>In fact, my statement on preprocessing is totally subjective, I have not tried any preprocessing.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 622013,
      "author_name": "Anna Novikova",
      "author_url": "",
      "post_date": "2019-09-09T07:01:59.367000",
      "content": "<p>Congratulations and thanks for your summary. Your approach seems amazing to me, first because of validation strategy, while everyone is complaining how different was public and private test sets, you relied on public LB and won! Then I thought that the winner would use ensemble of EfficientNets and something else and your solution doesn't use EfficientNets at all. Real surprise!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 621870,
      "author_name": "jayjhlin",
      "author_url": "",
      "post_date": "2019-09-09T03:53:55.457000",
      "content": "<p>Congrats and appreciate that sharing the best strategy!!!</p>\n\n<p>I have some little detailed questions, the first one is that since I treat this competition as classification problem, I am not pretty sure how to set the boundary like [0.5, 1.5, 2.5, 3.5] for regression problem, is that come from stage1 model predictions of each class group?</p>\n\n<p>Second is that how to set final prediction using the boundary(I guess setting 0 for prediction smaller than 0.5; 1 for prediction in [0.5, 1.5); and so on...)?</p>\n\n<p>Last one is that i am confused when you give the example about bounding the outlier for the Messidor dataset,  where is 0.5 in \"2.2-0.5=1.7\" come from?</p>\n\n<p>Thanks!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 622166,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-09T10:56:36.677000",
          "content": "<ul>\n<li>I'm not sure how to adjust the predicted probabilities in a scientific way, given I know the target label distribution bias. There might be some way in the literature.</li>\n<li>Yes</li>\n<li>0.5 is a dispersion value I manually set. It may not be optimal.</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 630561,
      "author_name": "Kaushik Ramachandran",
      "author_url": "",
      "post_date": "2019-09-20T11:58:20.540000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>  Congrats on your solution !! It's amazing. As a beginner to Kaggle, have some questions.</p>\n\n<ol>\n<li>You said you were using 2 x RTX titan instances. How do I get access to these instances? Should I go for AWS/Azure services, which I pay on demand?</li>\n<li>\"My final submission was a simple average of the following eight models.\" - Is it that you do ensemble, and take the average score of all the models?</li>\n<li>How do you change the learning rate after some epochs? Is it like you save the weights, and retrain in another session with a different learning rate.</li>\n<li>As a beginner, I don't have the patience to wait till my commit finishes and I get my score. Do you guys do it differently? Or do you do it offline in your own higher end GPUs?</li>\n</ol>",
      "votes": 2,
      "replies": [
        {
          "id": 630856,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-20T21:27:10.263000",
          "content": "<ol>\n<li>I bought the graphics cards myself.</li>\n<li>Yes</li>\n<li>Train until the validation results (the public LB for me, in this competition) are not increasing, then reduce lr. </li>\n<li>I did most of the hyperparameter tuning on kernels with a smaller CNN model, until Kaggle enforced the new kernel usage limit. The final ensemble was trained on my own machines. </li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 621767,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-09-08T23:09:41.577000",
      "content": "<p>Thanks for sharing! I have always seen you alone go straight to the top for every single competition we have. Not sure how can you do it, looking forward to learning from you again in the next comp.!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 621697,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2019-09-08T20:48:00.507000",
      "content": "<p>Big congratz, you are always doing amazing work. Fitting public LB done right I guess :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 621713,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-08T21:10:29.093000",
          "content": "<p>Thanks. \nYour 4 gold medals in 5 uncorrelated competitions is not legit😄 </p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 621819,
      "author_name": "HeVi27",
      "author_url": "",
      "post_date": "2019-09-09T01:34:43.737000",
      "content": "<p>大佬大佬恭喜恭喜</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3098129,
      "author_name": "RAVI BHUSHAN DIXIT",
      "author_url": "",
      "post_date": "2025-01-16T05:52:38.603000",
      "content": "<p>i want code can anyone share</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2508113,
      "author_name": "Ayushman Raghuvanshi",
      "author_url": "",
      "post_date": "2023-11-01T13:37:47.923000",
      "content": "<p>Great Work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1276158,
      "author_name": "Romel Correa",
      "author_url": "",
      "post_date": "2021-04-17T07:20:20.947000",
      "content": "<p>Hello Guanshuo Xu. It may be too late already, but congratulations for this accomplishment. I am new to machine learning and kaggle, and I have been exploring around. I read some of the solutions you posted on the competitions where  you ranked 1st, and although I cannot understand most of the details (like the solutions themselves), I was deeply amazed on how machine learning can be used to solve real world problems. I learned on some of your posts that building a successful ML model requires critical thinking, resourcefulness (like using other people's kernel) and sometimes even luck. Can you please give me some advise on how I can be as successful as you are now? I hope I can be like you someday. :-)<br>\nThanks for reading and have a nice day.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1104711,
      "author_name": "Jeremy Reeve Kurniawan",
      "author_url": "",
      "post_date": "2020-12-07T07:24:53.507000",
      "content": "<p>Thanks for sharing, this is awesome and will be helpful for a beginner like me</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821865,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-04-26T13:26:48.027000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a>  this was great summary i read after long time\n1) if you have git hub repo could u point me to ?\n2)which library did u  use for this augs\n<code>\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n</code>\n2) what do you think could u been a major contributor</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 636669,
      "author_name": "Joyce Chidi",
      "author_url": "",
      "post_date": "2019-09-30T02:41:12.260000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> Congratulations on winning Gold. I have read through your comments to various questions on your strategy. It's just amazing. I'm very new to Kaggle and data science in general.  I have definitely picked up a lot of valuable knowledgefrom your explanations. Let me go back to read the comments. 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 634929,
      "author_name": "henry",
      "author_url": "",
      "post_date": "2019-09-27T00:49:46.963000",
      "content": "<p>Congratulation!  BTW  may I ask you to share your kernel,  after all the competition is finished, so everyone can learn so much from your great kernel.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 630509,
      "author_name": "Alina Vladimirova",
      "author_url": "",
      "post_date": "2019-09-20T09:44:44.610000",
      "content": "<p>Good job!))and my congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 629734,
      "author_name": "Horus",
      "author_url": "",
      "post_date": "2019-09-19T05:40:55.027000",
      "content": "<p>谢分享！ 恭喜</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 629283,
      "author_name": "Ruslan Zabrodin",
      "author_url": "",
      "post_date": "2019-09-18T15:46:22.620000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 627873,
      "author_name": "Virilo Tejedor Aguilera",
      "author_url": "",
      "post_date": "2019-09-16T14:01:59.147000",
      "content": "<p>Congratulations for your first place <a href=\"/wowfattie\">@wowfattie</a> and for becoming 2nd in the global rank! Well done! </p>\n\n<p>And thanks a lot for sharing.</p>\n\n<p>Did you use any crop?</p>\n\n<p>How did you deal with the unbalanced classes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 628823,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-18T00:51:30.250000",
          "content": "<p>No, I did not perform any cropping.\nI did nothing to deal with the class imbalance</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 626666,
      "author_name": "Shize Su",
      "author_url": "",
      "post_date": "2019-09-14T16:39:34.587000",
      "content": "<p><a href=\"/wowfattie\">@wowfattie</a> a wonderful job，Guanshuo! Congrats for the 1st place winning as well as for achieving the 2nd overall kaggle user ranking!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 627179,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2019-09-15T14:34:34.827000",
          "content": "<p>Thank you, Shize</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 626466,
      "author_name": "Sagar Jatana",
      "author_url": "",
      "post_date": "2019-09-14T11:26:30.920000",
      "content": "<p>Congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624835,
      "author_name": "Seán McMahon",
      "author_url": "",
      "post_date": "2019-09-12T12:19:40.217000",
      "content": "<p>Congratulations on your win!</p>\n\n<p>Probably a silly question sorry, (new here!)  What does LB/ CV mean?</p>\n\n<p>S</p>",
      "votes": 0,
      "replies": [
        {
          "id": 624847,
          "author_name": "Rahul T P",
          "author_url": "",
          "post_date": "2019-09-12T12:39:17.013000",
          "content": "<p>LB -&gt; Leaderboard\nCV -&gt; Cross Validation</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 624712,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-12T09:58:17.187000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624621,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-12T08:07:28.467000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624579,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-12T07:33:37.457000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624245,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T21:01:37.950000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 624163,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T18:51:26.790000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 624210,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-11T20:03:45.457000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 624537,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-12T07:08:29.263000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 623746,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T09:02:22.900000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623735,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T08:52:26.913000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623693,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T08:02:20.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623246,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T16:25:01.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622948,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T09:45:01.230000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622859,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T07:06:48.690000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622772,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T04:39:20.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622769,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T04:36:22.617000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622753,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T03:39:00.803000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622709,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T01:47:31.693000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622697,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T01:15:04.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622547,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T19:53:10.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622410,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T16:05:10.630000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622092,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T09:13:04.253000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 622175,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-09T11:06:14.643000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 622284,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-09T13:13:37.127000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 622063,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T08:25:58.470000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 622172,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-09T11:03:30.770000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621911,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T04:56:42.177000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621879,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T04:12:18.283000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621863,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T03:35:48.423000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621822,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T01:45:30.013000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621820,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T01:40:20.027000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621815,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T01:30:01.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621725,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T21:20:19.653000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621723,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T21:17:40.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621722,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T21:16:38.237000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621688,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T20:38:59.930000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 621692,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-08T20:43:25.153000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621704,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-08T20:53:36.457000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621680,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T20:32:26.743000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 625619,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-13T08:27:21.710000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621677,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T20:30:52.957000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 621693,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-08T20:44:36.890000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 621700,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-08T20:51:06.140000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 621715,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-08T21:11:25.223000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 981235,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-22T10:02:13.553000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 673088,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-14T13:58:00.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 631361,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-21T22:06:05.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 626746,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-14T19:24:24.167000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 627182,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-15T14:36:44.623000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 626483,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-14T12:12:34.640000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 626510,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-14T12:50:45.900000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 626576,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-14T14:08:00.607000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 626582,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-14T14:11:33.920000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 626615,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-14T15:01:14.103000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 626800,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-14T21:41:41.303000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 632447,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-23T15:43:53.513000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 632689,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-23T23:15:56.667000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 633072,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-24T12:13:52.157000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 622667,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T00:26:46.040000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 622182,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T11:15:26.850000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621896,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T04:41:11.280000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621818,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T01:33:01.777000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 621825,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-09-09T01:54:52.093000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 621771,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-08T23:33:05.467000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1201788,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-15T16:57:07.410000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623691,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-11T08:01:05.080000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 623119,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-10T13:45:56.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 621795,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-09-09T00:44:31.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 823037,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-27T10:59:33.057000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2657843": "Could you please provide me with your answer submission file? I need it to train a model and test it on the test data.",
    "2402441": "Is the source code available somewhere? I wanted to integrate the model with some hardware for an on the go product as a part of my project semester",
    "621672": "Thanks to APTOS and Kaggle for hosting this interesting competition. \nI also would like to thank those who generously contributed in the kernels and discussions in this competition, and to the top teams that shared solutions and findings in the 2015 competition. Those findings and solutions have greatly impacted my strategy.\n\n**Validation Strategy**\n\nOne of the most popular topics in every competition are proper validation strategies. During the early stage, I tried using 2015 data (both train and test) as train set, and 2019 train data as validation set. Unfortunately, the validation results and public LB were very different, I was not able to make their performance correlate well. In some discussions and kernels, other participants were also reporting inconsistent performance between CV and LB. Because of this, I did not know how to move forward, so I shifted to some other competitions for a few weeks. When I came back to this competition, I made the decision to combine the whole 2015 and 2019 data as train set, and solely relied on public LB for validation.\n\n**Preprocessing**\n\nI don't think it's necessary to preprocess images to help with the modelling, the image qualities are perfect as input for deep neural networks. So, no special preprocessing, just plain resizing.\n\n**Models and Input sizes**\n\nMy final submission was a simple average of the following eight models. Inceptions and ResNets usually blend well. If I could have two more weeks I would definitely add some EfficientNets. \n\n```\n2 x inception_resnet_v2, input size 512\n2 x inception_v4, input size 512\n2 x seresnext50, input size 512\n2 x seresnext101, input size 384\n```\n\nThe input size was mainly determined by observations in the 2015 competition that larger input size brought better performance. Even though I did not find a lot beneficial to go beyond 384 based on the public LB feedback, I still push to the extreme the input size because the private set might benefit.\n\n**Loss, Augmentations, Pooling**\n\nI used only *nn.SmoothL1Loss()* as the loss function. Other loss functions may work well too. I sticked to this single loss just to simplify the emsembling process.\n\nFor augmentations, the following were helpful\n\n```\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n```\n\nFor the last pooling layer, I found the generalized mean pooling (https://arxiv.org/pdf/1711.02512.pdf) better than the original average pooling. Code copied from https://github.com/filipradenovic/cnnimageretrieval-pytorch.\n\n```\nfrom torch.nn.parameter import Parameter\ndef gem(x, p=3, eps=1e-6):\n    return F.avg_pool2d(x.clamp(min=eps).pow(p), (x.size(-2), x.size(-1))).pow(1./p)\nclass GeM(nn.Module):\n    def __init__(self, p=3, eps=1e-6):\n        super(GeM,self).__init__()\n        self.p = Parameter(torch.ones(1)*p)\n        self.eps = eps\n    def forward(self, x):\n        return gem(x, p=self.p, eps=self.eps)       \n    def __repr__(self):\n        return self.__class__.__name__ + '(' + 'p=' + '{:.4f}'.format(self.p.data.tolist()[0]) + ', ' + 'eps=' + str(self.eps) + ')'\nmodel = se_resnet50(num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = GeM()\n```\n\n**Training and Testing**\n\nThe training process can be divided into two stages. In the first stage, I routinely trained the eight models and validated each of them on the public LB. To get more stable results, models were evaluated in pairs (with different seeds), that's why I have 2x for each type of model. When probing LB, I tried to reduce the degree of freedom of hyperparemeters to alleviate overfitting, for example, to determine the best number of epochs for training I used a step size of five. The following are the optimized results after stage1 training:\n\ninception_resnet_v2   public: 0.831 private: 0.927\ninception_v4            public: 0.826 private: 0.924\nseresnext50            public: 0.826 private: 0.931\nseresnext101          public: 0.819 private: 0.923   (the 2nd best result, missing the best)\nensemble                public: 0.844 private: 0.934\n\nIn the second stage of training, I added pseudo-labelled (soft version) public test data and two additional external data - the Idrid and the Messidor dataset, to the stage1 trainset. The labels of Idrid also have five levels, to mitigate the labeling bias (my guess), the Idrid labels were smoothed by averaging the provided labels with the predicted labels from stage1 models. For the Messidor dataset which we only have four levels, I grouped the stage1 predicted soft labels by the provided groundtruth labels, calculated the mean of each group, and bounded the outliers using the mean values. For example, if the mean value of a group is 2.2, and within the same group an image has stage1 prediction of 1.1, then the label is adjusted to 2.2-0.5=1.7. After the preparation of all the labels and data, each stage1 model were trained for 10 more epoch. Finally, the LB improved from public:0.850 and private:0.935. Hours before the deadline, I made the last shot by changing the qwk thresholds from [0.5, 1.5, 2.5, 3.5] to [0.7, 1.5, 2.5, 3.5] and private improved to 0.936 ...",
    "752006": "Tq u for sharing Awesome ",
    "625604": "Congratulations! And it's really amazing that you became the solo winner while being short of time.... Just to summarize:\n1. A good loss function. I think smooth l1 loss is indeed better than mse, for it deals better with mis-labelled outliers. \n2. More data. You used all the train data, two external datasets and pseudo-labelling. Also the external data were carefully labelled. This should help quite a lot.\nOther techniques include heavy augmentations, generalized mean pooling, ensembling various of networks, validating with public lb and trying another threshold, I don't know which ones of these help more, but augmentations should help quite a lot(maybe this can be included as part of \"more data\").\n\nFrom your writeup, I believe 0.936 isn't the upper bound. This model might be further improved by:\n1. Adding some ensembles with preprocessing. This can at least add some variety to your models.\n2. The smooth l1 loss. Perhaps it can be furthered improved by e.g. setting the gradient with respect to outliers to be less than 1. Of course this hasn't been tried out yet.\n3. Adding efficientnets to the ensembles. For many of us, efficient is significantly better than all other conv nets, had you tried it out, the score could have been even higher...",
    "623667": "Thank you very much for your insights, this is very informative!\n1. Could you please elaborate on how did you train the models? What lr/batchsize did you use? Which optimizer? Did you use any lr scheduling?\n2. I was eagerly waiting for the end of this competition to learn about the validation strategy from the top solutions. And then you say that you didn't use any tricks - you even didn't use absolutely anything special, just use a leaderboard score for validation. So this is what blows my mind. How could you be so sure that you wouldn't be tragically affected by shakeup? You certainly can believe that your models generalize well, but how could you be so sure about their performance on private data?\n3. A similar question is about the thresholding. It seemed that one careless step in the thresholding could lead to a catastrophic overfitting. So this is why your success in switching 0.5 -&gt; 0.7 is interesting. I am wondering what made you do this? Why didn't you change other numbers, why did you change this one?\n4. I know many people who made a mistake by selecting not the best submission. So what did you choose for the final 2 submissions? What was your strategy of selecting submissions?",
    "622000": "I am amazed how abandoning common rules like \"never trust public LB alone\" can pay off, when done in proper circumstances. Congrats on your first win! Great job! And thanks for sharing your solution in details.. as you always do 😁 ",
    "621849": "Wow, congratulation!\nYour solution is totally out of my imagination 😂 No EfficientNet, No image preprocessing!",
    "628246": "Congrats for the top 1, Guanshuo, my I ask a question about the idea of Generalized Mean Pooling? How could you tune the p parameter? I am adapting this idea, but it is quite unstable to train, i.e it produce gradient overflow after just a few epochs. My current approach is to set initial p=1 (so that GeM acts like average pooling) and use longer warmup to stabilized the training process, but this setting lower my model performance. ",
    "622044": "I'm glad the winner also found out that preprocessing strategies didn't improve model performance. All the cropping and equalization didn't really add much wrt plain resizing. What makes the difference is larger input size and ensembling. Congratulations on your first place!",
    "622013": "Congratulations and thanks for your summary. Your approach seems amazing to me, first because of validation strategy, while everyone is complaining how different was public and private test sets, you relied on public LB and won! Then I thought that the winner would use ensemble of EfficientNets and something else and your solution doesn't use EfficientNets at all. Real surprise!",
    "621870": "Congrats and appreciate that sharing the best strategy!!!\n\nI have some little detailed questions, the first one is that since I treat this competition as classification problem, I am not pretty sure how to set the boundary like [0.5, 1.5, 2.5, 3.5] for regression problem, is that come from stage1 model predictions of each class group?\n\nSecond is that how to set final prediction using the boundary(I guess setting 0 for prediction smaller than 0.5; 1 for prediction in [0.5, 1.5); and so on...)?\n\nLast one is that i am confused when you give the example about bounding the outlier for the Messidor dataset,  where is 0.5 in \"2.2-0.5=1.7\" come from?\n\nThanks!!",
    "630561": "@wowfattie  Congrats on your solution !! It's amazing. As a beginner to Kaggle, have some questions.\n\n1. You said you were using 2 x RTX titan instances. How do I get access to these instances? Should I go for AWS/Azure services, which I pay on demand?\n2. \"My final submission was a simple average of the following eight models.\" - Is it that you do ensemble, and take the average score of all the models?\n3. How do you change the learning rate after some epochs? Is it like you save the weights, and retrain in another session with a different learning rate.\n4. As a beginner, I don't have the patience to wait till my commit finishes and I get my score. Do you guys do it differently? Or do you do it offline in your own higher end GPUs?",
    "621767": "Thanks for sharing! I have always seen you alone go straight to the top for every single competition we have. Not sure how can you do it, looking forward to learning from you again in the next comp.!",
    "621697": "Big congratz, you are always doing amazing work. Fitting public LB done right I guess :)",
    "621819": "大佬大佬恭喜恭喜",
    "3098129": "i want code can anyone share\n",
    "2508113": "Great Work!",
    "1276158": "Hello Guanshuo Xu. It may be too late already, but congratulations for this accomplishment. I am new to machine learning and kaggle, and I have been exploring around. I read some of the solutions you posted on the competitions where  you ranked 1st, and although I cannot understand most of the details (like the solutions themselves), I was deeply amazed on how machine learning can be used to solve real world problems. I learned on some of your posts that building a successful ML model requires critical thinking, resourcefulness (like using other people's kernel) and sometimes even luck. Can you please give me some advise on how I can be as successful as you are now? I hope I can be like you someday. :-)\nThanks for reading and have a nice day.",
    "1104711": "Thanks for sharing, this is awesome and will be helpful for a beginner like me",
    "821865": "@wowfattie  this was great summary i read after long time\n1) if you have git hub repo could u point me to ?\n2)which library did u  use for this augs\n```\ncontrast_range=0.2,\nbrightness_range=20.,\nhue_range=10.,\nsaturation_range=20.,\nblur_and_sharpen=True,\nrotate_range=180.,\nscale_range=0.2,\nshear_range=0.2,\nshift_range=0.2,\ndo_mirror=True,\n```\n2) what do you think could u been a major contributor\n",
    "636669": "@wowfattie Congratulations on winning Gold. I have read through your comments to various questions on your strategy. It's just amazing. I'm very new to Kaggle and data science in general.  I have definitely picked up a lot of valuable knowledgefrom your explanations. Let me go back to read the comments. 👍 ",
    "634929": "Congratulation!  BTW  may I ask you to share your kernel,  after all the competition is finished, so everyone can learn so much from your great kernel.",
    "630509": "Good job!))and my congratulations!",
    "629734": "谢分享！ 恭喜",
    "629283": "Congrats!",
    "627873": "Congratulations for your first place @wowfattie and for becoming 2nd in the global rank! Well done! \n\nAnd thanks a lot for sharing.\n\nDid you use any crop?\n\nHow did you deal with the unbalanced classes?",
    "626666": "@wowfattie a wonderful job，Guanshuo! Congrats for the 1st place winning as well as for achieving the 2nd overall kaggle user ranking!",
    "626466": "Congrats!",
    "624835": "Congratulations on your win!\n\nProbably a silly question sorry, (new here!)  What does LB/ CV mean?\n\nS",
    "624712": "Congrats!",
    "624621": "Congratulations! Great Work!",
    "624579": "Wow, congratulations!",
    "624245": "Congratulations!!",
    "624163": "Congrats! What augmentation module or package do you use? Or can you post the augmentation code? Thanks!",
    "623746": "Congratulations ",
    "623735": "Congratulations...and thanks for sharing.",
    "623693": "Many congratulations.  Thank you for sharing important insights.  Ensemble is next i need to work on. ",
    "623246": "Congratulation!!!!!",
    "622948": "Congrates. Impressive.\nI can feel your excitement to jump from rank 6 to rank 1.\n\nLearn a lot from this competition. Thanks for sharing.",
    "622859": "Congrats on your Win. Great job.",
    "622772": "Congratulations on your win!\n\nDo you plan on sharing your code?",
    "622769": "I bet on Seresnext too, but didn`t do my best(\nThanks for sharing! Good job",
    "622753": "Congratulations! Great Work!",
    "622709": "Congratulations",
    "622697": "Congratulations!!!",
    "622547": "Congratulations @wowfattie for a well deserved win. Thanks for sharing.",
    "622410": "congrats on your Win.",
    "622092": "Congratulations for your great results!\n\nI was taking a look at your strategy and I have a question: as the EyePACS dataset is so imbalanced, did you use some balancing or class weighting strategy on training time?",
    "622063": "Congrats! Did you apply TTA when ensemble?",
    "621911": "Congratulations",
    "621879": "Congratulations, Thank you for writing solution summary.",
    "621863": "Thank you for sharing your superb solution !\nWithout any efficient-nets &amp; pre-processing, you achieved this result. I am wondering, what will be the score if you add them.. \n",
    "621822": "Thanks for sharing and congratz for 1st place!!!!",
    "621820": "Congratulations!",
    "621815": "wow.. congrats.. just fitting the LB, amazing work😏 ",
    "621725": "Congratulations!",
    "621723": "Congratulations! \nI like your brave validation strategy a lot, it reminds me of our validation approach 😄 ",
    "621722": "Big Congrats!!! 👍 ",
    "621688": "Congratulations! Thank you for sharing your solution.\nI would like to ask you one thing.\nWhy did you decide to make the last threshold change?",
    "621680": "Congratulations! Finally a top solution without EfficientNets :p ",
    "621677": "Congrats and Thank you @wowfattie for summarizing strategy. Few queries:\n\n1. What was the size of the training set from the 2015 dataset you used?\n2. What machine configuration you used to train models? Model, Epochs and Training time?\n\nThanks,\nHimanshu",
    "981235": "",
    "673088": "",
    "631361": "",
    "626746": "",
    "626483": "",
    "622667": "",
    "622182": "Congratulations!!",
    "621896": "",
    "621818": "",
    "621771": "",
    "1201788": "Thanks for sharing\n",
    "623691": "Thanks for sharing and Congratulation!!",
    "623119": "Thanks for your contributions. 😁 ",
    "621795": "Big congrats! Thanks for sharing! ",
    "823037": "Thanks for sharing!!!"
  }
}