{
  "id": 175655,
  "title": "85th description - random label smoothing",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175655",
  "author_name": "Sam Klein",
  "post_date": "2020-08-18T23:37:58.601000",
  "votes": 11,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My approach was to build an ensemble of EfficientNets, I settled on this because it won last year and it did not seem that there were any new approaches performing better. I ran a lot of experiments using <a href=\"https://www.kaggle.com/cdeottes\" target=\"_blank\">@cdeottes</a> notebook with image size 128 and EfficientNet-B6 to find hyperparameters that increased my CV, then I made the assumption that these parameters would also work for larger image sizes and different EfficientNets. </p>\n<p>I included the meta data into the classifier by concatenating it with the output of an MLP on the EfficientNet output. I also experimented with different amounts of label smoothing and class weights, finding that a class weight of 5 and label smoothing of 0.05 worked best. I also experimented with assigning random label smoothing to the data as described at the end of this discussion. My final ensemble also included models trained on meta data only and on image embeddings combined with meta data, the ensembling was done using and explicit grid search on the out of fold scores for all models.</p>\n<p>For ensembling my models the best method I found was to find groups of predictions that were highly correlated (typically different EfficientNets trained on the same image size were highly correlated) and take a non-linear average that increased the out of fold cross validation score for each group. I then performed a grid search for the best weights to ensemble these averaged predictions. I found this approach to be superior to performing an explicit grid search without first averaging.</p>\n<p>To give an explicit example of this ensembling. Take all of the models trained on image size NxN, take the geometric mean (and/or median) and check that this has a better cross validation score than the best single model. Repeat this process for all image sizes that have models trained on them, taking either the geometric mean or the best single model in each case. Blend the averaged predictions for all image sizes. This can be applied to any type of grouping where correlations between predictions are found.</p>\n<p>The rationale behind this is that highly correlated models can still be included in the final ensemble, instead of ignoring models that took a long time to train.</p>\n<p>The only part of my final submission that is significantly different from what I have seen posted elsewhere is that I trained some models where the amount of label smoothing for each example was drawn randomly from a distribution. So each time the model saw an example, the label was different. As far as I am aware this is the first time this method has been employed.</p>\n<p>I trained multiple models with this strategy, experimenting with different distributions (for sampling the label smoothing), and then I ensembled all of these experiments (geometric mean and median both increased CV). This ensemble had a weight of 0.25 in my final submission.</p>\n<p>This approach produced the lowest difference between CV, private LB and public LB for a single model that I have trained. The scores for a single model trained in this way, with labels drawn from a gaussian with mean 0.05 and std 0.025, are:</p>\n<p>CV (2020 data only) 0.927<br>\nprivate LB 0.9319<br>\npublic LB 0.9389.</p>\n<p>Which can be compared with the original <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a> scores of:</p>\n<p>CV (2020 data only) 0.904<br>\nprivate LB 0.9262<br>\npublic LB 0.9454.</p>\n<p>To be completely clear on what this entails I used the following loss:</p>\n<pre><code>class sls_loss(tf.keras.losses.Loss):\n      def call(self, y_true, y_pred):\n            ls = tf.abs(tf.random.normal([1], 0, 0.025, tf.float32))\n            y_true = y_true * (1.0 - ls) + 0.5 * ls\n            bce = tf.keras.losses.BinaryCrossentropy(label_smoothing=0, \n                    reduction=tf.keras.losses.Reduction.NONE)\n            per_example_loss = bce(y_true,y_pred)\n            return tf.nn.compute_average_loss(per_example_loss, \n                    global_batch_size=GLOBAL_BATCH_SIZE)\n</code></pre>\n<p>Note this is slightly different to what I described as it is per batch.</p>",
  "messages": [
    {
      "id": 976526,
      "postDate": "2020-08-18T23:37:58.600Z",
      "content": "<p>My approach was to build an ensemble of EfficientNets, I settled on this because it won last year and it did not seem that there were any new approaches performing better. I ran a lot of experiments using <a href=\"https://www.kaggle.com/cdeottes\" target=\"_blank\">@cdeottes</a> notebook with image size 128 and EfficientNet-B6 to find hyperparameters that increased my CV, then I made the assumption that these parameters would also work for larger image sizes and different EfficientNets. </p>\n<p>I included the meta data into the classifier by concatenating it with the output of an MLP on the EfficientNet output. I also experimented with different amounts of label smoothing and class weights, finding that a class weight of 5 and label smoothing of 0.05 worked best. I also experimented with assigning random label smoothing to the data as described at the end of this discussion. My final ensemble also included models trained on meta data only and on image embeddings combined with meta data, the ensembling was done using and explicit grid search on the out of fold scores for all models.</p>\n<p>For ensembling my models the best method I found was to find groups of predictions that were highly correlated (typically different EfficientNets trained on the same image size were highly correlated) and take a non-linear average that increased the out of fold cross validation score for each group. I then performed a grid search for the best weights to ensemble these averaged predictions. I found this approach to be superior to performing an explicit grid search without first averaging.</p>\n<p>To give an explicit example of this ensembling. Take all of the models trained on image size NxN, take the geometric mean (and/or median) and check that this has a better cross validation score than the best single model. Repeat this process for all image sizes that have models trained on them, taking either the geometric mean or the best single model in each case. Blend the averaged predictions for all image sizes. This can be applied to any type of grouping where correlations between predictions are found.</p>\n<p>The rationale behind this is that highly correlated models can still be included in the final ensemble, instead of ignoring models that took a long time to train.</p>\n<p>The only part of my final submission that is significantly different from what I have seen posted elsewhere is that I trained some models where the amount of label smoothing for each example was drawn randomly from a distribution. So each time the model saw an example, the label was different. As far as I am aware this is the first time this method has been employed.</p>\n<p>I trained multiple models with this strategy, experimenting with different distributions (for sampling the label smoothing), and then I ensembled all of these experiments (geometric mean and median both increased CV). This ensemble had a weight of 0.25 in my final submission.</p>\n<p>This approach produced the lowest difference between CV, private LB and public LB for a single model that I have trained. The scores for a single model trained in this way, with labels drawn from a gaussian with mean 0.05 and std 0.025, are:</p>\n<p>CV (2020 data only) 0.927<br>\nprivate LB 0.9319<br>\npublic LB 0.9389.</p>\n<p>Which can be compared with the original <a href=\"https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords\" target=\"_blank\">notebook</a> scores of:</p>\n<p>CV (2020 data only) 0.904<br>\nprivate LB 0.9262<br>\npublic LB 0.9454.</p>\n<p>To be completely clear on what this entails I used the following loss:</p>\n<pre><code>class sls_loss(tf.keras.losses.Loss):\n      def call(self, y_true, y_pred):\n            ls = tf.abs(tf.random.normal([1], 0, 0.025, tf.float32))\n            y_true = y_true * (1.0 - ls) + 0.5 * ls\n            bce = tf.keras.losses.BinaryCrossentropy(label_smoothing=0, \n                    reduction=tf.keras.losses.Reduction.NONE)\n            per_example_loss = bce(y_true,y_pred)\n            return tf.nn.compute_average_loss(per_example_loss, \n                    global_batch_size=GLOBAL_BATCH_SIZE)\n</code></pre>\n<p>Note this is slightly different to what I described as it is per batch.</p>",
      "rawMarkdown": "My approach was to build an ensemble of EfficientNets, I settled on this because it won last year and it did not seem that there were any new approaches performing better. I ran a lot of experiments using @cdeottes notebook with image size 128 and EfficientNet-B6 to find hyperparameters that increased my CV, then I made the assumption that these parameters would also work for larger image sizes and different EfficientNets. \n\nI included the meta data into the classifier by concatenating it with the output of an MLP on the EfficientNet output. I also experimented with different amounts of label smoothing and class weights, finding that a class weight of 5 and label smoothing of 0.05 worked best. I also experimented with assigning random label smoothing to the data as described at the end of this discussion. My final ensemble also included models trained on meta data only and on image embeddings combined with meta data, the ensembling was done using and explicit grid search on the out of fold scores for all models.\n\nFor ensembling my models the best method I found was to find groups of predictions that were highly correlated (typically different EfficientNets trained on the same image size were highly correlated) and take a non-linear average that increased the out of fold cross validation score for each group. I then performed a grid search for the best weights to ensemble these averaged predictions. I found this approach to be superior to performing an explicit grid search without first averaging.\n\nTo give an explicit example of this ensembling. Take all of the models trained on image size NxN, take the geometric mean (and/or median) and check that this has a better cross validation score than the best single model. Repeat this process for all image sizes that have models trained on them, taking either the geometric mean or the best single model in each case. Blend the averaged predictions for all image sizes. This can be applied to any type of grouping where correlations between predictions are found.\n\nThe rationale behind this is that highly correlated models can still be included in the final ensemble, instead of ignoring models that took a long time to train.\n\nThe only part of my final submission that is significantly different from what I have seen posted elsewhere is that I trained some models where the amount of label smoothing for each example was drawn randomly from a distribution. So each time the model saw an example, the label was different. As far as I am aware this is the first time this method has been employed.\n\nI trained multiple models with this strategy, experimenting with different distributions (for sampling the label smoothing), and then I ensembled all of these experiments (geometric mean and median both increased CV). This ensemble had a weight of 0.25 in my final submission.\n\nThis approach produced the lowest difference between CV, private LB and public LB for a single model that I have trained. The scores for a single model trained in this way, with labels drawn from a gaussian with mean 0.05 and std 0.025, are:\n\nCV (2020 data only) 0.927\nprivate LB 0.9319\npublic LB 0.9389.\n\nWhich can be compared with the original [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) scores of:\n\nCV (2020 data only) 0.904\nprivate LB 0.9262\npublic LB 0.9454.\n\nTo be completely clear on what this entails I used the following loss:\n\n```\nclass sls_loss(tf.keras.losses.Loss):\n      def call(self, y_true, y_pred):\n            ls = tf.abs(tf.random.normal([1], 0, 0.025, tf.float32))\n            y_true = y_true * (1.0 - ls) + 0.5 * ls\n            bce = tf.keras.losses.BinaryCrossentropy(label_smoothing=0, \n                    reduction=tf.keras.losses.Reduction.NONE)\n            per_example_loss = bce(y_true,y_pred)\n            return tf.nn.compute_average_loss(per_example_loss, \n                    global_batch_size=GLOBAL_BATCH_SIZE)\n```\n\nNote this is slightly different to what I described as it is per batch.",
      "votes": 10
    },
    {
      "id": 985243,
      "postDate": "2020-08-25T15:21:47.547Z",
      "content": "<p>Great!!! i learned the nice thing !!</p>",
      "rawMarkdown": "Great!!! i learned the nice thing !!",
      "votes": 1
    },
    {
      "id": 976547,
      "postDate": "2020-08-19T00:07:24.687Z",
      "content": "<p>Interesting idea Sam. I'm going to try this out next comp. Thanks for sharing. Congrats on your strong solo silver finish.</p>",
      "rawMarkdown": "Interesting idea Sam. I'm going to try this out next comp. Thanks for sharing. Congrats on your strong solo silver finish.",
      "votes": 1,
      "replies": [
        {
          "id": 976601,
          "postDate": "2020-08-19T01:17:49.943Z",
          "content": "<p>Cool! I will be interested to see how that works out. The method is a little bit fiddly because it takes the single label smoothing parameter and makes many more. The results depend on the type of distribution (gauss vs uniform) as well as their shape. I suspect that I have not fully leveraged this idea, and I would like to explore what it means for the decision boundary that the network learns, especially wrt adversarial examples.</p>\n<p>Also thank you very much for the datasets and kernels that you published! I learned a lot from them.</p>",
          "rawMarkdown": "Cool! I will be interested to see how that works out. The method is a little bit fiddly because it takes the single label smoothing parameter and makes many more. The results depend on the type of distribution (gauss vs uniform) as well as their shape. I suspect that I have not fully leveraged this idea, and I would like to explore what it means for the decision boundary that the network learns, especially wrt adversarial examples.\n\nAlso thank you very much for the datasets and kernels that you published! I learned a lot from them.",
          "votes": 1
        }
      ]
    },
    {
      "id": 978271,
      "postDate": "2020-08-20T04:24:51.087Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 978288,
          "postDate": "2020-08-20T04:50:07.503Z",
          "content": "<p>Thanks! I hope it works out for you</p>",
          "rawMarkdown": "Thanks! I hope it works out for you"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 985243,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2020-08-25T15:21:47.547000",
      "content": "<p>Great!!! i learned the nice thing !!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976547,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-19T00:07:24.687000",
      "content": "<p>Interesting idea Sam. I'm going to try this out next comp. Thanks for sharing. Congrats on your strong solo silver finish.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976601,
          "author_name": "Sam Klein",
          "author_url": "",
          "post_date": "2020-08-19T01:17:49.943000",
          "content": "<p>Cool! I will be interested to see how that works out. The method is a little bit fiddly because it takes the single label smoothing parameter and makes many more. The results depend on the type of distribution (gauss vs uniform) as well as their shape. I suspect that I have not fully leveraged this idea, and I would like to explore what it means for the decision boundary that the network learns, especially wrt adversarial examples.</p>\n<p>Also thank you very much for the datasets and kernels that you published! I learned a lot from them.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 978271,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-20T04:24:51.087000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 978288,
          "author_name": "Sam Klein",
          "author_url": "",
          "post_date": "2020-08-20T04:50:07.503000",
          "content": "<p>Thanks! I hope it works out for you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "976526": "My approach was to build an ensemble of EfficientNets, I settled on this because it won last year and it did not seem that there were any new approaches performing better. I ran a lot of experiments using @cdeottes notebook with image size 128 and EfficientNet-B6 to find hyperparameters that increased my CV, then I made the assumption that these parameters would also work for larger image sizes and different EfficientNets. \n\nI included the meta data into the classifier by concatenating it with the output of an MLP on the EfficientNet output. I also experimented with different amounts of label smoothing and class weights, finding that a class weight of 5 and label smoothing of 0.05 worked best. I also experimented with assigning random label smoothing to the data as described at the end of this discussion. My final ensemble also included models trained on meta data only and on image embeddings combined with meta data, the ensembling was done using and explicit grid search on the out of fold scores for all models.\n\nFor ensembling my models the best method I found was to find groups of predictions that were highly correlated (typically different EfficientNets trained on the same image size were highly correlated) and take a non-linear average that increased the out of fold cross validation score for each group. I then performed a grid search for the best weights to ensemble these averaged predictions. I found this approach to be superior to performing an explicit grid search without first averaging.\n\nTo give an explicit example of this ensembling. Take all of the models trained on image size NxN, take the geometric mean (and/or median) and check that this has a better cross validation score than the best single model. Repeat this process for all image sizes that have models trained on them, taking either the geometric mean or the best single model in each case. Blend the averaged predictions for all image sizes. This can be applied to any type of grouping where correlations between predictions are found.\n\nThe rationale behind this is that highly correlated models can still be included in the final ensemble, instead of ignoring models that took a long time to train.\n\nThe only part of my final submission that is significantly different from what I have seen posted elsewhere is that I trained some models where the amount of label smoothing for each example was drawn randomly from a distribution. So each time the model saw an example, the label was different. As far as I am aware this is the first time this method has been employed.\n\nI trained multiple models with this strategy, experimenting with different distributions (for sampling the label smoothing), and then I ensembled all of these experiments (geometric mean and median both increased CV). This ensemble had a weight of 0.25 in my final submission.\n\nThis approach produced the lowest difference between CV, private LB and public LB for a single model that I have trained. The scores for a single model trained in this way, with labels drawn from a gaussian with mean 0.05 and std 0.025, are:\n\nCV (2020 data only) 0.927\nprivate LB 0.9319\npublic LB 0.9389.\n\nWhich can be compared with the original [notebook](https://www.kaggle.com/cdeotte/triple-stratified-kfold-with-tfrecords) scores of:\n\nCV (2020 data only) 0.904\nprivate LB 0.9262\npublic LB 0.9454.\n\nTo be completely clear on what this entails I used the following loss:\n\n```\nclass sls_loss(tf.keras.losses.Loss):\n      def call(self, y_true, y_pred):\n            ls = tf.abs(tf.random.normal([1], 0, 0.025, tf.float32))\n            y_true = y_true * (1.0 - ls) + 0.5 * ls\n            bce = tf.keras.losses.BinaryCrossentropy(label_smoothing=0, \n                    reduction=tf.keras.losses.Reduction.NONE)\n            per_example_loss = bce(y_true,y_pred)\n            return tf.nn.compute_average_loss(per_example_loss, \n                    global_batch_size=GLOBAL_BATCH_SIZE)\n```\n\nNote this is slightly different to what I described as it is per batch.",
    "985243": "Great!!! i learned the nice thing !!",
    "976547": "Interesting idea Sam. I'm going to try this out next comp. Thanks for sharing. Congrats on your strong solo silver finish.",
    "978271": ""
  }
}