{
  "id": 206454,
  "title": "Supervised Contrastive Learning",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/206454",
  "author_name": "",
  "post_date": "2020-12-24T17:55:09.602459600Z",
  "votes": 40,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Since I saw the concept of contrastive learning I became very interested in the subject, and a few weeks ago I saw an implementation on the <a href=\"https://keras.io/examples/vision/supervised-contrastive-learning/\" target=\"_blank\">Keras tutorials page</a>, so I decided to try it at this competition, I have spent most of my TPU quota from the last two weeks experimenting with it, and I think it has some real potential, so I want to share my finding, maybe some people can share ideas, suggestions and further improve the results.</p>\n<p>Update: I wrote an article about this <a href=\"https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966\" target=\"_blank\">Supervised Contrastive Learning for Cassava Leaf Disease Classification</a>.</p>\n<h2>Supervised Contrastive Learning</h2>\n<p><img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Cassava%20Leaf%20Disease%20Classification/sscl_vs_scl.png\" alt=\"\"></p>\n<h4>In short, this is how it works</h4>\n<blockquote>\n  <p>Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes.</p>\n</blockquote>\n<p>The idea is that training models using the supervised contrastive learning can make the model encoder learn better class representation from the samples, this should lead to better generalization and robustness to image and label corruption</p>\n<p><strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning</a><br>\n<strong>Inference notebook:</strong> <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference</a></p>\n<hr>\n<h3>Experiments summary</h3>\n<p><strong>What improved performance</strong><br>\nLarger batch size<br>\nAverage encoder size<br>\nData augmentation<br>\nLager image resolution<br>\nUsing pre-trained weights, I tried training from scratch but the results were very poor.<br>\nOversampling the minority classes from the dataset, gave a little help, but I expected more.</p>\n<hr>\n<p><strong>What made no difference</strong><br>\nTemperature parameter, according to the paper, lower temperature can benefit from longer training, since the data here is limited, tweaking this parameter maybe not worth it.<br>\nDifferent projection heads, I experimented a little with the number of neurons but got no relevant improvements.<br>\nDifferent classifier heads, I experimented a little with the number of neurons but got no relevant improvements.</p>\n<hr>\n<p><strong>Next steps</strong><br>\nDifferent pooling heads (AVG+MAX, Multi-sample, etc).<br>\nDifferent optimizers.<br>\nMore elaborated data augmentation schedules (MixUp, CutMix, etc).<br>\nUse external data.<br>\nIt may be interesting to try doing 2nd stage training but fine-tuning the complete network.</p>\n<hr>\n<p>An interesting way to evaluate the learned representation is to visualize the output of the trained embedding, at the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning#Visualize-embeddings-outputs\" target=\"_blank\">training notebook</a> I plot the predictions from the model trained using <code>cross-entropy</code> and <code>supervised contrastive learning</code></p>\n<h4>Embedding trained with cross-entropy</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F95d2c3c2071b70bf5a78b842181b584d%2FScreenshot%20from%202020-12-24%2014-51-18.png?generation=1608832461429961&amp;alt=media\" alt=\"\"></p>\n<h4>Embedding trained with supervised contrastive learning</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F433c6150356e7a709efc5698cc90e0c6%2FScreenshot%20from%202020-12-24%2014-51-31.png?generation=1608832477794000&amp;alt=media\" alt=\"\"></p>\n<p>You can see that the representation from the model trained with <code>supervised contrastive learning</code> clusters much more than the model trained with <code>cross-entropy</code>.</p>",
  "messages": [
    {
      "id": "1125436",
      "postDate": "12/24/2020 17:55:09",
      "content": "<p>Since I saw the concept of contrastive learning I became very interested in the subject, and a few weeks ago I saw an implementation on the <a href=\"https://keras.io/examples/vision/supervised-contrastive-learning/\" target=\"_blank\">Keras tutorials page</a>, so I decided to try it at this competition, I have spent most of my TPU quota from the last two weeks experimenting with it, and I think it has some real potential, so I want to share my finding, maybe some people can share ideas, suggestions and further improve the results.</p>\n<p>Update: I wrote an article about this <a href=\"https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966\" target=\"_blank\">Supervised Contrastive Learning for Cassava Leaf Disease Classification</a>.</p>\n<h2>Supervised Contrastive Learning</h2>\n<p><img src=\"https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Cassava%20Leaf%20Disease%20Classification/sscl_vs_scl.png\" alt=\"\"></p>\n<h4>In short, this is how it works</h4>\n<blockquote>\n  <p>Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes.</p>\n</blockquote>\n<p>The idea is that training models using the supervised contrastive learning can make the model encoder learn better class representation from the samples, this should lead to better generalization and robustness to image and label corruption</p>\n<p><strong>Training notebook:</strong> <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning</a><br>\n<strong>Inference notebook:</strong> <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference\" target=\"_blank\">https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference</a></p>\n<hr>\n<h3>Experiments summary</h3>\n<p><strong>What improved performance</strong><br>\nLarger batch size<br>\nAverage encoder size<br>\nData augmentation<br>\nLager image resolution<br>\nUsing pre-trained weights, I tried training from scratch but the results were very poor.<br>\nOversampling the minority classes from the dataset, gave a little help, but I expected more.</p>\n<hr>\n<p><strong>What made no difference</strong><br>\nTemperature parameter, according to the paper, lower temperature can benefit from longer training, since the data here is limited, tweaking this parameter maybe not worth it.<br>\nDifferent projection heads, I experimented a little with the number of neurons but got no relevant improvements.<br>\nDifferent classifier heads, I experimented a little with the number of neurons but got no relevant improvements.</p>\n<hr>\n<p><strong>Next steps</strong><br>\nDifferent pooling heads (AVG+MAX, Multi-sample, etc).<br>\nDifferent optimizers.<br>\nMore elaborated data augmentation schedules (MixUp, CutMix, etc).<br>\nUse external data.<br>\nIt may be interesting to try doing 2nd stage training but fine-tuning the complete network.</p>\n<hr>\n<p>An interesting way to evaluate the learned representation is to visualize the output of the trained embedding, at the <a href=\"https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning#Visualize-embeddings-outputs\" target=\"_blank\">training notebook</a> I plot the predictions from the model trained using <code>cross-entropy</code> and <code>supervised contrastive learning</code></p>\n<h4>Embedding trained with cross-entropy</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F95d2c3c2071b70bf5a78b842181b584d%2FScreenshot%20from%202020-12-24%2014-51-18.png?generation=1608832461429961&amp;alt=media\" alt=\"\"></p>\n<h4>Embedding trained with supervised contrastive learning</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F433c6150356e7a709efc5698cc90e0c6%2FScreenshot%20from%202020-12-24%2014-51-31.png?generation=1608832477794000&amp;alt=media\" alt=\"\"></p>\n<p>You can see that the representation from the model trained with <code>supervised contrastive learning</code> clusters much more than the model trained with <code>cross-entropy</code>.</p>",
      "rawMarkdown": "Since I saw the concept of contrastive learning I became very interested in the subject, and a few weeks ago I saw an implementation on the [Keras tutorials page](https://keras.io/examples/vision/supervised-contrastive-learning/), so I decided to try it at this competition, I have spent most of my TPU quota from the last two weeks experimenting with it, and I think it has some real potential, so I want to share my finding, maybe some people can share ideas, suggestions and further improve the results.\n\nUpdate: I wrote an article about this [Supervised Contrastive Learning for Cassava Leaf Disease Classification](https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966).\n\n## Supervised Contrastive Learning\n\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Cassava%20Leaf%20Disease%20Classification/sscl_vs_scl.png)\n\n#### In short, this is how it works\n> Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes.\n\nThe idea is that training models using the supervised contrastive learning can make the model encoder learn better class representation from the samples, this should lead to better generalization and robustness to image and label corruption\n\n\n**Training notebook:** https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning\n**Inference notebook:** https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference\n\n___\n\n### Experiments summary\n\n**What improved performance**\nLarger batch size\nAverage encoder size\nData augmentation\nLager image resolution\nUsing pre-trained weights, I tried training from scratch but the results were very poor.\nOversampling the minority classes from the dataset, gave a little help, but I expected more.\n\n___\n\n**What made no difference**\nTemperature parameter, according to the paper, lower temperature can benefit from longer training, since the data here is limited, tweaking this parameter maybe not worth it.\nDifferent projection heads, I experimented a little with the number of neurons but got no relevant improvements.\nDifferent classifier heads, I experimented a little with the number of neurons but got no relevant improvements.\n\n___\n\n**Next steps**\nDifferent pooling heads (AVG+MAX, Multi-sample, etc).\nDifferent optimizers.\nMore elaborated data augmentation schedules (MixUp, CutMix, etc).\nUse external data.\nIt may be interesting to try doing 2nd stage training but fine-tuning the complete network.\n\n___\n\nAn interesting way to evaluate the learned representation is to visualize the output of the trained embedding, at the [training notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning#Visualize-embeddings-outputs) I plot the predictions from the model trained using `cross-entropy` and `supervised contrastive learning`\n\n#### Embedding trained with cross-entropy\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F95d2c3c2071b70bf5a78b842181b584d%2FScreenshot%20from%202020-12-24%2014-51-18.png?generation=1608832461429961&alt=media)\n\n#### Embedding trained with supervised contrastive learning\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F433c6150356e7a709efc5698cc90e0c6%2FScreenshot%20from%202020-12-24%2014-51-31.png?generation=1608832477794000&alt=media)\n\nYou can see that the representation from the model trained with `supervised contrastive learning` clusters much more than the model trained with `cross-entropy`.",
      "votes": null
    },
    {
      "id": "1125445",
      "postDate": "12/24/2020 18:00:24",
      "content": "<p>About the model performance, the models trained with <code>supervised contrastive learning</code> consistently performed better (higher LB score) than the models trained with <code>cross-entropy</code>,  I ran out of TPU quota, but I am sure I could push it easily to <code>0.900</code> tweaking the parameters a little more.</p>",
      "rawMarkdown": "About the model performance, the models trained with `supervised contrastive learning` consistently performed better (higher LB score) than the models trained with `cross-entropy`,  I ran out of TPU quota, but I am sure I could push it easily to `0.900` tweaking the parameters a little more.",
      "votes": null
    },
    {
      "id": "1125617",
      "postDate": "12/24/2020 21:37:18",
      "content": "<p>very interesting</p>",
      "rawMarkdown": "very interesting",
      "votes": null
    },
    {
      "id": "1126216",
      "postDate": "12/25/2020 12:43:01",
      "content": "<p>The approach is definetly interesting but with one caveat. Our data has mislabelled examples (noisy labels) and I am afraid that Supervised Contrastive loss is not robust to noisy labels. It will be interesting to see results with just plain SimCLR(self supervision). </p>",
      "rawMarkdown": "The approach is definetly interesting but with one caveat. Our data has mislabelled examples (noisy labels) and I am afraid that Supervised Contrastive loss is not robust to noisy labels. It will be interesting to see results with just plain SimCLR(self supervision).",
      "votes": null
    },
    {
      "id": "1126758",
      "postDate": "12/25/2020 22:39:23",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/atharvap329\" target=\"_blank\">@atharvap329</a> , according to the paper, supervised contrastive learning performed better than cross-entropy on a benchmark that evaluate performance of different levels of corruption on imagenet, I am not sure the kind of corruption, but it might be  worth it to git it a try.</p>",
      "rawMarkdown": "Hi @atharvap329 , according to the paper, supervised contrastive learning performed better than cross-entropy on a benchmark that evaluate performance of different levels of corruption on imagenet, I am not sure the kind of corruption, but it might be  worth it to git it a try.",
      "votes": null
    },
    {
      "id": "1126818",
      "postDate": "12/26/2020 01:54:22",
      "content": "<p>They mention the corruption of images due to image images but its worth a try for sure it may happen that supervised contrastive loss may inherently be actually robust to noise(which will be a new finding :p). Looking forward to your results. </p>",
      "rawMarkdown": "They mention the corruption of images due to image images but its worth a try for sure it may happen that supervised contrastive loss may inherently be actually robust to noise(which will be a new finding :p). Looking forward to your results.",
      "votes": null
    },
    {
      "id": "1127635",
      "postDate": "12/26/2020 17:47:49",
      "content": "<p>Interesting. Need to be looked up :)</p>",
      "rawMarkdown": "Interesting. Need to be looked up :)",
      "votes": null
    },
    {
      "id": "1128194",
      "postDate": "12/27/2020 08:46:01",
      "content": "<p>Intersting. (Y)</p>",
      "rawMarkdown": "Intersting. (Y)",
      "votes": null
    },
    {
      "id": "1171609",
      "postDate": "01/27/2021 01:43:43",
      "content": "<p>For those interested in this method I just wrote <a href=\"https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966\" target=\"_blank\">an article about it</a>, using this competition's data as a use case</p>",
      "rawMarkdown": "For those interested in this method I just wrote [an article about it](https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966), using this competition's data as a use case",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1125445,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "12/24/2020 18:00:24",
      "content": "<p>About the model performance, the models trained with <code>supervised contrastive learning</code> consistently performed better (higher LB score) than the models trained with <code>cross-entropy</code>,  I ran out of TPU quota, but I am sure I could push it easily to <code>0.900</code> tweaking the parameters a little more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1125617,
      "author_name": "helenjwang",
      "author_url": "",
      "post_date": "12/24/2020 21:37:18",
      "content": "<p>very interesting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1126216,
      "author_name": "atharvap329",
      "author_url": "",
      "post_date": "12/25/2020 12:43:01",
      "content": "<p>The approach is definetly interesting but with one caveat. Our data has mislabelled examples (noisy labels) and I am afraid that Supervised Contrastive loss is not robust to noisy labels. It will be interesting to see results with just plain SimCLR(self supervision). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1126758,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "12/25/2020 22:39:23",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/atharvap329\" target=\"_blank\">@atharvap329</a> , according to the paper, supervised contrastive learning performed better than cross-entropy on a benchmark that evaluate performance of different levels of corruption on imagenet, I am not sure the kind of corruption, but it might be  worth it to git it a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126818,
          "author_name": "atharvap329",
          "author_url": "",
          "post_date": "12/26/2020 01:54:22",
          "content": "<p>They mention the corruption of images due to image images but its worth a try for sure it may happen that supervised contrastive loss may inherently be actually robust to noise(which will be a new finding :p). Looking forward to your results. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1127635,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "12/26/2020 17:47:49",
      "content": "<p>Interesting. Need to be looked up :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1128194,
      "author_name": "kiranchaudharyy",
      "author_url": "",
      "post_date": "12/27/2020 08:46:01",
      "content": "<p>Intersting. (Y)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1171609,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "01/27/2021 01:43:43",
      "content": "<p>For those interested in this method I just wrote <a href=\"https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966\" target=\"_blank\">an article about it</a>, using this competition's data as a use case</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1125436": "Since I saw the concept of contrastive learning I became very interested in the subject, and a few weeks ago I saw an implementation on the [Keras tutorials page](https://keras.io/examples/vision/supervised-contrastive-learning/), so I decided to try it at this competition, I have spent most of my TPU quota from the last two weeks experimenting with it, and I think it has some real potential, so I want to share my finding, maybe some people can share ideas, suggestions and further improve the results.\n\nUpdate: I wrote an article about this [Supervised Contrastive Learning for Cassava Leaf Disease Classification](https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966).\n\n## Supervised Contrastive Learning\n\n![](https://raw.githubusercontent.com/dimitreOliveira/MachineLearning/master/Kaggle/Cassava%20Leaf%20Disease%20Classification/sscl_vs_scl.png)\n\n#### In short, this is how it works\n> Clusters of points belonging to the same class are pulled together in embedding space, while simultaneously pushing apart clusters of samples from different classes.\n\nThe idea is that training models using the supervised contrastive learning can make the model encoder learn better class representation from the samples, this should lead to better generalization and robustness to image and label corruption\n\n\n**Training notebook:** https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning\n**Inference notebook:** https://www.kaggle.com/dimitreoliveira/cassava-supervised-contrastive-learning-inference\n\n___\n\n### Experiments summary\n\n**What improved performance**\nLarger batch size\nAverage encoder size\nData augmentation\nLager image resolution\nUsing pre-trained weights, I tried training from scratch but the results were very poor.\nOversampling the minority classes from the dataset, gave a little help, but I expected more.\n\n___\n\n**What made no difference**\nTemperature parameter, according to the paper, lower temperature can benefit from longer training, since the data here is limited, tweaking this parameter maybe not worth it.\nDifferent projection heads, I experimented a little with the number of neurons but got no relevant improvements.\nDifferent classifier heads, I experimented a little with the number of neurons but got no relevant improvements.\n\n___\n\n**Next steps**\nDifferent pooling heads (AVG+MAX, Multi-sample, etc).\nDifferent optimizers.\nMore elaborated data augmentation schedules (MixUp, CutMix, etc).\nUse external data.\nIt may be interesting to try doing 2nd stage training but fine-tuning the complete network.\n\n___\n\nAn interesting way to evaluate the learned representation is to visualize the output of the trained embedding, at the [training notebook](https://www.kaggle.com/dimitreoliveira/cassava-leaf-supervised-contrastive-learning#Visualize-embeddings-outputs) I plot the predictions from the model trained using `cross-entropy` and `supervised contrastive learning`\n\n#### Embedding trained with cross-entropy\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F95d2c3c2071b70bf5a78b842181b584d%2FScreenshot%20from%202020-12-24%2014-51-18.png?generation=1608832461429961&alt=media)\n\n#### Embedding trained with supervised contrastive learning\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F433c6150356e7a709efc5698cc90e0c6%2FScreenshot%20from%202020-12-24%2014-51-31.png?generation=1608832477794000&alt=media)\n\nYou can see that the representation from the model trained with `supervised contrastive learning` clusters much more than the model trained with `cross-entropy`.",
    "1125445": "About the model performance, the models trained with `supervised contrastive learning` consistently performed better (higher LB score) than the models trained with `cross-entropy`,  I ran out of TPU quota, but I am sure I could push it easily to `0.900` tweaking the parameters a little more.",
    "1125617": "very interesting",
    "1126216": "The approach is definetly interesting but with one caveat. Our data has mislabelled examples (noisy labels) and I am afraid that Supervised Contrastive loss is not robust to noisy labels. It will be interesting to see results with just plain SimCLR(self supervision).",
    "1126758": "Hi @atharvap329 , according to the paper, supervised contrastive learning performed better than cross-entropy on a benchmark that evaluate performance of different levels of corruption on imagenet, I am not sure the kind of corruption, but it might be  worth it to git it a try.",
    "1126818": "They mention the corruption of images due to image images but its worth a try for sure it may happen that supervised contrastive loss may inherently be actually robust to noise(which will be a new finding :p). Looking forward to your results.",
    "1127635": "Interesting. Need to be looked up :)",
    "1128194": "Intersting. (Y)",
    "1171609": "For those interested in this method I just wrote [an article about it](https://medium.com/towards-artificial-intelligence/supervised-contrastive-learning-for-cassava-leaf-disease-classification-9dd47779a966), using this competition's data as a use case"
  },
  "source": "meta"
}