{
  "id": 307308,
  "title": "Class Weighting in Supervised Contrastive Learning?",
  "url": "/competitions/happy-whale-and-dolphin/discussion/307308",
  "author_name": "",
  "post_date": "2022-02-13T18:09:16.948312300Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all, I had a question and maybe you could help.</p>\n<p>Since many of us are utilizing supervised contrastive learning techniques like ArcFace/Loss for this challenge, I was curious if anyone had any insight into whether class weighting can be effectively leveraged.</p>\n<p>My intuition would be that you would need to flatten the weighting maybe? </p>\n<hr>\n<p>Also, I assume (but would look for confirmation) that individuals repeatedly occur within the test data (similar to the training data). </p>\n<p>i.e. Jeff the whale may appear multiple times within the test dataset and may have appeared multiple times within the training dataset.</p>\n<p>However, it is unclear if the same individual distribution would be maintained across from train to test.</p>\n<p>i.e. Jeff appears 80 times in the training data but only appears 2 times in the test dataset whereas Frank appears just once in the training data yet appears 11 times in the test dataset.</p>\n<p>If this is possible, and especially if it is common, I think that class weighting could be very helpful.</p>\n<hr>\n<p>I'll be running an experiment soon, but it obviously is very time-intensive to run good experiments. If anyone has any insights I would be very appreciative.</p>\n<p>Thank you in advance!</p>\n<hr>\n<blockquote>\n  <p>[UPDATE]</p>\n  <p>I read the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632\" target=\"_blank\">6th Place Solution to the Google Landmark Recognition Challenge</a> and it appears they leveraged class weights.</p>\n  <p>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).</p>\n</blockquote>",
  "messages": [
    {
      "id": "1688600",
      "postDate": "02/13/2022 18:09:16",
      "content": "<p>Hi all, I had a question and maybe you could help.</p>\n<p>Since many of us are utilizing supervised contrastive learning techniques like ArcFace/Loss for this challenge, I was curious if anyone had any insight into whether class weighting can be effectively leveraged.</p>\n<p>My intuition would be that you would need to flatten the weighting maybe? </p>\n<hr>\n<p>Also, I assume (but would look for confirmation) that individuals repeatedly occur within the test data (similar to the training data). </p>\n<p>i.e. Jeff the whale may appear multiple times within the test dataset and may have appeared multiple times within the training dataset.</p>\n<p>However, it is unclear if the same individual distribution would be maintained across from train to test.</p>\n<p>i.e. Jeff appears 80 times in the training data but only appears 2 times in the test dataset whereas Frank appears just once in the training data yet appears 11 times in the test dataset.</p>\n<p>If this is possible, and especially if it is common, I think that class weighting could be very helpful.</p>\n<hr>\n<p>I'll be running an experiment soon, but it obviously is very time-intensive to run good experiments. If anyone has any insights I would be very appreciative.</p>\n<p>Thank you in advance!</p>\n<hr>\n<blockquote>\n  <p>[UPDATE]</p>\n  <p>I read the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632\" target=\"_blank\">6th Place Solution to the Google Landmark Recognition Challenge</a> and it appears they leveraged class weights.</p>\n  <p>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).</p>\n</blockquote>",
      "rawMarkdown": "Hi all, I had a question and maybe you could help.\n\nSince many of us are utilizing supervised contrastive learning techniques like ArcFace/Loss for this challenge, I was curious if anyone had any insight into whether class weighting can be effectively leveraged.\n\nMy intuition would be that you would need to flatten the weighting maybe? \n\n---\n\nAlso, I assume (but would look for confirmation) that individuals repeatedly occur within the test data (similar to the training data). \n\ni.e. Jeff the whale may appear multiple times within the test dataset and may have appeared multiple times within the training dataset.\n\nHowever, it is unclear if the same individual distribution would be maintained across from train to test.\n\ni.e. Jeff appears 80 times in the training data but only appears 2 times in the test dataset whereas Frank appears just once in the training data yet appears 11 times in the test dataset.\n\nIf this is possible, and especially if it is common, I think that class weighting could be very helpful.\n\n---\n\nI'll be running an experiment soon, but it obviously is very time-intensive to run good experiments. If anyone has any insights I would be very appreciative.\n\nThank you in advance!\n\n---\n\n> [UPDATE]\n>\n> I read the [6th Place Solution to the Google Landmark Recognition Challenge](https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632) and it appears they leveraged class weights.\n>\n>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).",
      "votes": null
    },
    {
      "id": "1689085",
      "postDate": "02/14/2022 02:58:47",
      "content": "<p>I have been reading a bit before I throw a model at it. Mayi ask what \"flattening\" the weights (I assume class weights) mean in this context.</p>\n<p>Nice point I have been wondering the same. I have read that identify and weeding out the duplicates is very important for good results. Folks seem to have done that in the flukes competition</p>",
      "rawMarkdown": "I have been reading a bit before I throw a model at it. Mayi ask what \"flattening\" the weights (I assume class weights) mean in this context.\n\nNice point I have been wondering the same. I have read that identify and weeding out the duplicates is very important for good results. Folks seem to have done that in the flukes competition",
      "votes": null
    },
    {
      "id": "1689100",
      "postDate": "02/14/2022 03:21:37",
      "content": "<p>EDIT: I had incorrectly understood, please see below for correct explaination</p>",
      "rawMarkdown": "EDIT: I had incorrectly understood, please see below for correct explaination",
      "votes": null
    },
    {
      "id": "1689116",
      "postDate": "02/14/2022 03:47:18",
      "content": "<p>I was referring to using an exponential on the weight term to push it closer to 1. </p>\n<p>So you have a dictionary of classes as keys and class weights as values… but instead of just using min_count/count as the weight you use (min_count/count)^some_exp</p>\n<p>The smaller the value for some_exp the more the weights will be shifted or flattened towards a consistent weighting of 1.</p>\n<p>Not sure if there’s a name for this, but I do it sometimes. Hope that makes sense.</p>",
      "rawMarkdown": "I was referring to using an exponential on the weight term to push it closer to 1. \n\nSo you have a dictionary of classes as keys and class weights as values… but instead of just using min_count/count as the weight you use (min_count/count)^some_exp\n\nThe smaller the value for some_exp the more the weights will be shifted or flattened towards a consistent weighting of 1.\n\nNot sure if there’s a name for this, but I do it sometimes. Hope that makes sense.",
      "votes": null
    },
    {
      "id": "1689117",
      "postDate": "02/14/2022 03:47:41",
      "content": "<p>If it still doesn’t make sense I’ll make a little notebook and show what I mean.</p>",
      "rawMarkdown": "If it still doesn’t make sense I’ll make a little notebook and show what I mean.",
      "votes": null
    },
    {
      "id": "1689158",
      "postDate": "02/14/2022 04:17:32",
      "content": "<p>Thanks for the explanation <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>. that's very interesting, I think I got a little idea of it. Thanks!</p>",
      "rawMarkdown": "Thanks for the explanation @dschettler8845. that's very interesting, I think I got a little idea of it. Thanks!",
      "votes": null
    },
    {
      "id": "1689181",
      "postDate": "02/14/2022 04:30:01",
      "content": "<p>Yes, that could be a great resource, and you could probably give it a name too while you are at it <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> 😅 I am not sure if many are aware of such techniques, at least I am not. Thanks</p>",
      "rawMarkdown": "Yes, that could be a great resource, and you could probably give it a name too while you are at it @dschettler8845 😅 I am not sure if many are aware of such techniques, at least I am not. Thanks",
      "votes": null
    },
    {
      "id": "1690113",
      "postDate": "02/14/2022 17:50:17",
      "content": "<p>[UPDATE]</p>\n<p>I read the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632\" target=\"_blank\">6th Place Solution to the Google Landmark Recognition Challenge</a> and it appears they leveraged class weights.</p>\n<p>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).</p>",
      "rawMarkdown": "[UPDATE]\n\nI read the [6th Place Solution to the Google Landmark Recognition Challenge](https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632) and it appears they leveraged class weights.\n\nThis gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1689085,
      "author_name": "bsridatta",
      "author_url": "",
      "post_date": "02/14/2022 02:58:47",
      "content": "<p>I have been reading a bit before I throw a model at it. Mayi ask what \"flattening\" the weights (I assume class weights) mean in this context.</p>\n<p>Nice point I have been wondering the same. I have read that identify and weeding out the duplicates is very important for good results. Folks seem to have done that in the flukes competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 1689100,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/14/2022 03:21:37",
          "content": "<p>EDIT: I had incorrectly understood, please see below for correct explaination</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689116,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "02/14/2022 03:47:18",
          "content": "<p>I was referring to using an exponential on the weight term to push it closer to 1. </p>\n<p>So you have a dictionary of classes as keys and class weights as values… but instead of just using min_count/count as the weight you use (min_count/count)^some_exp</p>\n<p>The smaller the value for some_exp the more the weights will be shifted or flattened towards a consistent weighting of 1.</p>\n<p>Not sure if there’s a name for this, but I do it sometimes. Hope that makes sense.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689117,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "02/14/2022 03:47:41",
          "content": "<p>If it still doesn’t make sense I’ll make a little notebook and show what I mean.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689158,
          "author_name": "bsridatta",
          "author_url": "",
          "post_date": "02/14/2022 04:17:32",
          "content": "<p>Thanks for the explanation <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>. that's very interesting, I think I got a little idea of it. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1689181,
          "author_name": "bsridatta",
          "author_url": "",
          "post_date": "02/14/2022 04:30:01",
          "content": "<p>Yes, that could be a great resource, and you could probably give it a name too while you are at it <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> 😅 I am not sure if many are aware of such techniques, at least I am not. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1690113,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "02/14/2022 17:50:17",
      "content": "<p>[UPDATE]</p>\n<p>I read the <a href=\"https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632\" target=\"_blank\">6th Place Solution to the Google Landmark Recognition Challenge</a> and it appears they leveraged class weights.</p>\n<p>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1688600": "Hi all, I had a question and maybe you could help.\n\nSince many of us are utilizing supervised contrastive learning techniques like ArcFace/Loss for this challenge, I was curious if anyone had any insight into whether class weighting can be effectively leveraged.\n\nMy intuition would be that you would need to flatten the weighting maybe? \n\n---\n\nAlso, I assume (but would look for confirmation) that individuals repeatedly occur within the test data (similar to the training data). \n\ni.e. Jeff the whale may appear multiple times within the test dataset and may have appeared multiple times within the training dataset.\n\nHowever, it is unclear if the same individual distribution would be maintained across from train to test.\n\ni.e. Jeff appears 80 times in the training data but only appears 2 times in the test dataset whereas Frank appears just once in the training data yet appears 11 times in the test dataset.\n\nIf this is possible, and especially if it is common, I think that class weighting could be very helpful.\n\n---\n\nI'll be running an experiment soon, but it obviously is very time-intensive to run good experiments. If anyone has any insights I would be very appreciative.\n\nThank you in advance!\n\n---\n\n> [UPDATE]\n>\n> I read the [6th Place Solution to the Google Landmark Recognition Challenge](https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632) and it appears they leveraged class weights.\n>\n>This gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset).",
    "1689085": "I have been reading a bit before I throw a model at it. Mayi ask what \"flattening\" the weights (I assume class weights) mean in this context.\n\nNice point I have been wondering the same. I have read that identify and weeding out the duplicates is very important for good results. Folks seem to have done that in the flukes competition",
    "1689100": "EDIT: I had incorrectly understood, please see below for correct explaination",
    "1689116": "I was referring to using an exponential on the weight term to push it closer to 1. \n\nSo you have a dictionary of classes as keys and class weights as values… but instead of just using min_count/count as the weight you use (min_count/count)^some_exp\n\nThe smaller the value for some_exp the more the weights will be shifted or flattened towards a consistent weighting of 1.\n\nNot sure if there’s a name for this, but I do it sometimes. Hope that makes sense.",
    "1689117": "If it still doesn’t make sense I’ll make a little notebook and show what I mean.",
    "1689158": "Thanks for the explanation @dschettler8845. that's very interesting, I think I got a little idea of it. Thanks!",
    "1689181": "Yes, that could be a great resource, and you could probably give it a name too while you are at it @dschettler8845 😅 I am not sure if many are aware of such techniques, at least I am not. Thanks",
    "1690113": "[UPDATE]\n\nI read the [6th Place Solution to the Google Landmark Recognition Challenge](https://www.kaggle.com/c/landmark-retrieval-2021/discussion/276632) and it appears they leveraged class weights.\n\nThis gives me confidence that this will be required in this competition as well. Especially considering the manually altered distribution present in the test dataset (like the most common individual in the train dataset not being found in the test dataset)."
  },
  "source": "meta"
}