{
  "id": 483299,
  "title": "Data Drift Caution",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/483299",
  "author_name": "",
  "post_date": "2024-03-11T20:36:01.658751600Z",
  "votes": 38,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I have seen quite a bit of discussion about the potential for data drift or not in this competition. Many seem to suggest that there will be very little to no drift, while others say that there will be significant drift. I just wanted to highlight the differences and why so those who may not follow can make a proper decision on what they want to do. Note: There is no way to be sure, so anyone with a secure feeling on this is probably not one you should follow blindly. </p>\n<p>What causes the drift? The drift in this comp is caused by the number of votes for each id. If the private LB has a change in the frequency of high number of votes vs low number of votes then LB will drift significantly. This can be seen as many 2 stage models do really well on the LB, but poorly on CV. While models doing well on CV do slightly worse on LB. </p>\n<p>Requirements for little/no drift: The distribution of high number of votes vs low number of votes needs to be basically identical to the public LB. </p>\n<p>Requirements for data drift (possibly more similar to training data): The distribution of high number of votes vs low number of votes needs to be similar to the training data. This would be the classic CV approach. This will not preform at the gold level (unless you know more than I do) on the public LB, but could do well in the private LB. </p>\n<p>My personal note: Don’t base your choices off what someone else will say or even the public LB, because it really means nothing in the race for the best private LB. I find that usually on Kaggle best CV is what comps are designed for, but there is certainly plenty of examples of caveats or little tricks! Make your own choices and stick to them! Good luck!</p>",
  "messages": [
    {
      "id": "2692390",
      "postDate": "03/11/2024 20:36:01",
      "content": "<p>I have seen quite a bit of discussion about the potential for data drift or not in this competition. Many seem to suggest that there will be very little to no drift, while others say that there will be significant drift. I just wanted to highlight the differences and why so those who may not follow can make a proper decision on what they want to do. Note: There is no way to be sure, so anyone with a secure feeling on this is probably not one you should follow blindly. </p>\n<p>What causes the drift? The drift in this comp is caused by the number of votes for each id. If the private LB has a change in the frequency of high number of votes vs low number of votes then LB will drift significantly. This can be seen as many 2 stage models do really well on the LB, but poorly on CV. While models doing well on CV do slightly worse on LB. </p>\n<p>Requirements for little/no drift: The distribution of high number of votes vs low number of votes needs to be basically identical to the public LB. </p>\n<p>Requirements for data drift (possibly more similar to training data): The distribution of high number of votes vs low number of votes needs to be similar to the training data. This would be the classic CV approach. This will not preform at the gold level (unless you know more than I do) on the public LB, but could do well in the private LB. </p>\n<p>My personal note: Don’t base your choices off what someone else will say or even the public LB, because it really means nothing in the race for the best private LB. I find that usually on Kaggle best CV is what comps are designed for, but there is certainly plenty of examples of caveats or little tricks! Make your own choices and stick to them! Good luck!</p>",
      "rawMarkdown": "I have seen quite a bit of discussion about the potential for data drift or not in this competition. Many seem to suggest that there will be very little to no drift, while others say that there will be significant drift. I just wanted to highlight the differences and why so those who may not follow can make a proper decision on what they want to do. Note: There is no way to be sure, so anyone with a secure feeling on this is probably not one you should follow blindly. \n\nWhat causes the drift? The drift in this comp is caused by the number of votes for each id. If the private LB has a change in the frequency of high number of votes vs low number of votes then LB will drift significantly. This can be seen as many 2 stage models do really well on the LB, but poorly on CV. While models doing well on CV do slightly worse on LB. \n\nRequirements for little/no drift: The distribution of high number of votes vs low number of votes needs to be basically identical to the public LB. \n\nRequirements for data drift (possibly more similar to training data): The distribution of high number of votes vs low number of votes needs to be similar to the training data. This would be the classic CV approach. This will not preform at the gold level (unless you know more than I do) on the public LB, but could do well in the private LB. \n\nMy personal note: Don’t base your choices off what someone else will say or even the public LB, because it really means nothing in the race for the best private LB. I find that usually on Kaggle best CV is what comps are designed for, but there is certainly plenty of examples of caveats or little tricks! Make your own choices and stick to them! Good luck!",
      "votes": null
    },
    {
      "id": "2692533",
      "postDate": "03/11/2024 22:43:45",
      "content": "<p>Thanks for bringing this topic <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>. Let me expand on a few things on validation and its importance.</p>\n<p>Firstly, having a good validation scheme is necessary to:</p>\n<ol>\n<li>Estimate the model's generalization performance.</li>\n<li>Increase model performance</li>\n<li>Model selection</li>\n</ol>\n<p>We always work (or try to work) under the assumption that the test distribution and the train distribution are the same. This is related to IID (independent identically distributed random variables). So, going back to data drift, when we have this issue what essentially happens is that this assumption is violated.</p>\n<p>Why data drift may happen in this competition? Because from the train data we know that the number of voters have 2 distributions. When you have a low number of voters your target variable becomes \"hard\", that is, the labels are composed of 1s and 0s and makes your model overconfident. On the other hand, when you have a high number of voters your target variable is more likely to become \"softer\"/\"smoother\".</p>\n<p>The \"gamble\" here is to know from which of the 2 distributions the test data ground truth comes from. If you assume test data comes from \"high number of voters\" then give more importance to this distribution when training. If you assume test data comes from \"low number of voters\" then give more importance to that distribution when training.</p>\n<p>In summary, it all comes down to making the train and test data as similar as possible so the model generalizes to unseen samples better. Also, the metric in this competition in KL-divergence which compares how similar 2 distributions are. If your model learns a different distribution at train stage compared to test stage you are likely to be in big trouble.</p>",
      "rawMarkdown": "Thanks for bringing this topic @cody11null. Let me expand on a few things on validation and its importance.\n\nFirstly, having a good validation scheme is necessary to:\n1. Estimate the model's generalization performance.\n2. Increase model performance\n3. Model selection\n\nWe always work (or try to work) under the assumption that the test distribution and the train distribution are the same. This is related to IID (independent identically distributed random variables). So, going back to data drift, when we have this issue what essentially happens is that this assumption is violated.\n\nWhy data drift may happen in this competition? Because from the train data we know that the number of voters have 2 distributions. When you have a low number of voters your target variable becomes \"hard\", that is, the labels are composed of 1s and 0s and makes your model overconfident. On the other hand, when you have a high number of voters your target variable is more likely to become \"softer\"/\"smoother\".\n\nThe \"gamble\" here is to know from which of the 2 distributions the test data ground truth comes from. If you assume test data comes from \"high number of voters\" then give more importance to this distribution when training. If you assume test data comes from \"low number of voters\" then give more importance to that distribution when training.\n\nIn summary, it all comes down to making the train and test data as similar as possible so the model generalizes to unseen samples better. Also, the metric in this competition in KL-divergence which compares how similar 2 distributions are. If your model learns a different distribution at train stage compared to test stage you are likely to be in big trouble.",
      "votes": null
    },
    {
      "id": "2692538",
      "postDate": "03/11/2024 22:57:06",
      "content": "<p>Great insight, thank you for adding more context! I am shocked with just how many seem so sure of the outcome. Truly once “the assumption” is broken it is anyone’s guess! </p>",
      "rawMarkdown": "Great insight, thank you for adding more context! I am shocked with just how many seem so sure of the outcome. Truly once “the assumption” is broken it is anyone’s guess!",
      "votes": null
    },
    {
      "id": "2692698",
      "postDate": "03/12/2024 02:16:38",
      "content": "<p>As I found when I did my <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">EDA</a>, not only are there two distributions of voters, but the labels distribution for the two are completely different.  </p>\n<p>For several years after I retired I did list my occupation as \"Professional Texas Holdem Gambler\".  Not sure how to play this hand.</p>",
      "rawMarkdown": "As I found when I did my [EDA](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda), not only are there two distributions of voters, but the labels distribution for the two are completely different.  \n\nFor several years after I retired I did list my occupation as \"Professional Texas Holdem Gambler\".  Not sure how to play this hand.",
      "votes": null
    },
    {
      "id": "2693340",
      "postDate": "03/12/2024 11:15:50",
      "content": "<p>Thank you for the wonderful topic</p>",
      "rawMarkdown": "Thank you for the wonderful topic",
      "votes": null
    },
    {
      "id": "2694985",
      "postDate": "03/13/2024 11:40:03",
      "content": "<p>Thanks for posting <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> ! Great insight</p>",
      "rawMarkdown": "Thanks for posting @cody11null ! Great insight",
      "votes": null
    },
    {
      "id": "2695004",
      "postDate": "03/13/2024 11:58:04",
      "content": "<p>Thanks for your help!</p>",
      "rawMarkdown": "Thanks for your help!",
      "votes": null
    },
    {
      "id": "2695205",
      "postDate": "03/13/2024 15:01:02",
      "content": "<p>Hi Cody, since we have the option to choose 2 models for the private LB the question for me is how i will select the 2 models, my current ideas to avoid overcommitting to one idea are:</p>\n<ol>\n<li>Select the best Public LB and the best CV model for the entire train distribution</li>\n<li>Select the best CV for data with little votes and the best CV data with many votes</li>\n<li>Select the best CV for data with little votes and the best Public LB</li>\n</ol>\n<p>the question here is, which of those 3 approaches is the best, or if there is even a better/safer option to select 2 models in order to avoid 'gambling' for the final placement</p>",
      "rawMarkdown": "Hi Cody, since we have the option to choose 2 models for the private LB the question for me is how i will select the 2 models, my current ideas to avoid overcommitting to one idea are:\n\n1. Select the best Public LB and the best CV model for the entire train distribution\n2. Select the best CV for data with little votes and the best CV data with many votes\n3. Select the best CV for data with little votes and the best Public LB\n\nthe question here is, which of those 3 approaches is the best, or if there is even a better/safer option to select 2 models in order to avoid 'gambling' for the final placement",
      "votes": null
    },
    {
      "id": "2695213",
      "postDate": "03/13/2024 15:07:58",
      "content": "<p>Another good way to put it! Having that debate myself. Not sure which way I will lean. Though I do have some thoughts! </p>",
      "rawMarkdown": "Another good way to put it! Having that debate myself. Not sure which way I will lean. Though I do have some thoughts!",
      "votes": null
    },
    {
      "id": "2698245",
      "postDate": "03/15/2024 11:16:46",
      "content": "<p>Super Nice Topic.<br>\nI speculate that the 2-stage learning, which relies on one distribution (Number of Vote) seen in some discussions, is risky. <br>\nThis is because, as described in this discussion, there is a possibility that the Private datasets could be from a different distribution.</p>",
      "rawMarkdown": "Super Nice Topic.\nI speculate that the 2-stage learning, which relies on one distribution (Number of Vote) seen in some discussions, is risky. \nThis is because, as described in this discussion, there is a possibility that the Private datasets could be from a different distribution.",
      "votes": null
    },
    {
      "id": "2701733",
      "postDate": "03/17/2024 07:49:37",
      "content": "<p>For the two stage training, I used the approach of splitting training data based on KL loss from the discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461\" target=\"_blank\">here</a><br>\nThe split is not based on number of votes, so the second stage would train on all voting groups.</p>\n<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> Let me know what you think about this approach?</p>",
      "rawMarkdown": "For the two stage training, I used the approach of splitting training data based on KL loss from the discussion [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461)\nThe split is not based on number of votes, so the second stage would train on all voting groups.\n\n@cody11null Let me know what you think about this approach?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2692533,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "03/11/2024 22:43:45",
      "content": "<p>Thanks for bringing this topic <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>. Let me expand on a few things on validation and its importance.</p>\n<p>Firstly, having a good validation scheme is necessary to:</p>\n<ol>\n<li>Estimate the model's generalization performance.</li>\n<li>Increase model performance</li>\n<li>Model selection</li>\n</ol>\n<p>We always work (or try to work) under the assumption that the test distribution and the train distribution are the same. This is related to IID (independent identically distributed random variables). So, going back to data drift, when we have this issue what essentially happens is that this assumption is violated.</p>\n<p>Why data drift may happen in this competition? Because from the train data we know that the number of voters have 2 distributions. When you have a low number of voters your target variable becomes \"hard\", that is, the labels are composed of 1s and 0s and makes your model overconfident. On the other hand, when you have a high number of voters your target variable is more likely to become \"softer\"/\"smoother\".</p>\n<p>The \"gamble\" here is to know from which of the 2 distributions the test data ground truth comes from. If you assume test data comes from \"high number of voters\" then give more importance to this distribution when training. If you assume test data comes from \"low number of voters\" then give more importance to that distribution when training.</p>\n<p>In summary, it all comes down to making the train and test data as similar as possible so the model generalizes to unseen samples better. Also, the metric in this competition in KL-divergence which compares how similar 2 distributions are. If your model learns a different distribution at train stage compared to test stage you are likely to be in big trouble.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2692538,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "03/11/2024 22:57:06",
          "content": "<p>Great insight, thank you for adding more context! I am shocked with just how many seem so sure of the outcome. Truly once “the assumption” is broken it is anyone’s guess! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2692698,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "03/12/2024 02:16:38",
          "content": "<p>As I found when I did my <a href=\"https://www.kaggle.com/code/pcjimmmy/patient-variation-eda\" target=\"_blank\">EDA</a>, not only are there two distributions of voters, but the labels distribution for the two are completely different.  </p>\n<p>For several years after I retired I did list my occupation as \"Professional Texas Holdem Gambler\".  Not sure how to play this hand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2693340,
          "author_name": "gentlezdh",
          "author_url": "",
          "post_date": "03/12/2024 11:15:50",
          "content": "<p>Thank you for the wonderful topic</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2694985,
      "author_name": "octaviograu",
      "author_url": "",
      "post_date": "03/13/2024 11:40:03",
      "content": "<p>Thanks for posting <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> ! Great insight</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2695004,
      "author_name": "mengzezhao312",
      "author_url": "",
      "post_date": "03/13/2024 11:58:04",
      "content": "<p>Thanks for your help!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2695205,
      "author_name": "maxuhl98",
      "author_url": "",
      "post_date": "03/13/2024 15:01:02",
      "content": "<p>Hi Cody, since we have the option to choose 2 models for the private LB the question for me is how i will select the 2 models, my current ideas to avoid overcommitting to one idea are:</p>\n<ol>\n<li>Select the best Public LB and the best CV model for the entire train distribution</li>\n<li>Select the best CV for data with little votes and the best CV data with many votes</li>\n<li>Select the best CV for data with little votes and the best Public LB</li>\n</ol>\n<p>the question here is, which of those 3 approaches is the best, or if there is even a better/safer option to select 2 models in order to avoid 'gambling' for the final placement</p>",
      "votes": null,
      "replies": [
        {
          "id": 2695213,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "03/13/2024 15:07:58",
          "content": "<p>Another good way to put it! Having that debate myself. Not sure which way I will lean. Though I do have some thoughts! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2698245,
      "author_name": "haruki741",
      "author_url": "",
      "post_date": "03/15/2024 11:16:46",
      "content": "<p>Super Nice Topic.<br>\nI speculate that the 2-stage learning, which relies on one distribution (Number of Vote) seen in some discussions, is risky. <br>\nThis is because, as described in this discussion, there is a possibility that the Private datasets could be from a different distribution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2701733,
      "author_name": "nartaa",
      "author_url": "",
      "post_date": "03/17/2024 07:49:37",
      "content": "<p>For the two stage training, I used the approach of splitting training data based on KL loss from the discussion <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461\" target=\"_blank\">here</a><br>\nThe split is not based on number of votes, so the second stage would train on all voting groups.</p>\n<p><a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> Let me know what you think about this approach?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2692390": "I have seen quite a bit of discussion about the potential for data drift or not in this competition. Many seem to suggest that there will be very little to no drift, while others say that there will be significant drift. I just wanted to highlight the differences and why so those who may not follow can make a proper decision on what they want to do. Note: There is no way to be sure, so anyone with a secure feeling on this is probably not one you should follow blindly. \n\nWhat causes the drift? The drift in this comp is caused by the number of votes for each id. If the private LB has a change in the frequency of high number of votes vs low number of votes then LB will drift significantly. This can be seen as many 2 stage models do really well on the LB, but poorly on CV. While models doing well on CV do slightly worse on LB. \n\nRequirements for little/no drift: The distribution of high number of votes vs low number of votes needs to be basically identical to the public LB. \n\nRequirements for data drift (possibly more similar to training data): The distribution of high number of votes vs low number of votes needs to be similar to the training data. This would be the classic CV approach. This will not preform at the gold level (unless you know more than I do) on the public LB, but could do well in the private LB. \n\nMy personal note: Don’t base your choices off what someone else will say or even the public LB, because it really means nothing in the race for the best private LB. I find that usually on Kaggle best CV is what comps are designed for, but there is certainly plenty of examples of caveats or little tricks! Make your own choices and stick to them! Good luck!",
    "2692533": "Thanks for bringing this topic @cody11null. Let me expand on a few things on validation and its importance.\n\nFirstly, having a good validation scheme is necessary to:\n1. Estimate the model's generalization performance.\n2. Increase model performance\n3. Model selection\n\nWe always work (or try to work) under the assumption that the test distribution and the train distribution are the same. This is related to IID (independent identically distributed random variables). So, going back to data drift, when we have this issue what essentially happens is that this assumption is violated.\n\nWhy data drift may happen in this competition? Because from the train data we know that the number of voters have 2 distributions. When you have a low number of voters your target variable becomes \"hard\", that is, the labels are composed of 1s and 0s and makes your model overconfident. On the other hand, when you have a high number of voters your target variable is more likely to become \"softer\"/\"smoother\".\n\nThe \"gamble\" here is to know from which of the 2 distributions the test data ground truth comes from. If you assume test data comes from \"high number of voters\" then give more importance to this distribution when training. If you assume test data comes from \"low number of voters\" then give more importance to that distribution when training.\n\nIn summary, it all comes down to making the train and test data as similar as possible so the model generalizes to unseen samples better. Also, the metric in this competition in KL-divergence which compares how similar 2 distributions are. If your model learns a different distribution at train stage compared to test stage you are likely to be in big trouble.",
    "2692538": "Great insight, thank you for adding more context! I am shocked with just how many seem so sure of the outcome. Truly once “the assumption” is broken it is anyone’s guess!",
    "2692698": "As I found when I did my [EDA](https://www.kaggle.com/code/pcjimmmy/patient-variation-eda), not only are there two distributions of voters, but the labels distribution for the two are completely different.  \n\nFor several years after I retired I did list my occupation as \"Professional Texas Holdem Gambler\".  Not sure how to play this hand.",
    "2693340": "Thank you for the wonderful topic",
    "2694985": "Thanks for posting @cody11null ! Great insight",
    "2695004": "Thanks for your help!",
    "2695205": "Hi Cody, since we have the option to choose 2 models for the private LB the question for me is how i will select the 2 models, my current ideas to avoid overcommitting to one idea are:\n\n1. Select the best Public LB and the best CV model for the entire train distribution\n2. Select the best CV for data with little votes and the best CV data with many votes\n3. Select the best CV for data with little votes and the best Public LB\n\nthe question here is, which of those 3 approaches is the best, or if there is even a better/safer option to select 2 models in order to avoid 'gambling' for the final placement",
    "2695213": "Another good way to put it! Having that debate myself. Not sure which way I will lean. Though I do have some thoughts!",
    "2698245": "Super Nice Topic.\nI speculate that the 2-stage learning, which relies on one distribution (Number of Vote) seen in some discussions, is risky. \nThis is because, as described in this discussion, there is a possibility that the Private datasets could be from a different distribution.",
    "2701733": "For the two stage training, I used the approach of splitting training data based on KL loss from the discussion [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477461)\nThe split is not based on number of votes, so the second stage would train on all voting groups.\n\n@cody11null Let me know what you think about this approach?"
  },
  "source": "meta"
}