{
  "id": 477498,
  "title": "Bridging the Gap between Cross-Validation and Leaderboard Scores [CV 0.44 – LB 0.47]",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/477498",
  "author_name": "",
  "post_date": "2024-02-16T12:15:25.882953700Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Greetings, Everyone,</p>\n<p><code>train[TARGETS] = K*train[TARGETS] + 0.166666667 # K=1</code></p>\n<p>Introducing a single line of code has notably reduced the Cross-Validation (CV) score, aligning it more closely with the Leaderboard (LB) scores. It's worth noting that the LB score itself remains unchanged. This observation suggests that during model initialization, a uniform probability distribution were predicted by the model at the begining, because it predicts classes poorly but equally, then eventually the model reduces the loss and learns the probability distributions of the classes.</p>\n<p>By incorporating an offset of 0.16667 (1/6) across all classes at the target probability distributions, the CV score decreases, approaching the learned distribution. The lack of change in the LB score potentially supports the hypothesis that this adjustment brings the learned probability closer to the target probability distribution. Consequently, the gap diminishes between the training loss and validation loss, while the underlying learned probability remains consistent. (further should be investigated)</p>\n<p>Here, the second row is the probability distribution of the target labels with the offset<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F6b9d655b354f3550f0123e512822a262%2Fhms.png?generation=1708084757579739&amp;alt=media\"></p>\n<p><em>Side Note: Adjusting the Coefficient K &lt; 1 for Voted Weighting: Possibly optimizing Model Confidence for Improved Leaderboard Scores</em><br>\nLet me know your thoughts.</p>",
  "messages": [
    {
      "id": "2654758",
      "postDate": "02/16/2024 12:15:25",
      "content": "<p>Greetings, Everyone,</p>\n<p><code>train[TARGETS] = K*train[TARGETS] + 0.166666667 # K=1</code></p>\n<p>Introducing a single line of code has notably reduced the Cross-Validation (CV) score, aligning it more closely with the Leaderboard (LB) scores. It's worth noting that the LB score itself remains unchanged. This observation suggests that during model initialization, a uniform probability distribution were predicted by the model at the begining, because it predicts classes poorly but equally, then eventually the model reduces the loss and learns the probability distributions of the classes.</p>\n<p>By incorporating an offset of 0.16667 (1/6) across all classes at the target probability distributions, the CV score decreases, approaching the learned distribution. The lack of change in the LB score potentially supports the hypothesis that this adjustment brings the learned probability closer to the target probability distribution. Consequently, the gap diminishes between the training loss and validation loss, while the underlying learned probability remains consistent. (further should be investigated)</p>\n<p>Here, the second row is the probability distribution of the target labels with the offset<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F6b9d655b354f3550f0123e512822a262%2Fhms.png?generation=1708084757579739&amp;alt=media\"></p>\n<p><em>Side Note: Adjusting the Coefficient K &lt; 1 for Voted Weighting: Possibly optimizing Model Confidence for Improved Leaderboard Scores</em><br>\nLet me know your thoughts.</p>",
      "rawMarkdown": "Greetings, Everyone,\n\n`train[TARGETS] = K*train[TARGETS] + 0.166666667 # K=1`\n\nIntroducing a single line of code has notably reduced the Cross-Validation (CV) score, aligning it more closely with the Leaderboard (LB) scores. It's worth noting that the LB score itself remains unchanged. This observation suggests that during model initialization, a uniform probability distribution were predicted by the model at the begining, because it predicts classes poorly but equally, then eventually the model reduces the loss and learns the probability distributions of the classes.\n\nBy incorporating an offset of 0.16667 (1/6) across all classes at the target probability distributions, the CV score decreases, approaching the learned distribution. The lack of change in the LB score potentially supports the hypothesis that this adjustment brings the learned probability closer to the target probability distribution. Consequently, the gap diminishes between the training loss and validation loss, while the underlying learned probability remains consistent. (further should be investigated)\n\nHere, the second row is the probability distribution of the target labels with the offset\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F6b9d655b354f3550f0123e512822a262%2Fhms.png?generation=1708084757579739&alt=media)\n\n*Side Note: Adjusting the Coefficient K < 1 for Voted Weighting: Possibly optimizing Model Confidence for Improved Leaderboard Scores*\nLet me know your thoughts.",
      "votes": null
    },
    {
      "id": "2654911",
      "postDate": "02/16/2024 14:50:33",
      "content": "<p>When you change this , do all the labels summation  for each row still remain 1 ?</p>",
      "rawMarkdown": "When you change this , do all the labels summation  for each row still remain 1 ?",
      "votes": null
    },
    {
      "id": "2654951",
      "postDate": "02/16/2024 15:37:38",
      "content": "<p>Yes, The summation of each row still remains 1.</p>\n<p>train[TARGETS] = 1*train[TARGETS] + 0.166666667<br>\ntrain[TARGETS] = train[TARGETS]/train[TARGETS].values.sum(axis=1,keepdims=True)</p>",
      "rawMarkdown": "Yes, The summation of each row still remains 1.\n\ntrain[TARGETS] = 1*train[TARGETS] + 0.166666667\ntrain[TARGETS] = train[TARGETS]/train[TARGETS].values.sum(axis=1,keepdims=True)",
      "votes": null
    },
    {
      "id": "2655490",
      "postDate": "02/16/2024 23:22:47",
      "content": "<p>just to confirm, the models are trained with the new targets right ? And not that just cv is recalculated with new targets. Although it might be worth looking at the second case too</p>",
      "rawMarkdown": "just to confirm, the models are trained with the new targets right ? And not that just cv is recalculated with new targets. Although it might be worth looking at the second case too",
      "votes": null
    },
    {
      "id": "2656010",
      "postDate": "02/17/2024 11:44:53",
      "content": "<p>Yes, the models are trained with the new targets. But I think the second case should give consistent results).</p>\n<p>Basically, the targets are started off with a uniform distribution i.e. ([0.16, 0.16, 0.16, 0.16, 0.16, 0.16]), then this distribution is reshaped with the provided votes for all classes, we can control how much the votes effect the distribution with the coefficient K.</p>",
      "rawMarkdown": "Yes, the models are trained with the new targets. But I think the second case should give consistent results).\n\nBasically, the targets are started off with a uniform distribution i.e. ([0.16, 0.16, 0.16, 0.16, 0.16, 0.16]), then this distribution is reshaped with the provided votes for all classes, we can control how much the votes effect the distribution with the coefficient K.",
      "votes": null
    },
    {
      "id": "2656020",
      "postDate": "02/17/2024 11:50:39",
      "content": "<p>What would you arbitrarily change the labels of validation set? The change in score should be reflected in the original validation, otherwise you're going to end up with unreliable results.</p>",
      "rawMarkdown": "What would you arbitrarily change the labels of validation set? The change in score should be reflected in the original validation, otherwise you're going to end up with unreliable results.",
      "votes": null
    },
    {
      "id": "2656043",
      "postDate": "02/17/2024 12:05:14",
      "content": "<p>The target labels are processed the same for all the training set once grouped by eeg_id (Chris's approach), that also includes validation sets. Let me know if I didn't answer your question</p>",
      "rawMarkdown": "The target labels are processed the same for all the training set once grouped by eeg_id (Chris's approach), that also includes validation sets. Let me know if I didn't answer your question",
      "votes": null
    },
    {
      "id": "2656046",
      "postDate": "02/17/2024 12:15:22",
      "content": "<p>just keep in mind that in order to have reliable CV results you need to do this after the train/val split only in train targets and keep valid untouched. So CV score is calculated on the original targets </p>\n<p>Think that in test time (submission) you won't have the test[TARGETS] to apply the same transformation.</p>\n<p>I'm guessing this is what also other guys above are trying to say. </p>",
      "rawMarkdown": "just keep in mind that in order to have reliable CV results you need to do this after the train/val split only in train targets and keep valid untouched. So CV score is calculated on the original targets \n\nThink that in test time (submission) you won't have the test[TARGETS] to apply the same transformation.\n\nI'm guessing this is what also other guys above are trying to say.",
      "votes": null
    },
    {
      "id": "2656140",
      "postDate": "02/17/2024 13:43:30",
      "content": "<p>I think this can go either way, I am seeing metric correlation either way.<br>\nIf we operate on the original targets for CV, the scores will be more familliar to us, but that will add additional complexity to the code (we would have to account for that in the data generator).</p>",
      "rawMarkdown": "I think this can go either way, I am seeing metric correlation either way.\nIf we operate on the original targets for CV, the scores will be more familliar to us, but that will add additional complexity to the code (we would have to account for that in the data generator).",
      "votes": null
    },
    {
      "id": "2656775",
      "postDate": "02/18/2024 02:27:04",
      "content": "<p>I agree with slime. I also think that more noise tends to less total votes for each patient, because it may be hard to diagose/predict and we should keep the original targets when calculating cv. Noise is also important features for these hard samples. keep hard samples/edge cases hard/edge, and don't change them.</p>",
      "rawMarkdown": "I agree with slime. I also think that more noise tends to less total votes for each patient, because it may be hard to diagose/predict and we should keep the original targets when calculating cv. Noise is also important features for these hard samples. keep hard samples/edge cases hard/edge, and don't change them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2654911,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "02/16/2024 14:50:33",
      "content": "<p>When you change this , do all the labels summation  for each row still remain 1 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2654951,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "02/16/2024 15:37:38",
          "content": "<p>Yes, The summation of each row still remains 1.</p>\n<p>train[TARGETS] = 1*train[TARGETS] + 0.166666667<br>\ntrain[TARGETS] = train[TARGETS]/train[TARGETS].values.sum(axis=1,keepdims=True)</p>",
          "votes": null,
          "replies": [
            {
              "id": 2656046,
              "author_name": "imeintanis",
              "author_url": "",
              "post_date": "02/17/2024 12:15:22",
              "content": "<p>just keep in mind that in order to have reliable CV results you need to do this after the train/val split only in train targets and keep valid untouched. So CV score is calculated on the original targets </p>\n<p>Think that in test time (submission) you won't have the test[TARGETS] to apply the same transformation.</p>\n<p>I'm guessing this is what also other guys above are trying to say. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2656140,
                  "author_name": "nartaa",
                  "author_url": "",
                  "post_date": "02/17/2024 13:43:30",
                  "content": "<p>I think this can go either way, I am seeing metric correlation either way.<br>\nIf we operate on the original targets for CV, the scores will be more familliar to us, but that will add additional complexity to the code (we would have to account for that in the data generator).</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2655490,
      "author_name": "nikhilmishradev",
      "author_url": "",
      "post_date": "02/16/2024 23:22:47",
      "content": "<p>just to confirm, the models are trained with the new targets right ? And not that just cv is recalculated with new targets. Although it might be worth looking at the second case too</p>",
      "votes": null,
      "replies": [
        {
          "id": 2656010,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "02/17/2024 11:44:53",
          "content": "<p>Yes, the models are trained with the new targets. But I think the second case should give consistent results).</p>\n<p>Basically, the targets are started off with a uniform distribution i.e. ([0.16, 0.16, 0.16, 0.16, 0.16, 0.16]), then this distribution is reshaped with the provided votes for all classes, we can control how much the votes effect the distribution with the coefficient K.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2656020,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "02/17/2024 11:50:39",
      "content": "<p>What would you arbitrarily change the labels of validation set? The change in score should be reflected in the original validation, otherwise you're going to end up with unreliable results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2656043,
          "author_name": "nartaa",
          "author_url": "",
          "post_date": "02/17/2024 12:05:14",
          "content": "<p>The target labels are processed the same for all the training set once grouped by eeg_id (Chris's approach), that also includes validation sets. Let me know if I didn't answer your question</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2656775,
          "author_name": "sweetyheehee",
          "author_url": "",
          "post_date": "02/18/2024 02:27:04",
          "content": "<p>I agree with slime. I also think that more noise tends to less total votes for each patient, because it may be hard to diagose/predict and we should keep the original targets when calculating cv. Noise is also important features for these hard samples. keep hard samples/edge cases hard/edge, and don't change them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2654758": "Greetings, Everyone,\n\n`train[TARGETS] = K*train[TARGETS] + 0.166666667 # K=1`\n\nIntroducing a single line of code has notably reduced the Cross-Validation (CV) score, aligning it more closely with the Leaderboard (LB) scores. It's worth noting that the LB score itself remains unchanged. This observation suggests that during model initialization, a uniform probability distribution were predicted by the model at the begining, because it predicts classes poorly but equally, then eventually the model reduces the loss and learns the probability distributions of the classes.\n\nBy incorporating an offset of 0.16667 (1/6) across all classes at the target probability distributions, the CV score decreases, approaching the learned distribution. The lack of change in the LB score potentially supports the hypothesis that this adjustment brings the learned probability closer to the target probability distribution. Consequently, the gap diminishes between the training loss and validation loss, while the underlying learned probability remains consistent. (further should be investigated)\n\nHere, the second row is the probability distribution of the target labels with the offset\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4495635%2F6b9d655b354f3550f0123e512822a262%2Fhms.png?generation=1708084757579739&alt=media)\n\n*Side Note: Adjusting the Coefficient K < 1 for Voted Weighting: Possibly optimizing Model Confidence for Improved Leaderboard Scores*\nLet me know your thoughts.",
    "2654911": "When you change this , do all the labels summation  for each row still remain 1 ?",
    "2654951": "Yes, The summation of each row still remains 1.\n\ntrain[TARGETS] = 1*train[TARGETS] + 0.166666667\ntrain[TARGETS] = train[TARGETS]/train[TARGETS].values.sum(axis=1,keepdims=True)",
    "2655490": "just to confirm, the models are trained with the new targets right ? And not that just cv is recalculated with new targets. Although it might be worth looking at the second case too",
    "2656010": "Yes, the models are trained with the new targets. But I think the second case should give consistent results).\n\nBasically, the targets are started off with a uniform distribution i.e. ([0.16, 0.16, 0.16, 0.16, 0.16, 0.16]), then this distribution is reshaped with the provided votes for all classes, we can control how much the votes effect the distribution with the coefficient K.",
    "2656020": "What would you arbitrarily change the labels of validation set? The change in score should be reflected in the original validation, otherwise you're going to end up with unreliable results.",
    "2656043": "The target labels are processed the same for all the training set once grouped by eeg_id (Chris's approach), that also includes validation sets. Let me know if I didn't answer your question",
    "2656046": "just keep in mind that in order to have reliable CV results you need to do this after the train/val split only in train targets and keep valid untouched. So CV score is calculated on the original targets \n\nThink that in test time (submission) you won't have the test[TARGETS] to apply the same transformation.\n\nI'm guessing this is what also other guys above are trying to say.",
    "2656140": "I think this can go either way, I am seeing metric correlation either way.\nIf we operate on the original targets for CV, the scores will be more familliar to us, but that will add additional complexity to the code (we would have to account for that in the data generator).",
    "2656775": "I agree with slime. I also think that more noise tends to less total votes for each patient, because it may be hard to diagose/predict and we should keep the original targets when calculating cv. Noise is also important features for these hard samples. keep hard samples/edge cases hard/edge, and don't change them."
  },
  "source": "meta"
}