{
  "id": 80674,
  "title": "Different distribution labels on test set?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/80674",
  "author_name": "Joop",
  "post_date": "2019-02-15T10:56:13.092000",
  "votes": 7,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I found it it's likely that the label positive/negative ratio is lower on the test set than on the training set. Does anyone concur?</p>\n\n<p>If so, what would be ways to deal with this?</p>\n\n<p>ps. the way I suspect this is because on my model, making the threshold higher gives me a significantly higher score. </p>",
  "messages": [
    {
      "id": 472536,
      "postDate": "2019-02-16T05:51:55.177Z",
      "content": "<p>It's hard to correlate total positive cases with lb score. Few lb, positive labels count and 5 fold cv score which i have achieved so far:\nLB         +ve Counts        CV 5 fold\n0.67            702                    0.7306\n0.689   639                    0.700\n0.707   723                    0.71243\n0.643   654                    0.729\n0.658   795                     0.7487\n0.677   699                     0.7268823\n0.644   807                      0.718493\n0.663   609                      0.7176142</p>",
      "rawMarkdown": "It's hard to correlate total positive cases with lb score. Few lb, positive labels count and 5 fold cv score which i have achieved so far:\nLB         +ve Counts        CV 5 fold\n0.67\t        702                    0.7306\n0.689\t639                    0.700\n0.707\t723                    0.71243\n0.643\t654                    0.729\n0.658\t795                     0.7487\n0.677\t699                     0.7268823\n0.644\t807                      0.718493\n0.663\t609                      0.7176142",
      "votes": 5,
      "replies": [
        {
          "id": 472857,
          "postDate": "2019-02-16T19:16:27.833Z",
          "content": "<p>Some of LB with related positives predicted labels:\n- LB 0.675 [762]\n- LB 0.688 [792]\n- LB 0.707 [675]\n- LB 0.712 [750]</p>\n\n<p>For most of my tests it looks it's more around 95-97% (no failure), 3-5% (failure) instead of 6% as in training dataset.</p>",
          "rawMarkdown": "Some of LB with related positives predicted labels:\n- LB 0.675 [762]\n- LB 0.688 [792]\n- LB 0.707 [675]\n- LB 0.712 [750]\n\nFor most of my tests it looks it's more around 95-97% (no failure), 3-5% (failure) instead of 6% as in training dataset.",
          "votes": 3
        }
      ]
    },
    {
      "id": 472103,
      "postDate": "2019-02-15T10:56:13.093Z",
      "content": "<p>I found it it's likely that the label positive/negative ratio is lower on the test set than on the training set. Does anyone concur?</p>\n\n<p>If so, what would be ways to deal with this?</p>\n\n<p>ps. the way I suspect this is because on my model, making the threshold higher gives me a significantly higher score. </p>",
      "rawMarkdown": "I found it it's likely that the label positive/negative ratio is lower on the test set than on the training set. Does anyone concur?\n\nIf so, what would be ways to deal with this?\n\nps. the way I suspect this is because on my model, making the threshold higher gives me a significantly higher score. ",
      "votes": 6
    },
    {
      "id": 472301,
      "postDate": "2019-02-15T16:52:31.463Z",
      "content": "<p>The file 'train' contains a division into 21 clusters according to the features that I used to train the ensemble of trees.\nIt can be seen that some clusters have no damage at all.\nThe file 'KZ' presents 3 types of damage in each of the clusters and 0 (no damage).\nThen the 'test' is divided into 21 clusters, you can visually see the sizes of the clusters and speculatively assume how much damage will be in each cluster and how they compare with the 'train'.</p>",
      "rawMarkdown": "\nThe file 'train' contains a division into 21 clusters according to the features that I used to train the ensemble of trees.\nIt can be seen that some clusters have no damage at all.\nThe file 'KZ' presents 3 types of damage in each of the clusters and 0 (no damage).\nThen the 'test' is divided into 21 clusters, you can visually see the sizes of the clusters and speculatively assume how much damage will be in each cluster and how they compare with the 'train'.",
      "votes": 4,
      "replies": [
        {
          "id": 472325,
          "postDate": "2019-02-15T17:38:32.743Z",
          "content": "<p>Awesome work.</p>",
          "rawMarkdown": "Awesome work."
        }
      ]
    },
    {
      "id": 472304,
      "postDate": "2019-02-15T16:54:13.227Z",
      "content": "<p>'KZ'</p>",
      "rawMarkdown": "'KZ'",
      "votes": 1
    },
    {
      "id": 472244,
      "postDate": "2019-02-15T15:07:23.773Z",
      "content": "<p>Determining the best threshold for positive percentage based on leader board performance, isn't this some kind of data leakage?</p>\n\n<p>What if you optimize for that and then the test set for the final leader board happens to have a different ratio altogether?</p>\n\n<p>I wonder if the organizers always change positive ratio on purpose to counter leader board hacking?</p>",
      "rawMarkdown": "Determining the best threshold for positive percentage based on leader board performance, isn't this some kind of data leakage?\n\nWhat if you optimize for that and then the test set for the final leader board happens to have a different ratio altogether?\n\nI wonder if the organizers always change positive ratio on purpose to counter leader board hacking?",
      "replies": [
        {
          "id": 482419,
          "postDate": "2019-03-03T00:16:47.077Z",
          "content": "<p>Of course public LB is, in many ways, a leak! Every time you take your LB score to make a decision, you actually use some statistic about the test labels.</p>\n\n<p>The question is: could this kind of leak be used to hack the competition results instead of achieving the best real-world solution to ML problem. Note, that only the latter is a win/win for both the competitors and stakeholders.</p>\n\n<p>Nevertheless, I would discourage anyone messing with the class ratio on evaluation data. And here is why. </p>\n\n<p>All our theory is built on the assumption that all available data (whether it is public, or not) are i.i.d. from <em>the same</em> unknown distribution. Changing this distribution on purpose will break the i.i.d. assumption with the following side effects: \na) algorithms may evolve into the wrong direction, since we all more or less use LB scores to decide between alternative approaches;\nb) organizers' choice of top N algorithms will be biased, especially if the final evaluation dataset is not drawn from the real-world distribution (from which our training data is ought to be sampled!).</p>\n\n<p>So, I hope the organizers don't do it otherwise it would be shot in the foot (IMHO). To protect the competition form leaks it would be much better to use smarter metrics along with good train/test/eval proportions (which, I guess, they already do).</p>",
          "rawMarkdown": "Of course public LB is, in many ways, a leak! Every time you take your LB score to make a decision, you actually use some statistic about the test labels.\n\nThe question is: could this kind of leak be used to hack the competition results instead of achieving the best real-world solution to ML problem. Note, that only the latter is a win/win for both the competitors and stakeholders.\n\nNevertheless, I would discourage anyone messing with the class ratio on evaluation data. And here is why. \n\nAll our theory is built on the assumption that all available data (whether it is public, or not) are i.i.d. from *the same* unknown distribution. Changing this distribution on purpose will break the i.i.d. assumption with the following side effects: \na) algorithms may evolve into the wrong direction, since we all more or less use LB scores to decide between alternative approaches;\nb) organizers' choice of top N algorithms will be biased, especially if the final evaluation dataset is not drawn from the real-world distribution (from which our training data is ought to be sampled!).\n\nSo, I hope the organizers don't do it otherwise it would be shot in the foot (IMHO). To protect the competition form leaks it would be much better to use smarter metrics along with good train/test/eval proportions (which, I guess, they already do).",
          "votes": 1
        },
        {
          "id": 499980,
          "postDate": "2019-03-25T13:15:57.633Z",
          "content": "<p>But it seems that they have done it. For one TP, we got MCC 0.060 on the public leaderboard and 0.050 on the private leaderboard. If we have calculated correctly, this corresponds to a positive rate of 0.023 in the public test set and 0.044 in the private test set (compared to 0.060 in the training set). \nWe guess that this difference is one reason why our MCC went up from 0.460 to 0.623 in the final evaluation. </p>",
          "rawMarkdown": "But it seems that they have done it. For one TP, we got MCC 0.060 on the public leaderboard and 0.050 on the private leaderboard. If we have calculated correctly, this corresponds to a positive rate of 0.023 in the public test set and 0.044 in the private test set (compared to 0.060 in the training set). \nWe guess that this difference is one reason why our MCC went up from 0.460 to 0.623 in the final evaluation. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 472125,
      "postDate": "2019-02-15T11:54:48.627Z",
      "content": "<p>For an lb score of 0.659 (CV 0.72) I have 825 positives out of 20337 test signals.\nCant say for sure that they have to reduce further for better scores, but I am just guessing that may be the case.</p>",
      "rawMarkdown": "For an lb score of 0.659 (CV 0.72) I have 825 positives out of 20337 test signals.\nCant say for sure that they have to reduce further for better scores, but I am just guessing that may be the case.",
      "replies": [
        {
          "id": 472241,
          "postDate": "2019-02-15T15:04:12.543Z",
          "content": "<p>My best score was also when I changed the rounding threshold to 4% positives from the original 6%, but only .578.</p>",
          "rawMarkdown": "My best score was also when I changed the rounding threshold to 4% positives from the original 6%, but only .578."
        }
      ]
    },
    {
      "id": 486785,
      "postDate": "2019-03-09T12:13:09.563Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 472536,
      "author_name": "HarshitMehta",
      "author_url": "",
      "post_date": "2019-02-16T05:51:55.177000",
      "content": "<p>It's hard to correlate total positive cases with lb score. Few lb, positive labels count and 5 fold cv score which i have achieved so far:\nLB         +ve Counts        CV 5 fold\n0.67            702                    0.7306\n0.689   639                    0.700\n0.707   723                    0.71243\n0.643   654                    0.729\n0.658   795                     0.7487\n0.677   699                     0.7268823\n0.644   807                      0.718493\n0.663   609                      0.7176142</p>",
      "votes": 5,
      "replies": [
        {
          "id": 472857,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2019-02-16T19:16:27.833000",
          "content": "<p>Some of LB with related positives predicted labels:\n- LB 0.675 [762]\n- LB 0.688 [792]\n- LB 0.707 [675]\n- LB 0.712 [750]</p>\n\n<p>For most of my tests it looks it's more around 95-97% (no failure), 3-5% (failure) instead of 6% as in training dataset.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 472301,
      "author_name": "sin2pi ",
      "author_url": "",
      "post_date": "2019-02-15T16:52:31.463000",
      "content": "<p>The file 'train' contains a division into 21 clusters according to the features that I used to train the ensemble of trees.\nIt can be seen that some clusters have no damage at all.\nThe file 'KZ' presents 3 types of damage in each of the clusters and 0 (no damage).\nThen the 'test' is divided into 21 clusters, you can visually see the sizes of the clusters and speculatively assume how much damage will be in each cluster and how they compare with the 'train'.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 472325,
          "author_name": "Joop",
          "author_url": "",
          "post_date": "2019-02-15T17:38:32.743000",
          "content": "<p>Awesome work.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 472304,
      "author_name": "sin2pi ",
      "author_url": "",
      "post_date": "2019-02-15T16:54:13.227000",
      "content": "<p>'KZ'</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472244,
      "author_name": "Joop",
      "author_url": "",
      "post_date": "2019-02-15T15:07:23.773000",
      "content": "<p>Determining the best threshold for positive percentage based on leader board performance, isn't this some kind of data leakage?</p>\n\n<p>What if you optimize for that and then the test set for the final leader board happens to have a different ratio altogether?</p>\n\n<p>I wonder if the organizers always change positive ratio on purpose to counter leader board hacking?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 482419,
          "author_name": "M0nZDeRR",
          "author_url": "",
          "post_date": "2019-03-03T00:16:47.077000",
          "content": "<p>Of course public LB is, in many ways, a leak! Every time you take your LB score to make a decision, you actually use some statistic about the test labels.</p>\n\n<p>The question is: could this kind of leak be used to hack the competition results instead of achieving the best real-world solution to ML problem. Note, that only the latter is a win/win for both the competitors and stakeholders.</p>\n\n<p>Nevertheless, I would discourage anyone messing with the class ratio on evaluation data. And here is why. </p>\n\n<p>All our theory is built on the assumption that all available data (whether it is public, or not) are i.i.d. from <em>the same</em> unknown distribution. Changing this distribution on purpose will break the i.i.d. assumption with the following side effects: \na) algorithms may evolve into the wrong direction, since we all more or less use LB scores to decide between alternative approaches;\nb) organizers' choice of top N algorithms will be biased, especially if the final evaluation dataset is not drawn from the real-world distribution (from which our training data is ought to be sampled!).</p>\n\n<p>So, I hope the organizers don't do it otherwise it would be shot in the foot (IMHO). To protect the competition form leaks it would be much better to use smarter metrics along with good train/test/eval proportions (which, I guess, they already do).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 499980,
          "author_name": "oranges",
          "author_url": "",
          "post_date": "2019-03-25T13:15:57.633000",
          "content": "<p>But it seems that they have done it. For one TP, we got MCC 0.060 on the public leaderboard and 0.050 on the private leaderboard. If we have calculated correctly, this corresponds to a positive rate of 0.023 in the public test set and 0.044 in the private test set (compared to 0.060 in the training set). \nWe guess that this difference is one reason why our MCC went up from 0.460 to 0.623 in the final evaluation. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 472125,
      "author_name": "Subrahmanyam V",
      "author_url": "",
      "post_date": "2019-02-15T11:54:48.627000",
      "content": "<p>For an lb score of 0.659 (CV 0.72) I have 825 positives out of 20337 test signals.\nCant say for sure that they have to reduce further for better scores, but I am just guessing that may be the case.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 472241,
          "author_name": "Joop",
          "author_url": "",
          "post_date": "2019-02-15T15:04:12.543000",
          "content": "<p>My best score was also when I changed the rounding threshold to 4% positives from the original 6%, but only .578.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 486785,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-09T12:13:09.563000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "472536": "It's hard to correlate total positive cases with lb score. Few lb, positive labels count and 5 fold cv score which i have achieved so far:\nLB         +ve Counts        CV 5 fold\n0.67\t        702                    0.7306\n0.689\t639                    0.700\n0.707\t723                    0.71243\n0.643\t654                    0.729\n0.658\t795                     0.7487\n0.677\t699                     0.7268823\n0.644\t807                      0.718493\n0.663\t609                      0.7176142",
    "472103": "I found it it's likely that the label positive/negative ratio is lower on the test set than on the training set. Does anyone concur?\n\nIf so, what would be ways to deal with this?\n\nps. the way I suspect this is because on my model, making the threshold higher gives me a significantly higher score. ",
    "472301": "\nThe file 'train' contains a division into 21 clusters according to the features that I used to train the ensemble of trees.\nIt can be seen that some clusters have no damage at all.\nThe file 'KZ' presents 3 types of damage in each of the clusters and 0 (no damage).\nThen the 'test' is divided into 21 clusters, you can visually see the sizes of the clusters and speculatively assume how much damage will be in each cluster and how they compare with the 'train'.",
    "472304": "'KZ'",
    "472244": "Determining the best threshold for positive percentage based on leader board performance, isn't this some kind of data leakage?\n\nWhat if you optimize for that and then the test set for the final leader board happens to have a different ratio altogether?\n\nI wonder if the organizers always change positive ratio on purpose to counter leader board hacking?",
    "472125": "For an lb score of 0.659 (CV 0.72) I have 825 positives out of 20337 test signals.\nCant say for sure that they have to reduce further for better scores, but I am just guessing that may be the case.",
    "486785": ""
  }
}