{
  "id": 86831,
  "title": "On the (In)Convenience of having different train / test distributions for a Kaggle Competition",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/86831",
  "author_name": "",
  "post_date": "2019-03-27T03:58:16.382210100Z",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Dear all,</p>\n\n<p>Now that this competition is over, I was thinking about the convenience of using different distributions of train / test set for a Kaggle competition. </p>\n\n<p>One the one side, competitions with different train / test distributions require algorithms that generalize well (generalization to unknown domains is a big topic in machine learning). If the organizers of this competition expected Kagglers to provide with a model that can generalize well to samples that are not included in the train set (and possibly not included in the public LB test set), then I guess that the goal was achieved within this competition.</p>\n\n<p>However, from the description of the competition it seems as though the main goal of organizers was to achieve the highest  possible detection of partial discharges, with the aim of using this classification to build technology solutions: </p>\n\n<blockquote>\n  <p>Your challenge is to detect partial discharge patterns in signals acquired from these power lines with a new meter designed at the ENET Centre at VŠB. Effective classifiers using this data will make it possible to continuously monitor power lines for faults.\n  ENET Centre researches and develops renewable energy resources with the goal of reducing or eliminating harmful environmental impacts. Their efforts focus on developing technology solutions around transportation and processing of energy raw materials.\n  By developing a solution to detect partial discharge you’ll help reduce maintenance costs, and prevent power outages.&gt; </p>\n</blockquote>\n\n<p>Many problems in machine learning at the production stage can be addressed by increasing the sample size and curating the dataset (e.g., checking that the train / test distributions are as similar as possible, performing error analysis, etc; see, for example, Andrew NG courses or Andrej Karpathy talks). A serious problem arises when we have different distributions for train, public LB and private LB: there is no easy way to prevent overfitting the private LB (although other Kagglers have provided some very clever ideas on how to check these discrepancies between train / test sets in the discussion section of this competition, there is no guarantee they will always work). This seems like the reason why most users / teams (several Kaggle grandmaster) that were leading the public LB went down the rankings: they worked hard on models that had high CV scores on the train set and the public LB set, but overfitted the private LB set. In my personal opinion, this can be very frustrating and often, due to the lack of a metric to evaluate the progress of your model on the private LB, a waste of time. </p>\n\n<p>If the final goal of this competition was to obtain algorithms to build technology solutions to prevent power outages and reduce costs,  a dataset with  a higher number of samples with partial discharges and similar distribution of train / test samples may have allowed Kagglers to provide the organizers with better models for partial discharge detection. Please don’t get me wrong here, I am not undermining the work of those people who won this competition, I think they did a great job and I’ve read some very interesting / ingenious approaches on how to eliminate features that had low impact on the public LB scores. I am just asking here, wouldn’t a dataset with similar train / test ditributions allow Kagglers to improve their models for partial discharge detection and, in the end, provide a better solution to the organizers of the competition (I suppose that the question boils down to whether a Matthews correlation coefficient of ~0.71 is sufficient at production stage for a technology solution in this area)?</p>\n\n<p>In my view (this is just my personal opinion), I would prefer to know in advance from the organizers whether a competition is aimed at building models that generalize well to unknown / unseen samples / categories (e.g., one shot-learning, or a model that learns to play chess after it learns to play Go, etc) or at improving a metric for solving a problem within a specific technology domain area. </p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "501222",
      "postDate": "03/27/2019 03:58:16",
      "content": "<p>Dear all,</p>\n\n<p>Now that this competition is over, I was thinking about the convenience of using different distributions of train / test set for a Kaggle competition. </p>\n\n<p>One the one side, competitions with different train / test distributions require algorithms that generalize well (generalization to unknown domains is a big topic in machine learning). If the organizers of this competition expected Kagglers to provide with a model that can generalize well to samples that are not included in the train set (and possibly not included in the public LB test set), then I guess that the goal was achieved within this competition.</p>\n\n<p>However, from the description of the competition it seems as though the main goal of organizers was to achieve the highest  possible detection of partial discharges, with the aim of using this classification to build technology solutions: </p>\n\n<blockquote>\n  <p>Your challenge is to detect partial discharge patterns in signals acquired from these power lines with a new meter designed at the ENET Centre at VŠB. Effective classifiers using this data will make it possible to continuously monitor power lines for faults.\n  ENET Centre researches and develops renewable energy resources with the goal of reducing or eliminating harmful environmental impacts. Their efforts focus on developing technology solutions around transportation and processing of energy raw materials.\n  By developing a solution to detect partial discharge you’ll help reduce maintenance costs, and prevent power outages.&gt; </p>\n</blockquote>\n\n<p>Many problems in machine learning at the production stage can be addressed by increasing the sample size and curating the dataset (e.g., checking that the train / test distributions are as similar as possible, performing error analysis, etc; see, for example, Andrew NG courses or Andrej Karpathy talks). A serious problem arises when we have different distributions for train, public LB and private LB: there is no easy way to prevent overfitting the private LB (although other Kagglers have provided some very clever ideas on how to check these discrepancies between train / test sets in the discussion section of this competition, there is no guarantee they will always work). This seems like the reason why most users / teams (several Kaggle grandmaster) that were leading the public LB went down the rankings: they worked hard on models that had high CV scores on the train set and the public LB set, but overfitted the private LB set. In my personal opinion, this can be very frustrating and often, due to the lack of a metric to evaluate the progress of your model on the private LB, a waste of time. </p>\n\n<p>If the final goal of this competition was to obtain algorithms to build technology solutions to prevent power outages and reduce costs,  a dataset with  a higher number of samples with partial discharges and similar distribution of train / test samples may have allowed Kagglers to provide the organizers with better models for partial discharge detection. Please don’t get me wrong here, I am not undermining the work of those people who won this competition, I think they did a great job and I’ve read some very interesting / ingenious approaches on how to eliminate features that had low impact on the public LB scores. I am just asking here, wouldn’t a dataset with similar train / test ditributions allow Kagglers to improve their models for partial discharge detection and, in the end, provide a better solution to the organizers of the competition (I suppose that the question boils down to whether a Matthews correlation coefficient of ~0.71 is sufficient at production stage for a technology solution in this area)?</p>\n\n<p>In my view (this is just my personal opinion), I would prefer to know in advance from the organizers whether a competition is aimed at building models that generalize well to unknown / unseen samples / categories (e.g., one shot-learning, or a model that learns to play chess after it learns to play Go, etc) or at improving a metric for solving a problem within a specific technology domain area. </p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Dear all,\n\nNow that this competition is over, I was thinking about the convenience of using different distributions of train / test set for a Kaggle competition. \n\nOne the one side, competitions with different train / test distributions require algorithms that generalize well (generalization to unknown domains is a big topic in machine learning). If the organizers of this competition expected Kagglers to provide with a model that can generalize well to samples that are not included in the train set (and possibly not included in the public LB test set), then I guess that the goal was achieved within this competition.\n\nHowever, from the description of the competition it seems as though the main goal of organizers was to achieve the highest  possible detection of partial discharges, with the aim of using this classification to build technology solutions: \n\n&gt; Your challenge is to detect partial discharge patterns in signals acquired from these power lines with a new meter designed at the ENET Centre at VŠB. Effective classifiers using this data will make it possible to continuously monitor power lines for faults.\nENET Centre researches and develops renewable energy resources with the goal of reducing or eliminating harmful environmental impacts. Their efforts focus on developing technology solutions around transportation and processing of energy raw materials.\nBy developing a solution to detect partial discharge you’ll help reduce maintenance costs, and prevent power outages.&gt; \n\nMany problems in machine learning at the production stage can be addressed by increasing the sample size and curating the dataset (e.g., checking that the train / test distributions are as similar as possible, performing error analysis, etc; see, for example, Andrew NG courses or Andrej Karpathy talks). A serious problem arises when we have different distributions for train, public LB and private LB: there is no easy way to prevent overfitting the private LB (although other Kagglers have provided some very clever ideas on how to check these discrepancies between train / test sets in the discussion section of this competition, there is no guarantee they will always work). This seems like the reason why most users / teams (several Kaggle grandmaster) that were leading the public LB went down the rankings: they worked hard on models that had high CV scores on the train set and the public LB set, but overfitted the private LB set. In my personal opinion, this can be very frustrating and often, due to the lack of a metric to evaluate the progress of your model on the private LB, a waste of time. \n\nIf the final goal of this competition was to obtain algorithms to build technology solutions to prevent power outages and reduce costs,  a dataset with  a higher number of samples with partial discharges and similar distribution of train / test samples may have allowed Kagglers to provide the organizers with better models for partial discharge detection. Please don’t get me wrong here, I am not undermining the work of those people who won this competition, I think they did a great job and I’ve read some very interesting / ingenious approaches on how to eliminate features that had low impact on the public LB scores. I am just asking here, wouldn’t a dataset with similar train / test ditributions allow Kagglers to improve their models for partial discharge detection and, in the end, provide a better solution to the organizers of the competition (I suppose that the question boils down to whether a Matthews correlation coefficient of ~0.71 is sufficient at production stage for a technology solution in this area)?\n\nIn my view (this is just my personal opinion), I would prefer to know in advance from the organizers whether a competition is aimed at building models that generalize well to unknown / unseen samples / categories (e.g., one shot-learning, or a model that learns to play chess after it learns to play Go, etc) or at improving a metric for solving a problem within a specific technology domain area. \n\nCheers",
      "votes": null
    },
    {
      "id": "501276",
      "postDate": "03/27/2019 06:20:50",
      "content": "<p>Good point. Just add more note : </p>\n\n<p>In a training set class1 is around 6.1%, in a public set around 2.3%, and around 4.3% in a private set. Therefore, a model that takes into account (e.g. by threshold adjustment) the class1 ratio will be likely to overfit more. </p>\n\n<p>I also think that top public LB kagglers do not get enough credits as it is actually also very difficult to generalize to the public LB (which also a test set that we will see in a real-world application). <a href=\"/mathurinache\">@mathurinache</a> , <a href=\"/sergeifironov\">@sergeifironov</a> , <a href=\"/thisisalpha\">@thisisalpha</a> , <a href=\"/raf123\">@raf123</a> , <a href=\"/wrosinski\">@wrosinski</a> , it would be super grateful to know your solution. I am sure that your solutions will be appreciate by the community.</p>\n\n<p>In any cases, I also learned a lot from top private LB kagglers who share their solutions how can they try to prevent this overfitting.</p>",
      "rawMarkdown": "Good point. Just add more note : \n\nIn a training set class1 is around 6.1%, in a public set around 2.3%, and around 4.3% in a private set. Therefore, a model that takes into account (e.g. by threshold adjustment) the class1 ratio will be likely to overfit more. \n\nI also think that top public LB kagglers do not get enough credits as it is actually also very difficult to generalize to the public LB (which also a test set that we will see in a real-world application). @mathurinache , @sergeifironov , @thisisalpha , @raf123 , @wrosinski , it would be super grateful to know your solution. I am sure that your solutions will be appreciate by the community.\n\nIn any cases, I also learned a lot from top private LB kagglers who share their solutions how can they try to prevent this overfitting.",
      "votes": null
    },
    {
      "id": "501372",
      "postDate": "03/27/2019 09:07:42",
      "content": "<p>Thank you for the mention. We dropped from 7th to 60th partially due to the fact that we trusted LB feedback more than CV score. But the differences were so significant that it was hard not to ;).</p>\n\n<p>I will try to post a summary of our approach in a few days.</p>",
      "rawMarkdown": "Thank you for the mention. We dropped from 7th to 60th partially due to the fact that we trusted LB feedback more than CV score. But the differences were so significant that it was hard not to ;).\n\nI will try to post a summary of our approach in a few days.",
      "votes": null
    },
    {
      "id": "501593",
      "postDate": "03/27/2019 13:48:25",
      "content": "<p><a href=\"/wrosinski\">@wrosinski</a> I am looking forward to your write up!</p>",
      "rawMarkdown": "wrosinski I am looking forward to your write up!",
      "votes": null
    },
    {
      "id": "501917",
      "postDate": "03/28/2019 00:37:52",
      "content": "<p>I've read in the discussion several people saying something similar to \"we dropped in the rankings because we trusted LB feedback more than CV score\". </p>\n\n<p>I am new in Kaggle and really want to learn, and I have this question: under what circumstances would you trust your CV score over your public LB score? And if you trust your CV score over your public LB score, how can you tell whether you are overfitting the train set or generalizing better to the private LB?</p>\n\n<p>Am I missing something very obvious here?</p>",
      "rawMarkdown": "I've read in the discussion several people saying something similar to \"we dropped in the rankings because we trusted LB feedback more than CV score\". \n\nI am new in Kaggle and really want to learn, and I have this question: under what circumstances would you trust your CV score over your public LB score? And if you trust your CV score over your public LB score, how can you tell whether you are overfitting the train set or generalizing better to the private LB?\n\nAm I missing something very obvious here?",
      "votes": null
    },
    {
      "id": "502029",
      "postDate": "03/28/2019 04:38:42",
      "content": "<p>I think it's reasonable to expect the organizers to inform the participants when each contest begins about how the separations were made between train and test datasets.  For example, in some contests train and test data are drawn randomly from the same population.  In others, such as the \"Heritage Health Prize\" contest several years ago, the test data represented a different time period from the train.  When I was developing real world predictive analytics applications this information was certainly available.  As to the separation between public and private LB, I can't imagine any sensible method other than random selection, possibly with nuances such as label stratification.</p>",
      "rawMarkdown": "I think it's reasonable to expect the organizers to inform the participants when each contest begins about how the separations were made between train and test datasets.  For example, in some contests train and test data are drawn randomly from the same population.  In others, such as the \"Heritage Health Prize\" contest several years ago, the test data represented a different time period from the train.  When I was developing real world predictive analytics applications this information was certainly available.  As to the separation between public and private LB, I can't imagine any sensible method other than random selection, possibly with nuances such as label stratification.",
      "votes": null
    },
    {
      "id": "502186",
      "postDate": "03/28/2019 09:06:13",
      "content": "<p>I agree with you that in real problems you know how data are split. It would be very useful to have this information on any kaggle competition. </p>",
      "rawMarkdown": "I agree with you that in real problems you know how data are split. It would be very useful to have this information on any kaggle competition.",
      "votes": null
    },
    {
      "id": "502977",
      "postDate": "03/29/2019 10:28:45",
      "content": "<p><a href=\"/lolo5996\">@lolo5996</a> I am not an expert, but let me try my best to address your question. </p>\n\n<h2>LB vs. CV</h2>\n\n<p>I think CV vs. LB here is a bit misleading. In principle, we only need </p>\n\n<p>“a good metric that measure our performance to the future use” .  In general, “test set” should reflect our future use, and many times, public LB does a great job.</p>\n\n<p>In this competition, as <a href=\"/dslate\">@dslate</a> stated above, since we do not have any information to convince us that public and private test set will be different (we only have information that “train” and “test” set are different using adversarial validation), it totally makes sense to believe in public LB . </p>\n\n<p>(Note also that the number of data in public LB is a lot more than that of tranining data)</p>\n\n<h2>When not to believe in Public LB</h2>\n\n<p>There are other situation : \n- Sometimes Public LB is indeed not much reliable : this happen in the recent quora competition where although public and private test data may still come from the same distribution, the public data is quite small and we have a lot more traninig data, so we can make a good CV in that case.</p>\n\n<ul>\n<li>In Doodle image recognition problem, we have very large dataset, and both CV, public and private score are all in agreement.</li>\n</ul>\n\n<h2>Why private and public test data do not come  from the same distribution here?</h2>\n\n<p>As far as I know, we’ve never got an official explanation, but some participants hypothesize that public and private data may come from <strong>different locations</strong> (this <strong>spatial splitting</strong> may make sense regarding to the application just like we have <strong>time spliting</strong> for a time-series competition)</p>",
      "rawMarkdown": "lolo5996 I am not an expert, but let me try my best to address your question. \n\n## LB vs. CV\nI think CV vs. LB here is a bit misleading. In principle, we only need \n\n“a good metric that measure our performance to the future use” .  In general, “test set” should reflect our future use, and many times, public LB does a great job.\n\nIn this competition, as @dslate stated above, since we do not have any information to convince us that public and private test set will be different (we only have information that “train” and “test” set are different using adversarial validation), it totally makes sense to believe in public LB . \n\n(Note also that the number of data in public LB is a lot more than that of tranining data)\n\n## When not to believe in Public LB\nThere are other situation : \n- Sometimes Public LB is indeed not much reliable : this happen in the recent quora competition where although public and private test data may still come from the same distribution, the public data is quite small and we have a lot more traninig data, so we can make a good CV in that case.\n\n- In Doodle image recognition problem, we have very large dataset, and both CV, public and private score are all in agreement.\n\n## Why private and public test data do not come  from the same distribution here?\nAs far as I know, we’ve never got an official explanation, but some participants hypothesize that public and private data may come from **different locations** (this **spatial splitting** may make sense regarding to the application just like we have **time spliting** for a time-series competition)",
      "votes": null
    },
    {
      "id": "506257",
      "postDate": "04/03/2019 07:52:36",
      "content": "<p>Hi <a href=\"/ratthachat\">@ratthachat</a>,</p>\n\n<p>Your classification of the problem makes sense. However, the question still remains for the situation you call \"when not to believe in public LB\". </p>\n\n<p>If we cannot trust the public LB as a measure of improvement of your model, as we progress with the feature engineering and training of the model, how can we tell whether we are overfitting the training set or improving our model with good generalization?</p>\n\n<p>If we got a higher score in our CV in the training set, would that mean that we have a better model or that we are overfitting the overall test set (as it happened during this competition to many Kagglers)?</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Hi @ratthachat,\n\nYour classification of the problem makes sense. However, the question still remains for the situation you call \"when not to believe in public LB\". \n \nIf we cannot trust the public LB as a measure of improvement of your model, as we progress with the feature engineering and training of the model, how can we tell whether we are overfitting the training set or improving our model with good generalization?\n\nIf we got a higher score in our CV in the training set, would that mean that we have a better model or that we are overfitting the overall test set (as it happened during this competition to many Kagglers)?\n\nCheers",
      "votes": null
    },
    {
      "id": "506312",
      "postDate": "04/03/2019 09:34:20",
      "content": "<p>Hi <a href=\"/lolo5996\">@lolo5996</a> , Saldinor,</p>\n\n<p>I agree with you. The issue still remains (but see below). Please note that as far as I know (see many other opinions), this competition is quite special where</p>\n\n<p>train, public and private distributions are all quite different.</p>\n\n<p>That's why public <strong>does not</strong> generalize to private;\nstraightforward CV from training data also <strong>does not</strong> generalize to private either.\n(this doesn't mean that private data is the ultimate truth that we have to believe, however).</p>\n\n<p><em>However, there is a nice idea by <a href=\"/yukinkgwa\">@yukinkgwa</a> (9th place), where he tried to transform the three of them  to be more similar. And in that's case you can believe more on both publicLB and CV.</em></p>",
      "rawMarkdown": "Hi @lolo5996 , Saldinor,\n\nI agree with you. The issue still remains (but see below). Please note that as far as I know (see many other opinions), this competition is quite special where\n\ntrain, public and private distributions are all quite different.\n\nThat's why public **does not** generalize to private;\nstraightforward CV from training data also **does not** generalize to private either.\n(this doesn't mean that private data is the ultimate truth that we have to believe, however).\n\n*However, there is a nice idea by @yukinkgwa (9th place), where he tried to transform the three of them  to be more similar. And in that's case you can believe more on both publicLB and CV.*",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 501276,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/27/2019 06:20:50",
      "content": "<p>Good point. Just add more note : </p>\n\n<p>In a training set class1 is around 6.1%, in a public set around 2.3%, and around 4.3% in a private set. Therefore, a model that takes into account (e.g. by threshold adjustment) the class1 ratio will be likely to overfit more. </p>\n\n<p>I also think that top public LB kagglers do not get enough credits as it is actually also very difficult to generalize to the public LB (which also a test set that we will see in a real-world application). <a href=\"/mathurinache\">@mathurinache</a> , <a href=\"/sergeifironov\">@sergeifironov</a> , <a href=\"/thisisalpha\">@thisisalpha</a> , <a href=\"/raf123\">@raf123</a> , <a href=\"/wrosinski\">@wrosinski</a> , it would be super grateful to know your solution. I am sure that your solutions will be appreciate by the community.</p>\n\n<p>In any cases, I also learned a lot from top private LB kagglers who share their solutions how can they try to prevent this overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 501372,
          "author_name": "wrosinski",
          "author_url": "",
          "post_date": "03/27/2019 09:07:42",
          "content": "<p>Thank you for the mention. We dropped from 7th to 60th partially due to the fact that we trusted LB feedback more than CV score. But the differences were so significant that it was hard not to ;).</p>\n\n<p>I will try to post a summary of our approach in a few days.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 501593,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/27/2019 13:48:25",
          "content": "<p><a href=\"/wrosinski\">@wrosinski</a> I am looking forward to your write up!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 501917,
          "author_name": "lolo5996",
          "author_url": "",
          "post_date": "03/28/2019 00:37:52",
          "content": "<p>I've read in the discussion several people saying something similar to \"we dropped in the rankings because we trusted LB feedback more than CV score\". </p>\n\n<p>I am new in Kaggle and really want to learn, and I have this question: under what circumstances would you trust your CV score over your public LB score? And if you trust your CV score over your public LB score, how can you tell whether you are overfitting the train set or generalizing better to the private LB?</p>\n\n<p>Am I missing something very obvious here?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 502977,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/29/2019 10:28:45",
          "content": "<p><a href=\"/lolo5996\">@lolo5996</a> I am not an expert, but let me try my best to address your question. </p>\n\n<h2>LB vs. CV</h2>\n\n<p>I think CV vs. LB here is a bit misleading. In principle, we only need </p>\n\n<p>“a good metric that measure our performance to the future use” .  In general, “test set” should reflect our future use, and many times, public LB does a great job.</p>\n\n<p>In this competition, as <a href=\"/dslate\">@dslate</a> stated above, since we do not have any information to convince us that public and private test set will be different (we only have information that “train” and “test” set are different using adversarial validation), it totally makes sense to believe in public LB . </p>\n\n<p>(Note also that the number of data in public LB is a lot more than that of tranining data)</p>\n\n<h2>When not to believe in Public LB</h2>\n\n<p>There are other situation : \n- Sometimes Public LB is indeed not much reliable : this happen in the recent quora competition where although public and private test data may still come from the same distribution, the public data is quite small and we have a lot more traninig data, so we can make a good CV in that case.</p>\n\n<ul>\n<li>In Doodle image recognition problem, we have very large dataset, and both CV, public and private score are all in agreement.</li>\n</ul>\n\n<h2>Why private and public test data do not come  from the same distribution here?</h2>\n\n<p>As far as I know, we’ve never got an official explanation, but some participants hypothesize that public and private data may come from <strong>different locations</strong> (this <strong>spatial splitting</strong> may make sense regarding to the application just like we have <strong>time spliting</strong> for a time-series competition)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506257,
          "author_name": "lolo5996",
          "author_url": "",
          "post_date": "04/03/2019 07:52:36",
          "content": "<p>Hi <a href=\"/ratthachat\">@ratthachat</a>,</p>\n\n<p>Your classification of the problem makes sense. However, the question still remains for the situation you call \"when not to believe in public LB\". </p>\n\n<p>If we cannot trust the public LB as a measure of improvement of your model, as we progress with the feature engineering and training of the model, how can we tell whether we are overfitting the training set or improving our model with good generalization?</p>\n\n<p>If we got a higher score in our CV in the training set, would that mean that we have a better model or that we are overfitting the overall test set (as it happened during this competition to many Kagglers)?</p>\n\n<p>Cheers</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 506312,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "04/03/2019 09:34:20",
          "content": "<p>Hi <a href=\"/lolo5996\">@lolo5996</a> , Saldinor,</p>\n\n<p>I agree with you. The issue still remains (but see below). Please note that as far as I know (see many other opinions), this competition is quite special where</p>\n\n<p>train, public and private distributions are all quite different.</p>\n\n<p>That's why public <strong>does not</strong> generalize to private;\nstraightforward CV from training data also <strong>does not</strong> generalize to private either.\n(this doesn't mean that private data is the ultimate truth that we have to believe, however).</p>\n\n<p><em>However, there is a nice idea by <a href=\"/yukinkgwa\">@yukinkgwa</a> (9th place), where he tried to transform the three of them  to be more similar. And in that's case you can believe more on both publicLB and CV.</em></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 502029,
      "author_name": "dslate",
      "author_url": "",
      "post_date": "03/28/2019 04:38:42",
      "content": "<p>I think it's reasonable to expect the organizers to inform the participants when each contest begins about how the separations were made between train and test datasets.  For example, in some contests train and test data are drawn randomly from the same population.  In others, such as the \"Heritage Health Prize\" contest several years ago, the test data represented a different time period from the train.  When I was developing real world predictive analytics applications this information was certainly available.  As to the separation between public and private LB, I can't imagine any sensible method other than random selection, possibly with nuances such as label stratification.</p>",
      "votes": null,
      "replies": [
        {
          "id": 502186,
          "author_name": "oguiza",
          "author_url": "",
          "post_date": "03/28/2019 09:06:13",
          "content": "<p>I agree with you that in real problems you know how data are split. It would be very useful to have this information on any kaggle competition. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "501222": "Dear all,\n\nNow that this competition is over, I was thinking about the convenience of using different distributions of train / test set for a Kaggle competition. \n\nOne the one side, competitions with different train / test distributions require algorithms that generalize well (generalization to unknown domains is a big topic in machine learning). If the organizers of this competition expected Kagglers to provide with a model that can generalize well to samples that are not included in the train set (and possibly not included in the public LB test set), then I guess that the goal was achieved within this competition.\n\nHowever, from the description of the competition it seems as though the main goal of organizers was to achieve the highest  possible detection of partial discharges, with the aim of using this classification to build technology solutions: \n\n&gt; Your challenge is to detect partial discharge patterns in signals acquired from these power lines with a new meter designed at the ENET Centre at VŠB. Effective classifiers using this data will make it possible to continuously monitor power lines for faults.\nENET Centre researches and develops renewable energy resources with the goal of reducing or eliminating harmful environmental impacts. Their efforts focus on developing technology solutions around transportation and processing of energy raw materials.\nBy developing a solution to detect partial discharge you’ll help reduce maintenance costs, and prevent power outages.&gt; \n\nMany problems in machine learning at the production stage can be addressed by increasing the sample size and curating the dataset (e.g., checking that the train / test distributions are as similar as possible, performing error analysis, etc; see, for example, Andrew NG courses or Andrej Karpathy talks). A serious problem arises when we have different distributions for train, public LB and private LB: there is no easy way to prevent overfitting the private LB (although other Kagglers have provided some very clever ideas on how to check these discrepancies between train / test sets in the discussion section of this competition, there is no guarantee they will always work). This seems like the reason why most users / teams (several Kaggle grandmaster) that were leading the public LB went down the rankings: they worked hard on models that had high CV scores on the train set and the public LB set, but overfitted the private LB set. In my personal opinion, this can be very frustrating and often, due to the lack of a metric to evaluate the progress of your model on the private LB, a waste of time. \n\nIf the final goal of this competition was to obtain algorithms to build technology solutions to prevent power outages and reduce costs,  a dataset with  a higher number of samples with partial discharges and similar distribution of train / test samples may have allowed Kagglers to provide the organizers with better models for partial discharge detection. Please don’t get me wrong here, I am not undermining the work of those people who won this competition, I think they did a great job and I’ve read some very interesting / ingenious approaches on how to eliminate features that had low impact on the public LB scores. I am just asking here, wouldn’t a dataset with similar train / test ditributions allow Kagglers to improve their models for partial discharge detection and, in the end, provide a better solution to the organizers of the competition (I suppose that the question boils down to whether a Matthews correlation coefficient of ~0.71 is sufficient at production stage for a technology solution in this area)?\n\nIn my view (this is just my personal opinion), I would prefer to know in advance from the organizers whether a competition is aimed at building models that generalize well to unknown / unseen samples / categories (e.g., one shot-learning, or a model that learns to play chess after it learns to play Go, etc) or at improving a metric for solving a problem within a specific technology domain area. \n\nCheers",
    "501276": "Good point. Just add more note : \n\nIn a training set class1 is around 6.1%, in a public set around 2.3%, and around 4.3% in a private set. Therefore, a model that takes into account (e.g. by threshold adjustment) the class1 ratio will be likely to overfit more. \n\nI also think that top public LB kagglers do not get enough credits as it is actually also very difficult to generalize to the public LB (which also a test set that we will see in a real-world application). @mathurinache , @sergeifironov , @thisisalpha , @raf123 , @wrosinski , it would be super grateful to know your solution. I am sure that your solutions will be appreciate by the community.\n\nIn any cases, I also learned a lot from top private LB kagglers who share their solutions how can they try to prevent this overfitting.",
    "501372": "Thank you for the mention. We dropped from 7th to 60th partially due to the fact that we trusted LB feedback more than CV score. But the differences were so significant that it was hard not to ;).\n\nI will try to post a summary of our approach in a few days.",
    "501593": "wrosinski I am looking forward to your write up!",
    "501917": "I've read in the discussion several people saying something similar to \"we dropped in the rankings because we trusted LB feedback more than CV score\". \n\nI am new in Kaggle and really want to learn, and I have this question: under what circumstances would you trust your CV score over your public LB score? And if you trust your CV score over your public LB score, how can you tell whether you are overfitting the train set or generalizing better to the private LB?\n\nAm I missing something very obvious here?",
    "502029": "I think it's reasonable to expect the organizers to inform the participants when each contest begins about how the separations were made between train and test datasets.  For example, in some contests train and test data are drawn randomly from the same population.  In others, such as the \"Heritage Health Prize\" contest several years ago, the test data represented a different time period from the train.  When I was developing real world predictive analytics applications this information was certainly available.  As to the separation between public and private LB, I can't imagine any sensible method other than random selection, possibly with nuances such as label stratification.",
    "502186": "I agree with you that in real problems you know how data are split. It would be very useful to have this information on any kaggle competition.",
    "502977": "lolo5996 I am not an expert, but let me try my best to address your question. \n\n## LB vs. CV\nI think CV vs. LB here is a bit misleading. In principle, we only need \n\n“a good metric that measure our performance to the future use” .  In general, “test set” should reflect our future use, and many times, public LB does a great job.\n\nIn this competition, as @dslate stated above, since we do not have any information to convince us that public and private test set will be different (we only have information that “train” and “test” set are different using adversarial validation), it totally makes sense to believe in public LB . \n\n(Note also that the number of data in public LB is a lot more than that of tranining data)\n\n## When not to believe in Public LB\nThere are other situation : \n- Sometimes Public LB is indeed not much reliable : this happen in the recent quora competition where although public and private test data may still come from the same distribution, the public data is quite small and we have a lot more traninig data, so we can make a good CV in that case.\n\n- In Doodle image recognition problem, we have very large dataset, and both CV, public and private score are all in agreement.\n\n## Why private and public test data do not come  from the same distribution here?\nAs far as I know, we’ve never got an official explanation, but some participants hypothesize that public and private data may come from **different locations** (this **spatial splitting** may make sense regarding to the application just like we have **time spliting** for a time-series competition)",
    "506257": "Hi @ratthachat,\n\nYour classification of the problem makes sense. However, the question still remains for the situation you call \"when not to believe in public LB\". \n \nIf we cannot trust the public LB as a measure of improvement of your model, as we progress with the feature engineering and training of the model, how can we tell whether we are overfitting the training set or improving our model with good generalization?\n\nIf we got a higher score in our CV in the training set, would that mean that we have a better model or that we are overfitting the overall test set (as it happened during this competition to many Kagglers)?\n\nCheers",
    "506312": "Hi @lolo5996 , Saldinor,\n\nI agree with you. The issue still remains (but see below). Please note that as far as I know (see many other opinions), this competition is quite special where\n\ntrain, public and private distributions are all quite different.\n\nThat's why public **does not** generalize to private;\nstraightforward CV from training data also **does not** generalize to private either.\n(this doesn't mean that private data is the ultimate truth that we have to believe, however).\n\n*However, there is a nice idea by @yukinkgwa (9th place), where he tried to transform the three of them  to be more similar. And in that's case you can believe more on both publicLB and CV.*"
  },
  "source": "meta"
}