{
  "id": 85167,
  "title": "If you implement a real world application, how should we employ Public or Private LB information?",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85167",
  "author_name": "",
  "post_date": "2019-03-22T02:48:22.705047200Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>EDIT</strong>: from probing (credited @putalay <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868</a>) \n  the private test has 4% positive class, the public has around 2.3%-2.4% positive class, and our training set has around 6% positive class. It is quite inevitable that our model will get confuse somehow.</p>\n</blockquote>\n\n<hr>\n\n<p>First of all, congratulation to all winners!  Shake up like an earthquake!</p>\n\n<p>Many of us, including myself,  are still wondering how to make a proper validation AND not to overfit the LB in a competition like this.</p>\n\n<p>Nevertheless, think about it, I have another question in mind : <strong>since both public LB and private LB are in fact the test set, they should represent the real-world application very well.</strong></p>\n\n<p>Now public is 57% and private is 43% , ASSUMING that we do not probe the public LB, does it still make more sense to believe in public LB more? or in fact believe in the weighted average :</p>\n\n<p>$$0.57 \\times PublicLB + 0.43 \\times PrivateLB $$</p>\n\n<p>——-\nI mean, getting high score on private LB is great, but in real life, the publicLB data/score should be highly relevant? </p>\n\n<p>Or in other words, if we have to make a real-world application on this problem, what will be the final validation strategy for you if we can access all train/test data &amp; labels?</p>",
  "messages": [
    {
      "id": "496279",
      "postDate": "03/22/2019 02:48:22",
      "content": "<blockquote>\n  <p><strong>EDIT</strong>: from probing (credited @putalay <a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868</a>) \n  the private test has 4% positive class, the public has around 2.3%-2.4% positive class, and our training set has around 6% positive class. It is quite inevitable that our model will get confuse somehow.</p>\n</blockquote>\n\n<hr>\n\n<p>First of all, congratulation to all winners!  Shake up like an earthquake!</p>\n\n<p>Many of us, including myself,  are still wondering how to make a proper validation AND not to overfit the LB in a competition like this.</p>\n\n<p>Nevertheless, think about it, I have another question in mind : <strong>since both public LB and private LB are in fact the test set, they should represent the real-world application very well.</strong></p>\n\n<p>Now public is 57% and private is 43% , ASSUMING that we do not probe the public LB, does it still make more sense to believe in public LB more? or in fact believe in the weighted average :</p>\n\n<p>$$0.57 \\times PublicLB + 0.43 \\times PrivateLB $$</p>\n\n<p>——-\nI mean, getting high score on private LB is great, but in real life, the publicLB data/score should be highly relevant? </p>\n\n<p>Or in other words, if we have to make a real-world application on this problem, what will be the final validation strategy for you if we can access all train/test data &amp; labels?</p>",
      "rawMarkdown": "&gt; **EDIT**: from probing (credited @putalay https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868) \nthe private test has 4% positive class, the public has around 2.3%-2.4% positive class, and our training set has around 6% positive class. It is quite inevitable that our model will get confuse somehow.\n\n-----\nFirst of all, congratulation to all winners!  Shake up like an earthquake!\n\nMany of us, including myself,  are still wondering how to make a proper validation AND not to overfit the LB in a competition like this.\n\nNevertheless, think about it, I have another question in mind : **since both public LB and private LB are in fact the test set, they should represent the real-world application very well.**\n\nNow public is 57% and private is 43% , ASSUMING that we do not probe the public LB, does it still make more sense to believe in public LB more? or in fact believe in the weighted average :\n\n$$0.57 \\times PublicLB + 0.43 \\times PrivateLB $$\n\n——-\nI mean, getting high score on private LB is great, but in real life, the publicLB data/score should be highly relevant? \n\nOr in other words, if we have to make a real-world application on this problem, what will be the final validation strategy for you if we can access all train/test data &amp; labels?",
      "votes": null
    },
    {
      "id": "496298",
      "postDate": "03/22/2019 03:04:42",
      "content": "<p>(Issue2)\nAfter reading other discussion , I also have another related question in mind : \nmany of us said that “trust your CV”,</p>\n\n<p>however, what is the valid CV strategy here? Because it is quite evidence that straightforward KFold CV here give a different data distribution than the test distribution.</p>\n\n<p>Moreover, does it make more sense to actually believe in LB more than CV provided that we have a lot more data on LB ?   (this is unlike competition like Quora where public LB data was very small)</p>\n\n<p><strong>EDIT</strong> — <a href=\"/blackboards\">@blackboards</a> also gave his ideas on this issue in the very beginning of his topic : \n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163</a></p>",
      "rawMarkdown": "(Issue2)\nAfter reading other discussion , I also have another related question in mind : \nmany of us said that “trust your CV”,\n\nhowever, what is the valid CV strategy here? Because it is quite evidence that straightforward KFold CV here give a different data distribution than the test distribution.\n\nMoreover, does it make more sense to actually believe in LB more than CV provided that we have a lot more data on LB ?   (this is unlike competition like Quora where public LB data was very small)\n\n**EDIT** — @blackboards also gave his ideas on this issue in the very beginning of his topic : \nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163",
      "votes": null
    },
    {
      "id": "496337",
      "postDate": "03/22/2019 04:03:32",
      "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I don't think you should look at it as believing public LB more. If the distribution of the training data and the test data are the same then yes LB is a good guide otherwse a good CV is your best bet. That is why it is good to have a very good cross validation setup. The reason why no one have so far posted a reliable CV setup for this competition does not mean there is none.  Perharps we are not usind the right features or there are some really noisy features that need to be excluded, etec..  My point is there is a reason even if we do not know it for sure yet.</p>",
      "rawMarkdown": "ratthachat, I don't think you should look at it as believing public LB more. If the distribution of the training data and the test data are the same then yes LB is a good guide otherwse a good CV is your best bet. That is why it is good to have a very good cross validation setup. The reason why no one have so far posted a reliable CV setup for this competition does not mean there is none.  Perharps we are not usind the right features or there are some really noisy features that need to be excluded, etec..  My point is there is a reason even if we do not know it for sure yet.",
      "votes": null
    },
    {
      "id": "496360",
      "postDate": "03/22/2019 04:27:24",
      "content": "<p>Thanks YaGana for your opinion! </p>\n\n<p>Just a bit more discussion — BTW, I realize that I have only a few experience on Kaggle, so please take my question as a newbie’s one ;) : </p>\n\n<p>Regarding the difference between CV vs. LB, how do we know that private LB will not be similar to Public LB ?  (Does it make more sense that public and private should be similar [should we believe that they are random splitting from the whole test set?] )</p>\n\n<p>I mean when we make a CV strategy, we have to make some assumptions, but how can we know that the distribution in private will be according to our assumptions? (compared to the simple assumptions that public/private should be randomly split)</p>",
      "rawMarkdown": "Thanks YaGana for your opinion! \n\nJust a bit more discussion — BTW, I realize that I have only a few experience on Kaggle, so please take my question as a newbie’s one ;) : \n\nRegarding the difference between CV vs. LB, how do we know that private LB will not be similar to Public LB ?  (Does it make more sense that public and private should be similar [should we believe that they are random splitting from the whole test set?] )\n\nI mean when we make a CV strategy, we have to make some assumptions, but how can we know that the distribution in private will be according to our assumptions? (compared to the simple assumptions that public/private should be randomly split)",
      "votes": null
    },
    {
      "id": "499044",
      "postDate": "03/24/2019 08:40:07",
      "content": "<p>The answer to this question would be the ultimate answer to succeeding in Kaggle.</p>",
      "rawMarkdown": "The answer to this question would be the ultimate answer to succeeding in Kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496298,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 03:04:42",
      "content": "<p>(Issue2)\nAfter reading other discussion , I also have another related question in mind : \nmany of us said that “trust your CV”,</p>\n\n<p>however, what is the valid CV strategy here? Because it is quite evidence that straightforward KFold CV here give a different data distribution than the test distribution.</p>\n\n<p>Moreover, does it make more sense to actually believe in LB more than CV provided that we have a lot more data on LB ?   (this is unlike competition like Quora where public LB data was very small)</p>\n\n<p><strong>EDIT</strong> — <a href=\"/blackboards\">@blackboards</a> also gave his ideas on this issue in the very beginning of his topic : \n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 496337,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "03/22/2019 04:03:32",
          "content": "<p><a href=\"/ratthachat\">@ratthachat</a>, I don't think you should look at it as believing public LB more. If the distribution of the training data and the test data are the same then yes LB is a good guide otherwse a good CV is your best bet. That is why it is good to have a very good cross validation setup. The reason why no one have so far posted a reliable CV setup for this competition does not mean there is none.  Perharps we are not usind the right features or there are some really noisy features that need to be excluded, etec..  My point is there is a reason even if we do not know it for sure yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 496360,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/22/2019 04:27:24",
          "content": "<p>Thanks YaGana for your opinion! </p>\n\n<p>Just a bit more discussion — BTW, I realize that I have only a few experience on Kaggle, so please take my question as a newbie’s one ;) : </p>\n\n<p>Regarding the difference between CV vs. LB, how do we know that private LB will not be similar to Public LB ?  (Does it make more sense that public and private should be similar [should we believe that they are random splitting from the whole test set?] )</p>\n\n<p>I mean when we make a CV strategy, we have to make some assumptions, but how can we know that the distribution in private will be according to our assumptions? (compared to the simple assumptions that public/private should be randomly split)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 499044,
          "author_name": "tarunpaparaju",
          "author_url": "",
          "post_date": "03/24/2019 08:40:07",
          "content": "<p>The answer to this question would be the ultimate answer to succeeding in Kaggle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "496279": "&gt; **EDIT**: from probing (credited @putalay https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/82868) \nthe private test has 4% positive class, the public has around 2.3%-2.4% positive class, and our training set has around 6% positive class. It is quite inevitable that our model will get confuse somehow.\n\n-----\nFirst of all, congratulation to all winners!  Shake up like an earthquake!\n\nMany of us, including myself,  are still wondering how to make a proper validation AND not to overfit the LB in a competition like this.\n\nNevertheless, think about it, I have another question in mind : **since both public LB and private LB are in fact the test set, they should represent the real-world application very well.**\n\nNow public is 57% and private is 43% , ASSUMING that we do not probe the public LB, does it still make more sense to believe in public LB more? or in fact believe in the weighted average :\n\n$$0.57 \\times PublicLB + 0.43 \\times PrivateLB $$\n\n——-\nI mean, getting high score on private LB is great, but in real life, the publicLB data/score should be highly relevant? \n\nOr in other words, if we have to make a real-world application on this problem, what will be the final validation strategy for you if we can access all train/test data &amp; labels?",
    "496298": "(Issue2)\nAfter reading other discussion , I also have another related question in mind : \nmany of us said that “trust your CV”,\n\nhowever, what is the valid CV strategy here? Because it is quite evidence that straightforward KFold CV here give a different data distribution than the test distribution.\n\nMoreover, does it make more sense to actually believe in LB more than CV provided that we have a lot more data on LB ?   (this is unlike competition like Quora where public LB data was very small)\n\n**EDIT** — @blackboards also gave his ideas on this issue in the very beginning of his topic : \nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/85163",
    "496337": "ratthachat, I don't think you should look at it as believing public LB more. If the distribution of the training data and the test data are the same then yes LB is a good guide otherwse a good CV is your best bet. That is why it is good to have a very good cross validation setup. The reason why no one have so far posted a reliable CV setup for this competition does not mean there is none.  Perharps we are not usind the right features or there are some really noisy features that need to be excluded, etec..  My point is there is a reason even if we do not know it for sure yet.",
    "496360": "Thanks YaGana for your opinion! \n\nJust a bit more discussion — BTW, I realize that I have only a few experience on Kaggle, so please take my question as a newbie’s one ;) : \n\nRegarding the difference between CV vs. LB, how do we know that private LB will not be similar to Public LB ?  (Does it make more sense that public and private should be similar [should we believe that they are random splitting from the whole test set?] )\n\nI mean when we make a CV strategy, we have to make some assumptions, but how can we know that the distribution in private will be according to our assumptions? (compared to the simple assumptions that public/private should be randomly split)",
    "499044": "The answer to this question would be the ultimate answer to succeeding in Kaggle."
  },
  "source": "meta"
}