{
  "id": 84683,
  "title": "Public LB scores!",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/84683",
  "author_name": "",
  "post_date": "2019-03-19T00:47:09.531739200Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hello guys, I have a question:</p>\n\n<p>How much we can trust in the public LB score?\nI mean, we have a lot of people that are playing with 'seed' lucky to achieve good scores.</p>\n\n<p>This question come to me after I tried to make some ensembles and, at least for me, they are looking random in the public LB.</p>\n\n<p>What are your thoughts?</p>",
  "messages": [
    {
      "id": "493684",
      "postDate": "03/19/2019 00:47:09",
      "content": "<p>Hello guys, I have a question:</p>\n\n<p>How much we can trust in the public LB score?\nI mean, we have a lot of people that are playing with 'seed' lucky to achieve good scores.</p>\n\n<p>This question come to me after I tried to make some ensembles and, at least for me, they are looking random in the public LB.</p>\n\n<p>What are your thoughts?</p>",
      "rawMarkdown": "Hello guys, I have a question:\n\nHow much we can trust in the public LB score?\nI mean, we have a lot of people that are playing with 'seed' lucky to achieve good scores.\n\nThis question come to me after I tried to make some ensembles and, at least for me, they are looking random in the public LB.\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "493956",
      "postDate": "03/19/2019 09:57:20",
      "content": "<p>I tried ensembling too, but it just gives random scores on the LB (similar to your situation). I think the key is a more robust validation setup. But, I do not know how to setup a better validation for my model. Do you have any ideas ?</p>",
      "rawMarkdown": "I tried ensembling too, but it just gives random scores on the LB (similar to your situation). I think the key is a more robust validation setup. But, I do not know how to setup a better validation for my model. Do you have any ideas ?",
      "votes": null
    },
    {
      "id": "494188",
      "postDate": "03/19/2019 14:58:06",
      "content": "<p>Perhaps</p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675</a>\nor\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317</a>\ncould work? (provided that we can make it in a KFold setting)</p>",
      "rawMarkdown": "Perhaps\n\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675\nor\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\ncould work? (provided that we can make it in a KFold setting)",
      "votes": null
    },
    {
      "id": "494291",
      "postDate": "03/19/2019 16:47:31",
      "content": "<p>Thanks for the links, <a href=\"/ratthachat\">@ratthachat</a> ! But, I think adversarial validation can cause overfitting to the public LB because it tries to match the distribution of the validation data as closely as possible to the public test data (if I'm not wrong). This could fail miserably if the private test data even has a slightly different distribution from the public test data.</p>",
      "rawMarkdown": "Thanks for the links, @ratthachat ! But, I think adversarial validation can cause overfitting to the public LB because it tries to match the distribution of the validation data as closely as possible to the public test data (if I'm not wrong). This could fail miserably if the private test data even has a slightly different distribution from the public test data.",
      "votes": null
    },
    {
      "id": "494512",
      "postDate": "03/19/2019 23:28:58",
      "content": "<p>Hi Tarun <a href=\"/tarunpaparaju\">@tarunpaparaju</a> , I agree. I have hoped that they should not be that much difference, but of course that may not be the case ;) However, in that case the validation process perhaps does not make sense anymore ? </p>\n\n<p>The whole purpose of the validation is to simulate test distribution. And if public and private (which is unknown) are totally different, what can we do ? —- On another thought, in this competition we are able to probe some amount of private data, if we would had more time, maybe we could try to adversarial validate the private/public haha</p>",
      "rawMarkdown": "Hi Tarun @tarunpaparaju , I agree. I have hoped that they should not be that much difference, but of course that may not be the case ;) However, in that case the validation process perhaps does not make sense anymore ? \n\nThe whole purpose of the validation is to simulate test distribution. And if public and private (which is unknown) are totally different, what can we do ? —- On another thought, in this competition we are able to probe some amount of private data, if we would had more time, maybe we could try to adversarial validate the private/public haha",
      "votes": null
    },
    {
      "id": "494543",
      "postDate": "03/20/2019 01:25:28",
      "content": "<p>You're right <a href=\"/ratthachat\">@ratthachat</a>. There's nothing we can do if they're totally different. I think I should try this, but I'm running out of time and my school exams are going on ;)</p>",
      "rawMarkdown": "You're right @ratthachat. There's nothing we can do if they're totally different. I think I should try this, but I'm running out of time and my school exams are going on ;)",
      "votes": null
    },
    {
      "id": "495057",
      "postDate": "03/20/2019 15:23:11",
      "content": "<p>Haha <a href=\"/tarunpaparaju\">@tarunpaparaju</a> the best of luck for us, and for your exam as well!</p>",
      "rawMarkdown": "Haha @tarunpaparaju the best of luck for us, and for your exam as well!",
      "votes": null
    },
    {
      "id": "495131",
      "postDate": "03/20/2019 17:22:30",
      "content": "<p>Thanks ! Good luck to us !</p>",
      "rawMarkdown": "Thanks ! Good luck to us !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 493956,
      "author_name": "tarunpaparaju",
      "author_url": "",
      "post_date": "03/19/2019 09:57:20",
      "content": "<p>I tried ensembling too, but it just gives random scores on the LB (similar to your situation). I think the key is a more robust validation setup. But, I do not know how to setup a better validation for my model. Do you have any ideas ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 494188,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/19/2019 14:58:06",
          "content": "<p>Perhaps</p>\n\n<p><a href=\"https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675\">https://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675</a>\nor\n<a href=\"https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\">https://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317</a>\ncould work? (provided that we can make it in a KFold setting)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494291,
          "author_name": "tarunpaparaju",
          "author_url": "",
          "post_date": "03/19/2019 16:47:31",
          "content": "<p>Thanks for the links, <a href=\"/ratthachat\">@ratthachat</a> ! But, I think adversarial validation can cause overfitting to the public LB because it tries to match the distribution of the validation data as closely as possible to the public test data (if I'm not wrong). This could fail miserably if the private test data even has a slightly different distribution from the public test data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494512,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/19/2019 23:28:58",
          "content": "<p>Hi Tarun <a href=\"/tarunpaparaju\">@tarunpaparaju</a> , I agree. I have hoped that they should not be that much difference, but of course that may not be the case ;) However, in that case the validation process perhaps does not make sense anymore ? </p>\n\n<p>The whole purpose of the validation is to simulate test distribution. And if public and private (which is unknown) are totally different, what can we do ? —- On another thought, in this competition we are able to probe some amount of private data, if we would had more time, maybe we could try to adversarial validate the private/public haha</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 494543,
          "author_name": "tarunpaparaju",
          "author_url": "",
          "post_date": "03/20/2019 01:25:28",
          "content": "<p>You're right <a href=\"/ratthachat\">@ratthachat</a>. There's nothing we can do if they're totally different. I think I should try this, but I'm running out of time and my school exams are going on ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495057,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "03/20/2019 15:23:11",
          "content": "<p>Haha <a href=\"/tarunpaparaju\">@tarunpaparaju</a> the best of luck for us, and for your exam as well!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 495131,
          "author_name": "tarunpaparaju",
          "author_url": "",
          "post_date": "03/20/2019 17:22:30",
          "content": "<p>Thanks ! Good luck to us !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "493684": "Hello guys, I have a question:\n\nHow much we can trust in the public LB score?\nI mean, we have a lot of people that are playing with 'seed' lucky to achieve good scores.\n\nThis question come to me after I tried to make some ensembles and, at least for me, they are looking random in the public LB.\n\nWhat are your thoughts?",
    "493956": "I tried ensembling too, but it just gives random scores on the LB (similar to your situation). I think the key is a more robust validation setup. But, I do not know how to setup a better validation for my model. Do you have any ideas ?",
    "494188": "Perhaps\n\nhttps://www.kaggle.com/c/microsoft-malware-prediction/discussion/84675\nor\nhttps://www.kaggle.com/c/vsb-power-line-fault-detection/discussion/81317\ncould work? (provided that we can make it in a KFold setting)",
    "494291": "Thanks for the links, @ratthachat ! But, I think adversarial validation can cause overfitting to the public LB because it tries to match the distribution of the validation data as closely as possible to the public test data (if I'm not wrong). This could fail miserably if the private test data even has a slightly different distribution from the public test data.",
    "494512": "Hi Tarun @tarunpaparaju , I agree. I have hoped that they should not be that much difference, but of course that may not be the case ;) However, in that case the validation process perhaps does not make sense anymore ? \n\nThe whole purpose of the validation is to simulate test distribution. And if public and private (which is unknown) are totally different, what can we do ? —- On another thought, in this competition we are able to probe some amount of private data, if we would had more time, maybe we could try to adversarial validate the private/public haha",
    "494543": "You're right @ratthachat. There's nothing we can do if they're totally different. I think I should try this, but I'm running out of time and my school exams are going on ;)",
    "495057": "Haha @tarunpaparaju the best of luck for us, and for your exam as well!",
    "495131": "Thanks ! Good luck to us !"
  },
  "source": "meta"
}