{
  "id": 85205,
  "title": "Strategy of Choosing RIGHT Submit File",
  "url": "/competitions/vsb-power-line-fault-detection/discussion/85205",
  "author_name": "",
  "post_date": "2019-03-22T09:19:58.368685300Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Congratulations to all the winners.   : )</p>\n\n<p>It’s a hard competition, not only for feature engineering, but also for the perplexing mismatch between local CV and public/private LB. Lots of people are frustrated by the violent shakeup, and regret not choosing the right submit file. I saw some masters even drop from top 50 to 4-5 hundreds. </p>\n\n<p>Here’s two submit files of mine:\n| Model       | CV    | Public LB | Public Rank | Private LB | Private Rank |\n| ----------- | ----- | --------- | ----------- | ---------- | ------------ |\n| NN/Ensemble | 0.73x | 0.714     | 95          | 0.629      | 273          |\n| Tree Based  | 0.718 | 0.611     | 995         | 0.666      | 42           |</p>\n\n<p>What confused me is:\n+ People say <code>Trust your CV</code>, but NN model’s private score is much lower than Tree model;\n+ For NN model, there’s no big gap between CV(around 0.73x) and public LB(0.714).  I cannot say NN model is overfitted in advance;\n+ With a common CV(0.718) and a dramatically poor public score(0.611), it’s more likely the second model is marked as overfitted; \n+ Should I include the second model into ensemble? Even it seems overfitted? During the competition, you got no chance to know it has such a good private score.\n+ During competition, some kagglers pointed out trainset might be very different from testset. What can be done when there’s big difference between train set/public LB dataset/private LB dataset, as we encountered in this competition?</p>\n\n<p><code>If there’re some more strategies to address above issues, and choose the right submit file?</code></p>\n\n<p>I believe most of you may have the same question. It will be appreciated if you can share some thoughts.  </p>",
  "messages": [
    {
      "id": "496492",
      "postDate": "03/22/2019 09:19:58",
      "content": "<p>Congratulations to all the winners.   : )</p>\n\n<p>It’s a hard competition, not only for feature engineering, but also for the perplexing mismatch between local CV and public/private LB. Lots of people are frustrated by the violent shakeup, and regret not choosing the right submit file. I saw some masters even drop from top 50 to 4-5 hundreds. </p>\n\n<p>Here’s two submit files of mine:\n| Model       | CV    | Public LB | Public Rank | Private LB | Private Rank |\n| ----------- | ----- | --------- | ----------- | ---------- | ------------ |\n| NN/Ensemble | 0.73x | 0.714     | 95          | 0.629      | 273          |\n| Tree Based  | 0.718 | 0.611     | 995         | 0.666      | 42           |</p>\n\n<p>What confused me is:\n+ People say <code>Trust your CV</code>, but NN model’s private score is much lower than Tree model;\n+ For NN model, there’s no big gap between CV(around 0.73x) and public LB(0.714).  I cannot say NN model is overfitted in advance;\n+ With a common CV(0.718) and a dramatically poor public score(0.611), it’s more likely the second model is marked as overfitted; \n+ Should I include the second model into ensemble? Even it seems overfitted? During the competition, you got no chance to know it has such a good private score.\n+ During competition, some kagglers pointed out trainset might be very different from testset. What can be done when there’s big difference between train set/public LB dataset/private LB dataset, as we encountered in this competition?</p>\n\n<p><code>If there’re some more strategies to address above issues, and choose the right submit file?</code></p>\n\n<p>I believe most of you may have the same question. It will be appreciated if you can share some thoughts.  </p>",
      "rawMarkdown": "Congratulations to all the winners.   : )\n\nIt’s a hard competition, not only for feature engineering, but also for the perplexing mismatch between local CV and public/private LB. Lots of people are frustrated by the violent shakeup, and regret not choosing the right submit file. I saw some masters even drop from top 50 to 4-5 hundreds. \n\nHere’s two submit files of mine:\n| Model       | CV    | Public LB | Public Rank | Private LB | Private Rank |\n| ----------- | ----- | --------- | ----------- | ---------- | ------------ |\n| NN/Ensemble | 0.73x | 0.714     | 95          | 0.629      | 273          |\n| Tree Based  | 0.718 | 0.611     | 995         | 0.666      | 42           |\n\n\nWhat confused me is:\n+ People say `Trust your CV`, but NN model’s private score is much lower than Tree model;\n+ For NN model, there’s no big gap between CV(around 0.73x) and public LB(0.714).  I cannot say NN model is overfitted in advance;\n+ With a common CV(0.718) and a dramatically poor public score(0.611), it’s more likely the second model is marked as overfitted; \n+ Should I include the second model into ensemble? Even it seems overfitted? During the competition, you got no chance to know it has such a good private score.\n+ During competition, some kagglers pointed out trainset might be very different from testset. What can be done when there’s big difference between train set/public LB dataset/private LB dataset, as we encountered in this competition?\n\n`If there’re some more strategies to address above issues, and choose the right submit file? `\n\nI believe most of you may have the same question. It will be appreciated if you can share some thoughts.",
      "votes": null
    },
    {
      "id": "496495",
      "postDate": "03/22/2019 09:22:50",
      "content": "<p>Tree based Approaches (in hindsight) should have been favoured for their extra robustness to overfitting. As far as I understand many DL models fell prey to extreme unevenness between not just train and test but their subsegments</p>",
      "rawMarkdown": "Tree based Approaches (in hindsight) should have been favoured for their extra robustness to overfitting. As far as I understand many DL models fell prey to extreme unevenness between not just train and test but their subsegments",
      "votes": null
    },
    {
      "id": "496565",
      "postDate": "03/22/2019 10:47:12",
      "content": "<p>I'm also wondering if perhaps the public/private split is not random, but by location.   Along with this, neural nets generally perform somewhat better with the public set locations, tree-based on the private set (and overall more robust as George indicates), and the training data is more similar to the public set.    @Tomas @Sohier , would you be willing to share how the training/public/private splits were done?</p>",
      "rawMarkdown": "I'm also wondering if perhaps the public/private split is not random, but by location.   Along with this, neural nets generally perform somewhat better with the public set locations, tree-based on the private set (and overall more robust as George indicates), and the training data is more similar to the public set.    @Tomas @Sohier , would you be willing to share how the training/public/private splits were done?",
      "votes": null
    },
    {
      "id": "496632",
      "postDate": "03/22/2019 12:08:38",
      "content": "<p>Thank you, Russ. It's a very reasonable conjecture.  Seems there're big difference between public/private test set. </p>",
      "rawMarkdown": "Thank you, Russ. It's a very reasonable conjecture.  Seems there're big difference between public/private test set.",
      "votes": null
    },
    {
      "id": "496636",
      "postDate": "03/22/2019 12:18:17",
      "content": "<p>I agree. Tree based models have more robustness. But it's interesting that tree model performs so bad on public test set, while the private score is surprisingly good.  I think lots of participants are curious about the reason. </p>",
      "rawMarkdown": "I agree. Tree based models have more robustness. But it's interesting that tree model performs so bad on public test set, while the private score is surprisingly good.  I think lots of participants are curious about the reason.",
      "votes": null
    },
    {
      "id": "496744",
      "postDate": "03/22/2019 14:36:13",
      "content": "<p>I think 0.629 vs. 0.666 and 995 vs. 95 look huge. But in fact the real difference of correct prediction may not that much. Doesn’t want to advertise, but my hack on MCC estimate that the difference is due to just around 5-6 positive signal data (3 phases predicted together) ... <a href=\"https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc\">https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc</a></p>\n\n<p>I mean perhaps the actual predictive performances of the two may be not that much different than they look like.</p>",
      "rawMarkdown": "I think 0.629 vs. 0.666 and 995 vs. 95 look huge. But in fact the real difference of correct prediction may not that much. Doesn’t want to advertise, but my hack on MCC estimate that the difference is due to just around 5-6 positive signal data (3 phases predicted together) ... https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc\n\nI mean perhaps the actual predictive performances of the two may be not that much different than they look like.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 496495,
      "author_name": "george1986",
      "author_url": "",
      "post_date": "03/22/2019 09:22:50",
      "content": "<p>Tree based Approaches (in hindsight) should have been favoured for their extra robustness to overfitting. As far as I understand many DL models fell prey to extreme unevenness between not just train and test but their subsegments</p>",
      "votes": null,
      "replies": [
        {
          "id": 496636,
          "author_name": "gitshe11",
          "author_url": "",
          "post_date": "03/22/2019 12:18:17",
          "content": "<p>I agree. Tree based models have more robustness. But it's interesting that tree model performs so bad on public test set, while the private score is surprisingly good.  I think lots of participants are curious about the reason. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496565,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "03/22/2019 10:47:12",
      "content": "<p>I'm also wondering if perhaps the public/private split is not random, but by location.   Along with this, neural nets generally perform somewhat better with the public set locations, tree-based on the private set (and overall more robust as George indicates), and the training data is more similar to the public set.    @Tomas @Sohier , would you be willing to share how the training/public/private splits were done?</p>",
      "votes": null,
      "replies": [
        {
          "id": 496632,
          "author_name": "gitshe11",
          "author_url": "",
          "post_date": "03/22/2019 12:08:38",
          "content": "<p>Thank you, Russ. It's a very reasonable conjecture.  Seems there're big difference between public/private test set. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 496744,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "03/22/2019 14:36:13",
      "content": "<p>I think 0.629 vs. 0.666 and 995 vs. 95 look huge. But in fact the real difference of correct prediction may not that much. Doesn’t want to advertise, but my hack on MCC estimate that the difference is due to just around 5-6 positive signal data (3 phases predicted together) ... <a href=\"https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc\">https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc</a></p>\n\n<p>I mean perhaps the actual predictive performances of the two may be not that much different than they look like.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "496492": "Congratulations to all the winners.   : )\n\nIt’s a hard competition, not only for feature engineering, but also for the perplexing mismatch between local CV and public/private LB. Lots of people are frustrated by the violent shakeup, and regret not choosing the right submit file. I saw some masters even drop from top 50 to 4-5 hundreds. \n\nHere’s two submit files of mine:\n| Model       | CV    | Public LB | Public Rank | Private LB | Private Rank |\n| ----------- | ----- | --------- | ----------- | ---------- | ------------ |\n| NN/Ensemble | 0.73x | 0.714     | 95          | 0.629      | 273          |\n| Tree Based  | 0.718 | 0.611     | 995         | 0.666      | 42           |\n\n\nWhat confused me is:\n+ People say `Trust your CV`, but NN model’s private score is much lower than Tree model;\n+ For NN model, there’s no big gap between CV(around 0.73x) and public LB(0.714).  I cannot say NN model is overfitted in advance;\n+ With a common CV(0.718) and a dramatically poor public score(0.611), it’s more likely the second model is marked as overfitted; \n+ Should I include the second model into ensemble? Even it seems overfitted? During the competition, you got no chance to know it has such a good private score.\n+ During competition, some kagglers pointed out trainset might be very different from testset. What can be done when there’s big difference between train set/public LB dataset/private LB dataset, as we encountered in this competition?\n\n`If there’re some more strategies to address above issues, and choose the right submit file? `\n\nI believe most of you may have the same question. It will be appreciated if you can share some thoughts.",
    "496495": "Tree based Approaches (in hindsight) should have been favoured for their extra robustness to overfitting. As far as I understand many DL models fell prey to extreme unevenness between not just train and test but their subsegments",
    "496565": "I'm also wondering if perhaps the public/private split is not random, but by location.   Along with this, neural nets generally perform somewhat better with the public set locations, tree-based on the private set (and overall more robust as George indicates), and the training data is more similar to the public set.    @Tomas @Sohier , would you be willing to share how the training/public/private splits were done?",
    "496632": "Thank you, Russ. It's a very reasonable conjecture.  Seems there're big difference between public/private test set.",
    "496636": "I agree. Tree based models have more robustness. But it's interesting that tree model performs so bad on public test set, while the private score is surprisingly good.  I think lots of participants are curious about the reason.",
    "496744": "I think 0.629 vs. 0.666 and 995 vs. 95 look huge. But in fact the real difference of correct prediction may not that much. Doesn’t want to advertise, but my hack on MCC estimate that the difference is due to just around 5-6 positive signal data (3 phases predicted together) ... https://www.kaggle.com/ratthachat/a-heuristic-to-understand-your-mcc\n\nI mean perhaps the actual predictive performances of the two may be not that much different than they look like."
  },
  "source": "meta"
}