{
  "id": 80646,
  "title": "Was local val score a good metric to follow?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/80646",
  "author_name": "",
  "post_date": "2019-02-15T04:36:16.082393400Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>During the competition, a lot of people (including me) were using their local validation score (cv or ensembling) as a way to figure out if their models had improved. However, after the competition ended, my best model with a cv score &gt;0.692 performed worse on the stage 2 test set than a public kernel that 230 people used as their final submission which got a cv score of 0.68.</p>\n\n<p>Since the competition is over, is anyone willing to share their best model's cv score and the corresponding private lb score? This might give everyone more insight into how to set up a reliable local cv to use in other competitions with small public test sets.</p>\n\n<p>Thanks</p>\n\n<p>My best kernel:</p>\n\n<p>local cv: 0.6927 <br>\npublic lb: 0.6962 <br>\nprivate lb: 0.70213  </p>\n\n<p>local roc-auc: 0.9697 <br>\ncorrelations between test predictions from each fold: 0.956</p>",
  "messages": [
    {
      "id": "471914",
      "postDate": "02/15/2019 04:36:16",
      "content": "<p>Hi,</p>\n\n<p>During the competition, a lot of people (including me) were using their local validation score (cv or ensembling) as a way to figure out if their models had improved. However, after the competition ended, my best model with a cv score &gt;0.692 performed worse on the stage 2 test set than a public kernel that 230 people used as their final submission which got a cv score of 0.68.</p>\n\n<p>Since the competition is over, is anyone willing to share their best model's cv score and the corresponding private lb score? This might give everyone more insight into how to set up a reliable local cv to use in other competitions with small public test sets.</p>\n\n<p>Thanks</p>\n\n<p>My best kernel:</p>\n\n<p>local cv: 0.6927 <br>\npublic lb: 0.6962 <br>\nprivate lb: 0.70213  </p>\n\n<p>local roc-auc: 0.9697 <br>\ncorrelations between test predictions from each fold: 0.956</p>",
      "rawMarkdown": "Hi,\n\nDuring the competition, a lot of people (including me) were using their local validation score (cv or ensembling) as a way to figure out if their models had improved. However, after the competition ended, my best model with a cv score &gt;0.692 performed worse on the stage 2 test set than a public kernel that 230 people used as their final submission which got a cv score of 0.68.\n\nSince the competition is over, is anyone willing to share their best model's cv score and the corresponding private lb score? This might give everyone more insight into how to set up a reliable local cv to use in other competitions with small public test sets.\n\nThanks\n\nMy best kernel:\n\nlocal cv: 0.6927  \npublic lb: 0.6962  \nprivate lb: 0.70213  \n\nlocal roc-auc: 0.9697  \ncorrelations between test predictions from each fold: 0.956",
      "votes": null
    },
    {
      "id": "471939",
      "postDate": "02/15/2019 05:47:39",
      "content": "<p>highest scoring model on private LB (which is of lower CV/PublicLB in the two submitted models)\nCV - 0.6910\nPublic LB - 0.69242\nPrivate LB - 0.70381\ncorrelation between fold predictions  is  around 0.95</p>\n\n<p>I have a feeling that the randomness of CUDNN layers has a pretty large impact on the scores. \nThe randomness/fluctuations of scores makes me feel really hard to know what is working and what is not.</p>",
      "rawMarkdown": "highest scoring model on private LB (which is of lower CV/PublicLB in the two submitted models)\nCV - 0.6910\nPublic LB - 0.69242\nPrivate LB - 0.70381\ncorrelation between fold predictions  is  around 0.95\n\nI have a feeling that the randomness of CUDNN layers has a pretty large impact on the scores. \nThe randomness/fluctuations of scores makes me feel really hard to know what is working and what is not.",
      "votes": null
    },
    {
      "id": "472023",
      "postDate": "02/15/2019 08:24:26",
      "content": "<p>Similar situation: my kernel with a cv (5fold) of ~0.69 also scored worse than the public kernel with a cv of 0.68. I used keras an didn't check for correlations between folds though. \nlocal cv: 0.69034\npublic lb: 0.69294\nprivate lb: 0.70196</p>",
      "rawMarkdown": "Similar situation: my kernel with a cv (5fold) of ~0.69 also scored worse than the public kernel with a cv of 0.68. I used keras an didn't check for correlations between folds though. \nlocal cv: 0.69034\npublic lb: 0.69294\nprivate lb: 0.70196",
      "votes": null
    },
    {
      "id": "472355",
      "postDate": "02/15/2019 18:57:06",
      "content": "<p>I think that it was my big mistake to trust the cv score fully. I have cv score 0.6938 and 0.689 public LB and 0.69935 private LB.  I had other kernel whose cv score was 0.692 and it scores 0.69980  in private LB(i used late submission to know this score).</p>\n\n<p>Recently I forked a public kernel whose cv score of 0.6804 and it gets 0.697 public LB and 0.70235 in private part(as you mentioned above).</p>\n\n<p>The current reason I think is that our real target is the output of the other model(quora's model) and due to that we are seeing this unpredicted behavior.</p>",
      "rawMarkdown": "I think that it was my big mistake to trust the cv score fully. I have cv score 0.6938 and 0.689 public LB and 0.69935 private LB.  I had other kernel whose cv score was 0.692 and it scores 0.69980  in private LB(i used late submission to know this score).\n\nRecently I forked a public kernel whose cv score of 0.6804 and it gets 0.697 public LB and 0.70235 in private part(as you mentioned above).\n\nThe current reason I think is that our real target is the output of the other model(quora's model) and due to that we are seeing this unpredicted behavior.",
      "votes": null
    },
    {
      "id": "473552",
      "postDate": "02/18/2019 07:10:46",
      "content": "<p>Highest local cv score - 0.687 which also get me the best private LB 0.705. I say trust your CV with some conditions. First of all, check if the training loss is lower than the val loss. I observed that the val loss is lower than the training loss for the first few epochs. If by \"trust your cv\" you mean to stop there, say epoch 3, then it's wrong.</p>",
      "rawMarkdown": "Highest local cv score - 0.687 which also get me the best private LB 0.705. I say trust your CV with some conditions. First of all, check if the training loss is lower than the val loss. I observed that the val loss is lower than the training loss for the first few epochs. If by \"trust your cv\" you mean to stop there, say epoch 3, then it's wrong.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471939,
      "author_name": "m7catsue",
      "author_url": "",
      "post_date": "02/15/2019 05:47:39",
      "content": "<p>highest scoring model on private LB (which is of lower CV/PublicLB in the two submitted models)\nCV - 0.6910\nPublic LB - 0.69242\nPrivate LB - 0.70381\ncorrelation between fold predictions  is  around 0.95</p>\n\n<p>I have a feeling that the randomness of CUDNN layers has a pretty large impact on the scores. \nThe randomness/fluctuations of scores makes me feel really hard to know what is working and what is not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472023,
      "author_name": "schlers",
      "author_url": "",
      "post_date": "02/15/2019 08:24:26",
      "content": "<p>Similar situation: my kernel with a cv (5fold) of ~0.69 also scored worse than the public kernel with a cv of 0.68. I used keras an didn't check for correlations between folds though. \nlocal cv: 0.69034\npublic lb: 0.69294\nprivate lb: 0.70196</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472355,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "02/15/2019 18:57:06",
      "content": "<p>I think that it was my big mistake to trust the cv score fully. I have cv score 0.6938 and 0.689 public LB and 0.69935 private LB.  I had other kernel whose cv score was 0.692 and it scores 0.69980  in private LB(i used late submission to know this score).</p>\n\n<p>Recently I forked a public kernel whose cv score of 0.6804 and it gets 0.697 public LB and 0.70235 in private part(as you mentioned above).</p>\n\n<p>The current reason I think is that our real target is the output of the other model(quora's model) and due to that we are seeing this unpredicted behavior.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 473552,
      "author_name": "wenrui29",
      "author_url": "",
      "post_date": "02/18/2019 07:10:46",
      "content": "<p>Highest local cv score - 0.687 which also get me the best private LB 0.705. I say trust your CV with some conditions. First of all, check if the training loss is lower than the val loss. I observed that the val loss is lower than the training loss for the first few epochs. If by \"trust your cv\" you mean to stop there, say epoch 3, then it's wrong.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471914": "Hi,\n\nDuring the competition, a lot of people (including me) were using their local validation score (cv or ensembling) as a way to figure out if their models had improved. However, after the competition ended, my best model with a cv score &gt;0.692 performed worse on the stage 2 test set than a public kernel that 230 people used as their final submission which got a cv score of 0.68.\n\nSince the competition is over, is anyone willing to share their best model's cv score and the corresponding private lb score? This might give everyone more insight into how to set up a reliable local cv to use in other competitions with small public test sets.\n\nThanks\n\nMy best kernel:\n\nlocal cv: 0.6927  \npublic lb: 0.6962  \nprivate lb: 0.70213  \n\nlocal roc-auc: 0.9697  \ncorrelations between test predictions from each fold: 0.956",
    "471939": "highest scoring model on private LB (which is of lower CV/PublicLB in the two submitted models)\nCV - 0.6910\nPublic LB - 0.69242\nPrivate LB - 0.70381\ncorrelation between fold predictions  is  around 0.95\n\nI have a feeling that the randomness of CUDNN layers has a pretty large impact on the scores. \nThe randomness/fluctuations of scores makes me feel really hard to know what is working and what is not.",
    "472023": "Similar situation: my kernel with a cv (5fold) of ~0.69 also scored worse than the public kernel with a cv of 0.68. I used keras an didn't check for correlations between folds though. \nlocal cv: 0.69034\npublic lb: 0.69294\nprivate lb: 0.70196",
    "472355": "I think that it was my big mistake to trust the cv score fully. I have cv score 0.6938 and 0.689 public LB and 0.69935 private LB.  I had other kernel whose cv score was 0.692 and it scores 0.69980  in private LB(i used late submission to know this score).\n\nRecently I forked a public kernel whose cv score of 0.6804 and it gets 0.697 public LB and 0.70235 in private part(as you mentioned above).\n\nThe current reason I think is that our real target is the output of the other model(quora's model) and due to that we are seeing this unpredicted behavior.",
    "473552": "Highest local cv score - 0.687 which also get me the best private LB 0.705. I say trust your CV with some conditions. First of all, check if the training loss is lower than the val loss. I observed that the val loss is lower than the training loss for the first few epochs. If by \"trust your cv\" you mean to stop there, say epoch 3, then it's wrong."
  },
  "source": "meta"
}