{
  "id": 73663,
  "title": "score between local cv and LB",
  "url": "/competitions/quora-insincere-questions-classification/discussion/73663",
  "author_name": "",
  "post_date": "2018-12-04T18:17:03.058741800Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>the question is: </p>\n\n<p>someone:  local cv score is much lower than LB, like cv (0.6771, 0.6740, 0.6812, 0.6751, ...) and got LB (0.693)</p>\n\n<p>others: cv (0.6779, 0.6821, 0.6860, 0.6883, 0.6817), only got LB (0.683)</p>\n\n<p>and I want to know:</p>\n\n<ul>\n<li>why low cv score can get high LB score?</li>\n<li>why high cv score get low LB score?</li>\n</ul>",
  "messages": [
    {
      "id": "433127",
      "postDate": "12/04/2018 18:17:03",
      "content": "<p>the question is: </p>\n\n<p>someone:  local cv score is much lower than LB, like cv (0.6771, 0.6740, 0.6812, 0.6751, ...) and got LB (0.693)</p>\n\n<p>others: cv (0.6779, 0.6821, 0.6860, 0.6883, 0.6817), only got LB (0.683)</p>\n\n<p>and I want to know:</p>\n\n<ul>\n<li>why low cv score can get high LB score?</li>\n<li>why high cv score get low LB score?</li>\n</ul>",
      "rawMarkdown": "the question is: \n\nsomeone:  local cv score is much lower than LB, like cv (0.6771, 0.6740, 0.6812, 0.6751, ...) and got LB (0.693)\n\nothers: cv (0.6779, 0.6821, 0.6860, 0.6883, 0.6817), only got LB (0.683)\n\nand I want to know:\n\n- why low cv score can get high LB score?\n- why high cv score get low LB score?",
      "votes": null
    },
    {
      "id": "433237",
      "postDate": "12/04/2018 21:58:33",
      "content": "<p>Well I think there are many factors. One is variance in CuDNN implementation and the others can possibly be differences between train and test sets (as in different distributions etc)</p>",
      "rawMarkdown": "Well I think there are many factors. One is variance in CuDNN implementation and the others can possibly be differences between train and test sets (as in different distributions etc)",
      "votes": null
    },
    {
      "id": "433265",
      "postDate": "12/04/2018 23:44:35",
      "content": "<p>This phenomenon sometimes happen on my model too. I guess the distribution in train set is not totally same to the test distribution. And I think the threshold of F1 plays an important role. High cv score but low threshold may get lower LB score.  </p>",
      "rawMarkdown": "This phenomenon sometimes happen on my model too. I guess the distribution in train set is not totally same to the test distribution. And I think the threshold of F1 plays an important role. High cv score but low threshold may get lower LB score.",
      "votes": null
    },
    {
      "id": "433375",
      "postDate": "12/05/2018 02:29:41",
      "content": "<p>Thank you and I'm going to set the seed, make the val set consistent on every runs to see whether the score boost.</p>",
      "rawMarkdown": "Thank you and I'm going to set the seed, make the val set consistent on every runs to see whether the score boost.",
      "votes": null
    },
    {
      "id": "433376",
      "postDate": "12/05/2018 02:29:47",
      "content": "<p>Thank you. I will have a try.</p>",
      "rawMarkdown": "Thank you. I will have a try.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 433237,
      "author_name": "rasvob",
      "author_url": "",
      "post_date": "12/04/2018 21:58:33",
      "content": "<p>Well I think there are many factors. One is variance in CuDNN implementation and the others can possibly be differences between train and test sets (as in different distributions etc)</p>",
      "votes": null,
      "replies": [
        {
          "id": 433375,
          "author_name": "syhens",
          "author_url": "",
          "post_date": "12/05/2018 02:29:41",
          "content": "<p>Thank you and I'm going to set the seed, make the val set consistent on every runs to see whether the score boost.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 433265,
      "author_name": "kaggleczs",
      "author_url": "",
      "post_date": "12/04/2018 23:44:35",
      "content": "<p>This phenomenon sometimes happen on my model too. I guess the distribution in train set is not totally same to the test distribution. And I think the threshold of F1 plays an important role. High cv score but low threshold may get lower LB score.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 433376,
          "author_name": "syhens",
          "author_url": "",
          "post_date": "12/05/2018 02:29:47",
          "content": "<p>Thank you. I will have a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "433127": "the question is: \n\nsomeone:  local cv score is much lower than LB, like cv (0.6771, 0.6740, 0.6812, 0.6751, ...) and got LB (0.693)\n\nothers: cv (0.6779, 0.6821, 0.6860, 0.6883, 0.6817), only got LB (0.683)\n\nand I want to know:\n\n- why low cv score can get high LB score?\n- why high cv score get low LB score?",
    "433237": "Well I think there are many factors. One is variance in CuDNN implementation and the others can possibly be differences between train and test sets (as in different distributions etc)",
    "433265": "This phenomenon sometimes happen on my model too. I guess the distribution in train set is not totally same to the test distribution. And I think the threshold of F1 plays an important role. High cv score but low threshold may get lower LB score.",
    "433375": "Thank you and I'm going to set the seed, make the val set consistent on every runs to see whether the score boost.",
    "433376": "Thank you. I will have a try."
  },
  "source": "meta"
}