{
  "id": 77268,
  "title": "Why does local cv score differ so much from lb score?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77268",
  "author_name": "Jeffery Chen",
  "post_date": "2019-01-11T02:43:34.067000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I tried to stack some well-performing model and consistently reached a f1 score of about 0.70 on multiple different holdout data that my model has never seen before. However, my score for lb test data is consistently about 0.687. Is there anything about lb test data that differ from the labeled data?</p>\n\n<p><a href=\"https://www.kaggle.com/jf2333/how-should-we-validate-our-models\">https://www.kaggle.com/jf2333/how-should-we-validate-our-models</a></p>",
  "messages": [
    {
      "id": 454005,
      "postDate": "2019-01-11T02:43:34.067Z",
      "content": "<p>I tried to stack some well-performing model and consistently reached a f1 score of about 0.70 on multiple different holdout data that my model has never seen before. However, my score for lb test data is consistently about 0.687. Is there anything about lb test data that differ from the labeled data?</p>\n\n<p><a href=\"https://www.kaggle.com/jf2333/how-should-we-validate-our-models\">https://www.kaggle.com/jf2333/how-should-we-validate-our-models</a></p>",
      "rawMarkdown": "I tried to stack some well-performing model and consistently reached a f1 score of about 0.70 on multiple different holdout data that my model has never seen before. However, my score for lb test data is consistently about 0.687. Is there anything about lb test data that differ from the labeled data?\n\nhttps://www.kaggle.com/jf2333/how-should-we-validate-our-models",
      "votes": 4
    },
    {
      "id": 454543,
      "postDate": "2019-01-11T20:05:40.757Z",
      "content": "<p>The lb score is definitely an outlier from the 20 holdout sets I tested. Maybe I should test on more splits but I still doubt the public test data is somehow different.</p>",
      "rawMarkdown": "The lb score is definitely an outlier from the 20 holdout sets I tested. Maybe I should test on more splits but I still doubt the public test data is somehow different.",
      "votes": 1
    },
    {
      "id": 454245,
      "postDate": "2019-01-11T09:54:38.580Z",
      "content": "<p>The holdout sets (several realizations of random splits) should have the same size of 56 k rows as the public test set. Then you can assess the variance of f1 from your holdout runs and see whether your LB score is within the same range. Did you fix size of holdout to 56 k rows ?</p>",
      "rawMarkdown": "The holdout sets (several realizations of random splits) should have the same size of 56 k rows as the public test set. Then you can assess the variance of f1 from your holdout runs and see whether your LB score is within the same range. Did you fix size of holdout to 56 k rows ?",
      "isDeleted": true,
      "replies": [
        {
          "id": 454537,
          "postDate": "2019-01-11T19:57:35.363Z",
          "content": "<p>Thank you! Yes, I controlled the holdout set to be within range of around 52k rows and ran the model on multiple different holdout sets, but the results are not really promising. I will try to calculate the variance to see if its within range. Also maybe I should try 56k rows :)</p>",
          "rawMarkdown": "Thank you! Yes, I controlled the holdout set to be within range of around 52k rows and ran the model on multiple different holdout sets, but the results are not really promising. I will try to calculate the variance to see if its within range. Also maybe I should try 56k rows :)"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 454543,
      "author_name": "Jeffery Chen",
      "author_url": "",
      "post_date": "2019-01-11T20:05:40.757000",
      "content": "<p>The lb score is definitely an outlier from the 20 holdout sets I tested. Maybe I should test on more splits but I still doubt the public test data is somehow different.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 454245,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-11T09:54:38.580000",
      "content": "<p>The holdout sets (several realizations of random splits) should have the same size of 56 k rows as the public test set. Then you can assess the variance of f1 from your holdout runs and see whether your LB score is within the same range. Did you fix size of holdout to 56 k rows ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454537,
          "author_name": "Jeffery Chen",
          "author_url": "",
          "post_date": "2019-01-11T19:57:35.363000",
          "content": "<p>Thank you! Yes, I controlled the holdout set to be within range of around 52k rows and ran the model on multiple different holdout sets, but the results are not really promising. I will try to calculate the variance to see if its within range. Also maybe I should try 56k rows :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "454005": "I tried to stack some well-performing model and consistently reached a f1 score of about 0.70 on multiple different holdout data that my model has never seen before. However, my score for lb test data is consistently about 0.687. Is there anything about lb test data that differ from the labeled data?\n\nhttps://www.kaggle.com/jf2333/how-should-we-validate-our-models",
    "454543": "The lb score is definitely an outlier from the 20 holdout sets I tested. Maybe I should test on more splits but I still doubt the public test data is somehow different.",
    "454245": "The holdout sets (several realizations of random splits) should have the same size of 56 k rows as the public test set. Then you can assess the variance of f1 from your holdout runs and see whether your LB score is within the same range. Did you fix size of holdout to 56 k rows ?"
  }
}