{
  "id": 20852,
  "title": "Local score highly underestimates loss",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20852",
  "author_name": "",
  "post_date": "2016-05-11T07:58:50.907Z",
  "votes": null,
  "comment_count": 2,
  "views": 668,
  "content": "<p>To estimate the loss locally before uploading any model I splitted test and train sets taking care of the fact that there should not be the same people in those sets, to do that I put all p064, p66, p72, p75, p81 images in the test set and everything else in the train set. Now, I've been able to achieve a loss of 0.54 with a model that locally achieved just 1.4, what might the reasons be?</p>\n\n<hr>\n\n<p>Edit: Sorry for the title, I meant &quot;overestimates&quot;</p>",
  "messages": [
    {
      "id": "119532",
      "postDate": "05/11/2016 07:58:50",
      "content": "<p>To estimate the loss locally before uploading any model I splitted test and train sets taking care of the fact that there should not be the same people in those sets, to do that I put all p064, p66, p72, p75, p81 images in the test set and everything else in the train set. Now, I've been able to achieve a loss of 0.54 with a model that locally achieved just 1.4, what might the reasons be?</p>\n\n<hr>\n\n<p>Edit: Sorry for the title, I meant &quot;overestimates&quot;</p>",
      "rawMarkdown": "To estimate the loss locally before uploading any model I splitted test and train sets taking care of the fact that there should not be the same people in those sets, to do that I put all p064, p66, p72, p75, p81 images in the test set and everything else in the train set. Now, I've been able to achieve a loss of 0.54 with a model that locally achieved just 1.4, what might the reasons be?\r\n_______________\r\nEdit: Sorry for the title, I meant \"overestimates\"",
      "votes": null
    },
    {
      "id": "120998",
      "postDate": "05/22/2016 15:32:09",
      "content": "<p>I think one reason it could be overestimating the loss is that the validation set is too hard.  To quote myself earlier: While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar.  That is, in the small training set, the drivers look more distinct to one another, but in the large test dataset, some drivers look similar to those in the training set.</p>\n\n<p>I personally don't just get overestimates, but just general variability, as mentioned here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss</a></p>",
      "rawMarkdown": "I think one reason it could be overestimating the loss is that the validation set is too hard.  To quote myself earlier: While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar.  That is, in the small training set, the drivers look more distinct to one another, but in the large test dataset, some drivers look similar to those in the training set.\r\n\r\nI personally don't just get overestimates, but just general variability, as mentioned here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss",
      "votes": null
    },
    {
      "id": "121049",
      "postDate": "05/23/2016 09:03:04",
      "content": "<p>If you splitted only training data for validation and got 1.4 then it makes sense.... During CV test set and train set drivers were distinct but when you predicted on actual test set, there were drivers from training set that were in test set thus you got lower LB than CV</p>",
      "rawMarkdown": "If you splitted only training data for validation and got 1.4 then it makes sense.... During CV test set and train set drivers were distinct but when you predicted on actual test set, there were drivers from training set that were in test set thus you got lower LB than CV",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 120998,
      "author_name": "jonathankchang",
      "author_url": "",
      "post_date": "05/22/2016 15:32:09",
      "content": "<p>I think one reason it could be overestimating the loss is that the validation set is too hard.  To quote myself earlier: While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar.  That is, in the small training set, the drivers look more distinct to one another, but in the large test dataset, some drivers look similar to those in the training set.</p>\n\n<p>I personally don't just get overestimates, but just general variability, as mentioned here:</p>\n\n<p><a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 121049,
      "author_name": "shahnawazakhtar",
      "author_url": "",
      "post_date": "05/23/2016 09:03:04",
      "content": "<p>If you splitted only training data for validation and got 1.4 then it makes sense.... During CV test set and train set drivers were distinct but when you predicted on actual test set, there were drivers from training set that were in test set thus you got lower LB than CV</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119532": "To estimate the loss locally before uploading any model I splitted test and train sets taking care of the fact that there should not be the same people in those sets, to do that I put all p064, p66, p72, p75, p81 images in the test set and everything else in the train set. Now, I've been able to achieve a loss of 0.54 with a model that locally achieved just 1.4, what might the reasons be?\r\n_______________\r\nEdit: Sorry for the title, I meant \"overestimates\"",
    "120998": "I think one reason it could be overestimating the loss is that the validation set is too hard.  To quote myself earlier: While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar.  That is, in the small training set, the drivers look more distinct to one another, but in the large test dataset, some drivers look similar to those in the training set.\r\n\r\nI personally don't just get overestimates, but just general variability, as mentioned here:\r\n\r\nhttps://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20936/high-instability-in-local-validation-loss",
    "121049": "If you splitted only training data for validation and got 1.4 then it makes sense.... During CV test set and train set drivers were distinct but when you predicted on actual test set, there were drivers from training set that were in test set thus you got lower LB than CV"
  },
  "source": "meta"
}