{
  "id": 21713,
  "title": "Overfitted Models",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/21713",
  "author_name": "",
  "post_date": "2016-06-16T03:15:09.517Z",
  "votes": null,
  "comment_count": 3,
  "views": 716,
  "content": "<p>Hi:</p>\n\n<p>I see a pattern in my results that I can't fully explain. I appreciate if someone can provide an explanation. When I train a model on the whole training dataset, I am able to get a very good accuracy and low log loss (~0.1-0.2) which is the evaluation metric of the competition but when I submit my results on the test dataset, I get much higher log loss (~2). I don't fully understand why. I expect some level of overfitting and worse log loss on the test dataset but since both training and test datasets are large, I don't expect to see this much difference between the loss for the two datasets.</p>\n\n<p>Am I missing something? Is this expected?</p>\n\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "124186",
      "postDate": "06/16/2016 03:15:09",
      "content": "<p>Hi:</p>\n\n<p>I see a pattern in my results that I can't fully explain. I appreciate if someone can provide an explanation. When I train a model on the whole training dataset, I am able to get a very good accuracy and low log loss (~0.1-0.2) which is the evaluation metric of the competition but when I submit my results on the test dataset, I get much higher log loss (~2). I don't fully understand why. I expect some level of overfitting and worse log loss on the test dataset but since both training and test datasets are large, I don't expect to see this much difference between the loss for the two datasets.</p>\n\n<p>Am I missing something? Is this expected?</p>\n\n<p>Thanks.</p>",
      "rawMarkdown": "Hi:\r\n\r\nI see a pattern in my results that I can't fully explain. I appreciate if someone can provide an explanation. When I train a model on the whole training dataset, I am able to get a very good accuracy and low log loss (~0.1-0.2) which is the evaluation metric of the competition but when I submit my results on the test dataset, I get much higher log loss (~2). I don't fully understand why. I expect some level of overfitting and worse log loss on the test dataset but since both training and test datasets are large, I don't expect to see this much difference between the loss for the two datasets.\r\n\r\nAm I missing something? Is this expected?\r\n\r\nThanks.",
      "votes": null
    },
    {
      "id": "124188",
      "postDate": "06/16/2016 03:24:24",
      "content": "<p>That is exactly overfitting. If you are evaluating your model on your training set, the log loss you get is a very biased estimation. What happens when you train your model on 25 of the drivers and evaluate on the remaining 1?</p>",
      "rawMarkdown": "That is exactly overfitting. If you are evaluating your model on your training set, the log loss you get is a very biased estimation. What happens when you train your model on 25 of the drivers and evaluate on the remaining 1?",
      "votes": null
    },
    {
      "id": "124291",
      "postDate": "06/16/2016 20:28:54",
      "content": "<p>I have actually tried what you mentioned: basically using a portion of photos in the training dataset for training a model and using the rest for evaluation (e.g. 4000 for training and 2000 for evaluation and 10000 for training and 2000 for evaluation) and I don't see such a discrepancy in loss values. So I still don't know what the issue is here.  </p>",
      "rawMarkdown": "I have actually tried what you mentioned: basically using a portion of photos in the training dataset for training a model and using the rest for evaluation (e.g. 4000 for training and 2000 for evaluation and 10000 for training and 2000 for evaluation) and I don't see such a discrepancy in loss values. So I still don't know what the issue is here.",
      "votes": null
    },
    {
      "id": "124307",
      "postDate": "06/17/2016 01:33:06",
      "content": "<p>Thanks Lucian. Ah, this is indeed possible because we have very limited number of drivers. I hadn't realized this. This issue is also discussed in this thread: <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad</a> </p>",
      "rawMarkdown": "Thanks Lucian. Ah, this is indeed possible because we have very limited number of drivers. I hadn't realized this. This issue is also discussed in this thread: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 124188,
      "author_name": "ilucian",
      "author_url": "",
      "post_date": "06/16/2016 03:24:24",
      "content": "<p>That is exactly overfitting. If you are evaluating your model on your training set, the log loss you get is a very biased estimation. What happens when you train your model on 25 of the drivers and evaluate on the remaining 1?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124291,
      "author_name": "rombuk74",
      "author_url": "",
      "post_date": "06/16/2016 20:28:54",
      "content": "<p>I have actually tried what you mentioned: basically using a portion of photos in the training dataset for training a model and using the rest for evaluation (e.g. 4000 for training and 2000 for evaluation and 10000 for training and 2000 for evaluation) and I don't see such a discrepancy in loss values. So I still don't know what the issue is here.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 124307,
      "author_name": "rombuk74",
      "author_url": "",
      "post_date": "06/17/2016 01:33:06",
      "content": "<p>Thanks Lucian. Ah, this is indeed possible because we have very limited number of drivers. I hadn't realized this. This issue is also discussed in this thread: <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad\">https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "124186": "Hi:\r\n\r\nI see a pattern in my results that I can't fully explain. I appreciate if someone can provide an explanation. When I train a model on the whole training dataset, I am able to get a very good accuracy and low log loss (~0.1-0.2) which is the evaluation metric of the competition but when I submit my results on the test dataset, I get much higher log loss (~2). I don't fully understand why. I expect some level of overfitting and worse log loss on the test dataset but since both training and test datasets are large, I don't expect to see this much difference between the loss for the two datasets.\r\n\r\nAm I missing something? Is this expected?\r\n\r\nThanks.",
    "124188": "That is exactly overfitting. If you are evaluating your model on your training set, the log loss you get is a very biased estimation. What happens when you train your model on 25 of the drivers and evaluate on the remaining 1?",
    "124291": "I have actually tried what you mentioned: basically using a portion of photos in the training dataset for training a model and using the rest for evaluation (e.g. 4000 for training and 2000 for evaluation and 10000 for training and 2000 for evaluation) and I don't see such a discrepancy in loss values. So I still don't know what the issue is here.",
    "124307": "Thanks Lucian. Ah, this is indeed possible because we have very limited number of drivers. I hadn't realized this. This issue is also discussed in this thread: https://www.kaggle.com/c/state-farm-distracted-driver-detection/forums/t/20775/why-is-my-leaderboard-score-so-bad"
  },
  "source": "meta"
}