{
  "id": 29914,
  "title": "Loss increases drastically on submission, don't understand why",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/29914",
  "author_name": "",
  "post_date": "2017-03-11T14:06:24.160811900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>My loss for my validation set while training is always around 0.04 or something similar and when I submit my results on gaggle the loss jumps to 1.9. I don't really understand why this happens, if someone could explain that would be good. </p>\n\n<p>What exactly is wrong with the training? Is it overfitting? </p>",
  "messages": [
    {
      "id": "166854",
      "postDate": "03/11/2017 14:06:24",
      "content": "<p>My loss for my validation set while training is always around 0.04 or something similar and when I submit my results on gaggle the loss jumps to 1.9. I don't really understand why this happens, if someone could explain that would be good. </p>\n\n<p>What exactly is wrong with the training? Is it overfitting? </p>",
      "rawMarkdown": "My loss for my validation set while training is always around 0.04 or something similar and when I submit my results on gaggle the loss jumps to 1.9. I don't really understand why this happens, if someone could explain that would be good. \n\nWhat exactly is wrong with the training? Is it overfitting?",
      "votes": null
    },
    {
      "id": "166856",
      "postDate": "03/11/2017 14:10:24",
      "content": "<p>Do you use validation by drivers? \nIf you haven't try using leave one group out validation based on drivers - this will likely improve your validation process and will prevent your model from overfitting the training data</p>",
      "rawMarkdown": "Do you use validation by drivers? \nIf you haven't try using leave one group out validation based on drivers - this will likely improve your validation process and will prevent your model from overfitting the training data",
      "votes": null
    },
    {
      "id": "166857",
      "postDate": "03/11/2017 14:16:35",
      "content": "<p>No I haven't tried that. I will though; you mean k-fold validation? or you mean using drivers to make sure the data is randomized properly? \nBut overfitting is the probable reason for this right ? </p>",
      "rawMarkdown": "No I haven't tried that. I will though; you mean k-fold validation? or you mean using drivers to make sure the data is randomized properly? \nBut overfitting is the probable reason for this right ?",
      "votes": null
    },
    {
      "id": "166860",
      "postDate": "03/11/2017 14:32:35",
      "content": "<p>I mean using drivers to make sure the data is randomized properly. \nAnd yes overfitting is the probable reason for this.\nCheck out scikit-learn leave one group out for an easy implementation</p>",
      "rawMarkdown": "I mean using drivers to make sure the data is randomized properly. \nAnd yes overfitting is the probable reason for this.\nCheck out scikit-learn leave one group out for an easy implementation",
      "votes": null
    },
    {
      "id": "175974",
      "postDate": "04/18/2017 10:17:08",
      "content": "<p>I'm facing the same problem with you, do you solve this problem? Does Nathan's advice work?</p>",
      "rawMarkdown": "I'm facing the same problem with you, do you solve this problem? Does Nathan's advice work?",
      "votes": null
    },
    {
      "id": "333175",
      "postDate": "05/24/2018 14:54:50",
      "content": "<p>It's depend what kind of loss have you choose in your CNN - kaggle use logloss..</p>",
      "rawMarkdown": "It's depend what kind of loss have you choose in your CNN - kaggle use logloss..",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 166856,
      "author_name": "interesting",
      "author_url": "",
      "post_date": "03/11/2017 14:10:24",
      "content": "<p>Do you use validation by drivers? \nIf you haven't try using leave one group out validation based on drivers - this will likely improve your validation process and will prevent your model from overfitting the training data</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 166857,
      "author_name": "amyhass",
      "author_url": "",
      "post_date": "03/11/2017 14:16:35",
      "content": "<p>No I haven't tried that. I will though; you mean k-fold validation? or you mean using drivers to make sure the data is randomized properly? \nBut overfitting is the probable reason for this right ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 166860,
          "author_name": "interesting",
          "author_url": "",
          "post_date": "03/11/2017 14:32:35",
          "content": "<p>I mean using drivers to make sure the data is randomized properly. \nAnd yes overfitting is the probable reason for this.\nCheck out scikit-learn leave one group out for an easy implementation</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 175974,
      "author_name": "justinho",
      "author_url": "",
      "post_date": "04/18/2017 10:17:08",
      "content": "<p>I'm facing the same problem with you, do you solve this problem? Does Nathan's advice work?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 333175,
      "author_name": "",
      "author_url": "",
      "post_date": "05/24/2018 14:54:50",
      "content": "<p>It's depend what kind of loss have you choose in your CNN - kaggle use logloss..</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "166854": "My loss for my validation set while training is always around 0.04 or something similar and when I submit my results on gaggle the loss jumps to 1.9. I don't really understand why this happens, if someone could explain that would be good. \n\nWhat exactly is wrong with the training? Is it overfitting?",
    "166856": "Do you use validation by drivers? \nIf you haven't try using leave one group out validation based on drivers - this will likely improve your validation process and will prevent your model from overfitting the training data",
    "166857": "No I haven't tried that. I will though; you mean k-fold validation? or you mean using drivers to make sure the data is randomized properly? \nBut overfitting is the probable reason for this right ?",
    "166860": "I mean using drivers to make sure the data is randomized properly. \nAnd yes overfitting is the probable reason for this.\nCheck out scikit-learn leave one group out for an easy implementation",
    "175974": "I'm facing the same problem with you, do you solve this problem? Does Nathan's advice work?",
    "333175": "It's depend what kind of loss have you choose in your CNN - kaggle use logloss.."
  },
  "source": "meta"
}