{
  "id": 20936,
  "title": "High instability in local validation loss?",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/20936",
  "author_name": "",
  "post_date": "2016-05-13T14:45:27.777Z",
  "votes": 4,
  "comment_count": 2,
  "views": 913,
  "content": "<p>My validation loss per epoch jumps around a lot from epoch to epoch, though a low pass filtered version of it does seem to generally trend down. The training loss is very smooth.</p>\n\n<p>What does that signify?</p>\n\n<p>Ideas:</p>\n\n<ul>\n<li>Too small of a validation set (10%) </li>\n<li>Not enough data for the problem complexity </li>\n<li>Validation set is unnecessarily hard</li>\n<li>Overfitting</li>\n</ul>\n\n<p>Validation set details</p>\n\n<ul>\n<li>A cursory look at the train/test didn't show drivers in both the training set and test set</li>\n<li>each driver's class distribution is approximately equal in the training set</li>\n<li>To test generalization on people it has never seen for the real world, and to match the construction of the test set from the above details, validation loss is selected by randomly selecting drivers and moving them and all of their classes into train or validation</li>\n<li>dividing train/test randomly doesn't have this issue, but of course yields  is due to the correlations in train and test</li>\n<li>dividing randomly by driver + action (instead of just driver) would likely yield smoother validation loss but seems to not match the test set construction</li>\n</ul>\n\n<p>Other details</p>\n\n<ul>\n<li>Reducing the learning rate reduces the variability. Note that I am using an &quot;automatic&quot; learning rate optimizer (ADAM)</li>\n<li>Does this mean its overfitting? Though the general trend of the validation loss is still going down.</li>\n</ul>\n\n<p>Is the validation set unnecessarily hard?  While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar, and given the amount of training data, selecting a validation set by driver+action may be okay.</p>\n\n<p>I like to think of the public LB as something I should make few decisions on to prevent overfitting vs the private LB.  However, validating my validation set quantitatively may make sense here: I could submit evaluations of weights in neighboring epochs with large variability in validation loss, to see if its reflected in the LB test set.</p>",
  "messages": [
    {
      "id": "119881",
      "postDate": "05/13/2016 14:45:27",
      "content": "<p>My validation loss per epoch jumps around a lot from epoch to epoch, though a low pass filtered version of it does seem to generally trend down. The training loss is very smooth.</p>\n\n<p>What does that signify?</p>\n\n<p>Ideas:</p>\n\n<ul>\n<li>Too small of a validation set (10%) </li>\n<li>Not enough data for the problem complexity </li>\n<li>Validation set is unnecessarily hard</li>\n<li>Overfitting</li>\n</ul>\n\n<p>Validation set details</p>\n\n<ul>\n<li>A cursory look at the train/test didn't show drivers in both the training set and test set</li>\n<li>each driver's class distribution is approximately equal in the training set</li>\n<li>To test generalization on people it has never seen for the real world, and to match the construction of the test set from the above details, validation loss is selected by randomly selecting drivers and moving them and all of their classes into train or validation</li>\n<li>dividing train/test randomly doesn't have this issue, but of course yields  is due to the correlations in train and test</li>\n<li>dividing randomly by driver + action (instead of just driver) would likely yield smoother validation loss but seems to not match the test set construction</li>\n</ul>\n\n<p>Other details</p>\n\n<ul>\n<li>Reducing the learning rate reduces the variability. Note that I am using an &quot;automatic&quot; learning rate optimizer (ADAM)</li>\n<li>Does this mean its overfitting? Though the general trend of the validation loss is still going down.</li>\n</ul>\n\n<p>Is the validation set unnecessarily hard?  While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar, and given the amount of training data, selecting a validation set by driver+action may be okay.</p>\n\n<p>I like to think of the public LB as something I should make few decisions on to prevent overfitting vs the private LB.  However, validating my validation set quantitatively may make sense here: I could submit evaluations of weights in neighboring epochs with large variability in validation loss, to see if its reflected in the LB test set.</p>",
      "rawMarkdown": "My validation loss per epoch jumps around a lot from epoch to epoch, though a low pass filtered version of it does seem to generally trend down. The training loss is very smooth.\r\n\r\nWhat does that signify?\r\n\r\nIdeas:\r\n\r\n - Too small of a validation set (10%) \r\n - Not enough data for the problem complexity \r\n - Validation set is unnecessarily hard\r\n - Overfitting\r\n\r\nValidation set details\r\n\r\n - A cursory look at the train/test didn't show drivers in both the training set and test set\r\n - each driver's class distribution is approximately equal in the training set\r\n - To test generalization on people it has never seen for the real world, and to match the construction of the test set from the above details, validation loss is selected by randomly selecting drivers and moving them and all of their classes into train or validation\r\n - dividing train/test randomly doesn't have this issue, but of course yields  is due to the correlations in train and test\r\n - dividing randomly by driver + action (instead of just driver) would likely yield smoother validation loss but seems to not match the test set construction\r\n\r\nOther details\r\n\r\n - Reducing the learning rate reduces the variability. Note that I am using an \"automatic\" learning rate optimizer (ADAM)\r\n - Does this mean its overfitting? Though the general trend of the validation loss is still going down.\r\n \r\nIs the validation set unnecessarily hard?  While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar, and given the amount of training data, selecting a validation set by driver+action may be okay.\r\n\r\nI like to think of the public LB as something I should make few decisions on to prevent overfitting vs the private LB.  However, validating my validation set quantitatively may make sense here: I could submit evaluations of weights in neighboring epochs with large variability in validation loss, to see if its reflected in the LB test set.",
      "votes": null
    },
    {
      "id": "119887",
      "postDate": "05/13/2016 15:10:25",
      "content": "<p>@JonathanKChang you are in 22nd place what seems to be the problem here :)... on a side note what kind of validation accuracy are you getting on your validation? Which framework are you using?  I have been training for 6+hrs now no pre-trained network and my  most recent epoch</p>\n\n<p>training loss:                1.340084</p>\n\n<p>validation loss:              1.175541</p>\n\n<p>validation accuracy:          55.95 %</p>\n\n<p>Progress seems to be slow at this point and I am wondering if my net is too small.  What validation accuracy are the people at the top of the leaderboard getting for single models.</p>",
      "rawMarkdown": "JonathanKChang you are in 22nd place what seems to be the problem here :)... on a side note what kind of validation accuracy are you getting on your validation? Which framework are you using?  I have been training for 6+hrs now no pre-trained network and my  most recent epoch\r\n\r\n training loss:                1.340084\r\n\r\n  validation loss:              1.175541\r\n\r\n  validation accuracy:          55.95 %\r\n\r\n\r\nProgress seems to be slow at this point and I am wondering if my net is too small.  What validation accuracy are the people at the top of the leaderboard getting for single models.",
      "votes": null
    },
    {
      "id": "120055",
      "postDate": "05/15/2016 00:42:45",
      "content": "<p>I'm getting validation losses in the 0.2 to 1.0 range. Sometimes it agrees with the leaderboard, other times it disagrees in either direction.  100% random validation losses are of course near 0, along with the training loss.</p>\n\n<p>I'm using theano and lasagne</p>\n\n<p>Given your training loss, I'd say your net is too small and needs more capacity.  You should be able to overfit your training set pretty well! You should verify you can do that, then you'll want to combat overfitting.</p>",
      "rawMarkdown": "I'm getting validation losses in the 0.2 to 1.0 range. Sometimes it agrees with the leaderboard, other times it disagrees in either direction.  100% random validation losses are of course near 0, along with the training loss.\r\n\r\nI'm using theano and lasagne\r\n\r\nGiven your training loss, I'd say your net is too small and needs more capacity.  You should be able to overfit your training set pretty well! You should verify you can do that, then you'll want to combat overfitting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 119887,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "05/13/2016 15:10:25",
      "content": "<p>@JonathanKChang you are in 22nd place what seems to be the problem here :)... on a side note what kind of validation accuracy are you getting on your validation? Which framework are you using?  I have been training for 6+hrs now no pre-trained network and my  most recent epoch</p>\n\n<p>training loss:                1.340084</p>\n\n<p>validation loss:              1.175541</p>\n\n<p>validation accuracy:          55.95 %</p>\n\n<p>Progress seems to be slow at this point and I am wondering if my net is too small.  What validation accuracy are the people at the top of the leaderboard getting for single models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 120055,
      "author_name": "jonathankchang",
      "author_url": "",
      "post_date": "05/15/2016 00:42:45",
      "content": "<p>I'm getting validation losses in the 0.2 to 1.0 range. Sometimes it agrees with the leaderboard, other times it disagrees in either direction.  100% random validation losses are of course near 0, along with the training loss.</p>\n\n<p>I'm using theano and lasagne</p>\n\n<p>Given your training loss, I'd say your net is too small and needs more capacity.  You should be able to overfit your training set pretty well! You should verify you can do that, then you'll want to combat overfitting.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "119881": "My validation loss per epoch jumps around a lot from epoch to epoch, though a low pass filtered version of it does seem to generally trend down. The training loss is very smooth.\r\n\r\nWhat does that signify?\r\n\r\nIdeas:\r\n\r\n - Too small of a validation set (10%) \r\n - Not enough data for the problem complexity \r\n - Validation set is unnecessarily hard\r\n - Overfitting\r\n\r\nValidation set details\r\n\r\n - A cursory look at the train/test didn't show drivers in both the training set and test set\r\n - each driver's class distribution is approximately equal in the training set\r\n - To test generalization on people it has never seen for the real world, and to match the construction of the test set from the above details, validation loss is selected by randomly selecting drivers and moving them and all of their classes into train or validation\r\n - dividing train/test randomly doesn't have this issue, but of course yields  is due to the correlations in train and test\r\n - dividing randomly by driver + action (instead of just driver) would likely yield smoother validation loss but seems to not match the test set construction\r\n\r\nOther details\r\n\r\n - Reducing the learning rate reduces the variability. Note that I am using an \"automatic\" learning rate optimizer (ADAM)\r\n - Does this mean its overfitting? Though the general trend of the validation loss is still going down.\r\n \r\nIs the validation set unnecessarily hard?  While I didn't see the same drivers in the training set vs test set, perhaps there drivers (and the way they dress) that look very similar, and given the amount of training data, selecting a validation set by driver+action may be okay.\r\n\r\nI like to think of the public LB as something I should make few decisions on to prevent overfitting vs the private LB.  However, validating my validation set quantitatively may make sense here: I could submit evaluations of weights in neighboring epochs with large variability in validation loss, to see if its reflected in the LB test set.",
    "119887": "JonathanKChang you are in 22nd place what seems to be the problem here :)... on a side note what kind of validation accuracy are you getting on your validation? Which framework are you using?  I have been training for 6+hrs now no pre-trained network and my  most recent epoch\r\n\r\n training loss:                1.340084\r\n\r\n  validation loss:              1.175541\r\n\r\n  validation accuracy:          55.95 %\r\n\r\n\r\nProgress seems to be slow at this point and I am wondering if my net is too small.  What validation accuracy are the people at the top of the leaderboard getting for single models.",
    "120055": "I'm getting validation losses in the 0.2 to 1.0 range. Sometimes it agrees with the leaderboard, other times it disagrees in either direction.  100% random validation losses are of course near 0, along with the training loss.\r\n\r\nI'm using theano and lasagne\r\n\r\nGiven your training loss, I'd say your net is too small and needs more capacity.  You should be able to overfit your training set pretty well! You should verify you can do that, then you'll want to combat overfitting."
  },
  "source": "meta"
}