{
  "id": 44814,
  "title": "Validation accuracy greater than training accuracy",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44814",
  "author_name": "",
  "post_date": "2017-12-02T21:40:39.075474100Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am following the tensorflow tutorial. While training, I am consistently observing validation accuracy (calculated every 400 steps) greater than training accuracy. Why?</p>\n\n<p>Samples are assigned to sets according to <code>which_set</code> function which ensures same speakers are in same set.</p>\n\n<p>Is it because we are modifying (almost 80% of) training set (by adding noise and time stretching etc)?</p>",
  "messages": [
    {
      "id": "252412",
      "postDate": "12/02/2017 21:40:39",
      "content": "<p>I am following the tensorflow tutorial. While training, I am consistently observing validation accuracy (calculated every 400 steps) greater than training accuracy. Why?</p>\n\n<p>Samples are assigned to sets according to <code>which_set</code> function which ensures same speakers are in same set.</p>\n\n<p>Is it because we are modifying (almost 80% of) training set (by adding noise and time stretching etc)?</p>",
      "rawMarkdown": "I am following the tensorflow tutorial. While training, I am consistently observing validation accuracy (calculated every 400 steps) greater than training accuracy. Why?\n\nSamples are assigned to sets according to ```which_set``` function which ensures same speakers are in same set.\n\nIs it because we are modifying (almost 80% of) training set (by adding noise and time stretching etc)?",
      "votes": null
    },
    {
      "id": "252421",
      "postDate": "12/02/2017 22:11:19",
      "content": "<p>You are right, its because regularization and data augmentation are only applied for the train.\nAnother possible reason is that the eval performance is computed at the end of the epoch when the model became better than when it was applied to train.</p>",
      "rawMarkdown": "You are right, its because regularization and data augmentation are only applied for the train.\nAnother possible reason is that the eval performance is computed at the end of the epoch when the model became better than when it was applied to train.",
      "votes": null
    },
    {
      "id": "252493",
      "postDate": "12/03/2017 02:45:41",
      "content": "<p>Actually I figured out the issue. In the script - <a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py</a></p>\n\n<p>on line 243, <code>validation_writer.add_summary(validation_summary, training_step)</code>, runs inside a loop. This loop runs (validation set size/ batch size) number of times. But the value of <code>training_step</code> is same  in the loop. So instead of writing average over all batches, the summary writer is writing validation accuracies for each validation batch at same training step.</p>\n\n<p>On tensorboard, when you set a particular value of smoothing, you see as if validation accuracy is higher than training accuracy with these spikes as shown in attached screen shot. But if you remove the smoothing you will see that both accuracies are similar. </p>",
      "rawMarkdown": "Actually I figured out the issue. In the script - https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py\n\non line 243, ```validation_writer.add_summary(validation_summary, training_step)```, runs inside a loop. This loop runs (validation set size/ batch size) number of times. But the value of ```training_step``` is same  in the loop. So instead of writing average over all batches, the summary writer is writing validation accuracies for each validation batch at same training step.\n\nOn tensorboard, when you set a particular value of smoothing, you see as if validation accuracy is higher than training accuracy with these spikes as shown in attached screen shot. But if you remove the smoothing you will see that both accuracies are similar.",
      "votes": null
    },
    {
      "id": "253042",
      "postDate": "12/04/2017 11:10:09",
      "content": "<p>The script is not wrong. It's a good practice to evaluate the validation set after every epoch (or after every few training steps). That's why the validation line on tensorboard is not as smoothing as the train one.</p>",
      "rawMarkdown": "The script is not wrong. It's a good practice to evaluate the validation set after every epoch (or after every few training steps). That's why the validation line on tensorboard is not as smoothing as the train one.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 252421,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "12/02/2017 22:11:19",
      "content": "<p>You are right, its because regularization and data augmentation are only applied for the train.\nAnother possible reason is that the eval performance is computed at the end of the epoch when the model became better than when it was applied to train.</p>",
      "votes": null,
      "replies": [
        {
          "id": 252493,
          "author_name": "tendolkar3",
          "author_url": "",
          "post_date": "12/03/2017 02:45:41",
          "content": "<p>Actually I figured out the issue. In the script - <a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py</a></p>\n\n<p>on line 243, <code>validation_writer.add_summary(validation_summary, training_step)</code>, runs inside a loop. This loop runs (validation set size/ batch size) number of times. But the value of <code>training_step</code> is same  in the loop. So instead of writing average over all batches, the summary writer is writing validation accuracies for each validation batch at same training step.</p>\n\n<p>On tensorboard, when you set a particular value of smoothing, you see as if validation accuracy is higher than training accuracy with these spikes as shown in attached screen shot. But if you remove the smoothing you will see that both accuracies are similar. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253042,
          "author_name": "barbolo",
          "author_url": "",
          "post_date": "12/04/2017 11:10:09",
          "content": "<p>The script is not wrong. It's a good practice to evaluate the validation set after every epoch (or after every few training steps). That's why the validation line on tensorboard is not as smoothing as the train one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "252412": "I am following the tensorflow tutorial. While training, I am consistently observing validation accuracy (calculated every 400 steps) greater than training accuracy. Why?\n\nSamples are assigned to sets according to ```which_set``` function which ensures same speakers are in same set.\n\nIs it because we are modifying (almost 80% of) training set (by adding noise and time stretching etc)?",
    "252421": "You are right, its because regularization and data augmentation are only applied for the train.\nAnother possible reason is that the eval performance is computed at the end of the epoch when the model became better than when it was applied to train.",
    "252493": "Actually I figured out the issue. In the script - https://github.com/tensorflow/tensorflow/blob/master/tensorflow/examples/speech_commands/train.py\n\non line 243, ```validation_writer.add_summary(validation_summary, training_step)```, runs inside a loop. This loop runs (validation set size/ batch size) number of times. But the value of ```training_step``` is same  in the loop. So instead of writing average over all batches, the summary writer is writing validation accuracies for each validation batch at same training step.\n\nOn tensorboard, when you set a particular value of smoothing, you see as if validation accuracy is higher than training accuracy with these spikes as shown in attached screen shot. But if you remove the smoothing you will see that both accuracies are similar.",
    "253042": "The script is not wrong. It's a good practice to evaluate the validation set after every epoch (or after every few training steps). That's why the validation line on tensorboard is not as smoothing as the train one."
  },
  "source": "meta"
}