{
  "id": 101997,
  "title": "pytorch model.eval seems suboptimal",
  "url": "/competitions/recursion-cellular-image-classification/discussion/101997",
  "author_name": "",
  "post_date": "2019-07-30T10:22:46.005002700Z",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>just found out, that using model.eval() almost halves my testset accuracy.. \naccording to the discussions in the interwebs, this could be due to the fact, that the batch norm layers use some sort of moving average over the testset distributions when in eval.mode, \nwhile in train mode the average of the current batch is used.\nWith small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode. \nIs anyone else seeing this? or is it a sign that my model is unstable?\nare you all just using train mode for predictions? <br>\nI thought eval mode should be better, but apparently not...</p>",
  "messages": [
    {
      "id": "588243",
      "postDate": "07/30/2019 10:22:46",
      "content": "<p>just found out, that using model.eval() almost halves my testset accuracy.. \naccording to the discussions in the interwebs, this could be due to the fact, that the batch norm layers use some sort of moving average over the testset distributions when in eval.mode, \nwhile in train mode the average of the current batch is used.\nWith small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode. \nIs anyone else seeing this? or is it a sign that my model is unstable?\nare you all just using train mode for predictions? <br>\nI thought eval mode should be better, but apparently not...</p>",
      "rawMarkdown": "just found out, that using model.eval() almost halves my testset accuracy.. \naccording to the discussions in the interwebs, this could be due to the fact, that the batch norm layers use some sort of moving average over the testset distributions when in eval.mode, \nwhile in train mode the average of the current batch is used.\nWith small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode. \nIs anyone else seeing this? or is it a sign that my model is unstable?\nare you all just using train mode for predictions?  \nI thought eval mode should be better, but apparently not...",
      "votes": null
    },
    {
      "id": "588300",
      "postDate": "07/30/2019 11:52:49",
      "content": "<p>interesting finding 4nna</p>\n\n<blockquote>\n  <p>With small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode.</p>\n</blockquote>\n\n<p>if you think that's the case you could reduce the momentum of your batch norm so it takes more batches into consideration:\n<a href=\"https://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html\">https://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html</a></p>",
      "rawMarkdown": "interesting finding 4nna\n&gt; With small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode.\n\nif you think that's the case you could reduce the momentum of your batch norm so it takes more batches into consideration:\nhttps://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html",
      "votes": null
    },
    {
      "id": "588369",
      "postDate": "07/30/2019 13:24:00",
      "content": "<p>On explanation could be, that if you are doing backpropagation and the experiments in your test set are ordered one after each other (no shuffle), you adapt the BN parameters to the experiment statistics during the first batches of the experiment data and use the optimized BN parameters for the classification. (BN statistics are calculated in training mode for each minibatch according to the pytorch docs.) When you are at the next experiment in the test data the BN parameters get adapted to the next experiment and so on.</p>\n\n<p>If you are not doing backpropagation then it is maybe caused by the BN statistics based on the mini-batches which seem to help the normalization of the unshuffled data.</p>\n\n<p>However, if the test dataset is shuffled this should not work anymore (if the above theory is true).</p>\n\n<p>Indeed a interesting finding!\nI wonder how that finding can be utilized for model training optimization?</p>",
      "rawMarkdown": "On explanation could be, that if you are doing backpropagation and the experiments in your test set are ordered one after each other (no shuffle), you adapt the BN parameters to the experiment statistics during the first batches of the experiment data and use the optimized BN parameters for the classification. (BN statistics are calculated in training mode for each minibatch according to the pytorch docs.) When you are at the next experiment in the test data the BN parameters get adapted to the next experiment and so on.\n\nIf you are not doing backpropagation then it is maybe caused by the BN statistics based on the mini-batches which seem to help the normalization of the unshuffled data.\n\nHowever, if the test dataset is shuffled this should not work anymore (if the above theory is true).\n\nIndeed a interesting finding!\nI wonder how that finding can be utilized for model training optimization?",
      "votes": null
    },
    {
      "id": "588402",
      "postDate": "07/30/2019 14:28:35",
      "content": "<p>hi <a href=\"/hmendonca\">@hmendonca</a>, is it possible to do only reduce momentum during evaluation? The <code>eval</code> function doesnt take any arguments: <a href=\"https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval\">https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval</a></p>\n\n<p>Found a workaround: <a href=\"https://github.com/pytorch/pytorch/issues/4741\">https://github.com/pytorch/pytorch/issues/4741</a></p>",
      "rawMarkdown": "hi @hmendonca, is it possible to do only reduce momentum during evaluation? The `eval` function doesnt take any arguments: https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval\n\nFound a workaround: https://github.com/pytorch/pytorch/issues/4741",
      "votes": null
    },
    {
      "id": "588638",
      "postDate": "07/30/2019 21:06:33",
      "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> BN uses momentum 0 during eval. I.e. it doesn't update the statistics doing validation.</p>",
      "rawMarkdown": "wjshenggggg BN uses momentum 0 during eval. I.e. it doesn't update the statistics doing validation.",
      "votes": null
    },
    {
      "id": "588737",
      "postDate": "07/31/2019 01:12:49",
      "content": "<p>Someone from the forum recommended setting <code>track_running_stats</code> to False</p>",
      "rawMarkdown": "Someone from the forum recommended setting `track_running_stats` to False",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 588300,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "07/30/2019 11:52:49",
      "content": "<p>interesting finding 4nna</p>\n\n<blockquote>\n  <p>With small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode.</p>\n</blockquote>\n\n<p>if you think that's the case you could reduce the momentum of your batch norm so it takes more batches into consideration:\n<a href=\"https://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html\">https://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 588402,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/30/2019 14:28:35",
          "content": "<p>hi <a href=\"/hmendonca\">@hmendonca</a>, is it possible to do only reduce momentum during evaluation? The <code>eval</code> function doesnt take any arguments: <a href=\"https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval\">https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval</a></p>\n\n<p>Found a workaround: <a href=\"https://github.com/pytorch/pytorch/issues/4741\">https://github.com/pytorch/pytorch/issues/4741</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588638,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "07/30/2019 21:06:33",
          "content": "<p><a href=\"/wjshenggggg\">@wjshenggggg</a> BN uses momentum 0 during eval. I.e. it doesn't update the statistics doing validation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588737,
          "author_name": "wjshenggggg",
          "author_url": "",
          "post_date": "07/31/2019 01:12:49",
          "content": "<p>Someone from the forum recommended setting <code>track_running_stats</code> to False</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 588369,
      "author_name": "micpie",
      "author_url": "",
      "post_date": "07/30/2019 13:24:00",
      "content": "<p>On explanation could be, that if you are doing backpropagation and the experiments in your test set are ordered one after each other (no shuffle), you adapt the BN parameters to the experiment statistics during the first batches of the experiment data and use the optimized BN parameters for the classification. (BN statistics are calculated in training mode for each minibatch according to the pytorch docs.) When you are at the next experiment in the test data the BN parameters get adapted to the next experiment and so on.</p>\n\n<p>If you are not doing backpropagation then it is maybe caused by the BN statistics based on the mini-batches which seem to help the normalization of the unshuffled data.</p>\n\n<p>However, if the test dataset is shuffled this should not work anymore (if the above theory is true).</p>\n\n<p>Indeed a interesting finding!\nI wonder how that finding can be utilized for model training optimization?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "588243": "just found out, that using model.eval() almost halves my testset accuracy.. \naccording to the discussions in the interwebs, this could be due to the fact, that the batch norm layers use some sort of moving average over the testset distributions when in eval.mode, \nwhile in train mode the average of the current batch is used.\nWith small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode. \nIs anyone else seeing this? or is it a sign that my model is unstable?\nare you all just using train mode for predictions?  \nI thought eval mode should be better, but apparently not...",
    "588300": "interesting finding 4nna\n&gt; With small batch-sizes and slightly unstable(?) distributions this can lead to lower testset accuracy when in eval-mode.\n\nif you think that's the case you could reduce the momentum of your batch norm so it takes more batches into consideration:\nhttps://pytorch.org/docs/stable/_modules/torch/nn/modules/batchnorm.html",
    "588369": "On explanation could be, that if you are doing backpropagation and the experiments in your test set are ordered one after each other (no shuffle), you adapt the BN parameters to the experiment statistics during the first batches of the experiment data and use the optimized BN parameters for the classification. (BN statistics are calculated in training mode for each minibatch according to the pytorch docs.) When you are at the next experiment in the test data the BN parameters get adapted to the next experiment and so on.\n\nIf you are not doing backpropagation then it is maybe caused by the BN statistics based on the mini-batches which seem to help the normalization of the unshuffled data.\n\nHowever, if the test dataset is shuffled this should not work anymore (if the above theory is true).\n\nIndeed a interesting finding!\nI wonder how that finding can be utilized for model training optimization?",
    "588402": "hi @hmendonca, is it possible to do only reduce momentum during evaluation? The `eval` function doesnt take any arguments: https://pytorch.org/docs/master/_modules/torch/nn/modules/module.html#Module.eval\n\nFound a workaround: https://github.com/pytorch/pytorch/issues/4741",
    "588638": "wjshenggggg BN uses momentum 0 during eval. I.e. it doesn't update the statistics doing validation.",
    "588737": "Someone from the forum recommended setting `track_running_stats` to False"
  },
  "source": "meta"
}