{
  "id": 193360,
  "title": "Training Loss vs Eval Loss deviation issue - the simplest solution",
  "url": "/competitions/lyft-motion-prediction-autonomous-vehicles/discussion/193360",
  "author_name": "",
  "post_date": "2020-10-26T17:27:21.398085100Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>This is a seemingly trivial point regarding the train/eval loss deviation that was not really raised in previous discussions regarding this issue:</p>\n<p>If anyone applied and understood the given points in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">this discussion</a>, is using chopped dataset, has good evaluation/LB correlation, but is still getting large training loss deviation, check how are you logging the loss:</p>\n<ul>\n<li>loss per batch/iteration, <code>loss.item()</code></li>\n<li>mean loss; append loss to a list, eg <code>training_loss = []</code> every iteration, calculate mean loss <code>np.mean(training_loss)</code>, log the mean loss value</li>\n</ul>\n<p>If you are only logging loss per batch (<code>loss.item()</code>), you are not seeing the mean loss during training which of course will result in a much smaller value than you will see during evaluation and LB submission.</p>\n<p>This is a simple point and probably implemented by many, but a number of people seem to be missing it and are digging a bit deeper than actually required for this specific issue.</p>",
  "messages": [
    {
      "id": "1061002",
      "postDate": "10/26/2020 17:27:21",
      "content": "<p>This is a seemingly trivial point regarding the train/eval loss deviation that was not really raised in previous discussions regarding this issue:</p>\n<p>If anyone applied and understood the given points in <a href=\"https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762\" target=\"_blank\">this discussion</a>, is using chopped dataset, has good evaluation/LB correlation, but is still getting large training loss deviation, check how are you logging the loss:</p>\n<ul>\n<li>loss per batch/iteration, <code>loss.item()</code></li>\n<li>mean loss; append loss to a list, eg <code>training_loss = []</code> every iteration, calculate mean loss <code>np.mean(training_loss)</code>, log the mean loss value</li>\n</ul>\n<p>If you are only logging loss per batch (<code>loss.item()</code>), you are not seeing the mean loss during training which of course will result in a much smaller value than you will see during evaluation and LB submission.</p>\n<p>This is a simple point and probably implemented by many, but a number of people seem to be missing it and are digging a bit deeper than actually required for this specific issue.</p>",
      "rawMarkdown": "This is a seemingly trivial point regarding the train/eval loss deviation that was not really raised in previous discussions regarding this issue:\n\nIf anyone applied and understood the given points in [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762), is using chopped dataset, has good evaluation/LB correlation, but is still getting large training loss deviation, check how are you logging the loss:\n\n- loss per batch/iteration, `loss.item()`\n- mean loss; append loss to a list, eg `training_loss = []` every iteration, calculate mean loss `np.mean(training_loss)`, log the mean loss value\n\nIf you are only logging loss per batch (`loss.item()`), you are not seeing the mean loss during training which of course will result in a much smaller value than you will see during evaluation and LB submission.\n\nThis is a simple point and probably implemented by many, but a number of people seem to be missing it and are digging a bit deeper than actually required for this specific issue.",
      "votes": null
    },
    {
      "id": "1061020",
      "postDate": "10/26/2020 17:39:49",
      "content": "<p>imho, it's good to do something like np.mean(training_loss[-N:]), where N is 1K or 10K.<br>\nit should give much better estimate of both val and LB when experimenting with chunked dataset (theoretically :) ) </p>",
      "rawMarkdown": "imho, it's good to do something like np.mean(training_loss[-N:]), where N is 1K or 10K.\nit should give much better estimate of both val and LB when experimenting with chunked dataset (theoretically :) )",
      "votes": null
    },
    {
      "id": "1061024",
      "postDate": "10/26/2020 17:42:46",
      "content": "<p>I was looking at the pandas dataframe.Describe() of my train and validation losses and realized something similar. I was recording mean of batches on train but the error on individual samples on validation. Took me a while to figure out why the max and std of my train was so much lower. </p>",
      "rawMarkdown": "I was looking at the pandas dataframe.Describe() of my train and validation losses and realized something similar. I was recording mean of batches on train but the error on individual samples on validation. Took me a while to figure out why the max and std of my train was so much lower.",
      "votes": null
    },
    {
      "id": "1061040",
      "postDate": "10/26/2020 17:52:26",
      "content": "<p>Yes, also took me a while to realize that. In the end, it ended up being that simple</p>",
      "rawMarkdown": "Yes, also took me a while to realize that. In the end, it ended up being that simple",
      "votes": null
    },
    {
      "id": "1061046",
      "postDate": "10/26/2020 17:55:41",
      "content": "<p>I was thinking something similar, since overall mean loss is not 100% accurate due to the large values at initial iterations. But I also haven't tried it - not a lot of time left for experimentation and I'm trying to keep my metric consistent:)</p>",
      "rawMarkdown": "I was thinking something similar, since overall mean loss is not 100% accurate due to the large values at initial iterations. But I also haven't tried it - not a lot of time left for experimentation and I'm trying to keep my metric consistent:)",
      "votes": null
    },
    {
      "id": "1061107",
      "postDate": "10/26/2020 18:44:01",
      "content": "<p>I am always doing it with large datasets when one or less than one full epoch is sufficient for some experimentation. Then, in addition to validation, training losses and metrics are equally good indications because your model sees the training data for the first time.</p>\n<p>Having little computing resources forces you to squeeze juice wherever you can.</p>",
      "rawMarkdown": "I am always doing it with large datasets when one or less than one full epoch is sufficient for some experimentation. Then, in addition to validation, training losses and metrics are equally good indications because your model sees the training data for the first time.\n\nHaving little computing resources forces you to squeeze juice wherever you can.",
      "votes": null
    },
    {
      "id": "1065220",
      "postDate": "10/31/2020 03:32:08",
      "content": "<p>In fact, you only have to log your loss data in tensorboard file, because you can adjust the smoothing parameter which is implemented with exponentially weighted averages(not pretty sure, but it has a similar function).</p>",
      "rawMarkdown": "In fact, you only have to log your loss data in tensorboard file, because you can adjust the smoothing parameter which is implemented with exponentially weighted averages(not pretty sure, but it has a similar function).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1061020,
      "author_name": "valanm",
      "author_url": "",
      "post_date": "10/26/2020 17:39:49",
      "content": "<p>imho, it's good to do something like np.mean(training_loss[-N:]), where N is 1K or 10K.<br>\nit should give much better estimate of both val and LB when experimenting with chunked dataset (theoretically :) ) </p>",
      "votes": null,
      "replies": [
        {
          "id": 1061046,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/26/2020 17:55:41",
          "content": "<p>I was thinking something similar, since overall mean loss is not 100% accurate due to the large values at initial iterations. But I also haven't tried it - not a lot of time left for experimentation and I'm trying to keep my metric consistent:)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1061107,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "10/26/2020 18:44:01",
          "content": "<p>I am always doing it with large datasets when one or less than one full epoch is sufficient for some experimentation. Then, in addition to validation, training losses and metrics are equally good indications because your model sees the training data for the first time.</p>\n<p>Having little computing resources forces you to squeeze juice wherever you can.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1065220,
          "author_name": "spicychicken38",
          "author_url": "",
          "post_date": "10/31/2020 03:32:08",
          "content": "<p>In fact, you only have to log your loss data in tensorboard file, because you can adjust the smoothing parameter which is implemented with exponentially weighted averages(not pretty sure, but it has a similar function).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1061024,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "10/26/2020 17:42:46",
      "content": "<p>I was looking at the pandas dataframe.Describe() of my train and validation losses and realized something similar. I was recording mean of batches on train but the error on individual samples on validation. Took me a while to figure out why the max and std of my train was so much lower. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1061040,
          "author_name": "indswetrust",
          "author_url": "",
          "post_date": "10/26/2020 17:52:26",
          "content": "<p>Yes, also took me a while to realize that. In the end, it ended up being that simple</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1061002": "This is a seemingly trivial point regarding the train/eval loss deviation that was not really raised in previous discussions regarding this issue:\n\nIf anyone applied and understood the given points in [this discussion](https://www.kaggle.com/c/lyft-motion-prediction-autonomous-vehicles/discussion/185762), is using chopped dataset, has good evaluation/LB correlation, but is still getting large training loss deviation, check how are you logging the loss:\n\n- loss per batch/iteration, `loss.item()`\n- mean loss; append loss to a list, eg `training_loss = []` every iteration, calculate mean loss `np.mean(training_loss)`, log the mean loss value\n\nIf you are only logging loss per batch (`loss.item()`), you are not seeing the mean loss during training which of course will result in a much smaller value than you will see during evaluation and LB submission.\n\nThis is a simple point and probably implemented by many, but a number of people seem to be missing it and are digging a bit deeper than actually required for this specific issue.",
    "1061020": "imho, it's good to do something like np.mean(training_loss[-N:]), where N is 1K or 10K.\nit should give much better estimate of both val and LB when experimenting with chunked dataset (theoretically :) )",
    "1061024": "I was looking at the pandas dataframe.Describe() of my train and validation losses and realized something similar. I was recording mean of batches on train but the error on individual samples on validation. Took me a while to figure out why the max and std of my train was so much lower.",
    "1061040": "Yes, also took me a while to realize that. In the end, it ended up being that simple",
    "1061046": "I was thinking something similar, since overall mean loss is not 100% accurate due to the large values at initial iterations. But I also haven't tried it - not a lot of time left for experimentation and I'm trying to keep my metric consistent:)",
    "1061107": "I am always doing it with large datasets when one or less than one full epoch is sufficient for some experimentation. Then, in addition to validation, training losses and metrics are equally good indications because your model sees the training data for the first time.\n\nHaving little computing resources forces you to squeeze juice wherever you can.",
    "1065220": "In fact, you only have to log your loss data in tensorboard file, because you can adjust the smoothing parameter which is implemented with exponentially weighted averages(not pretty sure, but it has a similar function)."
  },
  "source": "meta"
}