{
  "id": 131861,
  "title": "One Epoch Is All You Need(?)",
  "url": "/competitions/bengaliai-cv19/discussion/131861",
  "author_name": "",
  "post_date": "2020-02-22T07:36:23.529276500Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've been doing a lot of model tweaking and tuning. The process is really slow and annoying, but after having built out a decent number of experiments and letting them run to full convergence, I started wondering if there were a more expeditious means of identifying if the most recent 'tweak' was good to keep or should be rejected. What I've found is that, <em>for the most part</em>*, this behavior can be derived from the first epoch. There are some exceptions of course, for example increasing dropout or the amount of aug might in the short run degrade performance but then allow the model to consume more samples before over-fitting. But generally, what I've seen is that the first epoch's behavior is quite telling and it saves from having to run the full model out. Has anyone else observed similar behavior?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F933480%2F52f81735bd30ad6cfc8cd80104436a46%2Frun.png?generation=1582356972647833&amp;alt=media\" alt=\"minibatch losses\"></p>\n\n<p>As always, solid val is needed and the ability to reproduce results. Above are two runs, the blue one was a previous experiment and the brown one is a new experiment, the result of a single model tweak, currently running. What is logged are the minibatch gradients. Pictured are the train losses but what should be used per the above would be the holdout losses.</p>",
  "messages": [
    {
      "id": "753444",
      "postDate": "02/22/2020 07:36:23",
      "content": "<p>I've been doing a lot of model tweaking and tuning. The process is really slow and annoying, but after having built out a decent number of experiments and letting them run to full convergence, I started wondering if there were a more expeditious means of identifying if the most recent 'tweak' was good to keep or should be rejected. What I've found is that, <em>for the most part</em>*, this behavior can be derived from the first epoch. There are some exceptions of course, for example increasing dropout or the amount of aug might in the short run degrade performance but then allow the model to consume more samples before over-fitting. But generally, what I've seen is that the first epoch's behavior is quite telling and it saves from having to run the full model out. Has anyone else observed similar behavior?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F933480%2F52f81735bd30ad6cfc8cd80104436a46%2Frun.png?generation=1582356972647833&amp;alt=media\" alt=\"minibatch losses\"></p>\n\n<p>As always, solid val is needed and the ability to reproduce results. Above are two runs, the blue one was a previous experiment and the brown one is a new experiment, the result of a single model tweak, currently running. What is logged are the minibatch gradients. Pictured are the train losses but what should be used per the above would be the holdout losses.</p>",
      "rawMarkdown": "I've been doing a lot of model tweaking and tuning. The process is really slow and annoying, but after having built out a decent number of experiments and letting them run to full convergence, I started wondering if there were a more expeditious means of identifying if the most recent 'tweak' was good to keep or should be rejected. What I've found is that, _for the most part_*, this behavior can be derived from the first epoch. There are some exceptions of course, for example increasing dropout or the amount of aug might in the short run degrade performance but then allow the model to consume more samples before over-fitting. But generally, what I've seen is that the first epoch's behavior is quite telling and it saves from having to run the full model out. Has anyone else observed similar behavior?\n\n![minibatch losses](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F933480%2F52f81735bd30ad6cfc8cd80104436a46%2Frun.png?generation=1582356972647833&amp;alt=media)\n\nAs always, solid val is needed and the ability to reproduce results. Above are two runs, the blue one was a previous experiment and the brown one is a new experiment, the result of a single model tweak, currently running. What is logged are the minibatch gradients. Pictured are the train losses but what should be used per the above would be the holdout losses.",
      "votes": null
    },
    {
      "id": "753453",
      "postDate": "02/22/2020 07:57:51",
      "content": "<p>i have observed similar thing,if first epoch gets over 50% training accuracy then most of the times it does well in later epoches ,otherwise it  just keeps on overfitting no matter how long we train!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F2edaf17f481a46aba8dbbcb2299d43b9%2Ftweaking%20nn.gif?generation=1582358181898072&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "i have observed similar thing,if first epoch gets over 50% training accuracy then most of the times it does well in later epoches ,otherwise it  just keeps on overfitting no matter how long we train!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F2edaf17f481a46aba8dbbcb2299d43b9%2Ftweaking%20nn.gif?generation=1582358181898072&amp;alt=media)",
      "votes": null
    },
    {
      "id": "754791",
      "postDate": "02/24/2020 03:48:59",
      "content": "<p>With heavy augmentation in this competition, I'm afraid this observation doesn't work much. I used to compare a 50-epochs training performance. But once I've found that training for longer epochs make those seemly not-so-well experiments even much better.</p>",
      "rawMarkdown": "With heavy augmentation in this competition, I'm afraid this observation doesn't work much. I used to compare a 50-epochs training performance. But once I've found that training for longer epochs make those seemly not-so-well experiments even much better.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 753453,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "02/22/2020 07:57:51",
      "content": "<p>i have observed similar thing,if first epoch gets over 50% training accuracy then most of the times it does well in later epoches ,otherwise it  just keeps on overfitting no matter how long we train!</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F2edaf17f481a46aba8dbbcb2299d43b9%2Ftweaking%20nn.gif?generation=1582358181898072&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 754791,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "02/24/2020 03:48:59",
      "content": "<p>With heavy augmentation in this competition, I'm afraid this observation doesn't work much. I used to compare a 50-epochs training performance. But once I've found that training for longer epochs make those seemly not-so-well experiments even much better.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "753444": "I've been doing a lot of model tweaking and tuning. The process is really slow and annoying, but after having built out a decent number of experiments and letting them run to full convergence, I started wondering if there were a more expeditious means of identifying if the most recent 'tweak' was good to keep or should be rejected. What I've found is that, _for the most part_*, this behavior can be derived from the first epoch. There are some exceptions of course, for example increasing dropout or the amount of aug might in the short run degrade performance but then allow the model to consume more samples before over-fitting. But generally, what I've seen is that the first epoch's behavior is quite telling and it saves from having to run the full model out. Has anyone else observed similar behavior?\n\n![minibatch losses](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F933480%2F52f81735bd30ad6cfc8cd80104436a46%2Frun.png?generation=1582356972647833&amp;alt=media)\n\nAs always, solid val is needed and the ability to reproduce results. Above are two runs, the blue one was a previous experiment and the brown one is a new experiment, the result of a single model tweak, currently running. What is logged are the minibatch gradients. Pictured are the train losses but what should be used per the above would be the holdout losses.",
    "753453": "i have observed similar thing,if first epoch gets over 50% training accuracy then most of the times it does well in later epoches ,otherwise it  just keeps on overfitting no matter how long we train!\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F2edaf17f481a46aba8dbbcb2299d43b9%2Ftweaking%20nn.gif?generation=1582358181898072&amp;alt=media)",
    "754791": "With heavy augmentation in this competition, I'm afraid this observation doesn't work much. I used to compare a 50-epochs training performance. But once I've found that training for longer epochs make those seemly not-so-well experiments even much better."
  },
  "source": "meta"
}