{
  "id": 72932,
  "title": "Underestimated learning curves and bias/variance problem",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72932",
  "author_name": "BRad",
  "post_date": "2018-11-28T11:56:42.846000",
  "votes": 2,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many kernels concentrate on sophisticated NN architectures but a common pattern can be found in all notebooks: train on ~3 epochs and batch size between 512 and 2048. Looks like a wise setup because training longer leads to bias/variance issues\nHowever, it doesn't cover available kernel time span - training takes about 20 minutes. In further development a simple increment of epochs won't help.  That's what learning curves showed. \nNo one seems to check (in public kernels) bias/variance regarding train/validation set. </p>\n\n<p>What's your thoughts on that? Do you concentrate on filling available time and address overfit in your private kernels?</p>\n\n<p><img src=\"https://i.imgur.com/8N4Hxbn.png\" alt=\"learning curves\"></p>",
  "messages": [
    {
      "id": 429139,
      "postDate": "2018-11-28T11:56:42.847Z",
      "content": "<p>Many kernels concentrate on sophisticated NN architectures but a common pattern can be found in all notebooks: train on ~3 epochs and batch size between 512 and 2048. Looks like a wise setup because training longer leads to bias/variance issues\nHowever, it doesn't cover available kernel time span - training takes about 20 minutes. In further development a simple increment of epochs won't help.  That's what learning curves showed. \nNo one seems to check (in public kernels) bias/variance regarding train/validation set. </p>\n\n<p>What's your thoughts on that? Do you concentrate on filling available time and address overfit in your private kernels?</p>\n\n<p><img src=\"https://i.imgur.com/8N4Hxbn.png\" alt=\"learning curves\"></p>",
      "rawMarkdown": "Many kernels concentrate on sophisticated NN architectures but a common pattern can be found in all notebooks: train on ~3 epochs and batch size between 512 and 2048. Looks like a wise setup because training longer leads to bias/variance issues\nHowever, it doesn't cover available kernel time span - training takes about 20 minutes. In further development a simple increment of epochs won't help.  That's what learning curves showed. \nNo one seems to check (in public kernels) bias/variance regarding train/validation set. \n\nWhat's your thoughts on that? Do you concentrate on filling available time and address overfit in your private kernels?\n\n![learning curves][1]\n\n\n  [1]: https://i.imgur.com/8N4Hxbn.png",
      "votes": 2
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "429139": "Many kernels concentrate on sophisticated NN architectures but a common pattern can be found in all notebooks: train on ~3 epochs and batch size between 512 and 2048. Looks like a wise setup because training longer leads to bias/variance issues\nHowever, it doesn't cover available kernel time span - training takes about 20 minutes. In further development a simple increment of epochs won't help.  That's what learning curves showed. \nNo one seems to check (in public kernels) bias/variance regarding train/validation set. \n\nWhat's your thoughts on that? Do you concentrate on filling available time and address overfit in your private kernels?\n\n![learning curves][1]\n\n\n  [1]: https://i.imgur.com/8N4Hxbn.png"
  }
}