{
  "id": 75108,
  "title": "Possible problem with most kernels",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75108",
  "author_name": "",
  "post_date": "2018-12-18T16:22:52.777017200Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>I looked at some of the high scoring kernels and most of them use the model's weights at the last trained epoch instead of choosing the best performing epoch (lowest val loss) to find the best threshold and predict on the test set. </p>\n\n<p>By doing this, the model used to predict on the test set should be overfitting to the training set and should generalize worse to the val or test sets than choosing the epoch with the lowest validation loss. Also, the \"best threshold\" found wouldn't be the best threshold, but the best threshold to the validation set when overfitting to the train set.</p>\n\n<p>The kernels that I looked at are:</p>\n\n<p><a href=\"https://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora\">https://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora</a>\n<a href=\"https://www.kaggle.com/hengzheng/pytorch-starter\">https://www.kaggle.com/hengzheng/pytorch-starter</a>\n<a href=\"https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\">https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch</a>\n<a href=\"https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\">https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb</a>\n<a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a></p>",
  "messages": [
    {
      "id": "441412",
      "postDate": "12/18/2018 16:22:52",
      "content": "<p>Hi,</p>\n\n<p>I looked at some of the high scoring kernels and most of them use the model's weights at the last trained epoch instead of choosing the best performing epoch (lowest val loss) to find the best threshold and predict on the test set. </p>\n\n<p>By doing this, the model used to predict on the test set should be overfitting to the training set and should generalize worse to the val or test sets than choosing the epoch with the lowest validation loss. Also, the \"best threshold\" found wouldn't be the best threshold, but the best threshold to the validation set when overfitting to the train set.</p>\n\n<p>The kernels that I looked at are:</p>\n\n<p><a href=\"https://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora\">https://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora</a>\n<a href=\"https://www.kaggle.com/hengzheng/pytorch-starter\">https://www.kaggle.com/hengzheng/pytorch-starter</a>\n<a href=\"https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\">https://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch</a>\n<a href=\"https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\">https://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb</a>\n<a href=\"https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr\">https://www.kaggle.com/shujian/single-rnn-with-4-folds-clr</a></p>",
      "rawMarkdown": "Hi,\n\nI looked at some of the high scoring kernels and most of them use the model's weights at the last trained epoch instead of choosing the best performing epoch (lowest val loss) to find the best threshold and predict on the test set. \n\nBy doing this, the model used to predict on the test set should be overfitting to the training set and should generalize worse to the val or test sets than choosing the epoch with the lowest validation loss. Also, the \"best threshold\" found wouldn't be the best threshold, but the best threshold to the validation set when overfitting to the train set.\n\nThe kernels that I looked at are:\n\nhttps://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora\nhttps://www.kaggle.com/hengzheng/pytorch-starter\nhttps://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\nhttps://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\nhttps://www.kaggle.com/shujian/single-rnn-with-4-folds-clr",
      "votes": null
    },
    {
      "id": "441453",
      "postDate": "12/18/2018 17:11:59",
      "content": "<p>You are generally right. However, there are some recent papers that remark on the curious fact that NNs often improve their validation <strong>accuracy</strong> even after the validation <strong>loss</strong> has started increasing again. So training for 1 or 2 epochs past the \"overfit\" may actually be beneficial for the final performance.</p>\n\n<p>Anyway, apart from that, the art of this competition is to get great models in a short time. You should tune your models exactly so that in the given number of epochs, they end up being at their best. For off-line experimentation, training for many epochs and then selecting the best model may be fine. Here, it would be a waste of precious time.</p>\n\n<p>Since this tuning of the model and training process is hard work and a science in itself, the kernels posted here usually are not perfect. I believe you could probably get very far by taking one of the best kernels and then doing this tuning yourself.</p>",
      "rawMarkdown": "You are generally right. However, there are some recent papers that remark on the curious fact that NNs often improve their validation **accuracy** even after the validation **loss** has started increasing again. So training for 1 or 2 epochs past the \"overfit\" may actually be beneficial for the final performance.\n\nAnyway, apart from that, the art of this competition is to get great models in a short time. You should tune your models exactly so that in the given number of epochs, they end up being at their best. For off-line experimentation, training for many epochs and then selecting the best model may be fine. Here, it would be a waste of precious time.\n\nSince this tuning of the model and training process is hard work and a science in itself, the kernels posted here usually are not perfect. I believe you could probably get very far by taking one of the best kernels and then doing this tuning yourself.",
      "votes": null
    },
    {
      "id": "441713",
      "postDate": "12/19/2018 00:21:31",
      "content": "<p>Agreed.</p>",
      "rawMarkdown": "Agreed.",
      "votes": null
    },
    {
      "id": "442916",
      "postDate": "12/20/2018 17:39:19",
      "content": "<p>This is sounds counter intuitive to me. Isn't the loss supposed to act as a proxy for the accuracy? So far, whenever the loss decreased in my models, the accuracy increased (except of some small deviations). Could you share the paper(s) that explain the behavior you mentioned?</p>",
      "rawMarkdown": "This is sounds counter intuitive to me. Isn't the loss supposed to act as a proxy for the accuracy? So far, whenever the loss decreased in my models, the accuracy increased (except of some small deviations). Could you share the paper(s) that explain the behavior you mentioned?",
      "votes": null
    },
    {
      "id": "442926",
      "postDate": "12/20/2018 18:08:19",
      "content": "<p>Can you send a link to papers that claim that? Would be quite fun to read.</p>",
      "rawMarkdown": "Can you send a link to papers that claim that? Would be quite fun to read.",
      "votes": null
    },
    {
      "id": "443266",
      "postDate": "12/21/2018 10:14:45",
      "content": "<p>Unfortunately, I can't find any at the moment. I'm very certain to have read that in at least one paper recently... 🤔 Maybe someone else can help?</p>\n\n<p>EDIT: I definitely remember Andrew Ng saying in his DL course that he generally does not use early stopping. If your training process is tuned correctly, it should always be better to train longer.</p>",
      "rawMarkdown": "Unfortunately, I can't find any at the moment. I'm very certain to have read that in at least one paper recently... 🤔 Maybe someone else can help?\n\nEDIT: I definitely remember Andrew Ng saying in his DL course that he generally does not use early stopping. If your training process is tuned correctly, it should always be better to train longer.",
      "votes": null
    },
    {
      "id": "446370",
      "postDate": "12/28/2018 01:06:40",
      "content": "<p>I also remember him saying that, but I think it was said during arguing that Early Stopping is not his regularization method of choice. From that statement it doesn't really follow: \"If your training process is tuned correctly, it should always be better to train longer.\" If I remember correctly, what he tried to say is that he prefers using Dropout, L2 norm and other ways to regularize the network, rather than doing early stopping.</p>",
      "rawMarkdown": "I also remember him saying that, but I think it was said during arguing that Early Stopping is not his regularization method of choice. From that statement it doesn't really follow: \"If your training process is tuned correctly, it should always be better to train longer.\" If I remember correctly, what he tried to say is that he prefers using Dropout, L2 norm and other ways to regularize the network, rather than doing early stopping.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 441453,
      "author_name": "mschumacher",
      "author_url": "",
      "post_date": "12/18/2018 17:11:59",
      "content": "<p>You are generally right. However, there are some recent papers that remark on the curious fact that NNs often improve their validation <strong>accuracy</strong> even after the validation <strong>loss</strong> has started increasing again. So training for 1 or 2 epochs past the \"overfit\" may actually be beneficial for the final performance.</p>\n\n<p>Anyway, apart from that, the art of this competition is to get great models in a short time. You should tune your models exactly so that in the given number of epochs, they end up being at their best. For off-line experimentation, training for many epochs and then selecting the best model may be fine. Here, it would be a waste of precious time.</p>\n\n<p>Since this tuning of the model and training process is hard work and a science in itself, the kernels posted here usually are not perfect. I believe you could probably get very far by taking one of the best kernels and then doing this tuning yourself.</p>",
      "votes": null,
      "replies": [
        {
          "id": 441713,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "12/19/2018 00:21:31",
          "content": "<p>Agreed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442916,
          "author_name": "bminixhofer",
          "author_url": "",
          "post_date": "12/20/2018 17:39:19",
          "content": "<p>This is sounds counter intuitive to me. Isn't the loss supposed to act as a proxy for the accuracy? So far, whenever the loss decreased in my models, the accuracy increased (except of some small deviations). Could you share the paper(s) that explain the behavior you mentioned?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442926,
          "author_name": "ishitori",
          "author_url": "",
          "post_date": "12/20/2018 18:08:19",
          "content": "<p>Can you send a link to papers that claim that? Would be quite fun to read.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 443266,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/21/2018 10:14:45",
          "content": "<p>Unfortunately, I can't find any at the moment. I'm very certain to have read that in at least one paper recently... 🤔 Maybe someone else can help?</p>\n\n<p>EDIT: I definitely remember Andrew Ng saying in his DL course that he generally does not use early stopping. If your training process is tuned correctly, it should always be better to train longer.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446370,
          "author_name": "ishitori",
          "author_url": "",
          "post_date": "12/28/2018 01:06:40",
          "content": "<p>I also remember him saying that, but I think it was said during arguing that Early Stopping is not his regularization method of choice. From that statement it doesn't really follow: \"If your training process is tuned correctly, it should always be better to train longer.\" If I remember correctly, what he tried to say is that he prefers using Dropout, L2 norm and other ways to regularize the network, rather than doing early stopping.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "441412": "Hi,\n\nI looked at some of the high scoring kernels and most of them use the model's weights at the last trained epoch instead of choosing the best performing epoch (lowest val loss) to find the best threshold and predict on the test set. \n\nBy doing this, the model used to predict on the test set should be overfitting to the training set and should generalize worse to the val or test sets than choosing the epoch with the lowest validation loss. Also, the \"best threshold\" found wouldn't be the best threshold, but the best threshold to the validation set when overfitting to the train set.\n\nThe kernels that I looked at are:\n\nhttps://www.kaggle.com/ashishpatel26/nlp-text-analytics-solution-quora\nhttps://www.kaggle.com/hengzheng/pytorch-starter\nhttps://www.kaggle.com/bminixhofer/deterministic-neural-networks-using-pytorch\nhttps://www.kaggle.com/hung96ad/magic-numbers-is-all-you-need-0-696-lb\nhttps://www.kaggle.com/shujian/single-rnn-with-4-folds-clr",
    "441453": "You are generally right. However, there are some recent papers that remark on the curious fact that NNs often improve their validation **accuracy** even after the validation **loss** has started increasing again. So training for 1 or 2 epochs past the \"overfit\" may actually be beneficial for the final performance.\n\nAnyway, apart from that, the art of this competition is to get great models in a short time. You should tune your models exactly so that in the given number of epochs, they end up being at their best. For off-line experimentation, training for many epochs and then selecting the best model may be fine. Here, it would be a waste of precious time.\n\nSince this tuning of the model and training process is hard work and a science in itself, the kernels posted here usually are not perfect. I believe you could probably get very far by taking one of the best kernels and then doing this tuning yourself.",
    "441713": "Agreed.",
    "442916": "This is sounds counter intuitive to me. Isn't the loss supposed to act as a proxy for the accuracy? So far, whenever the loss decreased in my models, the accuracy increased (except of some small deviations). Could you share the paper(s) that explain the behavior you mentioned?",
    "442926": "Can you send a link to papers that claim that? Would be quite fun to read.",
    "443266": "Unfortunately, I can't find any at the moment. I'm very certain to have read that in at least one paper recently... 🤔 Maybe someone else can help?\n\nEDIT: I definitely remember Andrew Ng saying in his DL course that he generally does not use early stopping. If your training process is tuned correctly, it should always be better to train longer.",
    "446370": "I also remember him saying that, but I think it was said during arguing that Early Stopping is not his regularization method of choice. From that statement it doesn't really follow: \"If your training process is tuned correctly, it should always be better to train longer.\" If I remember correctly, what he tried to say is that he prefers using Dropout, L2 norm and other ways to regularize the network, rather than doing early stopping."
  },
  "source": "meta"
}