{
  "id": 76986,
  "title": "Why I overfit but get higher LB score?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/76986",
  "author_name": "",
  "post_date": "2019-01-08T14:01:14.807241Z",
  "votes": 8,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I check validation loss every epoch. If the loss is not decreased in three epoch, the training will stop. However, I find that if I don't early stop it and just continue to train three epoch more, the validation loss will increase, but the LB score will be higher. Early stopping is in every fold of five folds.\nIs that possible or just being something wrong in my code?</p>",
  "messages": [
    {
      "id": "452292",
      "postDate": "01/08/2019 14:01:14",
      "content": "<p>I check validation loss every epoch. If the loss is not decreased in three epoch, the training will stop. However, I find that if I don't early stop it and just continue to train three epoch more, the validation loss will increase, but the LB score will be higher. Early stopping is in every fold of five folds.\nIs that possible or just being something wrong in my code?</p>",
      "rawMarkdown": "I check validation loss every epoch. If the loss is not decreased in three epoch, the training will stop. However, I find that if I don't early stop it and just continue to train three epoch more, the validation loss will increase, but the LB score will be higher. Early stopping is in every fold of five folds.\nIs that possible or just being something wrong in my code?",
      "votes": null
    },
    {
      "id": "452378",
      "postDate": "01/08/2019 16:30:51",
      "content": "<p>For me also the LB scores are not that high even though the cv score is nearly the same as the public kernel. And according to me the difference between public kernels and me are I am saving the models with best validation score(using model checkpoint). I don't know what's the problem but this is worrying me a lot. </p>",
      "rawMarkdown": "For me also the LB scores are not that high even though the cv score is nearly the same as the public kernel. And according to me the difference between public kernels and me are I am saving the models with best validation score(using model checkpoint). I don't know what's the problem but this is worrying me a lot.",
      "votes": null
    },
    {
      "id": "452382",
      "postDate": "01/08/2019 16:42:00",
      "content": "<p>me too</p>",
      "rawMarkdown": "me too",
      "votes": null
    },
    {
      "id": "452435",
      "postDate": "01/08/2019 18:07:57",
      "content": "<p>It is possible, and there is an explanation. I'm working on a kernel on the topic. It will be published tomorrow :)</p>",
      "rawMarkdown": "It is possible, and there is an explanation. I'm working on a kernel on the topic. It will be published tomorrow :)",
      "votes": null
    },
    {
      "id": "452436",
      "postDate": "01/08/2019 18:08:36",
      "content": "<p>I tested my way of loading the models. It seems fine for me. This is the way I am doing </p>\n\n<pre><code>model=AttentionLSTM()\nmodel.float()\nmodel.cuda()  # pushing the model to gpu\nmodel.load_state_dict(torch.load(filename))\n# predicting on test data\n</code></pre>",
      "rawMarkdown": "I tested my way of loading the models. It seems fine for me. This is the way I am doing \n\n    model=AttentionLSTM()\n    model.float()\n    model.cuda()  # pushing the model to gpu\n    model.load_state_dict(torch.load(filename))\n    # predicting on test data",
      "votes": null
    },
    {
      "id": "452454",
      "postDate": "01/08/2019 18:40:49",
      "content": "<p>waiting for it...</p>",
      "rawMarkdown": "waiting for it...",
      "votes": null
    },
    {
      "id": "452550",
      "postDate": "01/08/2019 21:58:39",
      "content": "<p>I'm looking forward to see it :)</p>",
      "rawMarkdown": "I'm looking forward to see it :)",
      "votes": null
    },
    {
      "id": "452574",
      "postDate": "01/08/2019 23:00:06",
      "content": "<p>I also found this is true. Not to mention that local f1 CV score is NOT correlated with the LB score - my models with highest local CV score are not my highest scoring LB models. It's kind of annoying because Val loss and CV score are not good predictors of how it performs on the test set, so I'm never quite sure what to submit. I just submitted a few things and a model with local CV of .6836 scored .684 LB (at least the CV matched LB), but a model with CV .6768 scored .696 LB.</p>",
      "rawMarkdown": "I also found this is true. Not to mention that local f1 CV score is NOT correlated with the LB score - my models with highest local CV score are not my highest scoring LB models. It's kind of annoying because Val loss and CV score are not good predictors of how it performs on the test set, so I'm never quite sure what to submit. I just submitted a few things and a model with local CV of .6836 scored .684 LB (at least the CV matched LB), but a model with CV .6768 scored .696 LB.",
      "votes": null
    },
    {
      "id": "452594",
      "postDate": "01/08/2019 23:39:00",
      "content": "<p>Same as you! This is very annoying:(\nLB is so unstable</p>",
      "rawMarkdown": "Same as you! This is very annoying:(\nLB is so unstable",
      "votes": null
    },
    {
      "id": "452677",
      "postDate": "01/09/2019 03:01:49",
      "content": "<p>I load the model the way exactly same as yours</p>",
      "rawMarkdown": "I load the model the way exactly same as yours",
      "votes": null
    },
    {
      "id": "452680",
      "postDate": "01/09/2019 03:06:10",
      "content": "<p>looking forward to it</p>",
      "rawMarkdown": "looking forward to it",
      "votes": null
    },
    {
      "id": "452683",
      "postDate": "01/09/2019 03:17:14",
      "content": "<p>Yes, I also find that local f1 score is not correlated with LB score. Even if I separate a test set from the train set to replace the LB test set (I had planned to check my code), and run the code regarding my test set as the true test set, found that my test set score always matches the validation score.</p>",
      "rawMarkdown": "Yes, I also find that local f1 score is not correlated with LB score. Even if I separate a test set from the train set to replace the LB test set (I had planned to check my code), and run the code regarding my test set as the true test set, found that my test set score always matches the validation score.",
      "votes": null
    },
    {
      "id": "452900",
      "postDate": "01/09/2019 10:33:54",
      "content": "<p><a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">The kernel is now published.</a></p>",
      "rawMarkdown": "[The kernel is now published.](https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed)",
      "votes": null
    },
    {
      "id": "452984",
      "postDate": "01/09/2019 12:57:40",
      "content": "<p>The issue may be related to model selection. </p>\n\n<p>An extreme is the following: imagine every 100 steps you measure the F1 on CV, and save the model with best CV. After training, you load the model with best CV and test it. Your score at LB will be, most probably, lower, because you are \"overfitting to the validation set\". When you stop training after three epochs without the loss function not decreasing in some way you are choosing a \"best model\" which may be \"best\" in validation set, i.e. you are overfitting in that set.</p>",
      "rawMarkdown": "The issue may be related to model selection. \n\nAn extreme is the following: imagine every 100 steps you measure the F1 on CV, and save the model with best CV. After training, you load the model with best CV and test it. Your score at LB will be, most probably, lower, because you are \"overfitting to the validation set\". When you stop training after three epochs without the loss function not decreasing in some way you are choosing a \"best model\" which may be \"best\" in validation set, i.e. you are overfitting in that set.",
      "votes": null
    },
    {
      "id": "453988",
      "postDate": "01/11/2019 02:21:03",
      "content": "<p>That make sense</p>",
      "rawMarkdown": "That make sense",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 452378,
      "author_name": "suchith0312",
      "author_url": "",
      "post_date": "01/08/2019 16:30:51",
      "content": "<p>For me also the LB scores are not that high even though the cv score is nearly the same as the public kernel. And according to me the difference between public kernels and me are I am saving the models with best validation score(using model checkpoint). I don't know what's the problem but this is worrying me a lot. </p>",
      "votes": null,
      "replies": [
        {
          "id": 452436,
          "author_name": "suchith0312",
          "author_url": "",
          "post_date": "01/08/2019 18:08:36",
          "content": "<p>I tested my way of loading the models. It seems fine for me. This is the way I am doing </p>\n\n<pre><code>model=AttentionLSTM()\nmodel.float()\nmodel.cuda()  # pushing the model to gpu\nmodel.load_state_dict(torch.load(filename))\n# predicting on test data\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452677,
          "author_name": "wxytalent",
          "author_url": "",
          "post_date": "01/09/2019 03:01:49",
          "content": "<p>I load the model the way exactly same as yours</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452382,
      "author_name": "sshleifer",
      "author_url": "",
      "post_date": "01/08/2019 16:42:00",
      "content": "<p>me too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452435,
      "author_name": "bminixhofer",
      "author_url": "",
      "post_date": "01/08/2019 18:07:57",
      "content": "<p>It is possible, and there is an explanation. I'm working on a kernel on the topic. It will be published tomorrow :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 452454,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "01/08/2019 18:40:49",
          "content": "<p>waiting for it...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452550,
          "author_name": "francoisdubois",
          "author_url": "",
          "post_date": "01/08/2019 21:58:39",
          "content": "<p>I'm looking forward to see it :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452680,
          "author_name": "wxytalent",
          "author_url": "",
          "post_date": "01/09/2019 03:06:10",
          "content": "<p>looking forward to it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452900,
          "author_name": "bminixhofer",
          "author_url": "",
          "post_date": "01/09/2019 10:33:54",
          "content": "<p><a href=\"https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed\">The kernel is now published.</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452574,
      "author_name": "julius6",
      "author_url": "",
      "post_date": "01/08/2019 23:00:06",
      "content": "<p>I also found this is true. Not to mention that local f1 CV score is NOT correlated with the LB score - my models with highest local CV score are not my highest scoring LB models. It's kind of annoying because Val loss and CV score are not good predictors of how it performs on the test set, so I'm never quite sure what to submit. I just submitted a few things and a model with local CV of .6836 scored .684 LB (at least the CV matched LB), but a model with CV .6768 scored .696 LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 452594,
          "author_name": "vanche",
          "author_url": "",
          "post_date": "01/08/2019 23:39:00",
          "content": "<p>Same as you! This is very annoying:(\nLB is so unstable</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 452683,
          "author_name": "wxytalent",
          "author_url": "",
          "post_date": "01/09/2019 03:17:14",
          "content": "<p>Yes, I also find that local f1 score is not correlated with LB score. Even if I separate a test set from the train set to replace the LB test set (I had planned to check my code), and run the code regarding my test set as the true test set, found that my test set score always matches the validation score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 452984,
      "author_name": "manuelsh",
      "author_url": "",
      "post_date": "01/09/2019 12:57:40",
      "content": "<p>The issue may be related to model selection. </p>\n\n<p>An extreme is the following: imagine every 100 steps you measure the F1 on CV, and save the model with best CV. After training, you load the model with best CV and test it. Your score at LB will be, most probably, lower, because you are \"overfitting to the validation set\". When you stop training after three epochs without the loss function not decreasing in some way you are choosing a \"best model\" which may be \"best\" in validation set, i.e. you are overfitting in that set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 453988,
          "author_name": "wxytalent",
          "author_url": "",
          "post_date": "01/11/2019 02:21:03",
          "content": "<p>That make sense</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "452292": "I check validation loss every epoch. If the loss is not decreased in three epoch, the training will stop. However, I find that if I don't early stop it and just continue to train three epoch more, the validation loss will increase, but the LB score will be higher. Early stopping is in every fold of five folds.\nIs that possible or just being something wrong in my code?",
    "452378": "For me also the LB scores are not that high even though the cv score is nearly the same as the public kernel. And according to me the difference between public kernels and me are I am saving the models with best validation score(using model checkpoint). I don't know what's the problem but this is worrying me a lot.",
    "452382": "me too",
    "452435": "It is possible, and there is an explanation. I'm working on a kernel on the topic. It will be published tomorrow :)",
    "452436": "I tested my way of loading the models. It seems fine for me. This is the way I am doing \n\n    model=AttentionLSTM()\n    model.float()\n    model.cuda()  # pushing the model to gpu\n    model.load_state_dict(torch.load(filename))\n    # predicting on test data",
    "452454": "waiting for it...",
    "452550": "I'm looking forward to see it :)",
    "452574": "I also found this is true. Not to mention that local f1 CV score is NOT correlated with the LB score - my models with highest local CV score are not my highest scoring LB models. It's kind of annoying because Val loss and CV score are not good predictors of how it performs on the test set, so I'm never quite sure what to submit. I just submitted a few things and a model with local CV of .6836 scored .684 LB (at least the CV matched LB), but a model with CV .6768 scored .696 LB.",
    "452594": "Same as you! This is very annoying:(\nLB is so unstable",
    "452677": "I load the model the way exactly same as yours",
    "452680": "looking forward to it",
    "452683": "Yes, I also find that local f1 score is not correlated with LB score. Even if I separate a test set from the train set to replace the LB test set (I had planned to check my code), and run the code regarding my test set as the true test set, found that my test set score always matches the validation score.",
    "452900": "[The kernel is now published.](https://www.kaggle.com/bminixhofer/a-validation-framework-impact-of-the-random-seed)",
    "452984": "The issue may be related to model selection. \n\nAn extreme is the following: imagine every 100 steps you measure the F1 on CV, and save the model with best CV. After training, you load the model with best CV and test it. Your score at LB will be, most probably, lower, because you are \"overfitting to the validation set\". When you stop training after three epochs without the loss function not decreasing in some way you are choosing a \"best model\" which may be \"best\" in validation set, i.e. you are overfitting in that set.",
    "453988": "That make sense"
  },
  "source": "meta"
}