{
  "id": 358862,
  "title": "Why do we see a large difference between validation loss and test loss?",
  "url": "/competitions/tabular-playground-series-oct-2022/discussion/358862",
  "author_name": "",
  "post_date": "2022-10-09T20:31:22.145720300Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I, myself, have been shifting the focus from getting finding new features/models to trying to understand why most models see a validation loss around 0.18 but when submitting my test set to the LB, I get a loss that is in no way related to the validation loss. Does anyone have an idea why this happens? I assumed it had to do with class imbalance but that didn't seem to fix it. What do you all think it is? </p>",
  "messages": [
    {
      "id": "1979940",
      "postDate": "10/09/2022 20:31:22",
      "content": "<p>I, myself, have been shifting the focus from getting finding new features/models to trying to understand why most models see a validation loss around 0.18 but when submitting my test set to the LB, I get a loss that is in no way related to the validation loss. Does anyone have an idea why this happens? I assumed it had to do with class imbalance but that didn't seem to fix it. What do you all think it is? </p>",
      "rawMarkdown": "I, myself, have been shifting the focus from getting finding new features/models to trying to understand why most models see a validation loss around 0.18 but when submitting my test set to the LB, I get a loss that is in no way related to the validation loss. Does anyone have an idea why this happens? I assumed it had to do with class imbalance but that didn't seem to fix it. What do you all think it is?",
      "votes": null
    },
    {
      "id": "1979956",
      "postDate": "10/09/2022 20:57:35",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fabianbong\" target=\"_blank\">@fabianbong</a>, are you submitting probabilities?</p>",
      "rawMarkdown": "Hello @fabianbong, are you submitting probabilities?",
      "votes": null
    },
    {
      "id": "1979974",
      "postDate": "10/09/2022 21:59:56",
      "content": "<p>I am - I mean that's what the challenge asks for? And submitting classes (1 or 0) would give an even bigger loss for a wrong classification?</p>",
      "rawMarkdown": "I am - I mean that's what the challenge asks for? And submitting classes (1 or 0) would give an even bigger loss for a wrong classification?",
      "votes": null
    },
    {
      "id": "1979977",
      "postDate": "10/09/2022 22:11:35",
      "content": "<p>Hi, I think that maybe the fact that the public leaderboard is computed on 24% of the test data and the private leaderboard is computed on the remaining 76% is one important aspect to consider. I saw in many recent competition that the fact that public leaderboard was computed on a small part to test set lead to \"great shakeups\" in final leaderboard.</p>\n<p>There is a line of though around that says \"believe in your cross validation score rather than the public leaderboard\" especially when there are inbalances and smaller public sets than the private test set.</p>\n<p>My suggestion would be, take the leaderboard as an indicator:<br>\nIf you get 0.19 instead of 0.18 the model is not broken, maybe the public part of the test set is a harder part of the whole test set.<br>\nIf you get 0.25 start checking if your validation procedure is valid and you have no leakage, like validating on steps of games that are in your train set, in that case your validation loss would be significantly lower because of the leakage.</p>",
      "rawMarkdown": "Hi, I think that maybe the fact that the public leaderboard is computed on 24% of the test data and the private leaderboard is computed on the remaining 76% is one important aspect to consider. I saw in many recent competition that the fact that public leaderboard was computed on a small part to test set lead to \"great shakeups\" in final leaderboard.\n\nThere is a line of though around that says \"believe in your cross validation score rather than the public leaderboard\" especially when there are inbalances and smaller public sets than the private test set.\n\nMy suggestion would be, take the leaderboard as an indicator:\nIf you get 0.19 instead of 0.18 the model is not broken, maybe the public part of the test set is a harder part of the whole test set.\nIf you get 0.25 start checking if your validation procedure is valid and you have no leakage, like validating on steps of games that are in your train set, in that case your validation loss would be significantly lower because of the leakage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1979956,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "10/09/2022 20:57:35",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/fabianbong\" target=\"_blank\">@fabianbong</a>, are you submitting probabilities?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1979974,
          "author_name": "fabianbong",
          "author_url": "",
          "post_date": "10/09/2022 21:59:56",
          "content": "<p>I am - I mean that's what the challenge asks for? And submitting classes (1 or 0) would give an even bigger loss for a wrong classification?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1979977,
      "author_name": "pietromaldini1",
      "author_url": "",
      "post_date": "10/09/2022 22:11:35",
      "content": "<p>Hi, I think that maybe the fact that the public leaderboard is computed on 24% of the test data and the private leaderboard is computed on the remaining 76% is one important aspect to consider. I saw in many recent competition that the fact that public leaderboard was computed on a small part to test set lead to \"great shakeups\" in final leaderboard.</p>\n<p>There is a line of though around that says \"believe in your cross validation score rather than the public leaderboard\" especially when there are inbalances and smaller public sets than the private test set.</p>\n<p>My suggestion would be, take the leaderboard as an indicator:<br>\nIf you get 0.19 instead of 0.18 the model is not broken, maybe the public part of the test set is a harder part of the whole test set.<br>\nIf you get 0.25 start checking if your validation procedure is valid and you have no leakage, like validating on steps of games that are in your train set, in that case your validation loss would be significantly lower because of the leakage.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1979940": "I, myself, have been shifting the focus from getting finding new features/models to trying to understand why most models see a validation loss around 0.18 but when submitting my test set to the LB, I get a loss that is in no way related to the validation loss. Does anyone have an idea why this happens? I assumed it had to do with class imbalance but that didn't seem to fix it. What do you all think it is?",
    "1979956": "Hello @fabianbong, are you submitting probabilities?",
    "1979974": "I am - I mean that's what the challenge asks for? And submitting classes (1 or 0) would give an even bigger loss for a wrong classification?",
    "1979977": "Hi, I think that maybe the fact that the public leaderboard is computed on 24% of the test data and the private leaderboard is computed on the remaining 76% is one important aspect to consider. I saw in many recent competition that the fact that public leaderboard was computed on a small part to test set lead to \"great shakeups\" in final leaderboard.\n\nThere is a line of though around that says \"believe in your cross validation score rather than the public leaderboard\" especially when there are inbalances and smaller public sets than the private test set.\n\nMy suggestion would be, take the leaderboard as an indicator:\nIf you get 0.19 instead of 0.18 the model is not broken, maybe the public part of the test set is a harder part of the whole test set.\nIf you get 0.25 start checking if your validation procedure is valid and you have no leakage, like validating on steps of games that are in your train set, in that case your validation loss would be significantly lower because of the leakage."
  },
  "source": "meta"
}