{
  "id": 236218,
  "title": "F1 score and public score differing too much ",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/236218",
  "author_name": "",
  "post_date": "2021-05-03T11:25:28.297828900Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>This problem will be quite similar to <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234849?rvi=1\" target=\"_blank\">this topic</a>I have 80 percent f1 score both on train and val set but only 0.466 on public score i tried checking the distribution of my data but there is no problem in the data . How should i solve this problem i also would like to mention i have 0.0918 loss and 0.0875 (almost)loss on train and valid sets.</p>",
  "messages": [
    {
      "id": "1291834",
      "postDate": "05/03/2021 11:25:28",
      "content": "<p>This problem will be quite similar to <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234849?rvi=1\" target=\"_blank\">this topic</a>I have 80 percent f1 score both on train and val set but only 0.466 on public score i tried checking the distribution of my data but there is no problem in the data . How should i solve this problem i also would like to mention i have 0.0918 loss and 0.0875 (almost)loss on train and valid sets.</p>",
      "rawMarkdown": "This problem will be quite similar to [this topic](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234849?rvi=1)I have 80 percent f1 score both on train and val set but only 0.466 on public score i tried checking the distribution of my data but there is no problem in the data . How should i solve this problem i also would like to mention i have 0.0918 loss and 0.0875 (almost)loss on train and valid sets.",
      "votes": null
    },
    {
      "id": "1292958",
      "postDate": "05/04/2021 12:54:22",
      "content": "<p>I don't think you can solve that problem since we don't know how private test dataset differs from public training dataset. You can try to tune your metric though.</p>",
      "rawMarkdown": "I don't think you can solve that problem since we don't know how private test dataset differs from public training dataset. You can try to tune your metric though.",
      "votes": null
    },
    {
      "id": "1293076",
      "postDate": "05/04/2021 14:28:56",
      "content": "<p>I was also thinking the same, now its just so hard to tune the hyperparameters bad models often does better than good models . I think the distribution of images in hidden set and train set is quite different i should try more augmentations to change it works best as per my observation.</p>",
      "rawMarkdown": "I was also thinking the same, now its just so hard to tune the hyperparameters bad models often does better than good models . I think the distribution of images in hidden set and train set is quite different i should try more augmentations to change it works best as per my observation.",
      "votes": null
    },
    {
      "id": "1297114",
      "postDate": "05/07/2021 18:33:41",
      "content": "<p>I guess you are facing the problem of overfitting, it works too well for the known dataset, but not so for the hidden one. To begin with try to optimize on the metric they use, like in this case f1. Later you could try some regularization techniques to force the model to learn more rather than overfitting. Better learning will indeed help the model to generalize well. This might help to reduce the gap between your score and the final score.</p>",
      "rawMarkdown": "I guess you are facing the problem of overfitting, it works too well for the known dataset, but not so for the hidden one. To begin with try to optimize on the metric they use, like in this case f1. Later you could try some regularization techniques to force the model to learn more rather than overfitting. Better learning will indeed help the model to generalize well. This might help to reduce the gap between your score and the final score.",
      "votes": null
    },
    {
      "id": "1297487",
      "postDate": "05/08/2021 04:38:52",
      "content": "<p>Thank you i will try that out </p>",
      "rawMarkdown": "Thank you i will try that out",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1292958,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/04/2021 12:54:22",
      "content": "<p>I don't think you can solve that problem since we don't know how private test dataset differs from public training dataset. You can try to tune your metric though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1293076,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "05/04/2021 14:28:56",
          "content": "<p>I was also thinking the same, now its just so hard to tune the hyperparameters bad models often does better than good models . I think the distribution of images in hidden set and train set is quite different i should try more augmentations to change it works best as per my observation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1297114,
      "author_name": "jimitshah777",
      "author_url": "",
      "post_date": "05/07/2021 18:33:41",
      "content": "<p>I guess you are facing the problem of overfitting, it works too well for the known dataset, but not so for the hidden one. To begin with try to optimize on the metric they use, like in this case f1. Later you could try some regularization techniques to force the model to learn more rather than overfitting. Better learning will indeed help the model to generalize well. This might help to reduce the gap between your score and the final score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1297487,
          "author_name": "swaralipibose",
          "author_url": "",
          "post_date": "05/08/2021 04:38:52",
          "content": "<p>Thank you i will try that out </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1291834": "This problem will be quite similar to [this topic](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234849?rvi=1)I have 80 percent f1 score both on train and val set but only 0.466 on public score i tried checking the distribution of my data but there is no problem in the data . How should i solve this problem i also would like to mention i have 0.0918 loss and 0.0875 (almost)loss on train and valid sets.",
    "1292958": "I don't think you can solve that problem since we don't know how private test dataset differs from public training dataset. You can try to tune your metric though.",
    "1293076": "I was also thinking the same, now its just so hard to tune the hyperparameters bad models often does better than good models . I think the distribution of images in hidden set and train set is quite different i should try more augmentations to change it works best as per my observation.",
    "1297114": "I guess you are facing the problem of overfitting, it works too well for the known dataset, but not so for the hidden one. To begin with try to optimize on the metric they use, like in this case f1. Later you could try some regularization techniques to force the model to learn more rather than overfitting. Better learning will indeed help the model to generalize well. This might help to reduce the gap between your score and the final score.",
    "1297487": "Thank you i will try that out"
  },
  "source": "meta"
}