{
  "id": 234849,
  "title": "f1 score 90%",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/234849",
  "author_name": "AIT HAMMADI Abdellatif",
  "post_date": "2021-04-26T14:05:55.795000",
  "votes": 1,
  "comment_count": 8,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/aithammadiabdellatif/tpu-tf-efficientnetb7-with-extra-layer/\" target=\"_blank\">my notebook</a> i get 90% in f1 score but when i submit i get only 60%, someone have any explanation <br>\n<img src=\"https://abdo.ninja/images/90.png\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1285235,
      "postDate": "2021-04-26T17:52:33.790Z",
      "content": "<p>It seems that in this competition, the cv score and lb score are so inconsistent and uncorrelated. In my experients, many models with &gt;90% CV do poorly on the LB(~0.65). I geuss it may due to the <strong>Complex</strong> label, as the orgnizer said, </p>\n<blockquote>\n  <p>complex is best understood as \"<strong>too complex for an expert to classify</strong>\". You can think of it being used as a flag for cases where visual approaches aren't sufficient and further investigation with lab testing could be warranted. (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226411</a>) </p>\n</blockquote>\n<p>So the Complex class is really noisy for model to learn, and there may be many mislabeled data in this class. I'm still thinking of a practical solution and I post this problem here [<a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]</a>. Hoping there will be some insightful ideas.</p>",
      "rawMarkdown": "It seems that in this competition, the cv score and lb score are so inconsistent and uncorrelated. In my experients, many models with >90% CV do poorly on the LB(~0.65). I geuss it may due to the **Complex** label, as the orgnizer said, \n> complex is best understood as \"**too complex for an expert to classify**\". You can think of it being used as a flag for cases where visual approaches aren't sufficient and further investigation with lab testing could be warranted. ([https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226411](url)) \n\nSo the Complex class is really noisy for model to learn, and there may be many mislabeled data in this class. I'm still thinking of a practical solution and I post this problem here [https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]. Hoping there will be some insightful ideas.\n",
      "votes": 1
    },
    {
      "id": 1285028,
      "postDate": "2021-04-26T14:05:55.797Z",
      "content": "<p>In <a href=\"https://www.kaggle.com/aithammadiabdellatif/tpu-tf-efficientnetb7-with-extra-layer/\" target=\"_blank\">my notebook</a> i get 90% in f1 score but when i submit i get only 60%, someone have any explanation <br>\n<img src=\"https://abdo.ninja/images/90.png\" alt=\"\"></p>",
      "rawMarkdown": "In [my notebook](https://www.kaggle.com/aithammadiabdellatif/tpu-tf-efficientnetb7-with-extra-layer/) i get 90% in f1 score but when i submit i get only 60%, someone have any explanation \n![](https://abdo.ninja/images/90.png)",
      "votes": 1
    },
    {
      "id": 1285864,
      "postDate": "2021-04-27T10:15:27.880Z",
      "content": "<p>I've noticed that results heavily depend on chosen thresholds. I managed to get increase LB standing by ~0.2 by adjusting them.</p>",
      "rawMarkdown": "I've noticed that results heavily depend on chosen thresholds. I managed to get increase LB standing by ~0.2 by adjusting them.",
      "votes": 2,
      "replies": [
        {
          "id": 1289983,
          "postDate": "2021-05-01T14:33:30.183Z",
          "content": "<p>It heavily depends but keras uses 0.5 threshold and if f1 score is 90 percent using 0.5 threshold why 60 % with same threshold  ?</p>",
          "rawMarkdown": "It heavily depends but keras uses 0.5 threshold and if f1 score is 90 percent using 0.5 threshold why 60 % with same threshold  ?"
        },
        {
          "id": 1290007,
          "postDate": "2021-05-01T14:49:09.153Z",
          "content": "<p>Actually, I wrote that BEFORE Kaggle staff removed bad images from the test set. It should be much smoother now. As for choosing thresholds, see <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">this notebook</a> for how to calculate them to maximize your score based on validation set.</p>",
          "rawMarkdown": "Actually, I wrote that BEFORE Kaggle staff removed bad images from the test set. It should be much smoother now. As for choosing thresholds, see [this notebook](https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training) for how to calculate them to maximize your score based on validation set."
        },
        {
          "id": 1290017,
          "postDate": "2021-05-01T14:57:39.020Z",
          "content": "<p>Yeah it is indeed much smoother they should do something about the 'complex' class</p>",
          "rawMarkdown": "Yeah it is indeed much smoother they should do something about the 'complex' class"
        },
        {
          "id": 1290028,
          "postDate": "2021-05-01T15:06:25.503Z",
          "content": "<p>Read <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234936\" target=\"_blank\">this</a> for more info.</p>",
          "rawMarkdown": "Read [this](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234936) for more info."
        },
        {
          "id": 1290032,
          "postDate": "2021-05-01T15:08:39.653Z",
          "content": "<p>Oh yeah i have already read it i was talking about the complex class</p>",
          "rawMarkdown": "Oh yeah i have already read it i was talking about the complex class"
        }
      ]
    },
    {
      "id": 1289992,
      "postDate": "2021-05-01T14:40:33.940Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1285235,
      "author_name": "YTEP (Jiazhi Yang)",
      "author_url": "",
      "post_date": "2021-04-26T17:52:33.790000",
      "content": "<p>It seems that in this competition, the cv score and lb score are so inconsistent and uncorrelated. In my experients, many models with &gt;90% CV do poorly on the LB(~0.65). I geuss it may due to the <strong>Complex</strong> label, as the orgnizer said, </p>\n<blockquote>\n  <p>complex is best understood as \"<strong>too complex for an expert to classify</strong>\". You can think of it being used as a flag for cases where visual approaches aren't sufficient and further investigation with lab testing could be warranted. (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226411</a>) </p>\n</blockquote>\n<p>So the Complex class is really noisy for model to learn, and there may be many mislabeled data in this class. I'm still thinking of a practical solution and I post this problem here [<a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]\" target=\"_blank\">https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]</a>. Hoping there will be some insightful ideas.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1285864,
      "author_name": "Araik Tamazian",
      "author_url": "",
      "post_date": "2021-04-27T10:15:27.880000",
      "content": "<p>I've noticed that results heavily depend on chosen thresholds. I managed to get increase LB standing by ~0.2 by adjusting them.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1289983,
          "author_name": "Sayantan Mazumdar",
          "author_url": "",
          "post_date": "2021-05-01T14:33:30.183000",
          "content": "<p>It heavily depends but keras uses 0.5 threshold and if f1 score is 90 percent using 0.5 threshold why 60 % with same threshold  ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1290007,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2021-05-01T14:49:09.153000",
          "content": "<p>Actually, I wrote that BEFORE Kaggle staff removed bad images from the test set. It should be much smoother now. As for choosing thresholds, see <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-tpu-tf-training\" target=\"_blank\">this notebook</a> for how to calculate them to maximize your score based on validation set.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1290017,
          "author_name": "Sayantan Mazumdar",
          "author_url": "",
          "post_date": "2021-05-01T14:57:39.020000",
          "content": "<p>Yeah it is indeed much smoother they should do something about the 'complex' class</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1290028,
          "author_name": "Araik Tamazian",
          "author_url": "",
          "post_date": "2021-05-01T15:06:25.503000",
          "content": "<p>Read <a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234936\" target=\"_blank\">this</a> for more info.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1290032,
          "author_name": "Sayantan Mazumdar",
          "author_url": "",
          "post_date": "2021-05-01T15:08:39.653000",
          "content": "<p>Oh yeah i have already read it i was talking about the complex class</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1289992,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-01T14:40:33.940000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1285235": "It seems that in this competition, the cv score and lb score are so inconsistent and uncorrelated. In my experients, many models with >90% CV do poorly on the LB(~0.65). I geuss it may due to the **Complex** label, as the orgnizer said, \n> complex is best understood as \"**too complex for an expert to classify**\". You can think of it being used as a flag for cases where visual approaches aren't sufficient and further investigation with lab testing could be warranted. ([https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/226411](url)) \n\nSo the Complex class is really noisy for model to learn, and there may be many mislabeled data in this class. I'm still thinking of a practical solution and I post this problem here [https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/234886]. Hoping there will be some insightful ideas.\n",
    "1285028": "In [my notebook](https://www.kaggle.com/aithammadiabdellatif/tpu-tf-efficientnetb7-with-extra-layer/) i get 90% in f1 score but when i submit i get only 60%, someone have any explanation \n![](https://abdo.ninja/images/90.png)",
    "1285864": "I've noticed that results heavily depend on chosen thresholds. I managed to get increase LB standing by ~0.2 by adjusting them.",
    "1289992": ""
  }
}