{
  "id": 226587,
  "title": "Insights from last year's competition",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/226587",
  "author_name": "",
  "post_date": "2021-03-17T03:12:15.349891200Z",
  "votes": 9,
  "comment_count": 5,
  "views": 0,
  "content": "<p>After reading many top solutions from last year's <a href=\"https://www.kaggle.com/c/plant-pathology-2020-fgvc7\" target=\"_blank\">Plant Pathology competition</a>. Here are some insights:</p>\n<h4>1. Noisy Dataset</h4>\n<p>Same images with different labels. Some images in the training dataset were generated by one image, but they had different labels. Many top solutions used techniques like <strong>Knowledge Distillation</strong>, <strong>Label Smoothing</strong>, and <strong>Pseudo-Labelling</strong>.</p>\n<h4>2. Imbalance Dataset</h4>\n<p>Few people used <strong>upsampling</strong> to balance the dataset.</p>\n<p><strong>Advice:</strong> A wrong classification of a sample with multiple disease categories will have significant impact on the final results. So, don’t trust the public leaderboard, trust your CV.</p>\n<h4>3. Model architectures</h4>\n<p>People used different many different model but EfficientNet B(0-7) appeared most frequently. Model pre-trained on noisy-student dataset tend to perform better. And finally, larger images size (600+) tend to give better results. </p>\n<p>I will be testing all these ideas in the current competition. Stay tuned !</p>",
  "messages": [
    {
      "id": "1241412",
      "postDate": "03/17/2021 03:12:15",
      "content": "<p>After reading many top solutions from last year's <a href=\"https://www.kaggle.com/c/plant-pathology-2020-fgvc7\" target=\"_blank\">Plant Pathology competition</a>. Here are some insights:</p>\n<h4>1. Noisy Dataset</h4>\n<p>Same images with different labels. Some images in the training dataset were generated by one image, but they had different labels. Many top solutions used techniques like <strong>Knowledge Distillation</strong>, <strong>Label Smoothing</strong>, and <strong>Pseudo-Labelling</strong>.</p>\n<h4>2. Imbalance Dataset</h4>\n<p>Few people used <strong>upsampling</strong> to balance the dataset.</p>\n<p><strong>Advice:</strong> A wrong classification of a sample with multiple disease categories will have significant impact on the final results. So, don’t trust the public leaderboard, trust your CV.</p>\n<h4>3. Model architectures</h4>\n<p>People used different many different model but EfficientNet B(0-7) appeared most frequently. Model pre-trained on noisy-student dataset tend to perform better. And finally, larger images size (600+) tend to give better results. </p>\n<p>I will be testing all these ideas in the current competition. Stay tuned !</p>",
      "rawMarkdown": "After reading many top solutions from last year's [Plant Pathology competition](https://www.kaggle.com/c/plant-pathology-2020-fgvc7). Here are some insights:\n\n#### 1. Noisy Dataset\n\nSame images with different labels. Some images in the training dataset were generated by one image, but they had different labels. Many top solutions used techniques like **Knowledge Distillation**, **Label Smoothing**, and **Pseudo-Labelling**.\n\n#### 2. Imbalance Dataset\n\nFew people used **upsampling** to balance the dataset.\n\n**Advice:** A wrong classification of a sample with multiple disease categories will have significant impact on the final results. So, don’t trust the public leaderboard, trust your CV.\n\n#### 3. Model architectures\n\nPeople used different many different model but EfficientNet B(0-7) appeared most frequently. Model pre-trained on noisy-student dataset tend to perform better. And finally, larger images size (600+) tend to give better results. \n\nI will be testing all these ideas in the current competition. Stay tuned !",
      "votes": null
    },
    {
      "id": "1243322",
      "postDate": "03/18/2021 06:28:51",
      "content": "<p><a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a> Thanks so much for the curated insights Ankur . Very Helpful </p>",
      "rawMarkdown": "ankursingh12 Thanks so much for the curated insights Ankur . Very Helpful",
      "votes": null
    },
    {
      "id": "1244090",
      "postDate": "03/18/2021 17:33:18",
      "content": "<p>Trying my best to learn! Between, nice name!</p>",
      "rawMarkdown": "Trying my best to learn! Between, nice name!",
      "votes": null
    },
    {
      "id": "1248707",
      "postDate": "03/22/2021 19:12:52",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a>,  also believe last year we had a category for multiple diseases and single disease only, now we have each combination of disease which makes it a bit more challenging.</p>",
      "rawMarkdown": "Thanks @ankursingh12,  also believe last year we had a category for multiple diseases and single disease only, now we have each combination of disease which makes it a bit more challenging.",
      "votes": null
    },
    {
      "id": "1250307",
      "postDate": "03/23/2021 23:43:12",
      "content": "<p>Agreed, <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> ! </p>\n<p>It's actually, pretty complex. I have seen a lot of people approaching it as multi-class (12 classes ) classification problem. Technically, its a multi-label problem. But because of class imbalance, its much difficult to  get a good score. </p>\n<p>I think, multi-class approach is good for now. But as the competition progresses, it will start following short. What is your take?</p>",
      "rawMarkdown": "Agreed, @saurabhbagchi ! \n\nIt's actually, pretty complex. I have seen a lot of people approaching it as multi-class (12 classes ) classification problem. Technically, its a multi-label problem. But because of class imbalance, its much difficult to  get a good score. \n\nI think, multi-class approach is good for now. But as the competition progresses, it will start following short. What is your take?",
      "votes": null
    },
    {
      "id": "1250494",
      "postDate": "03/24/2021 04:53:57",
      "content": "<p>Yes class imbalance is difficult to tackle specially on unseen data, whoever handles it better through the code will prevail.</p>",
      "rawMarkdown": "Yes class imbalance is difficult to tackle specially on unseen data, whoever handles it better through the code will prevail.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1243322,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/18/2021 06:28:51",
      "content": "<p><a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a> Thanks so much for the curated insights Ankur . Very Helpful </p>",
      "votes": null,
      "replies": [
        {
          "id": 1244090,
          "author_name": "ankursingh12",
          "author_url": "",
          "post_date": "03/18/2021 17:33:18",
          "content": "<p>Trying my best to learn! Between, nice name!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1248707,
      "author_name": "saurabhbagchi",
      "author_url": "",
      "post_date": "03/22/2021 19:12:52",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/ankursingh12\" target=\"_blank\">@ankursingh12</a>,  also believe last year we had a category for multiple diseases and single disease only, now we have each combination of disease which makes it a bit more challenging.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1250307,
          "author_name": "ankursingh12",
          "author_url": "",
          "post_date": "03/23/2021 23:43:12",
          "content": "<p>Agreed, <a href=\"https://www.kaggle.com/saurabhbagchi\" target=\"_blank\">@saurabhbagchi</a> ! </p>\n<p>It's actually, pretty complex. I have seen a lot of people approaching it as multi-class (12 classes ) classification problem. Technically, its a multi-label problem. But because of class imbalance, its much difficult to  get a good score. </p>\n<p>I think, multi-class approach is good for now. But as the competition progresses, it will start following short. What is your take?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1250494,
          "author_name": "saurabhbagchi",
          "author_url": "",
          "post_date": "03/24/2021 04:53:57",
          "content": "<p>Yes class imbalance is difficult to tackle specially on unseen data, whoever handles it better through the code will prevail.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1241412": "After reading many top solutions from last year's [Plant Pathology competition](https://www.kaggle.com/c/plant-pathology-2020-fgvc7). Here are some insights:\n\n#### 1. Noisy Dataset\n\nSame images with different labels. Some images in the training dataset were generated by one image, but they had different labels. Many top solutions used techniques like **Knowledge Distillation**, **Label Smoothing**, and **Pseudo-Labelling**.\n\n#### 2. Imbalance Dataset\n\nFew people used **upsampling** to balance the dataset.\n\n**Advice:** A wrong classification of a sample with multiple disease categories will have significant impact on the final results. So, don’t trust the public leaderboard, trust your CV.\n\n#### 3. Model architectures\n\nPeople used different many different model but EfficientNet B(0-7) appeared most frequently. Model pre-trained on noisy-student dataset tend to perform better. And finally, larger images size (600+) tend to give better results. \n\nI will be testing all these ideas in the current competition. Stay tuned !",
    "1243322": "ankursingh12 Thanks so much for the curated insights Ankur . Very Helpful",
    "1244090": "Trying my best to learn! Between, nice name!",
    "1248707": "Thanks @ankursingh12,  also believe last year we had a category for multiple diseases and single disease only, now we have each combination of disease which makes it a bit more challenging.",
    "1250307": "Agreed, @saurabhbagchi ! \n\nIt's actually, pretty complex. I have seen a lot of people approaching it as multi-class (12 classes ) classification problem. Technically, its a multi-label problem. But because of class imbalance, its much difficult to  get a good score. \n\nI think, multi-class approach is good for now. But as the competition progresses, it will start following short. What is your take?",
    "1250494": "Yes class imbalance is difficult to tackle specially on unseen data, whoever handles it better through the code will prevail."
  },
  "source": "meta"
}