{
  "id": 143071,
  "title": "Test set has far more classes than WCS train data. How to train?",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/143071",
  "author_name": "",
  "post_date": "2020-04-13T17:21:28.487611100Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p><a href=\"/sbeery\">@sbeery</a> </p>\n\n<p>When I was exploring the WCS train data, I found that there are 216 unique classes in train data. But on contrary, sample_submission.csv file for test set has 676 unique classes. </p>\n\n<ol>\n<li>I was wondering how to train the model for rest of the classes in the test data? If there are no such classes in WCS data alone, then how can the model generalizes?</li>\n<li>Do I need to include other datasets as well (iNaturalist 2017, 2018, 2019)? Or WCS alone should be sufficient?</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1502092%2F2d6b8042f4b9591ea1927a5ec39ae12a%2FScreenshot%20from%202020-04-13%2018-01-5111.xcf?generation=1586798463866497&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "806396",
      "postDate": "04/13/2020 17:21:28",
      "content": "<p><a href=\"/sbeery\">@sbeery</a> </p>\n\n<p>When I was exploring the WCS train data, I found that there are 216 unique classes in train data. But on contrary, sample_submission.csv file for test set has 676 unique classes. </p>\n\n<ol>\n<li>I was wondering how to train the model for rest of the classes in the test data? If there are no such classes in WCS data alone, then how can the model generalizes?</li>\n<li>Do I need to include other datasets as well (iNaturalist 2017, 2018, 2019)? Or WCS alone should be sufficient?</li>\n</ol>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1502092%2F2d6b8042f4b9591ea1927a5ec39ae12a%2FScreenshot%20from%202020-04-13%2018-01-5111.xcf?generation=1586798463866497&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "sbeery \n\nWhen I was exploring the WCS train data, I found that there are 216 unique classes in train data. But on contrary, sample_submission.csv file for test set has 676 unique classes. \n\n1. I was wondering how to train the model for rest of the classes in the test data? If there are no such classes in WCS data alone, then how can the model generalizes?\n2. Do I need to include other datasets as well (iNaturalist 2017, 2018, 2019)? Or WCS alone should be sufficient?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1502092%2F2d6b8042f4b9591ea1927a5ec39ae12a%2FScreenshot%20from%202020-04-13%2018-01-5111.xcf?generation=1586798463866497&amp;alt=media)",
      "votes": null
    },
    {
      "id": "806403",
      "postDate": "04/13/2020 17:26:22",
      "content": "<p>There are no classes in the test set images that were not included in the training set (we explicitly removed images from non-training classes at the test locations), but we do encourage including other datasets to help with rare classes that have little training data.</p>",
      "rawMarkdown": "There are no classes in the test set images that were not included in the training set (we explicitly removed images from non-training classes at the test locations), but we do encourage including other datasets to help with rare classes that have little training data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 806403,
      "author_name": "sbeery",
      "author_url": "",
      "post_date": "04/13/2020 17:26:22",
      "content": "<p>There are no classes in the test set images that were not included in the training set (we explicitly removed images from non-training classes at the test locations), but we do encourage including other datasets to help with rare classes that have little training data.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "806396": "sbeery \n\nWhen I was exploring the WCS train data, I found that there are 216 unique classes in train data. But on contrary, sample_submission.csv file for test set has 676 unique classes. \n\n1. I was wondering how to train the model for rest of the classes in the test data? If there are no such classes in WCS data alone, then how can the model generalizes?\n2. Do I need to include other datasets as well (iNaturalist 2017, 2018, 2019)? Or WCS alone should be sufficient?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1502092%2F2d6b8042f4b9591ea1927a5ec39ae12a%2FScreenshot%20from%202020-04-13%2018-01-5111.xcf?generation=1586798463866497&amp;alt=media)",
    "806403": "There are no classes in the test set images that were not included in the training set (we explicitly removed images from non-training classes at the test locations), but we do encourage including other datasets to help with rare classes that have little training data."
  },
  "source": "meta"
}