{
  "id": 90965,
  "title": "How to do C/V  split?",
  "url": "/competitions/imet-2019-fgvc6/discussion/90965",
  "author_name": "",
  "post_date": "2019-04-29T15:13:14.743510700Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi guys, this is the first time I have to deal with such an unbalaced dataset, with more than 150 classes with less than 10 samples, and even 15 with only 1 sample. How do you do CV split so that you can do local cross validation? I can't evem trust my validation set with some labels with only 1 sample.\nThanks for your help in advance!</p>",
  "messages": [
    {
      "id": "524814",
      "postDate": "04/29/2019 15:13:14",
      "content": "<p>Hi guys, this is the first time I have to deal with such an unbalaced dataset, with more than 150 classes with less than 10 samples, and even 15 with only 1 sample. How do you do CV split so that you can do local cross validation? I can't evem trust my validation set with some labels with only 1 sample.\nThanks for your help in advance!</p>",
      "rawMarkdown": "Hi guys, this is the first time I have to deal with such an unbalaced dataset, with more than 150 classes with less than 10 samples, and even 15 with only 1 sample. How do you do CV split so that you can do local cross validation? I can't evem trust my validation set with some labels with only 1 sample.\nThanks for your help in advance!",
      "votes": null
    },
    {
      "id": "524938",
      "postDate": "04/29/2019 19:52:06",
      "content": "<p>Dear JayChen, have you tried SMOTE oversampling? I think you can fix this by SMOTE oversampling each minority class against all data not in that class. It will use k-means, and will create duplicates between an item and its neighbor. Hope this will help.</p>",
      "rawMarkdown": "Dear JayChen, have you tried SMOTE oversampling? I think you can fix this by SMOTE oversampling each minority class against all data not in that class. It will use k-means, and will create duplicates between an item and its neighbor. Hope this will help.",
      "votes": null
    },
    {
      "id": "525222",
      "postDate": "04/30/2019 13:17:37",
      "content": "<p>Thank you for your suggestion! I'll look into that!</p>",
      "rawMarkdown": "Thank you for your suggestion! I'll look into that!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 524938,
      "author_name": "karlentsatinyan",
      "author_url": "",
      "post_date": "04/29/2019 19:52:06",
      "content": "<p>Dear JayChen, have you tried SMOTE oversampling? I think you can fix this by SMOTE oversampling each minority class against all data not in that class. It will use k-means, and will create duplicates between an item and its neighbor. Hope this will help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 525222,
          "author_name": "marsyeti",
          "author_url": "",
          "post_date": "04/30/2019 13:17:37",
          "content": "<p>Thank you for your suggestion! I'll look into that!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "524814": "Hi guys, this is the first time I have to deal with such an unbalaced dataset, with more than 150 classes with less than 10 samples, and even 15 with only 1 sample. How do you do CV split so that you can do local cross validation? I can't evem trust my validation set with some labels with only 1 sample.\nThanks for your help in advance!",
    "524938": "Dear JayChen, have you tried SMOTE oversampling? I think you can fix this by SMOTE oversampling each minority class against all data not in that class. It will use k-means, and will create duplicates between an item and its neighbor. Hope this will help.",
    "525222": "Thank you for your suggestion! I'll look into that!"
  },
  "source": "meta"
}