{
  "id": 122629,
  "title": "Unequal proportion of subclasses in graphemes",
  "url": "/competitions/bengaliai-cv19/discussion/122629",
  "author_name": "",
  "post_date": "2019-12-21T15:04:15.955692600Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi all! I am new to machine learning and I am trying out this competition for learning.</p>\n\n<p>I notice that there are huge proportion differences within each components/graphemes, especially for consonant_diacritic. Below is the distribution of each of the graphemes for train_data_0. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2Fe20c515110aa3e7001fc1368c865cc0b%2Fdownload.png?generation=1576940353019216&amp;alt=media\" alt=\"\"></p>\n\n<p>In this case, how should I go about proceeding? As some subclasses are going to be trained more than the others. On the other hand, if the distribution is similar to the real-world distribution of subclasses within each graphemes, then will it be ok if we train them on unequal proportions?</p>",
  "messages": [
    {
      "id": "700179",
      "postDate": "12/21/2019 15:04:15",
      "content": "<p>Hi all! I am new to machine learning and I am trying out this competition for learning.</p>\n\n<p>I notice that there are huge proportion differences within each components/graphemes, especially for consonant_diacritic. Below is the distribution of each of the graphemes for train_data_0. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2Fe20c515110aa3e7001fc1368c865cc0b%2Fdownload.png?generation=1576940353019216&amp;alt=media\" alt=\"\"></p>\n\n<p>In this case, how should I go about proceeding? As some subclasses are going to be trained more than the others. On the other hand, if the distribution is similar to the real-world distribution of subclasses within each graphemes, then will it be ok if we train them on unequal proportions?</p>",
      "rawMarkdown": "Hi all! I am new to machine learning and I am trying out this competition for learning.\n\nI notice that there are huge proportion differences within each components/graphemes, especially for consonant_diacritic. Below is the distribution of each of the graphemes for train_data_0. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2Fe20c515110aa3e7001fc1368c865cc0b%2Fdownload.png?generation=1576940353019216&amp;alt=media)\n\nIn this case, how should I go about proceeding? As some subclasses are going to be trained more than the others. On the other hand, if the distribution is similar to the real-world distribution of subclasses within each graphemes, then will it be ok if we train them on unequal proportions?",
      "votes": null
    },
    {
      "id": "700934",
      "postDate": "12/22/2019 20:52:16",
      "content": "<p>I think there are multiple interesting approaches one could take for this:\n- Either try to balance the datasets by simply removing some images in the classes over represented to have a more balanced dataset. But this isn't a great approach since we actively would remove some data, and a lot of it.\n- Give a different weight to the prediction of each different class depending on it's frequency: the less a given class appears, the more penalizing the score could be. Definitely not perfect, but could be looked into.\n- Some unbalanced data augmentation might be interesting: \"make\" more artificial images for the ones the least represented.</p>\n\n<p>Those are simply a few very quick ideas. This is an interesting competition for this exact problem. I am pretty curious to see how people will deal with this!</p>",
      "rawMarkdown": "I think there are multiple interesting approaches one could take for this:\n- Either try to balance the datasets by simply removing some images in the classes over represented to have a more balanced dataset. But this isn't a great approach since we actively would remove some data, and a lot of it.\n- Give a different weight to the prediction of each different class depending on it's frequency: the less a given class appears, the more penalizing the score could be. Definitely not perfect, but could be looked into.\n- Some unbalanced data augmentation might be interesting: \"make\" more artificial images for the ones the least represented.\n\nThose are simply a few very quick ideas. This is an interesting competition for this exact problem. I am pretty curious to see how people will deal with this!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 700934,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "12/22/2019 20:52:16",
      "content": "<p>I think there are multiple interesting approaches one could take for this:\n- Either try to balance the datasets by simply removing some images in the classes over represented to have a more balanced dataset. But this isn't a great approach since we actively would remove some data, and a lot of it.\n- Give a different weight to the prediction of each different class depending on it's frequency: the less a given class appears, the more penalizing the score could be. Definitely not perfect, but could be looked into.\n- Some unbalanced data augmentation might be interesting: \"make\" more artificial images for the ones the least represented.</p>\n\n<p>Those are simply a few very quick ideas. This is an interesting competition for this exact problem. I am pretty curious to see how people will deal with this!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "700179": "Hi all! I am new to machine learning and I am trying out this competition for learning.\n\nI notice that there are huge proportion differences within each components/graphemes, especially for consonant_diacritic. Below is the distribution of each of the graphemes for train_data_0. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4022661%2Fe20c515110aa3e7001fc1368c865cc0b%2Fdownload.png?generation=1576940353019216&amp;alt=media)\n\nIn this case, how should I go about proceeding? As some subclasses are going to be trained more than the others. On the other hand, if the distribution is similar to the real-world distribution of subclasses within each graphemes, then will it be ok if we train them on unequal proportions?",
    "700934": "I think there are multiple interesting approaches one could take for this:\n- Either try to balance the datasets by simply removing some images in the classes over represented to have a more balanced dataset. But this isn't a great approach since we actively would remove some data, and a lot of it.\n- Give a different weight to the prediction of each different class depending on it's frequency: the less a given class appears, the more penalizing the score could be. Definitely not perfect, but could be looked into.\n- Some unbalanced data augmentation might be interesting: \"make\" more artificial images for the ones the least represented.\n\nThose are simply a few very quick ideas. This is an interesting competition for this exact problem. I am pretty curious to see how people will deal with this!"
  },
  "source": "meta"
}