{
  "id": 306510,
  "title": "Unsupervised clustering to facilitate bagging-ensemble classification on an imbalanced dataset",
  "url": "/competitions/quora-insincere-questions-classification/discussion/306510",
  "author_name": "Senthil Kumar R",
  "post_date": "2022-02-09T17:47:29.713000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm working on Kaggle's <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\" target=\"_blank\">Quora Insincere Questions Classification contest</a>. The dataset has ~1225k rows of <code>target=0 class</code> and ~88k of <code>target=1 class</code>. This is a binary classification problem where <code>one target=1</code> record occurs for every <code>fifteen target=0</code> records (93%:7% imbalance).</p>\n<p>In order to solve the data imbalance, I came up with the following idea:</p>\n<blockquote>\n  <p>STEP 1: Break <code>target=0</code> class into 15 subclasses using <strong>unsupervised K-means model</strong>. The model will auto-fit on the class to create 15 different subsets<br>\n  <br><br>\n  STEP 2: have <strong>15 ANN models</strong> that will be trained on the whole of <code>target=1</code> class with each <code>target=0</code> subclasses <em>(ie, all 15 models will get a particular subset of <code>target=0</code> class and all entries of <code>target=1</code> class)</em><br>\n  <br><br>\n  STEP 3: Ensemble these 15 models using <strong>max-voting</strong>/averaging<br>\n  <br></p>\n</blockquote>\n<p>The notion behind my idea is the fact that the K-means model could break the huge majority class into clusters, with each entry in a cluster having some similarity with all other entries in the same cluster <em>(ie, using K-means to group similar records together into a cluster)</em>. This, according to my noobish knowledge, would only help in better learning as each sub-model will now only get a properly grouped <code>target=0</code> subclass from which it may extract knowledge.</p>\n<p>I've seen a lot of videos where I am asked to prefer random sampling over K-means learning, but I really want to know why is it that so. What is the flaw with going ahead with this design?</p>",
  "messages": [
    {
      "id": 1683288,
      "postDate": "2022-02-09T17:47:29.713Z",
      "content": "<p>I'm working on Kaggle's <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification\" target=\"_blank\">Quora Insincere Questions Classification contest</a>. The dataset has ~1225k rows of <code>target=0 class</code> and ~88k of <code>target=1 class</code>. This is a binary classification problem where <code>one target=1</code> record occurs for every <code>fifteen target=0</code> records (93%:7% imbalance).</p>\n<p>In order to solve the data imbalance, I came up with the following idea:</p>\n<blockquote>\n  <p>STEP 1: Break <code>target=0</code> class into 15 subclasses using <strong>unsupervised K-means model</strong>. The model will auto-fit on the class to create 15 different subsets<br>\n  <br><br>\n  STEP 2: have <strong>15 ANN models</strong> that will be trained on the whole of <code>target=1</code> class with each <code>target=0</code> subclasses <em>(ie, all 15 models will get a particular subset of <code>target=0</code> class and all entries of <code>target=1</code> class)</em><br>\n  <br><br>\n  STEP 3: Ensemble these 15 models using <strong>max-voting</strong>/averaging<br>\n  <br></p>\n</blockquote>\n<p>The notion behind my idea is the fact that the K-means model could break the huge majority class into clusters, with each entry in a cluster having some similarity with all other entries in the same cluster <em>(ie, using K-means to group similar records together into a cluster)</em>. This, according to my noobish knowledge, would only help in better learning as each sub-model will now only get a properly grouped <code>target=0</code> subclass from which it may extract knowledge.</p>\n<p>I've seen a lot of videos where I am asked to prefer random sampling over K-means learning, but I really want to know why is it that so. What is the flaw with going ahead with this design?</p>",
      "rawMarkdown": "I'm working on Kaggle's [Quora Insincere Questions Classification contest](https://www.kaggle.com/c/quora-insincere-questions-classification). The dataset has ~1225k rows of `target=0 class` and ~88k of `target=1 class`. This is a binary classification problem where `one target=1` record occurs for every `fifteen target=0` records (93%:7% imbalance).\n\nIn order to solve the data imbalance, I came up with the following idea:\n> STEP 1: Break `target=0` class into 15 subclasses using **unsupervised K-means model**. The model will auto-fit on the class to create 15 different subsets\n<br>\n>STEP 2: have **15 ANN models** that will be trained on the whole of `target=1` class with each `target=0` subclasses *(ie, all 15 models will get a particular subset of `target=0` class and all entries of `target=1` class)*\n<br>\n>STEP 3: Ensemble these 15 models using **max-voting**/averaging\n<br>\n\nThe notion behind my idea is the fact that the K-means model could break the huge majority class into clusters, with each entry in a cluster having some similarity with all other entries in the same cluster *(ie, using K-means to group similar records together into a cluster)*. This, according to my noobish knowledge, would only help in better learning as each sub-model will now only get a properly grouped `target=0` subclass from which it may extract knowledge.\n\nI've seen a lot of videos where I am asked to prefer random sampling over K-means learning, but I really want to know why is it that so. What is the flaw with going ahead with this design?\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1683288": "I'm working on Kaggle's [Quora Insincere Questions Classification contest](https://www.kaggle.com/c/quora-insincere-questions-classification). The dataset has ~1225k rows of `target=0 class` and ~88k of `target=1 class`. This is a binary classification problem where `one target=1` record occurs for every `fifteen target=0` records (93%:7% imbalance).\n\nIn order to solve the data imbalance, I came up with the following idea:\n> STEP 1: Break `target=0` class into 15 subclasses using **unsupervised K-means model**. The model will auto-fit on the class to create 15 different subsets\n<br>\n>STEP 2: have **15 ANN models** that will be trained on the whole of `target=1` class with each `target=0` subclasses *(ie, all 15 models will get a particular subset of `target=0` class and all entries of `target=1` class)*\n<br>\n>STEP 3: Ensemble these 15 models using **max-voting**/averaging\n<br>\n\nThe notion behind my idea is the fact that the K-means model could break the huge majority class into clusters, with each entry in a cluster having some similarity with all other entries in the same cluster *(ie, using K-means to group similar records together into a cluster)*. This, according to my noobish knowledge, would only help in better learning as each sub-model will now only get a properly grouped `target=0` subclass from which it may extract knowledge.\n\nI've seen a lot of videos where I am asked to prefer random sampling over K-means learning, but I really want to know why is it that so. What is the flaw with going ahead with this design?\n"
  }
}