{
  "id": 308565,
  "title": "Why I don't think vanilla KNN is the best solution for embeddings classification",
  "url": "/competitions/happy-whale-and-dolphin/discussion/308565",
  "author_name": "",
  "post_date": "2022-02-19T08:20:21.438697700Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As we know that the class label assigned to a query vector is determined by a majority vote amongst the k-nearest neighbors of the query vector. For simplicity, let us assume that we are using Euclidean distance as the distance metric here and its a binary classification problem. The inductive bias of knn classification in this case assumes that the space of all samples can be partitioned into two disjoint subspaces. If training set is large enough and balanced, it'd seem as if the subspaces were sampled uniformly for generating the training set. If the training set were to be imbalanced, and if the uniform distribution assumption were to still hold, the probability that the k nearest neighbors of any random query point will belong to the class with more examples becomes higher. So, the closest neighbor of the query point may still belong to the class with less examples, but if rest (k-1) points belong to the other class (because of its higher density in the space), the point will get misclassified. <br>\nAs we can the case in our problem too where the individuals histogram shows us the unbalance of training samples<br>\n<img src=\"https://i.imgur.com/5NxHJOD.png\" alt=\"\"><br>\nSo, what can we do to compensate this ? Most easy to implement solutions are:</p>\n<ul>\n<li>weighting neighbors by the inverse of their class size converts neighbor counts into the fraction of each class that falls in your K nearest neighbors</li>\n<li>weighting neighbors by their distances</li>\n<li>using a radius-based rule for gathering neighbors instead of the K nearest ones (often implemented in KNN packages)</li>\n</ul>\n<p>But there are much more solutions and/or classifiers that can help. Below are 2 links with generical approaches when our targets are imbalanced</p>\n<ul>\n<li><a href=\"https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/\" target=\"_blank\">https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/</a></li>\n<li><a href=\"https://elitedatascience.com/imbalanced-classes\" target=\"_blank\">https://elitedatascience.com/imbalanced-classes</a></li>\n</ul>\n<p>Biography:<br>\n<a href=\"https://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf\" target=\"_blank\">https://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf</a><br>\n<a href=\"https://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes\" target=\"_blank\">https://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes</a><br>\n<a href=\"https://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance\" target=\"_blank\">https://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance</a></p>",
  "messages": [
    {
      "id": "1696935",
      "postDate": "02/19/2022 08:20:21",
      "content": "<p>As we know that the class label assigned to a query vector is determined by a majority vote amongst the k-nearest neighbors of the query vector. For simplicity, let us assume that we are using Euclidean distance as the distance metric here and its a binary classification problem. The inductive bias of knn classification in this case assumes that the space of all samples can be partitioned into two disjoint subspaces. If training set is large enough and balanced, it'd seem as if the subspaces were sampled uniformly for generating the training set. If the training set were to be imbalanced, and if the uniform distribution assumption were to still hold, the probability that the k nearest neighbors of any random query point will belong to the class with more examples becomes higher. So, the closest neighbor of the query point may still belong to the class with less examples, but if rest (k-1) points belong to the other class (because of its higher density in the space), the point will get misclassified. <br>\nAs we can the case in our problem too where the individuals histogram shows us the unbalance of training samples<br>\n<img src=\"https://i.imgur.com/5NxHJOD.png\" alt=\"\"><br>\nSo, what can we do to compensate this ? Most easy to implement solutions are:</p>\n<ul>\n<li>weighting neighbors by the inverse of their class size converts neighbor counts into the fraction of each class that falls in your K nearest neighbors</li>\n<li>weighting neighbors by their distances</li>\n<li>using a radius-based rule for gathering neighbors instead of the K nearest ones (often implemented in KNN packages)</li>\n</ul>\n<p>But there are much more solutions and/or classifiers that can help. Below are 2 links with generical approaches when our targets are imbalanced</p>\n<ul>\n<li><a href=\"https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/\" target=\"_blank\">https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/</a></li>\n<li><a href=\"https://elitedatascience.com/imbalanced-classes\" target=\"_blank\">https://elitedatascience.com/imbalanced-classes</a></li>\n</ul>\n<p>Biography:<br>\n<a href=\"https://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf\" target=\"_blank\">https://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf</a><br>\n<a href=\"https://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes\" target=\"_blank\">https://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes</a><br>\n<a href=\"https://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance\" target=\"_blank\">https://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance</a></p>",
      "rawMarkdown": "As we know that the class label assigned to a query vector is determined by a majority vote amongst the k-nearest neighbors of the query vector. For simplicity, let us assume that we are using Euclidean distance as the distance metric here and its a binary classification problem. The inductive bias of knn classification in this case assumes that the space of all samples can be partitioned into two disjoint subspaces. If training set is large enough and balanced, it'd seem as if the subspaces were sampled uniformly for generating the training set. If the training set were to be imbalanced, and if the uniform distribution assumption were to still hold, the probability that the k nearest neighbors of any random query point will belong to the class with more examples becomes higher. So, the closest neighbor of the query point may still belong to the class with less examples, but if rest (k-1) points belong to the other class (because of its higher density in the space), the point will get misclassified. \nAs we can the case in our problem too where the individuals histogram shows us the unbalance of training samples\n![](https://i.imgur.com/5NxHJOD.png)\nSo, what can we do to compensate this ? Most easy to implement solutions are:\n\n- weighting neighbors by the inverse of their class size converts neighbor counts into the fraction of each class that falls in your K nearest neighbors\n- weighting neighbors by their distances\n- using a radius-based rule for gathering neighbors instead of the K nearest ones (often implemented in KNN packages)\n\nBut there are much more solutions and/or classifiers that can help. Below are 2 links with generical approaches when our targets are imbalanced\n- https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/\n- https://elitedatascience.com/imbalanced-classes\n\nBiography:\nhttps://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf\nhttps://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes\nhttps://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance",
      "votes": null
    },
    {
      "id": "1697316",
      "postDate": "02/19/2022 14:28:41",
      "content": "<p>Can we just calculate the centroid of all clusters? Assuming that all clusters are spheres, will it bring the same result as a uniformly balanced dataset?</p>",
      "rawMarkdown": "Can we just calculate the centroid of all clusters? Assuming that all clusters are spheres, will it bring the same result as a uniformly balanced dataset?",
      "votes": null
    },
    {
      "id": "1697328",
      "postDate": "02/19/2022 14:42:07",
      "content": "<p>ICYMI: Grandmaster Awsaf has made a neat <a href=\"https://www.kaggle.com/awsaf49/happywhale-data-distribution\" target=\"_blank\">kernel</a> showcasing the distributions</p>",
      "rawMarkdown": "ICYMI: Grandmaster Awsaf has made a neat [kernel](https://www.kaggle.com/awsaf49/happywhale-data-distribution) showcasing the distributions",
      "votes": null
    },
    {
      "id": "1697376",
      "postDate": "02/19/2022 15:18:59",
      "content": "<p><a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a>  I am not sure that I understand but we may think of similar solutions. I am thinking of making just one data point for every individual_id by making for example the mean (for every dimension  of the 512) of all samples for that individual so in this way every individual will have just one point (considered the centroid) then for inference calculate the distance from the unknown embedding to each of the centroids and classify regarding that distance</p>",
      "rawMarkdown": "meowmeowmeowmeowmeow  I am not sure that I understand but we may think of similar solutions. I am thinking of making just one data point for every individual_id by making for example the mean (for every dimension  of the 512) of all samples for that individual so in this way every individual will have just one point (considered the centroid) then for inference calculate the distance from the unknown embedding to each of the centroids and classify regarding that distance",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1697316,
      "author_name": "meowmeowmeowmeowmeow",
      "author_url": "",
      "post_date": "02/19/2022 14:28:41",
      "content": "<p>Can we just calculate the centroid of all clusters? Assuming that all clusters are spheres, will it bring the same result as a uniformly balanced dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1697328,
          "author_name": "init27",
          "author_url": "",
          "post_date": "02/19/2022 14:42:07",
          "content": "<p>ICYMI: Grandmaster Awsaf has made a neat <a href=\"https://www.kaggle.com/awsaf49/happywhale-data-distribution\" target=\"_blank\">kernel</a> showcasing the distributions</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1697376,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/19/2022 15:18:59",
          "content": "<p><a href=\"https://www.kaggle.com/meowmeowmeowmeowmeow\" target=\"_blank\">@meowmeowmeowmeowmeow</a>  I am not sure that I understand but we may think of similar solutions. I am thinking of making just one data point for every individual_id by making for example the mean (for every dimension  of the 512) of all samples for that individual so in this way every individual will have just one point (considered the centroid) then for inference calculate the distance from the unknown embedding to each of the centroids and classify regarding that distance</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1696935": "As we know that the class label assigned to a query vector is determined by a majority vote amongst the k-nearest neighbors of the query vector. For simplicity, let us assume that we are using Euclidean distance as the distance metric here and its a binary classification problem. The inductive bias of knn classification in this case assumes that the space of all samples can be partitioned into two disjoint subspaces. If training set is large enough and balanced, it'd seem as if the subspaces were sampled uniformly for generating the training set. If the training set were to be imbalanced, and if the uniform distribution assumption were to still hold, the probability that the k nearest neighbors of any random query point will belong to the class with more examples becomes higher. So, the closest neighbor of the query point may still belong to the class with less examples, but if rest (k-1) points belong to the other class (because of its higher density in the space), the point will get misclassified. \nAs we can the case in our problem too where the individuals histogram shows us the unbalance of training samples\n![](https://i.imgur.com/5NxHJOD.png)\nSo, what can we do to compensate this ? Most easy to implement solutions are:\n\n- weighting neighbors by the inverse of their class size converts neighbor counts into the fraction of each class that falls in your K nearest neighbors\n- weighting neighbors by their distances\n- using a radius-based rule for gathering neighbors instead of the K nearest ones (often implemented in KNN packages)\n\nBut there are much more solutions and/or classifiers that can help. Below are 2 links with generical approaches when our targets are imbalanced\n- https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/\n- https://elitedatascience.com/imbalanced-classes\n\nBiography:\nhttps://www.site.uottawa.ca/~nat/Workshop2003/jzhang.pdf\nhttps://stats.stackexchange.com/questions/341/knn-and-unbalanced-classes\nhttps://www.quora.com/Why-does-knn-get-effected-by-the-class-imbalance",
    "1697316": "Can we just calculate the centroid of all clusters? Assuming that all clusters are spheres, will it bring the same result as a uniformly balanced dataset?",
    "1697328": "ICYMI: Grandmaster Awsaf has made a neat [kernel](https://www.kaggle.com/awsaf49/happywhale-data-distribution) showcasing the distributions",
    "1697376": "meowmeowmeowmeowmeow  I am not sure that I understand but we may think of similar solutions. I am thinking of making just one data point for every individual_id by making for example the mean (for every dimension  of the 512) of all samples for that individual so in this way every individual will have just one point (considered the centroid) then for inference calculate the distance from the unknown embedding to each of the centroids and classify regarding that distance"
  },
  "source": "meta"
}