{
  "id": 72164,
  "title": "Clustering for classification?  For class 99?",
  "url": "/competitions/PLAsTiCC-2018/discussion/72164",
  "author_name": "",
  "post_date": "2018-11-20T23:54:18.314643400Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I feel like there may be a way to use clustering possibly for classification and at least for class 99.  If we did kMeans with 14 clusters setting the initial centroids as the centroids of our 14 training set classes I would expect to end up with centroids that represent the classes.  Then we could measure the distance of the test points (using a non Euclidean metric).  Those that are outliers to every centroid could be put into class 99.</p>\n\n<p>Anyone with experience using clustering algorithms have thoughts on this?</p>",
  "messages": [
    {
      "id": "424960",
      "postDate": "11/20/2018 23:54:18",
      "content": "<p>I feel like there may be a way to use clustering possibly for classification and at least for class 99.  If we did kMeans with 14 clusters setting the initial centroids as the centroids of our 14 training set classes I would expect to end up with centroids that represent the classes.  Then we could measure the distance of the test points (using a non Euclidean metric).  Those that are outliers to every centroid could be put into class 99.</p>\n\n<p>Anyone with experience using clustering algorithms have thoughts on this?</p>",
      "rawMarkdown": "I feel like there may be a way to use clustering possibly for classification and at least for class 99.  If we did kMeans with 14 clusters setting the initial centroids as the centroids of our 14 training set classes I would expect to end up with centroids that represent the classes.  Then we could measure the distance of the test points (using a non Euclidean metric).  Those that are outliers to every centroid could be put into class 99.\n\nAnyone with experience using clustering algorithms have thoughts on this?",
      "votes": null
    },
    {
      "id": "425084",
      "postDate": "11/21/2018 05:33:11",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709</a></p>",
      "rawMarkdown": "See https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 425084,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/21/2018 05:33:11",
      "content": "<p>See <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709\">https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "424960": "I feel like there may be a way to use clustering possibly for classification and at least for class 99.  If we did kMeans with 14 clusters setting the initial centroids as the centroids of our 14 training set classes I would expect to end up with centroids that represent the classes.  Then we could measure the distance of the test points (using a non Euclidean metric).  Those that are outliers to every centroid could be put into class 99.\n\nAnyone with experience using clustering algorithms have thoughts on this?",
    "425084": "See https://www.kaggle.com/c/PLAsTiCC-2018/discussion/69376#latest-408709"
  },
  "source": "meta"
}