{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Visually Separating the Clusters\n\nIn this TPS competition, many of us are facing two challenges:\n1. We haven't yet found a good cross-validation metric.\n2. Visualizing seven clusters in 29 dimensions is quite difficult.\n\nThe notebook aims at dealing with these two challenges by simplifying the visualization task as follows:\n\n1. We look at only two clusters at a time: If we can separate every pair of two clusters, we have got what we need. This corresponds to what one-vs-one classifiers do in multiclass classification.\n2. We project the points of the two clusters from 29 dimensions to a single, well-chosen dimension, i.e. to a straight line. Of course the projection must be chosen so that the two clusters overlap as little as possible. It turns out that the ideal projection function is the decision function of a (linear) classifier.\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"import pandas as pd\nimport matplotlib.pyplot as plt\nimport numpy as np\nfrom sklearn.preprocessing import PowerTransformer, StandardScaler\nfrom sklearn.svm import LinearSVC, SVC\nfrom sklearn.discriminant_analysis import QuadraticDiscriminantAnalysis\nfrom sklearn.pipeline import make_pipeline\nfrom itertools import combinations","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-21T16:22:57.523153Z","iopub.execute_input":"2022-07-21T16:22:57.523618Z","iopub.status.idle":"2022-07-21T16:22:58.055741Z","shell.execute_reply.started":"2022-07-21T16:22:57.523522Z","shell.execute_reply":"2022-07-21T16:22:58.054509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We start by reading the data and @hiro5299834's clustering:","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv', index_col='id')\ngood_features = [f'f_{i:02}' for i in [7, 8, 9, 10, 11, 12, 13, 22, 23, 24, 25, 26, 27, 28]]\n#good_features = [f'f_{i:02}' for i in range(29)]\ndata.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T16:22:58.057770Z","iopub.execute_input":"2022-07-21T16:22:58.058135Z","iopub.status.idle":"2022-07-21T16:22:58.787842Z","shell.execute_reply.started":"2022-07-21T16:22:58.058104Z","shell.execute_reply":"2022-07-21T16:22:58.786607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clusters = pd.read_csv('../input/tps-jul-2022-unsupervised-and-supervised-learning/submission.csv', index_col='Id')\nclusters.head(2)","metadata":{"execution":{"iopub.status.busy":"2022-07-21T16:22:58.789651Z","iopub.execute_input":"2022-07-21T16:22:58.790004Z","iopub.status.idle":"2022-07-21T16:22:58.831181Z","shell.execute_reply.started":"2022-07-21T16:22:58.789972Z","shell.execute_reply":"2022-07-21T16:22:58.829988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we'll show the result of the projection for every pair of clusters and for three different classifiers:\n- A linear SVM\n- QuadraticDiscriminantAnalysis\n- A pipeline made from a PowerTransformer and a linear SVM\n\nYou can see that some clusters are separated quite well by the decision functions (e.g., clusters 0 and 3); for other pairs we haven't yet found a good decision function (e.g., clusters 0 and 1).","metadata":{}},{"cell_type":"code","source":"for ca, cb in list(combinations(np.unique(clusters.Predicted), 2)):\n    # Show the separation between clusters ca and cb\n    subset_mask = (clusters.Predicted == ca) | (clusters.Predicted == cb)\n    X = data[subset_mask][good_features].astype(np.float64)\n    y = clusters.Predicted[subset_mask]\n    \n    classifiers = [('Linear SVM', make_pipeline(StandardScaler(), \n                                                LinearSVC(max_iter=5000, random_state=1))),\n                   ('QDA', QuadraticDiscriminantAnalysis()),\n                   ('PowerTransformer + linear SVM', make_pipeline(PowerTransformer(),\n                                                                   StandardScaler(),\n                                                                   LinearSVC(max_iter=20000,random_state=1)))]\n    \n    _, axs = plt.subplots(1, len(classifiers), figsize=(12, 3))\n    for ax, (xlabel, clf) in zip(axs, classifiers):\n        clf.fit(X, y)\n        #print(clf.score(X, y))\n        ax.set_facecolor('#0057b8')\n        ax.hist(clf.decision_function(X), bins=100, density=True, color='#ffd700')\n        ax.vlines([0], 0, ax.get_ylim()[1], color='#a0ffa0')\n        ax.set_xlabel(f'{xlabel}')\n    plt.suptitle(f\"Separation for clusters {ca} and {cb}\", fontsize=16)\n    #plt.savefig(f'dec_{ca}_{cb}', bbox_inches='tight')\n    plt.show()\n    ","metadata":{"execution":{"iopub.status.busy":"2022-07-21T16:34:13.644682Z","iopub.execute_input":"2022-07-21T16:34:13.645155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The diagrams don't give us a cross-validation metric which measures the quality of a clustering in a single floating point number, but they give a good impression of the quality, and most important, they show which pairs of clusters need further investigation.\n\nWhat next?\n\n- Try out other classifiers with their decision functions, in particular for the cluster pairs which are not yet well separated\n- Tune the classifiers for optimal cluster separation by varying all available hyperparameters\n- Do the same experiment with eight or nine clusters and see what happens\n","metadata":{}}]}