{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### How to ensemble clustering algorithms?\n\n> - Since this competition is about clustering \n> - And since it is necessary to ensemble in order to achieve maximum Kaggling. \n\nOn this notebook we show how to ensemble multiple clustering algorithms.\n\n![](https://i.ibb.co/Sy5xgnQ/Scheme-of-ensemble-clustering-approach.png)\n\n[[source](https://www.researchgate.net/figure/Scheme-of-ensemble-clustering-approach_fig1_277589770)]\n\n\nThe approach is simple: We create a sparse matrix of all pairs of samples and assign 1 to any pair of samples that are on the same cluster.\n\nWe then simply take the mean/median of all sparse matrices, apply a threshold, and convert it back to cluster ids.\n\nThis way we can use multiple clustering algorithms at the same time. \n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-04T15:41:53.439899Z","iopub.execute_input":"2022-07-04T15:41:53.440308Z","iopub.status.idle":"2022-07-04T15:41:53.45155Z","shell.execute_reply.started":"2022-07-04T15:41:53.440276Z","shell.execute_reply":"2022-07-04T15:41:53.450577Z"}}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom tqdm import trange\nfrom sklearn.cluster import KMeans\nfrom sklearn.mixture import BayesianGaussianMixture\n\ndata = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv', index_col='id')\n\ndata = data[:5000] # Just for the fast demonstration. ","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:46:21.614725Z","iopub.execute_input":"2022-07-08T13:46:21.615064Z","iopub.status.idle":"2022-07-08T13:46:22.542458Z","shell.execute_reply.started":"2022-07-08T13:46:21.615038Z","shell.execute_reply":"2022-07-08T13:46:22.541143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **Removing 'useless' columns**\n\nUsless columns: Columns that have same mean for all the 7 clusters are refered to **'useless'** in certain kernels. And removing them slightly increases the scores, so it is being used here.","metadata":{}},{"cell_type":"code","source":"imp_col = ['f_07','f_08','f_09','f_10','f_11','f_12','f_13',\n          'f_22', 'f_23', 'f_24', 'f_25', 'f_26', 'f_27', 'f_28']\ndf = data[imp_col]","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:46:24.631973Z","iopub.execute_input":"2022-07-08T13:46:24.632343Z","iopub.status.idle":"2022-07-08T13:46:24.638791Z","shell.execute_reply.started":"2022-07-08T13:46:24.632313Z","shell.execute_reply":"2022-07-08T13:46:24.637810Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **Scale the columns**\n\nThis scales both 'int' and 'float' columns. Separating them could be useful.","metadata":{}},{"cell_type":"code","source":"from sklearn import preprocessing\n\nx = df.values\nrb_scalar = preprocessing.RobustScaler()\nx_scaled = rb_scalar.fit_transform(x)\ndata_scaled = pd.DataFrame(x_scaled, columns=df.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:46:28.485781Z","iopub.execute_input":"2022-07-08T13:46:28.486107Z","iopub.status.idle":"2022-07-08T13:46:28.500421Z","shell.execute_reply.started":"2022-07-08T13:46:28.486082Z","shell.execute_reply":"2022-07-08T13:46:28.499304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### **K-means and Bayesian Gaussian Mixture Models**","metadata":{}},{"cell_type":"code","source":"clusters_k_means = KMeans(n_clusters = 7).fit_predict(data_scaled)\nclusters_mixture = BayesianGaussianMixture(n_components = 7).fit_predict(data_scaled)\n\nprint(clusters_k_means.shape)\nprint(clusters_mixture.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:46:31.642482Z","iopub.execute_input":"2022-07-08T13:46:31.642839Z","iopub.status.idle":"2022-07-08T13:46:36.356178Z","shell.execute_reply.started":"2022-07-08T13:46:31.642809Z","shell.execute_reply":"2022-07-08T13:46:36.354027Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### We create a sparse matrix of all pairs of samples and assign 1 to any pair of samples that are on the same cluster.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"from scipy.sparse import csr_matrix\n\ndef create_sparse_matrix(clusters):\n    n = len(clusters)\n    data = []\n    row = []\n    col = []\n    for i in trange(n):\n        for j in range(i+1, n):\n            if clusters[i] == clusters[j]:\n                data.append(1)\n                row.append(i)\n                col.append(j)\n    return csr_matrix((data, (row, col)), shape=(n, n))\n\nsparse_matrix_k_means = create_sparse_matrix(clusters_k_means)\nsparse_matrix_mixture = create_sparse_matrix(clusters_mixture)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:46:42.730171Z","iopub.execute_input":"2022-07-08T13:46:42.730576Z","iopub.status.idle":"2022-07-08T13:46:58.110639Z","shell.execute_reply.started":"2022-07-08T13:46:42.730543Z","shell.execute_reply":"2022-07-08T13:46:58.109732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### We then simply take the mean/median of all sparse matrices, apply a threshold, and convert it back to cluster ids.\n\n![](https://i.ibb.co/84nnT31/1-s2-0-S153204641300021-X-gr1.jpg)\n[[source](https://www.sciencedirect.com/science/article/pii/S153204641300021X)]\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"sparse_matrix_mean = (sparse_matrix_k_means + sparse_matrix_mixture) / 2","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:46:58.112497Z","iopub.execute_input":"2022-07-08T13:46:58.113127Z","iopub.status.idle":"2022-07-08T13:46:58.156927Z","shell.execute_reply.started":"2022-07-08T13:46:58.113089Z","shell.execute_reply":"2022-07-08T13:46:58.155815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now we apply the threshold","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"sparse_matrix_mean[sparse_matrix_mean < 0.5] = 0\nsparse_matrix_mean[sparse_matrix_mean >= 0.5] = 1","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:46:58.158410Z","iopub.execute_input":"2022-07-08T13:46:58.159200Z","iopub.status.idle":"2022-07-08T13:47:05.900054Z","shell.execute_reply.started":"2022-07-08T13:46:58.159156Z","shell.execute_reply":"2022-07-08T13:47:05.898541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Now we reconstruct our clusters and return an integer column","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"sparse_matrix_mean = sparse_matrix_mean.toarray()\n\nclusters_final = np.zeros(len(df))\nclusters_final_next_id = 0\n\nfor i in range(len(df)):\n    if clusters_final[i] == 0:\n        clusters_final_current_id = clusters_final_next_id\n        clusters_final_next_id += 1\n        clusters_final[i] = clusters_final_current_id\n        for j in range(i+1, len(df)):\n            if sparse_matrix_mean[i, j] == 1:\n                clusters_final[j] = clusters_final_current_id\n\nprint(clusters_final)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-08T13:47:05.902866Z","iopub.execute_input":"2022-07-08T13:47:05.903152Z","iopub.status.idle":"2022-07-08T13:47:06.234549Z","shell.execute_reply.started":"2022-07-08T13:47:05.903124Z","shell.execute_reply":"2022-07-08T13:47:06.233497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Visualize with with tSNE**\nPlay around with the perplexity (currently 5% of the dataset)","metadata":{}},{"cell_type":"code","source":"from sklearn.manifold import TSNE\ntsne = TSNE(perplexity=250)\ndata_tsne_2d = tsne.fit_transform(data_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-08T13:55:36.863595Z","iopub.execute_input":"2022-07-08T13:55:36.863984Z","iopub.status.idle":"2022-07-08T13:56:35.905113Z","shell.execute_reply.started":"2022-07-08T13:55:36.863944Z","shell.execute_reply":"2022-07-08T13:56:35.904278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nplt.figure(figsize = (10,10))\nplt.scatter(data_tsne_2d[:,0], data_tsne_2d[:,1], cmap=\"Set1\", s=10, c=clusters_final)\nplt.axis('off')","metadata":{"execution":{"iopub.status.busy":"2022-07-08T14:00:29.468550Z","iopub.execute_input":"2022-07-08T14:00:29.468955Z","iopub.status.idle":"2022-07-08T14:00:29.685523Z","shell.execute_reply.started":"2022-07-08T14:00:29.468924Z","shell.execute_reply":"2022-07-08T14:00:29.684696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Export the results**\nNot done for trail runs","metadata":{}},{"cell_type":"code","source":"#ss = pd.read_csv('/kaggle/input/tabular-playground-series-jul-2022/sample_submission.csv', index_col='Id')\n#ss[:5000].head()","metadata":{"execution":{"iopub.status.busy":"2022-07-08T14:01:07.215577Z","iopub.execute_input":"2022-07-08T14:01:07.216398Z","iopub.status.idle":"2022-07-08T14:01:07.220582Z","shell.execute_reply.started":"2022-07-08T14:01:07.216370Z","shell.execute_reply":"2022-07-08T14:01:07.219771Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ss['Predicted'] = clusters_final\n#ss.to_csv(\"sample_submission_Kcluster_BGMM_C7_ImpCol.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-07-08T14:01:10.595745Z","iopub.execute_input":"2022-07-08T14:01:10.596101Z","iopub.status.idle":"2022-07-08T14:01:10.611459Z","shell.execute_reply.started":"2022-07-08T14:01:10.596072Z","shell.execute_reply":"2022-07-08T14:01:10.610378Z"},"trusted":true},"execution_count":null,"outputs":[]}]}