{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Getting Cluster_Ensembles to Work\n\nThere's been some some discussion of ensembling clusters, particularly in light of the sensitivity of results to seeds:\n- [@thedevastator](https://www.kaggle.com/thedevastator) posted a great notebook; [How to Ensemble Clustering algorithms?](https://www.kaggle.com/competitions/tabular-playground-series-jul-2022/discussion/335078#1843648). \n- [@adaubas](https://www.kaggle.com/adaubas) posted a couple of links [How to Ensemble Clustering Algorithms](https://towardsdatascience.com/how-to-ensemble-clustering-algorithms-bf78d7602265) and [Cluster_Ensembles](https://pypi.org/project/Cluster_Ensembles/)\n\nI think the notebook and the first link are the same algorithm, known as the __Cluster-based Similarity Partitioning Algorithm (CSPA)__ from a paper by [Strehl & Ghosh](http://www.strehl.com/download/strehl-aaai02.pdf). It's a very intuitive algorithm, but unfortunately it's computational and storage complexity are both $O(n^2).$ In our case we have ~100,000 observations, so $n^2$ is huge.\n\nI checked out the other suggestion; the Cluster_Ensembles package. It implements __CSPA__, as well as __HyperGraph Partitioning (HGPA)__ and __Meta-Clustering MCLA__, also from the Stehl & Ghosh paper. But this package is a little old and needs a little help to run...","metadata":{}},{"cell_type":"markdown","source":"#### Per the documentation, Cluster_Ensembles is based on the METIS hypergraph library. \n\nFirst we need to be able to run 32-bit executables.","metadata":{}},{"cell_type":"code","source":"!sudo apt -y install libc6-dev-i386","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:29:46.708299Z","iopub.execute_input":"2022-07-11T14:29:46.709390Z","iopub.status.idle":"2022-07-11T14:29:49.113508Z","shell.execute_reply.started":"2022-07-11T14:29:46.709345Z","shell.execute_reply":"2022-07-11T14:29:49.112022Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Install METIS","metadata":{}},{"cell_type":"code","source":"!sudo apt-get -y install metis","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:29:49.116121Z","iopub.execute_input":"2022-07-11T14:29:49.116601Z","iopub.status.idle":"2022-07-11T14:29:51.494968Z","shell.execute_reply.started":"2022-07-11T14:29:49.116560Z","shell.execute_reply":"2022-07-11T14:29:51.494035Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Install the Cluster_Ensembles package.","metadata":{}},{"cell_type":"code","source":"!pip install Cluster_Ensembles","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-11T14:29:51.496297Z","iopub.execute_input":"2022-07-11T14:29:51.496624Z","iopub.status.idle":"2022-07-11T14:30:02.177724Z","shell.execute_reply.started":"2022-07-11T14:29:51.496588Z","shell.execute_reply":"2022-07-11T14:30:02.176470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Resolve scikit-learn incompatibility\n\nWe need to replace a reference to `jaccard_similarity_score` with `jaccard_score` to be compatible with the most recent scikit-learn.","metadata":{}},{"cell_type":"code","source":"!curl https://raw.githubusercontent.com/GGiecold-zz/Cluster_Ensembles/master/src/Cluster_Ensembles/Cluster_Ensembles.py > /tmp/Cluster_Ensembles.py\n!sed -i s/jaccard_similarity_score/jaccard_score/g /tmp/Cluster_Ensembles.py\n!cp /tmp/Cluster_Ensembles.py /opt/conda/lib/python3.7/site-packages/Cluster_Ensembles/Cluster_Ensembles.py\n\n!grep jaccard /opt/conda/lib/python3.7/site-packages/Cluster_Ensembles/Cluster_Ensembles.py","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:02.180603Z","iopub.execute_input":"2022-07-11T14:30:02.181022Z","iopub.status.idle":"2022-07-11T14:30:05.548705Z","shell.execute_reply.started":"2022-07-11T14:30:02.180979Z","shell.execute_reply":"2022-07-11T14:30:05.547348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Ready to go! Let's give it a test run\n- Load and transform the data\n- Fit and save 3 clusterings using a GaussianMixture","metadata":{}},{"cell_type":"code","source":"import gc\nfrom pathlib import Path\n\nimport numpy as np \nimport pandas as pd\nimport random\n\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import MaxNLocator\nimport seaborn as sns\n\nfrom sklearn.mixture import GaussianMixture\nfrom sklearn.preprocessing import RobustScaler, PowerTransformer\n\nimport warnings\nwarnings.simplefilter('ignore')\n\npd.set_option('display.max_columns', None)\npd.set_option('display.float_format', '{:.3f}'.format)\n\nINPUT = Path('../input/tabular-playground-series-jul-2022')\nSEED = 420","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:05.551043Z","iopub.execute_input":"2022-07-11T14:30:05.551422Z","iopub.status.idle":"2022-07-11T14:30:05.560696Z","shell.execute_reply.started":"2022-07-11T14:30:05.551384Z","shell.execute_reply":"2022-07-11T14:30:05.559388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.read_csv(INPUT / 'data.csv',index_col='id')\ndata.info()\n\ndf = data[['f_07','f_08','f_09','f_10','f_11','f_12','f_13',\n           'f_22','f_23','f_24','f_25','f_26','f_27','f_28']]\n\nX = RobustScaler().fit_transform(df)\nX = pd.DataFrame(X, columns = df.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:05.562344Z","iopub.execute_input":"2022-07-11T14:30:05.563038Z","iopub.status.idle":"2022-07-11T14:30:06.528949Z","shell.execute_reply.started":"2022-07-11T14:30:05.562995Z","shell.execute_reply":"2022-07-11T14:30:06.527818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"n_iter = 3\nn_comp = 7\n\nclusterings = np.empty((n_iter, X.shape[0]))\nprint('fitting seeds: ',end='')\nfor i in range(n_iter):\n    seed = 100 * i\n    print(seed, end='...')\n    gmm = GaussianMixture(n_components=n_comp, \n                          covariance_type='full',\n                          max_iter=200,\n                          n_init=1,\n                          random_state=seed)\n    gmm.fit(X)\n    clusterings[i,:] = gmm.predict(X)    ","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:06.530282Z","iopub.execute_input":"2022-07-11T14:30:06.530608Z","iopub.status.idle":"2022-07-11T14:30:27.728285Z","shell.execute_reply.started":"2022-07-11T14:30:06.530578Z","shell.execute_reply":"2022-07-11T14:30:27.726929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"clusterings","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:27.730132Z","iopub.execute_input":"2022-07-11T14:30:27.731474Z","iopub.status.idle":"2022-07-11T14:30:27.741017Z","shell.execute_reply.started":"2022-07-11T14:30:27.731419Z","shell.execute_reply":"2022-07-11T14:30:27.739766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Generate a consensus clustering\n\nNote: Cluster_Ensembles detects that we have too much data to use CSPA.","metadata":{}},{"cell_type":"code","source":"import Cluster_Ensembles as CE\n\nconsensus_labels = CE.cluster_ensembles(clusterings, verbose = True, N_clusters_max = n_comp)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:27.742879Z","iopub.execute_input":"2022-07-11T14:30:27.749457Z","iopub.status.idle":"2022-07-11T14:30:35.611694Z","shell.execute_reply.started":"2022-07-11T14:30:27.749384Z","shell.execute_reply":"2022-07-11T14:30:35.610394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Sanity check that we did indeed get some kind of labeling","metadata":{}},{"cell_type":"code","source":"consensus_labels","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:35.615291Z","iopub.execute_input":"2022-07-11T14:30:35.615634Z","iopub.status.idle":"2022-07-11T14:30:35.623196Z","shell.execute_reply.started":"2022-07-11T14:30:35.615596Z","shell.execute_reply":"2022-07-11T14:30:35.621728Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Visually compare consensus clustering to initial clusterings\n\nIt does appear that the consensus clustering shows less fragmentation vs the initial clusterings then they do between themselves.\n\nBorrowed this cool routine from [@ambrosm](https://www.kaggle.com/ambrosm)\nhere: https://www.kaggle.com/code/ambrosm/tpsjul22-gaussian-mixture-cluster-analysis","metadata":{}},{"cell_type":"code","source":"def compare_clusterings(ax, y1, y2, title=''):\n    \"\"\"plot the two clusterings in color\"\"\"\n    n1 = y1.max() + 1\n    n2 = y2.max() + 1\n    argsort = np.argsort(y1*100 + y2) if n1 >= n2 else np.argsort(y2*100 + y1)\n    for i in range(6, 11):\n        ax.scatter(np.arange(len(y1)), np.full_like(y1, i), c=y1[argsort], s=1, cmap='tab10')\n    for i in range(5):\n        ax.scatter(np.arange(len(y2)), np.full_like(y2, i), c=y2[argsort], s=1, cmap='tab10')\n    ax.axis('off')\n    ax.set_title(title)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:35.625164Z","iopub.execute_input":"2022-07-11T14:30:35.625607Z","iopub.status.idle":"2022-07-11T14:30:35.636019Z","shell.execute_reply.started":"2022-07-11T14:30:35.625562Z","shell.execute_reply":"2022-07-11T14:30:35.634643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(3, 1, figsize=(16, 2.6))\ncompare_clusterings(ax[0], clusterings[0], clusterings[1], 'Clusterings: 0 vs 1')\ncompare_clusterings(ax[1], clusterings[0], clusterings[2], 'Clusterings: 1 vs 2')\ncompare_clusterings(ax[2], clusterings[1], clusterings[2], 'Clusterings: 0 vs 2')\nfig.subplots_adjust(hspace=0.7)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:30:35.638120Z","iopub.execute_input":"2022-07-11T14:30:35.638552Z","iopub.status.idle":"2022-07-11T14:31:08.230215Z","shell.execute_reply.started":"2022-07-11T14:30:35.638499Z","shell.execute_reply":"2022-07-11T14:31:08.228646Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, ax = plt.subplots(3, 1, figsize=(16, 2.6))\ncompare_clusterings(ax[0], clusterings[0], consensus_labels, 'Consensus vs 0')\ncompare_clusterings(ax[1], clusterings[1], consensus_labels, 'Consensus vs 1')\ncompare_clusterings(ax[2], clusterings[2], consensus_labels, 'Consensus vs 2')\nfig.subplots_adjust(hspace=0.7)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:31:08.232267Z","iopub.execute_input":"2022-07-11T14:31:08.233033Z","iopub.status.idle":"2022-07-11T14:31:40.834360Z","shell.execute_reply.started":"2022-07-11T14:31:08.232978Z","shell.execute_reply":"2022-07-11T14:31:40.833590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Use Mutual Information to compare clusterings\n\nWe can also use mutual information to compare the consensus clustering. Note that the consensus vs individual scores are higher than the individual vs individual scores.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import adjusted_mutual_info_score\n\nprint('clusters vs clusters')\nfor i in range(n_iter):\n    for j in range(i+1, n_iter):\n        print(adjusted_mutual_info_score(clusterings[i], clusterings[j]))\n        \nprint('\\nconsensus vs clusters')\nfor i in range(n_iter):\n    print(adjusted_mutual_info_score(clusterings[i], consensus_labels))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T14:31:40.835629Z","iopub.execute_input":"2022-07-11T14:31:40.836566Z","iopub.status.idle":"2022-07-11T14:31:41.408365Z","shell.execute_reply.started":"2022-07-11T14:31:40.836524Z","shell.execute_reply":"2022-07-11T14:31:41.407200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Have fun and good luck!","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}