{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Clustering on mass**\n\nDuring my own tests for this competition I wrote a series of **functions to make life easier.** \n\nHopefully they'll help you too. I won't be doing any analysis in *this* notebook.\n\nUsually in notebooks we tend to hide the code. In this case I will leave the majority of it on show.","metadata":{}},{"cell_type":"code","source":"# Libs\nimport numpy as np \nimport pandas as pd\nimport os\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom tqdm import tqdm, trange\n\nfrom sklearn.decomposition import PCA\nfrom sklearn.mixture import GaussianMixture, BayesianGaussianMixture\nfrom sklearn.metrics import davies_bouldin_score\nfrom sklearn.preprocessing import RobustScaler, StandardScaler, MinMaxScaler, MaxAbsScaler, MinMaxScaler, QuantileTransformer, PowerTransformer\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Custom function for time logging\nfrom datetime import datetime\n\ndef time_log(f):\n    def wrapper(*args, **kwargs):\n        tick = datetime.now()\n        results = f(*args, **kwargs)\n        tock = datetime.now()\n        print(f\"{f.__name__} function took {tock-tick}\")\n        return results\n    wrapper.unwrapped = f\n    return wrapper\n\n# For initial overview\n@time_log\ndef plot_overview(df, rows: int = 8, columns:int = 5, figsize=(20, 28)):\n    \n    plt.figure(figsize=(figsize), dpi=75)\n\n    for index, col in enumerate(df.columns):\n\n        ax = plt.subplot(rows, columns, index+1) \n        sns.kdeplot(x = df[col], shade=True, color='steelblue', ec='black', alpha=0.7)\n        ax.set(xticks=([]), yticks=([]), xlabel='', ylabel='')\n        for s in ['left', 'right', 'top']:\n            ax.spines[s].set_visible(False)\n        ax.spines['bottom'].set_linewidth(2)\n        ax.set_title(f'{col}', loc='left', weight='bold', fontsize=10)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-07-11T10:34:12.471377Z","iopub.execute_input":"2022-07-11T10:34:12.471828Z","iopub.status.idle":"2022-07-11T10:34:14.064220Z","shell.execute_reply.started":"2022-07-11T10:34:12.471735Z","shell.execute_reply":"2022-07-11T10:34:14.063011Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Read in the data**","metadata":{}},{"cell_type":"code","source":"def get_data(filepath:str, fraction_percentage:float = None) -> pd.DataFrame():\n    \"\"\"\n    Read in DataFrame from file path.\n    \n    fraction_percentage: The percentage of the data you want to read in.\n    Default is None which reads in entire dataset. \n    \n    e.g. fraction_percentage = 0.1 would read in randomized 10% of the data\n    \"\"\"\n    df = (pd.read_csv(filepath)\n      .astype({'id': 'category'})\n      .set_index(['id'])\n     )\n    \n    if fraction_percentage is not None:\n        df = df.sample(frac=fraction_percentage)\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:34:23.757766Z","iopub.execute_input":"2022-07-11T10:34:23.758200Z","iopub.status.idle":"2022-07-11T10:34:23.766217Z","shell.execute_reply.started":"2022-07-11T10:34:23.758163Z","shell.execute_reply":"2022-07-11T10:34:23.764933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"filepath = '/kaggle/input/tabular-playground-series-jul-2022/data.csv'\n\ndf = get_data(filepath, fraction_percentage=0.05)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:34:23.982446Z","iopub.execute_input":"2022-07-11T10:34:23.982821Z","iopub.status.idle":"2022-07-11T10:34:24.832283Z","shell.execute_reply.started":"2022-07-11T10:34:23.982791Z","shell.execute_reply":"2022-07-11T10:34:24.830970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Quick visual overview of the data**","metadata":{}},{"cell_type":"code","source":"plot_overview(df, figsize=(20, 15), columns=10)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:34:26.370308Z","iopub.execute_input":"2022-07-11T10:34:26.370746Z","iopub.status.idle":"2022-07-11T10:34:28.448421Z","shell.execute_reply.started":"2022-07-11T10:34:26.370712Z","shell.execute_reply":"2022-07-11T10:34:28.447104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Modelling on mass**\n\nSometimes it is valuable to run many tests in order to gauge performance for a variety of parameters.\n\nThese functions enable you to do that.\n\n**Add to the models, scalers, and transformers if you want to...** (just be sure you've imported them)\n\nWe'll define the number of components to test later on.","metadata":{}},{"cell_type":"code","source":"# Models & Transformers\ndef define_models(models=dict(), n_components_list=[]):\n    random_state=1\n    for number in n_components_list:\n        models[f'GaussianMixture_n{number}'] = GaussianMixture(n_components=number,\n                                                               random_state=random_state)\n        \n        models[f'BayesianGaussianMixture_n{number}'] = BayesianGaussianMixture(n_components=number,\n                                                                               random_state=random_state)\n        \n    return models\n    \ndef define_scalers(scalers=dict()):\n    scalers['robust'] = RobustScaler()\n    scalers['standard'] = StandardScaler()\n    scalers['maxabs'] = MaxAbsScaler()\n    scalers['minmax'] = MinMaxScaler()\n    \n    return scalers\n\ndef define_transformers(transformers=dict()):\n    transformers['power'] = PowerTransformer()\n    transformers['quantile'] = QuantileTransformer()\n    \n    return transformers","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:34:32.485078Z","iopub.execute_input":"2022-07-11T10:34:32.485477Z","iopub.status.idle":"2022-07-11T10:34:32.494520Z","shell.execute_reply.started":"2022-07-11T10:34:32.485445Z","shell.execute_reply":"2022-07-11T10:34:32.493568Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Implementation & evaluation \ndef scale_data(df, scaler, transformer):\n    X_scaled = scaler.fit(df).transform(df)\n    X_scaled = transformer.fit(X_scaled).transform(X_scaled)\n    X_scaled = pd.DataFrame(X_scaled, columns = df.columns)\n    return X_scaled\n\ndef evaluate_single_model(model, df, scaler, transformer):\n    X_scaled = scale_data(df, scaler, transformer)\n    labels = model.fit_predict(X_scaled)\n    proxy_score = davies_bouldin_score(df, labels)\n    return proxy_score, labels\n\n@time_log\ndef score_models(df, models, scalers, transformers):\n    results = dict()\n    outputs = dict()\n    for model_name, model in tqdm(models.items()):\n        for scaler_name, scaler in scalers.items():\n            for transformer_name, transformer in transformers.items():\n                    results[model_name, scaler_name, transformer_name], outputs[model_name, scaler_name, transformer_name] = evaluate_single_model(model, df, scaler, transformer)\n                \n    return results, outputs\n\ndef score_summary(results):\n    scores = [(k,v) for k,v in results.items()]\n    scores = sorted(scores, key=lambda x: x[1])\n    for name, score in scores:\n        print(f'Score={score:.3f}\\t{name}')\n        \ndef get_top_n(results, outputs, n:int=5):\n    top_n = dict(list(results.items())[:n])\n    top_outputs = [outputs[o] for o in top_n]\n    return top_outputs\n","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:44:46.377046Z","iopub.execute_input":"2022-07-11T10:44:46.377440Z","iopub.status.idle":"2022-07-11T10:44:46.391744Z","shell.execute_reply.started":"2022-07-11T10:44:46.377408Z","shell.execute_reply":"2022-07-11T10:44:46.390807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualising results\n@time_log       \ndef plot_results_pca(df, outputs, rows: int = 10, columns:int = 5, figsize=(20, 28)):\n    pca = PCA(n_components=2)\n    pca_transform = pca.fit_transform(df)\n    df_pca = pd.DataFrame({\"x\" : pca_transform[:,0],\n                           \"y\" : pca_transform[:,1]}\n                         )\n    \n    for i in outputs:\n        name = i \n        labels = outputs[(name)]\n        df_pca[f'{name}'] = labels\n\n    plt.figure(figsize=(figsize), dpi=75)\n    \n    for index, col in enumerate(df_pca.columns[2:]):\n        ax = plt.subplot(rows, columns, index+1) \n        sns.scatterplot(data=df_pca,\n                    x='x', \n                    y='y', \n                    hue=df_pca[col],\n                    alpha=0.7,\n                    ec='black',\n                    palette='mako',\n                    lw=0.5, \n                    s=5,\n                    legend=False, ax=ax)\n        \n        ax.set(xticks=([]), yticks=([]), xlabel='', ylabel='')\n        for s in ['left', 'right', 'top', 'bottom']:\n            ax.spines[s].set_visible(False)\n        ax.set_title(f'{col}', loc='center', weight='semibold', fontsize=10)\n        plt.tight_layout()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:35:45.593428Z","iopub.execute_input":"2022-07-11T10:35:45.593795Z","iopub.status.idle":"2022-07-11T10:35:45.607396Z","shell.execute_reply.started":"2022-07-11T10:35:45.593766Z","shell.execute_reply":"2022-07-11T10:35:45.606397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Implement all of the above\n\n@time_log\ndef run_experiments(df, n_components_range:list = [7, 8], summarize=True):\n    \n    results, outputs = score_models(df,\n                                       define_models(n_components_list=n_components_range),\n                                       define_scalers(),\n                                       define_transformers()\n                                   )\n    \n    if summarize:\n        score_summary(results)\n    else:\n        pass\n    \n    return results, outputs","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:35:45.991790Z","iopub.execute_input":"2022-07-11T10:35:45.992363Z","iopub.status.idle":"2022-07-11T10:35:45.997949Z","shell.execute_reply.started":"2022-07-11T10:35:45.992331Z","shell.execute_reply":"2022-07-11T10:35:45.996988Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Run models**\n\nDefine the number of components/centroids.etc. you want to test here.","metadata":{}},{"cell_type":"code","source":"results, outputs = run_experiments(df, n_components_range=[7, 8, 9], summarize=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:36:17.477565Z","iopub.execute_input":"2022-07-11T10:36:17.478638Z","iopub.status.idle":"2022-07-11T10:39:50.444145Z","shell.execute_reply.started":"2022-07-11T10:36:17.478594Z","shell.execute_reply":"2022-07-11T10:39:50.442904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The results are summarized above, but of course **we can quickly view our results separately too if we want.** \n\nAlthough in this case the scores will not necessarily mean a better competition standing as it is unsupervised and therefore we cannot verify our results!\n\nThe metric output is the **Davies Bouldin score**. I find this the easiest to interpret (lower is better).\n\nFrom the docs:\n\n\"*The score is defined as the average similarity measure of each cluster with its most similar cluster, where similarity is the ratio of within-cluster distances to between-cluster distances. Thus, clusters which are farther apart and less dispersed will result in a better score.*\n\n*The minimum score is zero, with lower values indicating better clustering.*\"\n\nMore info. can be found here:\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.metrics.davies_bouldin_score.html\n\nYou could tweak this to a different metric, or even include multiple if you wanted.","metadata":{}},{"cell_type":"code","source":"score_summary(results)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:39:50.448238Z","iopub.execute_input":"2022-07-11T10:39:50.450120Z","iopub.status.idle":"2022-07-11T10:39:50.460955Z","shell.execute_reply.started":"2022-07-11T10:39:50.450065Z","shell.execute_reply":"2022-07-11T10:39:50.459743Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And of course we can **plot the clusters** with the help of **PCA** within the function.","metadata":{}},{"cell_type":"code","source":"plot_results_pca(df, outputs, rows=12, columns=4, figsize=(20, 28))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:39:50.463603Z","iopub.execute_input":"2022-07-11T10:39:50.465339Z","iopub.status.idle":"2022-07-11T10:40:00.660748Z","shell.execute_reply.started":"2022-07-11T10:39:50.465295Z","shell.execute_reply":"2022-07-11T10:40:00.656732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And of course we can access our **outputs/labels**","metadata":{}},{"cell_type":"code","source":"#outputs","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:45:07.346665Z","iopub.execute_input":"2022-07-11T10:45:07.347096Z","iopub.status.idle":"2022-07-11T10:45:07.352554Z","shell.execute_reply.started":"2022-07-11T10:45:07.347059Z","shell.execute_reply":"2022-07-11T10:45:07.351283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To get a **specific output** just pass the name tuple","metadata":{}},{"cell_type":"code","source":"outputs[('GaussianMixture_n7', 'robust', 'power')]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:40:00.687622Z","iopub.execute_input":"2022-07-11T10:40:00.688593Z","iopub.status.idle":"2022-07-11T10:40:00.696591Z","shell.execute_reply.started":"2022-07-11T10:40:00.688553Z","shell.execute_reply":"2022-07-11T10:40:00.695295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"and to get your **top n** outputs...","metadata":{}},{"cell_type":"code","source":"get_top_n(results, outputs, 5)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T10:44:29.233485Z","iopub.execute_input":"2022-07-11T10:44:29.233868Z","iopub.status.idle":"2022-07-11T10:44:29.241862Z","shell.execute_reply.started":"2022-07-11T10:44:29.233837Z","shell.execute_reply":"2022-07-11T10:44:29.241017Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"You could then make an **ensemble of the outputs**","metadata":{}},{"cell_type":"markdown","source":"**Fin.**","metadata":{}}]}