{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# TPS JUL 22 - metrics vs LB?\n\nThere are a handful of metrics used to assess goodness-of-fit of potential cluster groupings. I wanted to see\nif any of them were at all predictive of performance on the leaderboard; in which case they could be used to \nhelp validate model performance without burning submission attempts.\n\nThis a half-baked experiment - I'm only going to burn 5 submission attempts on this. We've already seen\nthat these metrics are not definitive in selecting the optimal number of clusters. But perhaps for a given number of clusters\nthere is some correlation between one of the metrics and the LB score?\n\nSo, I'm going to:\n- fit several GaussianMixtureModels with n_components=7 and different seeds\n- ensure that these models are different from each other\n- compute several metrics for each model\n- submit the results of each model to the LB \n- compare metrics vs LB scores to see if any of them correlate\n\nI have absolutely no idea if this experiment has any real validity - happy to receive any and all feedback!\n\n__tldr__:\n- LB results are sensitive to seeds\n- I don't think any of the tested metrics are definitively predictive of LB improvement; but perhaps someone wiser than me can make something more out this!\n\n__Acknowledgements__:\n\n- GaussianMixture approach and parameters from https://www.kaggle.com/code/thedevastator/bruteforce-clustering \n- Thinking about how to apply metrics from https://www.kaggle.com/code/ambrosm/tpsjul22-gaussian-mixture-cluster-analysis\n\n","metadata":{}},{"cell_type":"code","source":"import gc\nfrom pathlib import Path\n\nimport numpy as np \nimport pandas as pd\nimport random\n\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\nsns.set_style('darkgrid')\n\nfrom sklearn.mixture import GaussianMixture, BayesianGaussianMixture\nfrom sklearn.preprocessing import RobustScaler, PowerTransformer, MinMaxScaler\nfrom sklearn.metrics import silhouette_score, calinski_harabasz_score, davies_bouldin_score\n\nimport warnings\nwarnings.simplefilter('ignore')\n\npd.set_option('display.max_columns', None)\npd.set_option('display.float_format', '{:.3f}'.format)\n\nINPUT = Path('../input/tabular-playground-series-jul-2022')\nSEED = 420\nN_COMPONENTS = 7","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","execution":{"iopub.status.busy":"2022-07-06T15:18:03.294251Z","iopub.execute_input":"2022-07-06T15:18:03.294731Z","iopub.status.idle":"2022-07-06T15:18:03.305456Z","shell.execute_reply.started":"2022-07-06T15:18:03.294696Z","shell.execute_reply":"2022-07-06T15:18:03.304249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Load data","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv(INPUT / 'data.csv',\n                   index_col='id')\ndata.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:03.312036Z","iopub.execute_input":"2022-07-06T15:18:03.313427Z","iopub.status.idle":"2022-07-06T15:18:04.701940Z","shell.execute_reply.started":"2022-07-06T15:18:03.313369Z","shell.execute_reply":"2022-07-06T15:18:04.700373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Scale and transform training data\n\nDropping some columns and scaling/transforming via `RobustScaler + PowerTransform` seems to be the current 'standard' approach.","metadata":{}},{"cell_type":"code","source":"df = data[['f_07','f_08','f_09','f_10','f_11','f_12','f_13',\n           'f_22','f_23','f_24','f_25','f_26','f_27','f_28']]\n\nX = RobustScaler().fit(df).transform(df)\nX = PowerTransformer().fit(X).transform(X)\nX = pd.DataFrame(X, columns = df.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:04.703892Z","iopub.execute_input":"2022-07-06T15:18:04.704358Z","iopub.status.idle":"2022-07-06T15:18:07.055459Z","shell.execute_reply.started":"2022-07-06T15:18:04.704321Z","shell.execute_reply":"2022-07-06T15:18:07.054361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Helper Functions","metadata":{}},{"cell_type":"code","source":"def fit_gmm(X, n_comp=N_COMPONENTS, seed=SEED):\n    \n    # only n_init once; we are looking for variance between models\n    gmm = GaussianMixture(n_components=n_comp, \n                          covariance_type='full',\n                          max_iter=200,\n                          n_init=1,\n                          random_state=seed)\n    gmm.fit(X)\n    return gmm\n    \ndef gmm_unique(old_gmms, new_gmm):\n    for gmm in old_gmms:\n        if np.allclose(sorted(gmm.weights_), sorted(new_gmm.weights_), rtol=.025):\n            return False\n    return True\n                      \ndef get_metrics(gmm, X):\n    labels = gmm.predict(X)\n    print('bic...', end=' ')\n    bic = gmm.bic(X)\n    print('silhouette...', end=' ')\n    ss = silhouette_score(X, labels, metric='euclidean')\n    print('calinski_harabasz...', end=' ')\n    chs = calinski_harabasz_score(X, labels)\n    print('davies_bouldin...')\n    dbs = davies_bouldin_score(X, labels)\n    return {'bic':bic, 'ss':ss, 'chs':chs, 'dbs':dbs}","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:07.056938Z","iopub.execute_input":"2022-07-06T15:18:07.057700Z","iopub.status.idle":"2022-07-06T15:18:07.068096Z","shell.execute_reply.started":"2022-07-06T15:18:07.057662Z","shell.execute_reply":"2022-07-06T15:18:07.067121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def make_submission(gmm, X, n_components, id):\n    labels = gmm.predict(X)\n    submission = pd.DataFrame(labels, index=data.index, columns=['Predicted'])\n    submission.to_csv(f'./submission_{n_components}_{id}.csv')\n    return submission","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:07.070728Z","iopub.execute_input":"2022-07-06T15:18:07.071566Z","iopub.status.idle":"2022-07-06T15:18:07.081531Z","shell.execute_reply.started":"2022-07-06T15:18:07.071518Z","shell.execute_reply":"2022-07-06T15:18:07.080142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fit some mixture models\n\nI use the `weights_` attribute to assess if a new model is distinguishable from retained models","metadata":{}},{"cell_type":"code","source":"gmms = {}\nseed = SEED\n\nwhile len(gmms) < 5:\n    seed = seed + 100\n    print(f'fitting seed: {seed}', end= ' - ')\n    gmm = fit_gmm(X, N_COMPONENTS, seed)\n    if gmm_unique(gmms.values(), gmm):\n        print('retaining model')\n        gmms[seed] = gmm\n    else:\n        print('bypassing model')","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:07.082659Z","iopub.execute_input":"2022-07-06T15:18:07.083101Z","iopub.status.idle":"2022-07-06T15:18:53.311960Z","shell.execute_reply.started":"2022-07-06T15:18:07.083067Z","shell.execute_reply":"2022-07-06T15:18:53.310518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compute metrics for each model","metadata":{}},{"cell_type":"code","source":"metrics = {}\nfor seed, gmm in gmms.items():\n    print(f'computing metrics {seed}:', end=' ')\n    metrics[seed]= get_metrics(gmm, X)\n    \nmetrics","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:18:53.313839Z","iopub.execute_input":"2022-07-06T15:18:53.314646Z","iopub.status.idle":"2022-07-06T15:27:15.899134Z","shell.execute_reply.started":"2022-07-06T15:18:53.314592Z","shell.execute_reply":"2022-07-06T15:27:15.897705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Save metrics and submissions together\n\nHere I manually submit the files one-by-one and record the scores in LB below\n\n__IMPORTANT__: use the `SAVE` flag to retain metrics to plot against LB scores","metadata":{}},{"cell_type":"code","source":"# use this flag to retain metrics and submission files\nSAVE=True\n\nif SAVE:\n    metrics_df = pd.DataFrame.from_dict(metrics, orient='index')\n    metrics_df.index.name = 'seed'\n    metrics_df.to_csv(\"./metrics.csv\")\n\n    for seed, gmm in gmms.items():\n        make_submission(gmm, X, N_COMPONENTS, seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:27:15.901195Z","iopub.execute_input":"2022-07-06T15:27:15.902018Z","iopub.status.idle":"2022-07-06T15:27:17.899450Z","shell.execute_reply.started":"2022-07-06T15:27:15.901965Z","shell.execute_reply":"2022-07-06T15:27:17.898432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Plot outputs\n\nI scale the metrics and LB scores and sort by LB. Are any metrics directionally correlated with the LB?","metadata":{}},{"cell_type":"code","source":"metrics_df = pd.read_csv('./metrics.csv', index_col='seed')\n\nLB = {520 : 0.50983,\n      620 : 0.50478,\n      720 : 0.52661,\n      820 : 0.58147,\n      920 : 0.57691}\n      \nLB_df = pd.DataFrame.from_dict(LB, orient='index', columns=['LB'])\nresults = metrics_df.join(LB_df)\nresults = pd.DataFrame(MinMaxScaler().fit_transform(results), index=results.index, columns=results.columns)\nresults.sort_values(by='LB', inplace=True)\nresults","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:27:17.900741Z","iopub.execute_input":"2022-07-06T15:27:17.901033Z","iopub.status.idle":"2022-07-06T15:27:17.935468Z","shell.execute_reply.started":"2022-07-06T15:27:17.901007Z","shell.execute_reply":"2022-07-06T15:27:17.934326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results.reset_index().drop(columns='seed').plot(figsize=(10,10), subplots=True, layout=(3,2))\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:27:17.936894Z","iopub.execute_input":"2022-07-06T15:27:17.937251Z","iopub.status.idle":"2022-07-06T15:27:18.806106Z","shell.execute_reply.started":"2022-07-06T15:27:17.937216Z","shell.execute_reply":"2022-07-06T15:27:18.804730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusions\n\n","metadata":{}},{"cell_type":"markdown","source":"- Note the large variance between seeds on the LB; .50478 to .58147!!\n- Are any of the metrics a sure thing in predicting LB improvement? 🤷‍♂️","metadata":{"execution":{"iopub.status.busy":"2022-07-06T15:27:18.809235Z","iopub.execute_input":"2022-07-06T15:27:18.809585Z","iopub.status.idle":"2022-07-06T15:27:18.817050Z","shell.execute_reply.started":"2022-07-06T15:27:18.809552Z","shell.execute_reply":"2022-07-06T15:27:18.815512Z"}}}]}