{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-10T06:01:49.312136Z","iopub.execute_input":"2022-07-10T06:01:49.313101Z","iopub.status.idle":"2022-07-10T06:01:49.347336Z","shell.execute_reply.started":"2022-07-10T06:01:49.313Z","shell.execute_reply":"2022-07-10T06:01:49.346256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is an improvement of the brute force clustering, found [here](https://www.kaggle.com/code/thedevastator/bruteforce-clustering). Please, give the man a vote, after you give me one, because he did a great job! I did some tweaks to his code and, at the moment of writing this, I am 26th overall. \n\nAlso, there is another guy which did brute force clustering [here](https://www.kaggle.com/code/plarmuseau/bruteforce-clustering/notebook?scriptVersionId=100033863). His version did slightly better than mine (he is 20th at the moment of writing this), so please give him a vote as well. My opinion is that my version is better, even though is 6 places below, because it's simpler.\n\nLet's see what is going on here:\n\n# 1. Library importing\n\nWe import the needed library, except for Pandas, which are already being imported.","metadata":{}},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.decomposition import PCA\nfrom sklearn.mixture import BayesianGaussianMixture\nfrom sklearn.preprocessing import PowerTransformer","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:49.34894Z","iopub.execute_input":"2022-07-10T06:01:49.349498Z","iopub.status.idle":"2022-07-10T06:01:51.20802Z","shell.execute_reply.started":"2022-07-10T06:01:49.349465Z","shell.execute_reply":"2022-07-10T06:01:51.206706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Style setting\n\nEverybody can play here with the default Seaborn styles. I found this to suit me the best.","metadata":{}},{"cell_type":"code","source":"sns.set_style('darkgrid')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:51.20956Z","iopub.execute_input":"2022-07-10T06:01:51.209905Z","iopub.status.idle":"2022-07-10T06:01:51.216136Z","shell.execute_reply.started":"2022-07-10T06:01:51.209874Z","shell.execute_reply":"2022-07-10T06:01:51.214873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Reading the data file \n\n...and displaying its main characteristics","metadata":{}},{"cell_type":"code","source":"data = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv')\nprint(data.head())\nprint(data.shape)\nprint(data.info())\nprint(data.describe().to_string())","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:51.219501Z","iopub.execute_input":"2022-07-10T06:01:51.22039Z","iopub.status.idle":"2022-07-10T06:01:52.743718Z","shell.execute_reply.started":"2022-07-10T06:01:51.220339Z","shell.execute_reply":"2022-07-10T06:01:52.742314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Reading the sample submission file\n\nIt will be used to build the final output file.","metadata":{}},{"cell_type":"code","source":"submission = pd.read_csv('../input/tabular-playground-series-jul-2022/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:52.74541Z","iopub.execute_input":"2022-07-10T06:01:52.745861Z","iopub.status.idle":"2022-07-10T06:01:52.775017Z","shell.execute_reply.started":"2022-07-10T06:01:52.745817Z","shell.execute_reply":"2022-07-10T06:01:52.773711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Displaying the correlation matrix:\n\n... in a graphical manner.","metadata":{}},{"cell_type":"code","source":"mask = np.triu(np.ones_like(data.corr(), dtype='bool'))\nf, ax = plt.subplots(figsize=(20, 20))\nsns.heatmap(data.corr(), mask=mask, annot=True, fmt='.2f')\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Dropping the ID column\n\nWe don't need it, right?","metadata":{}},{"cell_type":"code","source":"data = data.drop(columns='id')","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:52.776551Z","iopub.execute_input":"2022-07-10T06:01:52.777029Z","iopub.status.idle":"2022-07-10T06:01:52.789571Z","shell.execute_reply.started":"2022-07-10T06:01:52.776983Z","shell.execute_reply":"2022-07-10T06:01:52.788251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Saving only the correlated columns from the DataFrame:\n\nWe took these columns from the correlation matrix in step 5.","metadata":{}},{"cell_type":"code","source":"data = data[\n    ['f_07', 'f_08', 'f_09', 'f_10', 'f_11', 'f_12', 'f_13', 'f_22', 'f_23', 'f_24', 'f_25', 'f_26', 'f_27', 'f_28']]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. Saving the Dataframe columns\n\n... for further use.","metadata":{}},{"cell_type":"code","source":"cols = list(data.columns)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:52.791317Z","iopub.execute_input":"2022-07-10T06:01:52.791976Z","iopub.status.idle":"2022-07-10T06:01:52.800802Z","shell.execute_reply.started":"2022-07-10T06:01:52.791909Z","shell.execute_reply":"2022-07-10T06:01:52.799925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 9. Data scaling\n\nWe use the PowerTransformer class","metadata":{}},{"cell_type":"code","source":"X_scaled = PowerTransformer().fit_transform(data)\nX_scaled = pd.DataFrame(X_scaled, columns=cols)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:52.804785Z","iopub.execute_input":"2022-07-10T06:01:52.805157Z","iopub.status.idle":"2022-07-10T06:01:56.626026Z","shell.execute_reply.started":"2022-07-10T06:01:52.805121Z","shell.execute_reply":"2022-07-10T06:01:56.62458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 10. Using PCA to display the datapoints on an X-Y scale","metadata":{}},{"cell_type":"code","source":"pca = PCA(random_state=10, whiten=True)\nX_pca = pca.fit_transform(X_scaled)\nPCA_df = pd.DataFrame({'PCA_1': X_pca[:, 0], 'PCA_2': X_pca[:, 1]})\nplt.figure(figsize=(14, 14))\nsns.scatterplot(data=PCA_df, x='PCA_1', y='PCA_2', s=3)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:56.637562Z","iopub.execute_input":"2022-07-10T06:01:56.637937Z","iopub.status.idle":"2022-07-10T06:01:57.317014Z","shell.execute_reply.started":"2022-07-10T06:01:56.637895Z","shell.execute_reply":"2022-07-10T06:01:57.316047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 10. The actual brute force algorithm.\n\nInitially, this was being done by applying the *BayesianGaussianMixture* algorithm for a number of clusters between 1 and 10, according to the algorithm below:\n\n1. Applying the clustering algorithm, with the determined hyperparameters, thus obtaining the cluster number for each row;\n1. Applying PCA, similar to point 9, to display the clustered data points on an X-Y scale.\n\nThe main idea was to obtain the cluster number for each row iteratively from 1 to 10 component numbers and submitting each submission file, to get the score. By trial and error, I noticed that the optimum number of clusters is 7, thus the parameters of the *range* function.\n\nIn this improved version, I simply chose the best option: 7 clusters. The entire algorithm implementation is being found in the previous versions.","metadata":{}},{"cell_type":"code","source":"gmm = BayesianGaussianMixture(n_components=7, n_init=5, covariance_type='full')\npreds = gmm.fit_predict(X_scaled)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:01:57.31831Z","iopub.execute_input":"2022-07-10T06:01:57.318803Z","iopub.status.idle":"2022-07-10T06:07:42.22821Z","shell.execute_reply.started":"2022-07-10T06:01:57.318771Z","shell.execute_reply":"2022-07-10T06:07:42.226866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 11. Displaying the PCA:","metadata":{}},{"cell_type":"code","source":"pca = PCA(n_components=2)\nreduced_data = pca.fit_transform(X_scaled)\ndf = pd.DataFrame({'x': reduced_data[:, 0], 'y': reduced_data[:, 1], 'clusters': preds})\nplt.figure(figsize=(20, 10))\nsns.scatterplot(x=df['x'], y=df['y'], hue=df['clusters'])\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 12. Creating the submission file and submitting it\n\nIf the code got here, it exited the *for* loop above.","metadata":{}},{"cell_type":"code","source":"submission['Predicted'] = preds\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-07-10T06:07:42.229874Z","iopub.execute_input":"2022-07-10T06:07:42.230258Z","iopub.status.idle":"2022-07-10T06:07:42.392737Z","shell.execute_reply.started":"2022-07-10T06:07:42.230224Z","shell.execute_reply":"2022-07-10T06:07:42.39147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Conclusion: ** I found this setup to be the best, without changing the main philosophy behind it. I tried other algorithms and/or parameters and I found this setup to be the best.\n\nI am looking forward to your comments and questions. And votes.","metadata":{}}]}