{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkgreen;\">Tabular Playground Series - Jul 2022</h1>  \n</div>\n\n<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">C l u s t e r i n g - 3/3</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"#### In this notebook, we Ensembling the results of two public notebooks together. We assume that we don't have access to the predict_proba() file, so we first get these files using the BayesianGMMClassifier() library and then simply perform Ensembling. This method is our own initiative.\n\n#### Thanks to: @thedevastator @cabaxiom\n\n#### https://www.kaggle.com/code/thedevastator/the-fine-art-of-fine-tuning\n\n#### https://www.kaggle.com/code/cabaxiom/tps-jul-22-bgmm-semi-supervised\n\n-----\n\n#### In the **third version**, we gave proportional weight to the results. That is, we multiply the probabilities of the weaker notebook by 0.80\n\n#### In the **fourth version**, we use the results of another notebook to improve the final results. Of course, we can always use more notebooks, provided that we choose the appropriate coefficient (weight) for each notebook.\n\n#### Thanks to: @sanaaburrows\n\n#### https://www.kaggle.com/code/sanaaburrows/tps-2022jul-voting-classifier/notebook","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"import warnings # suppress warnings\nwarnings.filterwarnings('ignore')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport gc\nimport random\n\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\n\nfrom tqdm import tqdm\nfrom scipy import stats\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport plotly.figure_factory as ff\nimport plotly.express as px\n%matplotlib inline\n!ls ../input/*","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install sklego","metadata":{"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklego.mixture import BayesianGMMClassifier\nfrom sklearn.mixture import BayesianGaussianMixture\n\nfrom sklearn.preprocessing import MinMaxScaler, PowerTransformer, StandardScaler, RobustScaler, LabelEncoder","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Load Data & Preprocessing</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"DATA = pd.read_csv('../input/tabular-playground-series-jul-2022/data.csv')\nSAMPLE = pd.read_csv('../input/tabular-playground-series-jul-2022/sample_submission.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = DATA.copy()\ndf.drop(\"id\", axis=1, inplace=True)\ncols = list(df.columns)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols_select = []\nalpha = 0.005\n\nfor col in cols:\n    _, p_value = stats.shapiro(df[col])\n    \n    if (p_value <= alpha): \n        cols_select.append(col)       \nprint(cols_select)  ","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"color:darkgreen;\">Scaling</span>","metadata":{}},{"cell_type":"code","source":"dff = DATA[cols_select]\n\ndffs = dff.copy()\ndffs = PowerTransformer().fit_transform(dffs)\ndffs = pd.DataFrame(dffs, columns=cols_select)\ndffs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Ensembling with BayesianGMMClassifier</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"sub_prime = pd.read_csv('../input/tps22jul81580/submission.csv', index_col=[0])\nsub_prime['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_prime['Predicted'] += -1\n\nsub_prime['Predicted'].value_counts().plot(kind='bar')\nsub_prime['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"support1 = pd.read_csv('../input/tps22jul81661/submission.csv', index_col=[0])\nsupport1['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"support1['Predicted'] += -1\n\nsupport1['Predicted'].value_counts().plot(kind='bar')\nsupport1['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"support2 = pd.read_csv('../input/tps22jun81232/submission.csv', index_col=[0])\nsupport2['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"support2['Predicted'] += -1\n\nsupport2['Predicted'].value_counts().plot(kind='bar')\nsupport2['Predicted'].value_counts()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"X = np.array(dffs)\ny = np.array(sub_prime)\n\ns1 = np.array(support1)\ns2 = np.array(support2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bgm = BayesianGMMClassifier(n_components=7, random_state=1234, tol=0.001, max_iter=300, n_init=3, verbose=0)\n\nbgm.fit(X,y)\nproba = bgm.predict_proba(X)\n\nbgm.fit(X,s1)\nprobs1 = bgm.predict_proba(X)\n\nbgm.fit(X,s2)\nprobs2 = bgm.predict_proba(X)\n\nproba.shape, probs1.shape, probs2.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"markdown","source":"### <span style=\"color:darkgreen;\">We multiply the probabilities in the support notebooks by appropriate coefficients.</span>","metadata":{}},{"cell_type":"code","source":"prob = np.concatenate((proba, probs1*0.99), axis=1)\nprob = np.concatenate((prob,  probs2*0.93), axis=1)\nprob.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <span style=\"color:darkgreen;\">The final result is the highest probability (related to each of the algorithms).</span>","metadata":{}},{"cell_type":"code","source":"pred = np.argmax(prob, axis=1)\npred, min(pred), max(pred)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"clusters = np.zeros(shape=(7, 7), dtype=int)\nfor n1, n2 in zip(y, s1):\n    clusters[n1, n2] += 1\n    \nclusters","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_clusters = np.argmax(clusters, axis=0)\nmax_clusters","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(len(pred)):\n    \n    if (pred[i] == 7): \n        pred[i] = max_clusters[0]\n    if (pred[i] == 8): \n        pred[i] = max_clusters[1]\n    if (pred[i] == 9): \n        pred[i] = max_clusters[2]\n    if (pred[i] == 10): \n        pred[i] = max_clusters[3]\n    if (pred[i] == 11): \n        pred[i] = max_clusters[4]\n    if (pred[i] == 12): \n        pred[i] = max_clusters[5]\n    if (pred[i] == 13): \n        pred[i] = max_clusters[6]        \n\npred, min(pred), max(pred)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#for i in range(len(pred)):   \n    #for j in range(7, 14):\n        #if (pred[i] == j):  \n            #pred[i] = max_clusters[j-7]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"clusters = np.zeros(shape=(7, 7), dtype=int)\nfor n1, n2 in zip(y, s2):\n    clusters[n1, n2] += 1\n    \nclusters","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_clusters = np.argmax(clusters, axis=0)\nmax_clusters","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(len(pred)):\n    \n    if (pred[i] == 14): \n        pred[i] = max_clusters[0]\n    if (pred[i] == 15): \n        pred[i] = max_clusters[1]\n    if (pred[i] == 16): \n        pred[i] = max_clusters[2]\n    if (pred[i] == 17): \n        pred[i] = max_clusters[3]\n    if (pred[i] == 18): \n        pred[i] = max_clusters[4]\n    if (pred[i] == 19): \n        pred[i] = max_clusters[5]\n    if (pred[i] == 20): \n        pred[i] = max_clusters[6]        \n\npred, min(pred), max(pred)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div>\n    <h1 align=\"center\" style=\"color:darkgreen;\">Submission</h1>\n</div>\n\n<div class=\"alert alert-success\">  \n</div>","metadata":{}},{"cell_type":"code","source":"sub = SAMPLE.copy()\nsub['Predicted'] = pred","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"hist_data = [sub_prime.iloc[:, 0], pred]  \ngroup_labels = ['Sub_Prime', 'Submission']\n  \nfig = ff.create_distplot(hist_data, group_labels, bin_size=.2, show_hist=False, show_rug=False) \nfig.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub.to_csv(\"submission.csv\", index=False)\n!ls","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-success\">  \n    <h1 align=\"center\" style=\"color:darkgreen;\">Good Luck</h1>  \n</div>","metadata":{}}]}