{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Setup and Imports","metadata":{"execution":{"iopub.status.busy":"2022-07-13T13:25:16.199716Z","iopub.execute_input":"2022-07-13T13:25:16.200222Z","iopub.status.idle":"2022-07-13T13:25:16.206483Z","shell.execute_reply.started":"2022-07-13T13:25:16.200155Z","shell.execute_reply":"2022-07-13T13:25:16.204848Z"}}},{"cell_type":"code","source":"from IPython.display import clear_output\nfrom IPython.core.interactiveshell import InteractiveShell\nInteractiveShell.ast_node_interactivity = 'all'\n\nRS = 335566\n\nimport os\nfrom pathlib import Path\nimport datetime, time\nimport pickle\n\nimport pandas as pd\npd.options.display.max_columns = None\npd.options.display.max_colwidth = 999\npd.options.display.max_rows = 999\n\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.cluster import KMeans\nfrom sklearn.mixture import GaussianMixture, BayesianGaussianMixture\nfrom sklearn.metrics import calinski_harabasz_score, davies_bouldin_score, silhouette_score\n\nfrom umap import UMAP\nfrom sklearn.decomposition import PCA\nfrom sklearn.preprocessing import StandardScaler, Normalizer, PowerTransformer\n\nRS = 335577\ndata_dir = '/kaggle/input/tabular-playground-series-jul-2022/'\n# data_dir = 'data/'\n\ndf_data = pd.read_csv(f'{data_dir}data.csv', index_col='id')","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Intro","metadata":{}},{"cell_type":"markdown","source":"Hi All !!!\n\nIn this notebook I will present the results of a custom implementation of the <a href=https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html>SequentialFeatureSelector</a> from scikit-learn.\nThe actual idea behind my algorithm is the same, but I've added some helpful logging (controlled with ```verbose``` param), and the ability to resume (with use of base_features and num_features_to_select params).\n\nThe algorithm follows the below steps:\n- create a base dataframe with no features\n- score each feature individually using the ```score_func``` function\n- add the feature with the best score to base dataframe\n- repeat the process again","metadata":{}},{"cell_type":"code","source":"from collections import OrderedDict\ndef adv_sequential_feature_selection(X, y, score_func, base_features = [], num_features_to_select = -1, verbose= 0):\n    \"\"\"\n    Custom implementation of Sequential Feature Selection that allowes to select starting/base features, \n    prints the progress of finding features and allowes to pick up selecting features where you left off.\n\n    :param X: train df with features\n    :param y: train df with target\n    :param score_func: function that will be called with X, and y. Should return score in form of 1 float - this si your function with CV or test/train split\n    :param base_features: list of features to always include in scoring\n    :param num_features_to_select: number of featrues to select from (X.columns - base_features). Will start from best and finish when limit is reached\n    :param verbose: logging levels 1 - log when best featre is found, 2 - log every sequence trained\n    :return: tuple with: [1] series with best features, [2] df with train history\n    \"\"\"\n    \n    base_features = list(base_features)\n    features_to_select_from = list(set(X.columns) - set(base_features))\n    if verbose > 1:\n        print(f'features_to_select_from: {features_to_select_from}')\n        print(f'base_features: {base_features}')\n    \n    if num_features_to_select == -1:\n        num_features_to_select = len(features_to_select_from)\n    \n    scorred_features = OrderedDict()\n    pending_features_to_score = []\n    \n    df_verb_2_log = pd.DataFrame(index=features_to_select_from)\n    \n    for i in range(num_features_to_select):\n        scores = {}\n        for feature in features_to_select_from:\n            features_to_fit = base_features + list(scorred_features.keys()) + [feature]\n            returned_scores = score_func( X[features_to_fit], y)\n            score = returned_scores[0]\n\n            if verbose > 1:\n                print(f'Base: {base_features}   Selected: {list(scorred_features.keys())}')\n                print(f'Added: {feature}   Score all: {returned_scores[0]}   Score subset: {returned_scores[1]}')\n            scores[feature] = score\n        \n        best_feature = max(scores, key=scores.get)\n        best_score = scores[best_feature]\n        if verbose > 0:\n            print(f'Best feature: {best_feature}   Score: {best_score}')\n        \n        df_verb_2_log[best_feature] = pd.Series(scores)\n        \n        scorred_features[best_feature] = best_score\n        features_to_select_from.remove(best_feature)\n\n    return pd.Series(scorred_features), df_verb_2_log","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The below ```score_func``` return 2 values for ```calinski_harabasz_score```.\n\n\n- 'Score all' - returns a score of labels created on subset of features on the entire dataset. This shows how well currently selected features are representing all features.\n- 'Score subset' - returns a score of labels created on subset of features on this subset only. This shows how well those features cluster.\n\nIn early iterations, 'Score all' is low and 'Score subset' is high. \n'Score subset' is high at the beginning, as it is easier for algorithm to create good clustering with just 2-3 features, but it gets harder with 10.\n'Score all' is low, because clustering just based on 2-3 features is not representing well ","metadata":{}},{"cell_type":"markdown","source":"# GM with 4 clusters","metadata":{"execution":{"iopub.status.busy":"2022-07-13T15:48:57.485856Z","iopub.execute_input":"2022-07-13T15:48:57.486382Z","iopub.status.idle":"2022-07-13T15:48:57.493239Z","shell.execute_reply.started":"2022-07-13T15:48:57.486343Z","shell.execute_reply":"2022-07-13T15:48:57.491719Z"}}},{"cell_type":"code","source":"%%time\ndef score_func(X, y=None):\n    X_pow = PowerTransformer().fit_transform(X)\n    \n    model = GaussianMixture(random_state=RS, n_components=4)\n    lbls = model.fit_predict(X_pow)\n    return int(calinski_harabasz_score(df_, lbls)), int(calinski_harabasz_score(X_pow, lbls))\n\n# df_ = df_data.sample(100)\ndf_ = df_data.copy()\nall_results = adv_sequential_feature_selection(df_, None, score_func, verbose=100)\n\ndf_res = all_results[1]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_res.style.format(na_rep='-', precision=0).highlight_max(color = \"red\")","metadata":{"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_plt = df_res.max().reset_index(drop=True)\n\ncols = list(sns.color_palette(\"Blues\", 40))\ncols.reverse()\n\nfig, ax = plt.subplots(1, 1, figsize=(20, 15))\nbars = ax.barh(df_plt.index, df_plt.values, color=cols)\n_ = ax.bar_label(bars)\n_ = ax.set_facecolor('#EEE')\n_ = ax.set_ylabel('Number of features')\n_ = ax.set_xlabel('Calinski-Harabasz Score')\n_ = ax.set_title('How number of features affect the Calinski-Harabasz for GaussianMixture with 4 Clusters')","metadata":{"_kg_hide-input":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# GM with 7 clusters","metadata":{}},{"cell_type":"code","source":"%%time\ndef score_func(X, y=None):\n    X_pow = PowerTransformer().fit_transform(X)\n    \n    model = GaussianMixture(random_state=RS, n_components=7)\n    lbls = model.fit_predict(X_pow)\n    return int(calinski_harabasz_score(df_, lbls)), int(calinski_harabasz_score(X_pow, lbls))\n\n# df_ = df_data.sample(100)\ndf_ = df_data.copy()\nall_results = adv_sequential_feature_selection(df_, None, score_func, verbose=100)\n\ndf_res = all_results[1]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_res.style.format(na_rep='-', precision=0).highlight_max(color = \"red\")","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_plt = df_res.max().reset_index(drop=True)\n\ncols = list(sns.color_palette(\"Blues\", 40))\ncols.reverse()\n\nfig, ax = plt.subplots(1, 1, figsize=(20, 15))\nbars = ax.barh(df_plt.index, df_plt.values, color=cols)\n_ = ax.bar_label(bars)\n_ = ax.set_facecolor('#EEE')\n_ = ax.set_ylabel('Number of features')\n_ = ax.set_xlabel('Calinski-Harabasz Score')\n_ = ax.set_title('How number of features affect the Calinski-Harabasz for GaussianMixture with 7 Clusters')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Comments","metadata":{"execution":{"iopub.status.busy":"2022-07-13T15:50:20.152737Z","iopub.execute_input":"2022-07-13T15:50:20.153241Z","iopub.status.idle":"2022-07-13T15:50:20.15941Z","shell.execute_reply.started":"2022-07-13T15:50:20.153205Z","shell.execute_reply":"2022-07-13T15:50:20.158091Z"}}},{"cell_type":"markdown","source":"I've run this on combinations of data, models and params.\nIn most cases the score stops improving, in few it starts going down.","metadata":{}}]}