{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Telecom Churn Case Study(Machine learning -II)","metadata":{}},{"cell_type":"markdown","source":"# Problem Statement\n\n### Business problem overview\n\n. In the telecom industry, customers are able to choose from multiple service providers and actively switch from one operator to another. In this highly competitive market, the telecommunications industry experiences an average of 15-25% annual churn rate. Given the fact that it costs 5-10 times more to acquire a new customer than to retain an existing one, customer retention has now become even more important than customer acquisition.\n\n. For many incumbent operators, retaining high profitable customers is the number one business goal.\n\n. To reduce customer churn, telecom companies need to predict which customers are at high risk of churn.\n\n. In this project, we will analyse customer-level data of a leading telecom firm, build predictive models to identify customers at high risk of churn and identify the main indicators of churn.","metadata":{}},{"cell_type":"markdown","source":"### Definitions of churn\n. There are various ways to define churn, such as:\n\n### Revenue-based churn:\n. Customers who have not utilised any revenue-generating facilities such as mobile internet, outgoing calls, SMS etc. over a given period of time. One could also use aggregate metrics such as ‘customers who have generated less than INR 4 per month in total/average/median revenue’.\n\nThe main shortcoming of this definition is that there are customers who only receive calls/SMSes from their wage-earning counterparts, i.e. they don’t generate revenue but use the services. For example, many users in rural areas only receive calls from their wage-earning siblings in urban areas.\n\n### Usage-based churn:\nCustomers who have not done any usage, either incoming or outgoing - in terms of calls, internet etc. over a period of time.\n\nA potential shortcoming of this definition is that when the customer has stopped using the services for a while, it may be too late to take any corrective actions to retain them. For e.g., if you define churn based on a ‘two-months zero usage’ period, predicting churn could be useless since by that time the customer would have already switched to another operator.\n\nIn this project, we will use the usage-based definition to define churn.","metadata":{}},{"cell_type":"markdown","source":"# Objective\n- To Predict the customers who are about to churn from a telecom operator\n- Business Objective is to predict the High Value Customers only\n- We need to predict Churn on the basis of Action Period (Churn period data needs to be deleted after labelling)\n  Churn would be based on Usage\n\n### Requirement:\n\n- Churn Prediction Model\n- Best Predictor Variables","metadata":{}},{"cell_type":"markdown","source":"# Steps to Approach The  Best Solution For This Case Study\nThere are mainly 6 steps\n#### Step 1 :\n- Data reading\n- Data Understanding\n- Data Cleaning\n- Imputing missing values \n\n#### Step-2 :\nNeed to Filter high value customers\n\n#### Step-3 :\nDerive churn\n   need to Derive the Target Variable\n   \n#### Step-4 :\nData Preparation\n  - Derived variable\n  - EDA\n  - Split data in to train and test sets\n  - Performing Scaling\n \n#### Step-5 :\n- Handle class imbalance\n- Dimensionality Reduction using PCA\n- Classification models to predict Churn (Use various Models )\n\n#### Step-6 :\n- Model Evaluation\n- Prepare Model for Predictor variables selection (Prepare multiple models & choose the best one)\n\nFinally we need to give best Summarize to the company ","metadata":{}},{"cell_type":"markdown","source":"## Import  Libraries","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n# Import the logistic regression module\nfrom sklearn.linear_model import LogisticRegression\n\n# Importing 'variance_inflation_factor' or VIF\nfrom statsmodels.stats.outliers_influence import variance_inflation_factor\n\n# Import RFE for RFE selection\nfrom sklearn.feature_selection import RFE\n\n# Importing statsmodels\nimport statsmodels.api as sm\n\n# Importing the precision recall curve\nfrom sklearn.metrics import precision_recall_curve\n\n# Importing evaluation metrics from scikitlearn \nfrom sklearn import metrics\n\nfrom imblearn.over_sampling import SMOTE\n\nfrom sklearn.decomposition import IncrementalPCA\n\n# To suppress the warnings which will be raised\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Displaying all Columns without restrictions\npd.set_option('display.max_columns', None)\npd.set_option('display.max_rows', None)\npd.set_option('display.max_colwidth', -1)\n\n\n# import required libraries\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.decomposition import PCA\nfrom sklearn.preprocessing import MinMaxScaler\nfrom sklearn.pipeline import FeatureUnion\nfrom sklearn.base import BaseEstimator, TransformerMixin\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import roc_auc_score\nfrom imblearn.metrics import sensitivity_specificity_support\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom sklearn.svm import SVC","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-09T14:29:07.036756Z","iopub.execute_input":"2022-08-09T14:29:07.037252Z","iopub.status.idle":"2022-08-09T14:29:07.049359Z","shell.execute_reply.started":"2022-08-09T14:29:07.037213Z","shell.execute_reply":"2022-08-09T14:29:07.048417Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read data\nchurn = pd.read_csv(\"../input/telecom-churn-case-study-hackathon-38/train (1).csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:07.062980Z","iopub.execute_input":"2022-08-09T14:29:07.063588Z","iopub.status.idle":"2022-08-09T14:29:08.281307Z","shell.execute_reply.started":"2022-08-09T14:29:07.063552Z","shell.execute_reply":"2022-08-09T14:29:08.279986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:08.283727Z","iopub.execute_input":"2022-08-09T14:29:08.284173Z","iopub.status.idle":"2022-08-09T14:29:08.442131Z","shell.execute_reply.started":"2022-08-09T14:29:08.284131Z","shell.execute_reply":"2022-08-09T14:29:08.441070Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# create backup of data\noriginal = churn.copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:08.443237Z","iopub.execute_input":"2022-08-09T14:29:08.443526Z","iopub.status.idle":"2022-08-09T14:29:08.485548Z","shell.execute_reply.started":"2022-08-09T14:29:08.443499Z","shell.execute_reply":"2022-08-09T14:29:08.484643Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#look at the last 5 rows\nchurn.tail() ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:08.487809Z","iopub.execute_input":"2022-08-09T14:29:08.488242Z","iopub.status.idle":"2022-08-09T14:29:08.634831Z","shell.execute_reply.started":"2022-08-09T14:29:08.488211Z","shell.execute_reply":"2022-08-09T14:29:08.633592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the columns of data\nchurn.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:08.636917Z","iopub.execute_input":"2022-08-09T14:29:08.637351Z","iopub.status.idle":"2022-08-09T14:29:08.644191Z","shell.execute_reply.started":"2022-08-09T14:29:08.637307Z","shell.execute_reply":"2022-08-09T14:29:08.643385Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Checking the numerical columns data distribution statistics\nchurn.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:08.645519Z","iopub.execute_input":"2022-08-09T14:29:08.645941Z","iopub.status.idle":"2022-08-09T14:29:09.574446Z","shell.execute_reply.started":"2022-08-09T14:29:08.645906Z","shell.execute_reply":"2022-08-09T14:29:09.573676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check dataframe for null and datatype \nchurn.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.575470Z","iopub.execute_input":"2022-08-09T14:29:09.575784Z","iopub.status.idle":"2022-08-09T14:29:09.594933Z","shell.execute_reply.started":"2022-08-09T14:29:09.575750Z","shell.execute_reply":"2022-08-09T14:29:09.593570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# feature type summary\nchurn.info(verbose=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.596509Z","iopub.execute_input":"2022-08-09T14:29:09.596853Z","iopub.status.idle":"2022-08-09T14:29:09.616121Z","shell.execute_reply.started":"2022-08-09T14:29:09.596809Z","shell.execute_reply":"2022-08-09T14:29:09.615219Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking for null values\nchurn.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.617024Z","iopub.execute_input":"2022-08-09T14:29:09.617321Z","iopub.status.idle":"2022-08-09T14:29:09.680851Z","shell.execute_reply.started":"2022-08-09T14:29:09.617293Z","shell.execute_reply":"2022-08-09T14:29:09.679626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the null value percentage\nchurn.isna().sum()/churn.isna().count()*100","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.685691Z","iopub.execute_input":"2022-08-09T14:29:09.686155Z","iopub.status.idle":"2022-08-09T14:29:09.794429Z","shell.execute_reply.started":"2022-08-09T14:29:09.686122Z","shell.execute_reply":"2022-08-09T14:29:09.793393Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking for shape of a data set\nchurn.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.795738Z","iopub.execute_input":"2022-08-09T14:29:09.796718Z","iopub.status.idle":"2022-08-09T14:29:09.801559Z","shell.execute_reply.started":"2022-08-09T14:29:09.796682Z","shell.execute_reply":"2022-08-09T14:29:09.800887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking for the duplicates\nchurn.drop_duplicates(subset=None, inplace=True)\nchurn.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:09.802777Z","iopub.execute_input":"2022-08-09T14:29:09.803329Z","iopub.status.idle":"2022-08-09T14:29:10.275276Z","shell.execute_reply.started":"2022-08-09T14:29:09.803297Z","shell.execute_reply":"2022-08-09T14:29:10.274431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the size of data\nchurn.size","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.276892Z","iopub.execute_input":"2022-08-09T14:29:10.277420Z","iopub.status.idle":"2022-08-09T14:29:10.283105Z","shell.execute_reply.started":"2022-08-09T14:29:10.277388Z","shell.execute_reply":"2022-08-09T14:29:10.282341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the axes of data\nchurn.axes","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.284299Z","iopub.execute_input":"2022-08-09T14:29:10.285070Z","iopub.status.idle":"2022-08-09T14:29:10.299200Z","shell.execute_reply.started":"2022-08-09T14:29:10.285038Z","shell.execute_reply":"2022-08-09T14:29:10.298248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the dimensions of data\nchurn.ndim","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.300797Z","iopub.execute_input":"2022-08-09T14:29:10.301478Z","iopub.status.idle":"2022-08-09T14:29:10.309866Z","shell.execute_reply.started":"2022-08-09T14:29:10.301426Z","shell.execute_reply":"2022-08-09T14:29:10.309163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#check the values of data\nchurn.values","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.311438Z","iopub.execute_input":"2022-08-09T14:29:10.312115Z","iopub.status.idle":"2022-08-09T14:29:10.961504Z","shell.execute_reply.started":"2022-08-09T14:29:10.312073Z","shell.execute_reply":"2022-08-09T14:29:10.960127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#list of columns\npd.DataFrame(churn.columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.962667Z","iopub.execute_input":"2022-08-09T14:29:10.962990Z","iopub.status.idle":"2022-08-09T14:29:10.979928Z","shell.execute_reply.started":"2022-08-09T14:29:10.962962Z","shell.execute_reply":"2022-08-09T14:29:10.978716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# look at missing value ratio in each column\nchurn.isnull().sum()*100/churn.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:10.981155Z","iopub.execute_input":"2022-08-09T14:29:10.981501Z","iopub.status.idle":"2022-08-09T14:29:11.043648Z","shell.execute_reply.started":"2022-08-09T14:29:10.981471Z","shell.execute_reply":"2022-08-09T14:29:11.042892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# some recharge columns have minimum value of 1 while some don't have\nrecharge_cols = ['total_rech_data_6', 'total_rech_data_7', 'total_rech_data_8', \n                 'count_rech_2g_6', 'count_rech_2g_7', 'count_rech_2g_8', \n                 'count_rech_3g_6', 'count_rech_3g_7', 'count_rech_3g_8', \n                 'max_rech_data_6', 'max_rech_data_7', 'max_rech_data_8', \n                 'av_rech_amt_data_6', 'av_rech_amt_data_7', 'av_rech_amt_data_8', \n                 ]\n\nchurn[recharge_cols].describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.044953Z","iopub.execute_input":"2022-08-09T14:29:11.045485Z","iopub.status.idle":"2022-08-09T14:29:11.156075Z","shell.execute_reply.started":"2022-08-09T14:29:11.045454Z","shell.execute_reply":"2022-08-09T14:29:11.155218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" We can create new feature as total_rech_amt_data using total_rech_data and av_rech_amt_data to capture amount utilized by customer for data.\n\n Also as the minimum value is 1 we can impute the NA values by 0, Considering there were no recharges done by the customer.","metadata":{}},{"cell_type":"code","source":"# It is also observed that the recharge date and the recharge value are missing together which means the customer didn't recharge\nchurn.loc[churn.total_rech_data_6.isnull() & churn.date_of_last_rech_data_6.isnull(), [\"total_rech_data_6\", \"date_of_last_rech_data_6\"]].head(20)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.157540Z","iopub.execute_input":"2022-08-09T14:29:11.157830Z","iopub.status.idle":"2022-08-09T14:29:11.220166Z","shell.execute_reply.started":"2022-08-09T14:29:11.157803Z","shell.execute_reply":"2022-08-09T14:29:11.219437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the recharge variables where minumum value is 1, we can impute missing values with zeroes since it means customer didn't recharge their numbers that month.","metadata":{}},{"cell_type":"code","source":"# create a list of recharge columns where we will impute missing values with zeroes\nzero_impute = ['total_rech_data_6', 'total_rech_data_7', 'total_rech_data_8', \n        'av_rech_amt_data_6', 'av_rech_amt_data_7', 'av_rech_amt_data_8', \n        'max_rech_data_6', 'max_rech_data_7', 'max_rech_data_8'\n       ]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.221196Z","iopub.execute_input":"2022-08-09T14:29:11.222083Z","iopub.status.idle":"2022-08-09T14:29:11.227138Z","shell.execute_reply.started":"2022-08-09T14:29:11.222040Z","shell.execute_reply":"2022-08-09T14:29:11.226074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# impute missing values with 0\nchurn[zero_impute] = churn[zero_impute].apply(lambda x: x.fillna(0))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.228708Z","iopub.execute_input":"2022-08-09T14:29:11.229609Z","iopub.status.idle":"2022-08-09T14:29:11.252419Z","shell.execute_reply.started":"2022-08-09T14:29:11.229578Z","shell.execute_reply":"2022-08-09T14:29:11.251186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now we have to  make sure the values are imputed correctly for that we can check \"Missing value ratio\"\nchurn[zero_impute].isnull().sum()*100/churn.shape[1]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.253953Z","iopub.execute_input":"2022-08-09T14:29:11.254997Z","iopub.status.idle":"2022-08-09T14:29:11.266724Z","shell.execute_reply.started":"2022-08-09T14:29:11.254960Z","shell.execute_reply":"2022-08-09T14:29:11.265717Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now we can check the \"statistics Summary\"\nchurn[zero_impute].describe(include='all')","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.267980Z","iopub.execute_input":"2022-08-09T14:29:11.268487Z","iopub.status.idle":"2022-08-09T14:29:11.323766Z","shell.execute_reply.started":"2022-08-09T14:29:11.268455Z","shell.execute_reply":"2022-08-09T14:29:11.322622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now we can create some column name list by there types using description of columns\nid_cols = ['id', 'circle_id']\n\ndate_cols = ['last_date_of_month_6',\n             'last_date_of_month_7',\n             'last_date_of_month_8',             \n             'date_of_last_rech_6',\n             'date_of_last_rech_7',\n             'date_of_last_rech_8',             \n             'date_of_last_rech_data_6',\n             'date_of_last_rech_data_7',\n             'date_of_last_rech_data_8'             \n            ]\n\ncat_cols =  ['night_pck_user_6',\n             'night_pck_user_7',\n             'night_pck_user_8',             \n             'fb_user_6',\n             'fb_user_7',\n             'fb_user_8'             \n            ]\n\nnum_cols = [column for column in churn.columns if column not in id_cols + date_cols + cat_cols]\n\n# print the number of columns in each list\nprint(\"#ID cols: %d\\n#Date cols:%d\\n#Numeric cols:%d\\n#Category cols:%d\" % (len(id_cols), len(date_cols), len(num_cols), len(cat_cols)))\n\n# check if we have missed any column or not\nprint(len(id_cols) + len(date_cols) + len(num_cols) + len(cat_cols) == churn.shape[1])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.325383Z","iopub.execute_input":"2022-08-09T14:29:11.326077Z","iopub.status.idle":"2022-08-09T14:29:11.335559Z","shell.execute_reply.started":"2022-08-09T14:29:11.326033Z","shell.execute_reply":"2022-08-09T14:29:11.334355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# drop id and date columns\nchurn = churn.drop(id_cols + date_cols, axis=1)\n#check the shape again\nchurn.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.338876Z","iopub.execute_input":"2022-08-09T14:29:11.339505Z","iopub.status.idle":"2022-08-09T14:29:11.377957Z","shell.execute_reply.started":"2022-08-09T14:29:11.339475Z","shell.execute_reply":"2022-08-09T14:29:11.376706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# replace missing values with '-1' in categorical columns\nchurn[cat_cols] = churn[cat_cols].apply(lambda x: x.fillna(-1))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.380009Z","iopub.execute_input":"2022-08-09T14:29:11.380632Z","iopub.status.idle":"2022-08-09T14:29:11.396423Z","shell.execute_reply.started":"2022-08-09T14:29:11.380518Z","shell.execute_reply":"2022-08-09T14:29:11.395594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# missing value ratio\nchurn[cat_cols].isnull().sum()*100/churn.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.403841Z","iopub.execute_input":"2022-08-09T14:29:11.404381Z","iopub.status.idle":"2022-08-09T14:29:11.414101Z","shell.execute_reply.started":"2022-08-09T14:29:11.404339Z","shell.execute_reply":"2022-08-09T14:29:11.413378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Droping variables with more than 70% of missing values (we can call it as threshold )","metadata":{}},{"cell_type":"code","source":"initial_cols = churn.shape[1]\n\nMISSING_THRESHOLD = 0.7\n\ninclude_cols = list(churn.apply(lambda column: True if column.isnull().sum()/churn.shape[0] < MISSING_THRESHOLD else False))\n\ndrop_missing = pd.DataFrame({'features':churn.columns , 'include': include_cols})\ndrop_missing.loc[drop_missing.include == True,:]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.415382Z","iopub.execute_input":"2022-08-09T14:29:11.415703Z","iopub.status.idle":"2022-08-09T14:29:11.491231Z","shell.execute_reply.started":"2022-08-09T14:29:11.415676Z","shell.execute_reply":"2022-08-09T14:29:11.490425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# now we can drop  some more columns\nchurn = churn.loc[:, include_cols]\n\ndropped_cols = churn.shape[1] - initial_cols\ndropped_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.492390Z","iopub.execute_input":"2022-08-09T14:29:11.492857Z","iopub.status.idle":"2022-08-09T14:29:11.523947Z","shell.execute_reply.started":"2022-08-09T14:29:11.492815Z","shell.execute_reply":"2022-08-09T14:29:11.523152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#rechecking the shape of a dataframe\nchurn.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.525019Z","iopub.execute_input":"2022-08-09T14:29:11.525731Z","iopub.status.idle":"2022-08-09T14:29:11.536267Z","shell.execute_reply.started":"2022-08-09T14:29:11.525700Z","shell.execute_reply":"2022-08-09T14:29:11.535491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# rechecking the missing values for how many missing values has left\nchurn.isnull().sum()*100/churn.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.537469Z","iopub.execute_input":"2022-08-09T14:29:11.537991Z","iopub.status.idle":"2022-08-09T14:29:11.572174Z","shell.execute_reply.started":"2022-08-09T14:29:11.537960Z","shell.execute_reply":"2022-08-09T14:29:11.571150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_cols = [column for column in churn.columns if column not in id_cols + date_cols + cat_cols]\nnum_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.573410Z","iopub.execute_input":"2022-08-09T14:29:11.573707Z","iopub.status.idle":"2022-08-09T14:29:11.582527Z","shell.execute_reply.started":"2022-08-09T14:29:11.573679Z","shell.execute_reply":"2022-08-09T14:29:11.581626Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#imputing with meadian for num_cols\nchurn[num_cols] = churn[num_cols].apply(lambda x: x.fillna(x.median()))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.583933Z","iopub.execute_input":"2022-08-09T14:29:11.584220Z","iopub.status.idle":"2022-08-09T14:29:11.964291Z","shell.execute_reply.started":"2022-08-09T14:29:11.584193Z","shell.execute_reply":"2022-08-09T14:29:11.963217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#again checking for the missing values\nchurn.isnull().sum()*100/churn.shape[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:11.965968Z","iopub.execute_input":"2022-08-09T14:29:11.966352Z","iopub.status.idle":"2022-08-09T14:29:12.000367Z","shell.execute_reply.started":"2022-08-09T14:29:11.966319Z","shell.execute_reply":"2022-08-09T14:29:11.999175Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In churn prediction, we assume that there are three phases of customer lifecycle :\n\n- The ‘good & action’ phase [Month 6 & 7]\n- The ‘churn’ phase [Month 8]\nIn this case, since we are working over a three-month window, the first two months are the ‘good & action’ phase, the third month is the ‘churn’ phase.","metadata":{}},{"cell_type":"markdown","source":"# Step 2:\n\n# Filter high-value customers","metadata":{}},{"cell_type":"markdown","source":"Here we can take good phase ( it means month 6 and 7) data to get high value customers","metadata":{}},{"cell_type":"code","source":"# calculate the total data recharge amount for June and July --> number of recharges * average recharge amount\nchurn['total_data_rech_6'] = churn.total_rech_data_6 * churn.av_rech_amt_data_6\nchurn['total_data_rech_7'] = churn.total_rech_data_7 * churn.av_rech_amt_data_7","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.002150Z","iopub.execute_input":"2022-08-09T14:29:12.002477Z","iopub.status.idle":"2022-08-09T14:29:12.009995Z","shell.execute_reply.started":"2022-08-09T14:29:12.002446Z","shell.execute_reply":"2022-08-09T14:29:12.009189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"add total data recharge and total recharge to get total combined recharge amount for a month","metadata":{}},{"cell_type":"code","source":"# calculate total recharge amount for June and July --> call recharge amount + data recharge amount\nchurn['amt_data_6'] = churn.total_rech_amt_6 + churn.total_data_rech_6\nchurn['amt_data_7'] = churn.total_rech_amt_7 + churn.total_data_rech_7","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.011201Z","iopub.execute_input":"2022-08-09T14:29:12.011671Z","iopub.status.idle":"2022-08-09T14:29:12.022524Z","shell.execute_reply.started":"2022-08-09T14:29:12.011643Z","shell.execute_reply":"2022-08-09T14:29:12.021751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# calculate average recharge done by customer in June and July\nchurn['av_amt_data_6_7'] = (churn.amt_data_6 + churn.amt_data_7)/2","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.023674Z","iopub.execute_input":"2022-08-09T14:29:12.024255Z","iopub.status.idle":"2022-08-09T14:29:12.032785Z","shell.execute_reply.started":"2022-08-09T14:29:12.024218Z","shell.execute_reply":"2022-08-09T14:29:12.031855Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# look at the 70th percentile recharge amount\nprint(\"Recharge amount at 70th percentile: {0}\".format(churn.av_amt_data_6_7.quantile(0.7)))\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.033775Z","iopub.execute_input":"2022-08-09T14:29:12.034549Z","iopub.status.idle":"2022-08-09T14:29:12.046362Z","shell.execute_reply.started":"2022-08-09T14:29:12.034520Z","shell.execute_reply":"2022-08-09T14:29:12.045341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.047830Z","iopub.execute_input":"2022-08-09T14:29:12.048574Z","iopub.status.idle":"2022-08-09T14:29:12.180367Z","shell.execute_reply.started":"2022-08-09T14:29:12.048534Z","shell.execute_reply":"2022-08-09T14:29:12.179261Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# retain only those customers who have recharged their mobiles with more than or equal to 70th percentile amount\nchurn_filtered = churn.loc[churn.av_amt_data_6_7 >= churn.av_amt_data_6_7.quantile(0.7), :]\nchurn_filtered = churn_filtered.reset_index(drop=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.181997Z","iopub.execute_input":"2022-08-09T14:29:12.182626Z","iopub.status.idle":"2022-08-09T14:29:12.260555Z","shell.execute_reply.started":"2022-08-09T14:29:12.182584Z","shell.execute_reply":"2022-08-09T14:29:12.259303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.262068Z","iopub.execute_input":"2022-08-09T14:29:12.262794Z","iopub.status.idle":"2022-08-09T14:29:12.268637Z","shell.execute_reply.started":"2022-08-09T14:29:12.262758Z","shell.execute_reply":"2022-08-09T14:29:12.267724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# delete variables created to filter high-value customers\nchurn_filtered = churn_filtered.drop(['total_data_rech_6', 'total_data_rech_7',\n                                      'amt_data_6', 'amt_data_7', 'av_amt_data_6_7'], axis=1)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.269792Z","iopub.execute_input":"2022-08-09T14:29:12.270669Z","iopub.status.idle":"2022-08-09T14:29:12.285220Z","shell.execute_reply.started":"2022-08-09T14:29:12.270637Z","shell.execute_reply":"2022-08-09T14:29:12.284397Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.286558Z","iopub.execute_input":"2022-08-09T14:29:12.287447Z","iopub.status.idle":"2022-08-09T14:29:12.294304Z","shell.execute_reply.started":"2022-08-09T14:29:12.287417Z","shell.execute_reply":"2022-08-09T14:29:12.293601Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" hear we're left with 21,013 rows  and 149 columns after selecting the customers who have provided recharge value of more than or equal to the recharge value of the 70th percentile customer.","metadata":{}},{"cell_type":"markdown","source":"# Step 3:\n\n# Derive churn\n\nDerive churn means hear we are using 8 month(The ‘churn’ phase) data , To get the target variable(In this case stydy they did not provide any target variable we have to derive it from churn phase data)\nFor that, we need to find the derive churn variable using total_ic_mou_8,total_og_mou_8,vol_2g_mb_8 and vol_3g_mb_8 attributes","metadata":{}},{"cell_type":"code","source":"# Selecting the columns to define churn variable (i.e. TARGET Variable)\nchurn_col=['total_ic_mou_8','total_og_mou_8','vol_2g_mb_8','vol_3g_mb_8']\nchurn_filtered[churn_col].info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.295290Z","iopub.execute_input":"2022-08-09T14:29:12.296042Z","iopub.status.idle":"2022-08-09T14:29:12.312487Z","shell.execute_reply.started":"2022-08-09T14:29:12.296008Z","shell.execute_reply":"2022-08-09T14:29:12.311635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lets find out churn/non churn percentage\nprint((churn_filtered['churn_probability'].value_counts()/len(churn))*100)\n((churn_filtered['churn_probability'].value_counts()/len(churn))*100).plot(kind=\"pie\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.313594Z","iopub.execute_input":"2022-08-09T14:29:12.314174Z","iopub.status.idle":"2022-08-09T14:29:12.405274Z","shell.execute_reply.started":"2022-08-09T14:29:12.314144Z","shell.execute_reply":"2022-08-09T14:29:12.403918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### ***As we can see that 90% of the customers do not churn, there is a possibility of class imbalance*** \nSince this variable churn is the target variable, all the columns relating to this variable(i.e. all columns with suffix _8) can be dropped forn the dataset.\n","metadata":{}},{"cell_type":"markdown","source":"We can still clean the data by few possible columns relating to the good phase.\n\nAs we derived few columns in the good phase earlier, we can drop those related columns during creation.","metadata":{}},{"cell_type":"code","source":"#churn['total_rech_amt_data_6']=churn['av_rech_amt_data_6'] * churn['total_rech_data_6']\n# churn['total_rech_amt_data_7']=churn['av_rech_amt_data_7'] * churn['total_rech_data_7']\n\n# # Calculating the overall recharge amount for the months 6,7 and 8\n\n# churn['overall_rech_amt_6'] = churn['total_rech_amt_data_6'] + churn['total_rech_amt_6']\n# churn['overall_rech_amt_7'] = churn['total_rech_amt_data_7'] + churn['total_rech_amt_7']\n\nchurn_filtered.drop(['av_rech_amt_data_6',\n                   'total_rech_data_6','total_rech_amt_6',\n                  'av_rech_amt_data_7',\n                   'total_rech_data_7','total_rech_amt_7'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.407878Z","iopub.execute_input":"2022-08-09T14:29:12.408815Z","iopub.status.idle":"2022-08-09T14:29:12.425757Z","shell.execute_reply.started":"2022-08-09T14:29:12.408754Z","shell.execute_reply":"2022-08-09T14:29:12.424264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can also create new columns for the defining the good phase variables and drop the seperate 6th and 7 month variables.\n\nBefore proceding to check the remaining missing value handling, let us check the collineartity of the indepedent variables and try to understand their dependencies.","metadata":{}},{"cell_type":"code","source":"# creating a list of column names for each month\nmon_6_cols = [col for col in churn_filtered.columns if '_6' in col]\nmon_7_cols = [col for col in churn_filtered.columns if '_7' in col]\nmon_8_cols = [col for col in churn_filtered.columns if '_8' in col]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.427660Z","iopub.execute_input":"2022-08-09T14:29:12.428360Z","iopub.status.idle":"2022-08-09T14:29:12.439172Z","shell.execute_reply.started":"2022-08-09T14:29:12.428318Z","shell.execute_reply":"2022-08-09T14:29:12.437598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mon_7_cols","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.441903Z","iopub.execute_input":"2022-08-09T14:29:12.442900Z","iopub.status.idle":"2022-08-09T14:29:12.454401Z","shell.execute_reply.started":"2022-08-09T14:29:12.442794Z","shell.execute_reply":"2022-08-09T14:29:12.452935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# lets check the correlation amongst the independent variables, drop the highly correlated ones\nchurn_corr = churn_filtered.corr()\nchurn_corr.loc[:,:] = np.tril(churn_corr, k=-1)\nchurn_corr = churn_corr.stack()\nchurn_corr\nchurn_corr[(churn_corr > 0.80) | (churn_corr < -0.80)].sort_values(ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:12.457281Z","iopub.execute_input":"2022-08-09T14:29:12.458426Z","iopub.status.idle":"2022-08-09T14:29:13.594134Z","shell.execute_reply.started":"2022-08-09T14:29:12.458361Z","shell.execute_reply":"2022-08-09T14:29:13.593354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"col_to_drop=['fb_user_6','fb_user_7','total_ic_mou_6','total_ic_mou_7',               \n               'std_og_t2t_mou_7','std_og_t2t_mou_6' ,'std_og_t2m_mou_7','std_ic_mou_7',]\n\n# These columns can be dropped as they are highly collinered with other predictor variables.\n# criteria set is for collinearity of 85%\n\n#  dropping these column\nchurn_filtered.drop(col_to_drop, axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:13.595666Z","iopub.execute_input":"2022-08-09T14:29:13.595979Z","iopub.status.idle":"2022-08-09T14:29:13.607529Z","shell.execute_reply.started":"2022-08-09T14:29:13.595950Z","shell.execute_reply":"2022-08-09T14:29:13.606539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The curent dimension of the dataset after dropping few unwanted columns\nchurn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:13.609206Z","iopub.execute_input":"2022-08-09T14:29:13.609612Z","iopub.status.idle":"2022-08-09T14:29:13.622467Z","shell.execute_reply.started":"2022-08-09T14:29:13.609571Z","shell.execute_reply":"2022-08-09T14:29:13.621575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Step 4:\n# Data preparation\n\n# i.Deriving new variables to understand the data \n\n# ii.EDA","metadata":{}},{"cell_type":"code","source":"# We have a column called 'aon'\n\n# we can derive new variables from this to explain the data w.r.t churn.\n\n# creating a new variable 'tenure'\nchurn_filtered['tenure'] = (churn_filtered['aon']/30).round(0)\n\n# Since we derived a new column from 'aon', we can drop it\nchurn_filtered.drop('aon',axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:13.623747Z","iopub.execute_input":"2022-08-09T14:29:13.624773Z","iopub.status.idle":"2022-08-09T14:29:13.647357Z","shell.execute_reply.started":"2022-08-09T14:29:13.624739Z","shell.execute_reply":"2022-08-09T14:29:13.646410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the distribution of he tenure variable\n\nsns.distplot(churn_filtered['tenure'],bins=30)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:13.648522Z","iopub.execute_input":"2022-08-09T14:29:13.648791Z","iopub.status.idle":"2022-08-09T14:29:14.014055Z","shell.execute_reply.started":"2022-08-09T14:29:13.648766Z","shell.execute_reply":"2022-08-09T14:29:14.013221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tn_range = [0, 6, 12, 24, 60, 61]\ntn_label = [ '0-6 Months', '6-12 Months', '1-2 Yrs', '2-5 Yrs', '5 Yrs and above']\nchurn_filtered['tenure_range'] = pd.cut(churn_filtered['tenure'], tn_range, labels=tn_label)\nchurn_filtered['tenure_range'].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.016140Z","iopub.execute_input":"2022-08-09T14:29:14.016576Z","iopub.status.idle":"2022-08-09T14:29:14.029993Z","shell.execute_reply.started":"2022-08-09T14:29:14.016533Z","shell.execute_reply":"2022-08-09T14:29:14.028904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting a bar plot for tenure range\nplt.figure(figsize=[12,7])\nsns.barplot(x='tenure_range',y='churn_probability', data=churn_filtered)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.031246Z","iopub.execute_input":"2022-08-09T14:29:14.031667Z","iopub.status.idle":"2022-08-09T14:29:14.526561Z","shell.execute_reply.started":"2022-08-09T14:29:14.031626Z","shell.execute_reply":"2022-08-09T14:29:14.525406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It can be seen that the maximum churn rate happens within 0-6 month, but it gradually decreases as the customer retains in the network.\n\nThe average revenue per user is good phase of customer is given by arpu_6 and arpu_7. since we have two separate averages, lets take an average to these two and drop the other columns","metadata":{}},{"cell_type":"code","source":"churn_filtered[\"avg_arpu_6_7\"]= (churn_filtered['arpu_6']+churn_filtered['arpu_7'])/2\nchurn_filtered['avg_arpu_6_7'].head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.527950Z","iopub.execute_input":"2022-08-09T14:29:14.528270Z","iopub.status.idle":"2022-08-09T14:29:14.537907Z","shell.execute_reply.started":"2022-08-09T14:29:14.528240Z","shell.execute_reply":"2022-08-09T14:29:14.536788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lets drop the original columns as they are derived to a new column for better understanding of the data\n\nchurn_filtered.drop(['arpu_6','arpu_7'], axis=1, inplace=True)\n\n\n# The curent dimension of the dataset after dropping few unwanted columns\nchurn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.539080Z","iopub.execute_input":"2022-08-09T14:29:14.539372Z","iopub.status.idle":"2022-08-09T14:29:14.562927Z","shell.execute_reply.started":"2022-08-09T14:29:14.539339Z","shell.execute_reply":"2022-08-09T14:29:14.561724Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualizing the column created\nsns.distplot(churn_filtered['avg_arpu_6_7'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.564032Z","iopub.execute_input":"2022-08-09T14:29:14.564323Z","iopub.status.idle":"2022-08-09T14:29:14.993617Z","shell.execute_reply.started":"2022-08-09T14:29:14.564295Z","shell.execute_reply":"2022-08-09T14:29:14.992708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking Correlation between target variable(SalePrice) with the other variable in the dataset\nplt.figure(figsize=(10,50))\nheatmap_churn = sns.heatmap(churn_filtered.corr()[['churn_probability']].sort_values(ascending=False, by='churn_probability'),annot=True, \n                                cmap='summer')\nheatmap_churn.set_title(\"Features Correlating with Churn variable\", fontsize=15)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:14.994767Z","iopub.execute_input":"2022-08-09T14:29:14.995063Z","iopub.status.idle":"2022-08-09T14:29:18.863860Z","shell.execute_reply.started":"2022-08-09T14:29:14.995036Z","shell.execute_reply":"2022-08-09T14:29:18.862753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.columns","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:18.865049Z","iopub.execute_input":"2022-08-09T14:29:18.865361Z","iopub.status.idle":"2022-08-09T14:29:18.873132Z","shell.execute_reply.started":"2022-08-09T14:29:18.865334Z","shell.execute_reply":"2022-08-09T14:29:18.872281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Avg Outgoing Calls & calls on roaming for 6th & 7th months are positively correlated with churn.\n- Avg Revenue, No. of Recharge for 8th month has negative correlation with churn.","metadata":{}},{"cell_type":"code","source":"# lets now draw a scatter plot between total recharge and avg revenue for the 8th month\nchurn_filtered[['total_rech_num_8', 'arpu_8']].plot.scatter(x = 'total_rech_num_8',\n                                                              y='arpu_8')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:18.874894Z","iopub.execute_input":"2022-08-09T14:29:18.875201Z","iopub.status.idle":"2022-08-09T14:29:19.064318Z","shell.execute_reply.started":"2022-08-09T14:29:18.875173Z","shell.execute_reply":"2022-08-09T14:29:19.063254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating categories for month 8 column totalrecharge and their count\nchurn_filtered['total_rech_data_group_8']=pd.cut(churn_filtered['total_rech_data_8'],[-1,0,10,25,100],labels=[\"No_Recharge\",\"<=10_Recharges\",\"10-25_Recharges\",\">25_Recharges\"])\nchurn_filtered['total_rech_num_group_8']=pd.cut(churn_filtered['total_rech_num_8'],[-1,0,10,25,1000],labels=[\"No_Recharge\",\"<=10_Recharges\",\"10-25_Recharges\",\">25_Recharges\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.065925Z","iopub.execute_input":"2022-08-09T14:29:19.066356Z","iopub.status.idle":"2022-08-09T14:29:19.079680Z","shell.execute_reply.started":"2022-08-09T14:29:19.066314Z","shell.execute_reply":"2022-08-09T14:29:19.078888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting the results\n\nplt.figure(figsize=[12,4])\nsns.countplot(data=churn_filtered,x=\"total_rech_data_group_8\",hue=\"churn_probability\")\nprint(\"\\t\\t\\t\\t\\tDistribution of total_rech_data_8 variable\\n\",churn_filtered['total_rech_data_group_8'].value_counts())\nplt.show()\nplt.figure(figsize=[12,4])\nsns.countplot(data=churn_filtered,x=\"total_rech_num_group_8\",hue=\"churn_probability\")\nprint(\"\\t\\t\\t\\t\\tDistribution of total_rech_num_8 variable\\n\",churn_filtered['total_rech_num_group_8'].value_counts())\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.081131Z","iopub.execute_input":"2022-08-09T14:29:19.081881Z","iopub.status.idle":"2022-08-09T14:29:19.467152Z","shell.execute_reply.started":"2022-08-09T14:29:19.081811Z","shell.execute_reply":"2022-08-09T14:29:19.466085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As the number of recharge rate increases, the churn rate decreases clearly.","metadata":{}},{"cell_type":"code","source":"churn_filtered.drop(['av_rech_amt_data_8','total_rech_data_8','sachet_2g_6','sachet_2g_7','sachet_3g_6',\n              'sachet_3g_7','sachet_3g_8','last_day_rch_amt_6','last_day_rch_amt_7',\n              'last_day_rch_amt_8',], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.468448Z","iopub.execute_input":"2022-08-09T14:29:19.469041Z","iopub.status.idle":"2022-08-09T14:29:19.480284Z","shell.execute_reply.started":"2022-08-09T14:29:19.469007Z","shell.execute_reply":"2022-08-09T14:29:19.479208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.drop(['loc_og_t2o_mou', 'std_og_t2o_mou', 'loc_ic_t2o_mou','roam_ic_mou_6', 'roam_ic_mou_7', 'roam_ic_mou_8', \n         'roam_og_mou_6', 'roam_og_mou_7', 'roam_og_mou_8', 'loc_og_t2t_mou_6', 'loc_og_t2t_mou_7', 'loc_og_t2t_mou_8',\n         'loc_og_t2m_mou_6', 'loc_og_t2m_mou_7', 'loc_og_t2m_mou_8', 'loc_og_t2f_mou_6', 'loc_og_t2f_mou_7', 'loc_og_t2f_mou_8',\n         'loc_og_t2c_mou_6', 'loc_og_t2c_mou_7', 'loc_og_t2c_mou_8', 'loc_og_mou_6', 'loc_og_mou_7', 'loc_og_mou_8', \n         'std_og_t2m_mou_6', 'std_og_t2f_mou_6', 'std_og_t2f_mou_7', 'std_og_t2f_mou_8', 'std_og_t2c_mou_6', 'std_og_t2c_mou_7',\n         'std_og_t2c_mou_8', 'std_og_mou_6', 'std_og_mou_7', 'std_og_mou_8', 'isd_og_mou_6', 'isd_og_mou_7', 'spl_og_mou_6',\n         'spl_og_mou_7', 'spl_og_mou_8','total_og_mou_6', 'loc_ic_t2t_mou_6', 'loc_ic_t2t_mou_7', 'loc_ic_t2t_mou_8', \n         'loc_ic_t2m_mou_6', 'loc_ic_t2m_mou_7', 'loc_ic_t2m_mou_8', 'loc_ic_t2f_mou_6', 'loc_ic_t2f_mou_7', 'loc_ic_t2f_mou_8',\n         'loc_ic_mou_6', 'loc_ic_mou_7', 'loc_ic_mou_8', 'std_ic_t2t_mou_6', 'std_ic_t2t_mou_7', 'std_ic_t2t_mou_8', \n         'std_ic_t2m_mou_6', 'std_ic_t2m_mou_7', 'std_ic_t2m_mou_8', 'std_ic_t2f_mou_6', 'std_ic_t2f_mou_7', 'std_ic_t2f_mou_8',\n         'std_ic_t2o_mou_6', 'std_ic_t2o_mou_7', 'std_ic_t2o_mou_8', 'std_ic_mou_6', 'spl_ic_mou_6', 'spl_ic_mou_7',\n         'spl_ic_mou_8', 'isd_ic_mou_6', 'isd_ic_mou_7', 'isd_ic_mou_8',], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.481649Z","iopub.execute_input":"2022-08-09T14:29:19.482372Z","iopub.status.idle":"2022-08-09T14:29:19.496625Z","shell.execute_reply.started":"2022-08-09T14:29:19.482341Z","shell.execute_reply":"2022-08-09T14:29:19.495145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.499714Z","iopub.execute_input":"2022-08-09T14:29:19.500753Z","iopub.status.idle":"2022-08-09T14:29:19.512348Z","shell.execute_reply.started":"2022-08-09T14:29:19.500709Z","shell.execute_reply":"2022-08-09T14:29:19.511221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (50, 50))\nsns.heatmap(churn_filtered.corr())\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:19.514852Z","iopub.execute_input":"2022-08-09T14:29:19.515936Z","iopub.status.idle":"2022-08-09T14:29:21.997393Z","shell.execute_reply.started":"2022-08-09T14:29:19.515895Z","shell.execute_reply":"2022-08-09T14:29:21.996242Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:21.999160Z","iopub.execute_input":"2022-08-09T14:29:21.999585Z","iopub.status.idle":"2022-08-09T14:29:22.022447Z","shell.execute_reply.started":"2022-08-09T14:29:21.999552Z","shell.execute_reply":"2022-08-09T14:29:22.021388Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.drop(['total_rech_data_group_8','total_rech_num_group_8',] , axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.023901Z","iopub.execute_input":"2022-08-09T14:29:22.024756Z","iopub.status.idle":"2022-08-09T14:29:22.033031Z","shell.execute_reply.started":"2022-08-09T14:29:22.024711Z","shell.execute_reply":"2022-08-09T14:29:22.031967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.034367Z","iopub.execute_input":"2022-08-09T14:29:22.034662Z","iopub.status.idle":"2022-08-09T14:29:22.046459Z","shell.execute_reply.started":"2022-08-09T14:29:22.034633Z","shell.execute_reply":"2022-08-09T14:29:22.045636Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.047888Z","iopub.execute_input":"2022-08-09T14:29:22.048544Z","iopub.status.idle":"2022-08-09T14:29:22.072333Z","shell.execute_reply.started":"2022-08-09T14:29:22.048502Z","shell.execute_reply":"2022-08-09T14:29:22.071133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.drop(['tenure_range'] , axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.075522Z","iopub.execute_input":"2022-08-09T14:29:22.075993Z","iopub.status.idle":"2022-08-09T14:29:22.084521Z","shell.execute_reply.started":"2022-08-09T14:29:22.075956Z","shell.execute_reply":"2022-08-09T14:29:22.083400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_filtered.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.085781Z","iopub.execute_input":"2022-08-09T14:29:22.086101Z","iopub.status.idle":"2022-08-09T14:29:22.108443Z","shell.execute_reply.started":"2022-08-09T14:29:22.086073Z","shell.execute_reply":"2022-08-09T14:29:22.107492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_rate = (sum(churn_filtered[\"churn_probability\"])/len(churn_filtered[\"churn_probability\"].index))*100\nchurn_rate","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.109586Z","iopub.execute_input":"2022-08-09T14:29:22.109938Z","iopub.status.idle":"2022-08-09T14:29:22.118590Z","shell.execute_reply.started":"2022-08-09T14:29:22.109907Z","shell.execute_reply":"2022-08-09T14:29:22.117404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# v.Split Data Into Train and Test Data","metadata":{}},{"cell_type":"code","source":"churn_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.120444Z","iopub.execute_input":"2022-08-09T14:29:22.121227Z","iopub.status.idle":"2022-08-09T14:29:22.129145Z","shell.execute_reply.started":"2022-08-09T14:29:22.121196Z","shell.execute_reply":"2022-08-09T14:29:22.128391Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# divide data into train and test\nX = churn_filtered.drop(\"churn_probability\", axis = 1)\ny = churn_filtered.churn_probability\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size = 0.25, random_state = 4, stratify = y)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.130062Z","iopub.execute_input":"2022-08-09T14:29:22.130516Z","iopub.status.idle":"2022-08-09T14:29:22.158796Z","shell.execute_reply.started":"2022-08-09T14:29:22.130488Z","shell.execute_reply":"2022-08-09T14:29:22.157831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print shapes of train and test sets\nX_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.162606Z","iopub.execute_input":"2022-08-09T14:29:22.162918Z","iopub.status.idle":"2022-08-09T14:29:22.168101Z","shell.execute_reply.started":"2022-08-09T14:29:22.162889Z","shell.execute_reply":"2022-08-09T14:29:22.167342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(y_train.shape)\nprint(X_test.shape)\nprint(y_test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.169062Z","iopub.execute_input":"2022-08-09T14:29:22.169786Z","iopub.status.idle":"2022-08-09T14:29:22.179890Z","shell.execute_reply.started":"2022-08-09T14:29:22.169750Z","shell.execute_reply":"2022-08-09T14:29:22.178629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# vi.Perform Scaling","metadata":{}},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.181099Z","iopub.execute_input":"2022-08-09T14:29:22.181854Z","iopub.status.idle":"2022-08-09T14:29:22.234909Z","shell.execute_reply.started":"2022-08-09T14:29:22.181815Z","shell.execute_reply":"2022-08-09T14:29:22.233866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.238223Z","iopub.execute_input":"2022-08-09T14:29:22.238505Z","iopub.status.idle":"2022-08-09T14:29:22.255127Z","shell.execute_reply.started":"2022-08-09T14:29:22.238479Z","shell.execute_reply":"2022-08-09T14:29:22.254076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_col = X_train.select_dtypes(include = ['int64','float64']).columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.271597Z","iopub.execute_input":"2022-08-09T14:29:22.272006Z","iopub.status.idle":"2022-08-09T14:29:22.278859Z","shell.execute_reply.started":"2022-08-09T14:29:22.271974Z","shell.execute_reply":"2022-08-09T14:29:22.278192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# apply scaling on the dataset\nfrom sklearn import preprocessing\nfrom sklearn.preprocessing import MinMaxScaler\n\nscaler = MinMaxScaler()\nX_train[num_col] = scaler.fit_transform(X_train[num_col])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.279752Z","iopub.execute_input":"2022-08-09T14:29:22.280221Z","iopub.status.idle":"2022-08-09T14:29:22.311109Z","shell.execute_reply.started":"2022-08-09T14:29:22.280191Z","shell.execute_reply":"2022-08-09T14:29:22.310044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.312664Z","iopub.execute_input":"2022-08-09T14:29:22.313019Z","iopub.status.idle":"2022-08-09T14:29:22.360773Z","shell.execute_reply.started":"2022-08-09T14:29:22.312988Z","shell.execute_reply":"2022-08-09T14:29:22.360002Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As there are many variables we will start the process of dropping variables after doing the RFE","metadata":{}},{"cell_type":"markdown","source":"# Data Modeling and Model Evaluation and Prepare Model for Predictor variables selection\n","metadata":{}},{"cell_type":"markdown","source":"## Data Imbalance Handling\nUsing SMOTE method, we can balance the data w.r.t. churn variable and proceed further","metadata":{}},{"cell_type":"code","source":"smote = SMOTE(random_state=42)\nX_train_sm,y_train_sm = smote.fit_resample(X_train,y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.361792Z","iopub.execute_input":"2022-08-09T14:29:22.362452Z","iopub.status.idle":"2022-08-09T14:29:22.477652Z","shell.execute_reply.started":"2022-08-09T14:29:22.362419Z","shell.execute_reply":"2022-08-09T14:29:22.475670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Dimension of X_train_sm Shape:\", X_train_sm.shape)\nprint(\"Dimension of y_train_sm Shape:\", y_train_sm.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.480013Z","iopub.execute_input":"2022-08-09T14:29:22.480860Z","iopub.status.idle":"2022-08-09T14:29:22.488821Z","shell.execute_reply.started":"2022-08-09T14:29:22.480792Z","shell.execute_reply":"2022-08-09T14:29:22.487377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logistic Regression","metadata":{}},{"cell_type":"code","source":"# Logistic regression model\nlogm1 = sm.GLM(y_train_sm,(sm.add_constant(X_train_sm)), family = sm.families.Binomial())\nlogm1.fit().summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:22.491211Z","iopub.execute_input":"2022-08-09T14:29:22.492191Z","iopub.status.idle":"2022-08-09T14:29:26.084734Z","shell.execute_reply.started":"2022-08-09T14:29:22.492000Z","shell.execute_reply":"2022-08-09T14:29:26.083236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logistic Regression using Feature Selection (RFE method)\n","metadata":{}},{"cell_type":"code","source":"logreg = LogisticRegression()\n\nfrom sklearn.feature_selection import RFE\n\n# running RFE with 20 variables as output\nrfe = RFE(logreg,  n_features_to_select= 20)             \nrfe = rfe.fit(X_train_sm, y_train_sm)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:26.087057Z","iopub.execute_input":"2022-08-09T14:29:26.088002Z","iopub.status.idle":"2022-08-09T14:29:41.417010Z","shell.execute_reply.started":"2022-08-09T14:29:26.087942Z","shell.execute_reply":"2022-08-09T14:29:41.415452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rfe.support_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.419410Z","iopub.execute_input":"2022-08-09T14:29:41.420450Z","iopub.status.idle":"2022-08-09T14:29:41.430068Z","shell.execute_reply.started":"2022-08-09T14:29:41.420392Z","shell.execute_reply":"2022-08-09T14:29:41.428644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rfe_columns=X_train_sm.columns[rfe.support_]\nprint(\"The selected columns by RFE for modelling are: \\n\\n\",rfe_columns)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.432665Z","iopub.execute_input":"2022-08-09T14:29:41.433777Z","iopub.status.idle":"2022-08-09T14:29:41.442621Z","shell.execute_reply.started":"2022-08-09T14:29:41.433716Z","shell.execute_reply":"2022-08-09T14:29:41.441290Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list(zip(X_train_sm.columns, rfe.support_, rfe.ranking_))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.444827Z","iopub.execute_input":"2022-08-09T14:29:41.445759Z","iopub.status.idle":"2022-08-09T14:29:41.465940Z","shell.execute_reply.started":"2022-08-09T14:29:41.445705Z","shell.execute_reply":"2022-08-09T14:29:41.464576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Assessing the model with StatsModels","metadata":{}},{"cell_type":"code","source":"X_train_SM = sm.add_constant(X_train_sm[rfe_columns])\nlogm2 = sm.GLM(y_train_sm,X_train_SM, family = sm.families.Binomial())\nres = logm2.fit()\nres.summary()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.468429Z","iopub.execute_input":"2022-08-09T14:29:41.469594Z","iopub.status.idle":"2022-08-09T14:29:41.752829Z","shell.execute_reply.started":"2022-08-09T14:29:41.469533Z","shell.execute_reply":"2022-08-09T14:29:41.751351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Getting the predicted values on the train set\ny_train_sm_pred = res.predict(X_train_SM)\ny_train_sm_pred = y_train_sm_pred.values.reshape(-1)\ny_train_sm_pred[:10]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.756073Z","iopub.execute_input":"2022-08-09T14:29:41.757226Z","iopub.status.idle":"2022-08-09T14:29:41.772172Z","shell.execute_reply.started":"2022-08-09T14:29:41.757164Z","shell.execute_reply":"2022-08-09T14:29:41.770645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a dataframe with the actual churn flag and the predicted probabilities\ny_train_sm_pred_final = pd.DataFrame({'Converted':y_train_sm.values, 'Converted_prob':y_train_sm_pred})\ny_train_sm_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.775049Z","iopub.execute_input":"2022-08-09T14:29:41.776164Z","iopub.status.idle":"2022-08-09T14:29:41.795863Z","shell.execute_reply.started":"2022-08-09T14:29:41.776106Z","shell.execute_reply":"2022-08-09T14:29:41.794134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating new column 'churn_pred' with 1 if Churn_Prob > 0.8 else 0","metadata":{}},{"cell_type":"code","source":"y_train_sm_pred_final['churn_pred'] = y_train_sm_pred_final.Converted_prob.map(lambda x: 1 if x > 0.5 else 0)\n\n# Viewing the prediction results\ny_train_sm_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.798441Z","iopub.execute_input":"2022-08-09T14:29:41.799553Z","iopub.status.idle":"2022-08-09T14:29:41.857485Z","shell.execute_reply.started":"2022-08-09T14:29:41.799491Z","shell.execute_reply":"2022-08-09T14:29:41.855251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Confusion matrix \nconfusion = metrics.confusion_matrix(y_train_sm_pred_final.Converted, y_train_sm_pred_final.churn_pred )\nprint(confusion)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.860235Z","iopub.execute_input":"2022-08-09T14:29:41.861316Z","iopub.status.idle":"2022-08-09T14:29:41.882626Z","shell.execute_reply.started":"2022-08-09T14:29:41.861258Z","shell.execute_reply":"2022-08-09T14:29:41.881135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Confusion matrix\n\n# Predicted     not_churn    churn\n# Actual\n\n# not_churn     11630           2825\n                    \n# churn             2238            12217  ","metadata":{}},{"cell_type":"code","source":"# Checking the overall accuracy.\nprint(\"The overall accuracy of the model is:\",metrics.accuracy_score(y_train_sm_pred_final.Converted, y_train_sm_pred_final.churn_pred))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.884865Z","iopub.execute_input":"2022-08-09T14:29:41.885447Z","iopub.status.idle":"2022-08-09T14:29:41.893795Z","shell.execute_reply.started":"2022-08-09T14:29:41.885416Z","shell.execute_reply":"2022-08-09T14:29:41.892962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Check for the VIF values of the feature variables\n","metadata":{}},{"cell_type":"code","source":"# Create a dataframe that will contain the names of all the feature variables and their respective VIFs\nvif = pd.DataFrame()\nvif['Features'] = X_train_sm[rfe_columns].columns\nvif['VIF'] = [variance_inflation_factor(X_train_sm[rfe_columns].values, i) for i in range(X_train_sm[rfe_columns].shape[1])]\nvif['VIF'] = round(vif['VIF'], 2)\nvif = vif.sort_values(by = \"VIF\", ascending = False)\nvif","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:41.895732Z","iopub.execute_input":"2022-08-09T14:29:41.896490Z","iopub.status.idle":"2022-08-09T14:29:42.967354Z","shell.execute_reply.started":"2022-08-09T14:29:41.896449Z","shell.execute_reply":"2022-08-09T14:29:42.966001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Metrics beyond simply accuracy\n","metadata":{}},{"cell_type":"code","source":"TP = confusion[1,1] # true positive \nTN = confusion[0,0] # true negatives\nFP = confusion[0,1] # false positives\nFN = confusion[1,0] # false negatives","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:42.969623Z","iopub.execute_input":"2022-08-09T14:29:42.970517Z","iopub.status.idle":"2022-08-09T14:29:42.978418Z","shell.execute_reply.started":"2022-08-09T14:29:42.970459Z","shell.execute_reply":"2022-08-09T14:29:42.977031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's see the sensitivity of our logistic regression model\nprint(\"Sensitivity = \",TP / float(TP+FN))\n\n# Let us calculate specificity\nprint(\"Specificity = \",TN / float(TN+FP))\n\n# Calculate false postive rate - predicting churn when customer does not have churned\nprint(\"False Positive Rate = \",FP/ float(TN+FP))\n\n# positive predictive value \nprint (\"Precision = \",TP / float(TP+FP))\n\n# Negative predictive value\nprint (\"True Negative Prediction Rate = \",TN / float(TN+ FN))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:42.980667Z","iopub.execute_input":"2022-08-09T14:29:42.981659Z","iopub.status.idle":"2022-08-09T14:29:42.994374Z","shell.execute_reply.started":"2022-08-09T14:29:42.981598Z","shell.execute_reply":"2022-08-09T14:29:42.992956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Plotting the ROC Curve","metadata":{}},{"cell_type":"code","source":"# Defining a function to plot the roc curve\ndef draw_roc( actual, probs ):\n    fpr, tpr, thresholds = metrics.roc_curve( actual, probs,\n                                              drop_intermediate = False )\n    auc_score = metrics.roc_auc_score( actual, probs )\n    plt.figure(figsize=(5, 5))\n    plt.plot( fpr, tpr, label='ROC curve (area = %0.2f)' % auc_score )\n    plt.plot([0, 1], [0, 1], 'k--')\n    plt.xlim([0.0, 1.0])\n    plt.ylim([0.0, 1.05])\n    plt.xlabel('False Positive Rate or [1 - True Negative Prediction Rate]')\n    plt.ylabel('True Positive Rate')\n    plt.title('Receiver operating characteristic example')\n    plt.legend(loc=\"lower right\")\n    plt.show()\n\n    return None","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:42.997292Z","iopub.execute_input":"2022-08-09T14:29:42.998346Z","iopub.status.idle":"2022-08-09T14:29:43.013171Z","shell.execute_reply.started":"2022-08-09T14:29:42.998283Z","shell.execute_reply":"2022-08-09T14:29:43.011547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Defining the variables to plot the curve\nfpr, tpr, thresholds = metrics.roc_curve( y_train_sm_pred_final.Converted, y_train_sm_pred_final.Converted_prob, drop_intermediate = False )","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.015778Z","iopub.execute_input":"2022-08-09T14:29:43.017056Z","iopub.status.idle":"2022-08-09T14:29:43.039697Z","shell.execute_reply.started":"2022-08-09T14:29:43.016994Z","shell.execute_reply":"2022-08-09T14:29:43.038145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Plotting the curve for the obtained metrics\ndraw_roc(y_train_sm_pred_final.Converted, y_train_sm_pred_final.Converted_prob)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.042378Z","iopub.execute_input":"2022-08-09T14:29:43.043434Z","iopub.status.idle":"2022-08-09T14:29:43.273113Z","shell.execute_reply.started":"2022-08-09T14:29:43.043374Z","shell.execute_reply":"2022-08-09T14:29:43.271995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Finding Optimal Cutoff Point\n","metadata":{}},{"cell_type":"code","source":"# Let's create columns with different probability cutoffs \nnumbers = [float(x)/10 for x in range(10)]\nfor i in numbers:\n    y_train_sm_pred_final[i]= y_train_sm_pred_final.Converted_prob.map(lambda x: 1 if x > i else 0)\ny_train_sm_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.274354Z","iopub.execute_input":"2022-08-09T14:29:43.274650Z","iopub.status.idle":"2022-08-09T14:29:43.430905Z","shell.execute_reply.started":"2022-08-09T14:29:43.274622Z","shell.execute_reply":"2022-08-09T14:29:43.429366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now let's calculate accuracy sensitivity and specificity for various probability cutoffs.\ncutoff_df = pd.DataFrame( columns = ['probability','accuracy','sensitivity','specificity'])\nfrom sklearn.metrics import confusion_matrix\n\n# TP = confusion[1,1] # true positive \n# TN = confusion[0,0] # true negatives\n# FP = confusion[0,1] # false positives\n# FN = confusion[1,0] # false negatives\n\nnum = [0.0,0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9]\nfor i in num:\n    cm1 = metrics.confusion_matrix(y_train_sm_pred_final.Converted, y_train_sm_pred_final[i] )\n    total1=sum(sum(cm1))\n    accuracy = (cm1[0,0]+cm1[1,1])/total1\n    \n    specificity = cm1[0,0]/(cm1[0,0]+cm1[0,1])\n    sensitivity = cm1[1,1]/(cm1[1,0]+cm1[1,1])\n    cutoff_df.loc[i] =[ i ,accuracy,sensitivity,specificity]\nprint(cutoff_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.432885Z","iopub.execute_input":"2022-08-09T14:29:43.433200Z","iopub.status.idle":"2022-08-09T14:29:43.518894Z","shell.execute_reply.started":"2022-08-09T14:29:43.433171Z","shell.execute_reply":"2022-08-09T14:29:43.517884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plotting accuracy sensitivity and specificity for various probabilities calculated above.\ncutoff_df.plot.line(x='probability', y=['accuracy','sensitivity','specificity'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.520460Z","iopub.execute_input":"2022-08-09T14:29:43.520750Z","iopub.status.idle":"2022-08-09T14:29:43.733999Z","shell.execute_reply.started":"2022-08-09T14:29:43.520722Z","shell.execute_reply":"2022-08-09T14:29:43.732883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Initially we selected the optimum point of classification as 0.5.<br><br>From the above graph, we can see the optimum cutoff is slightly higher than 0.5 but lies lower than 0.6. So lets tweek a little more within this range.**","metadata":{}},{"cell_type":"code","source":"# Let's create columns with refined probability cutoffs \nnumbers = [0.50,0.51,0.52,0.53,0.54,0.55,0.56,0.57,0.58,0.59]\nfor i in numbers:\n    y_train_sm_pred_final[i]= y_train_sm_pred_final.Converted_prob.map(lambda x: 1 if x > i else 0)\ny_train_sm_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.735862Z","iopub.execute_input":"2022-08-09T14:29:43.736757Z","iopub.status.idle":"2022-08-09T14:29:43.880211Z","shell.execute_reply.started":"2022-08-09T14:29:43.736711Z","shell.execute_reply":"2022-08-09T14:29:43.878920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Now let's calculate accuracy sensitivity and specificity for various probability cutoffs.\ncutoff_df = pd.DataFrame( columns = ['probability','accuracy','sensitivity','specificity'])\nfrom sklearn.metrics import confusion_matrix\n\n# TP = confusion[1,1] # true positive \n# TN = confusion[0,0] # true negatives\n# FP = confusion[0,1] # false positives\n# FN = confusion[1,0] # false negatives\n\nnum = [0.50,0.51,0.52,0.53,0.54,0.55,0.56,0.57,0.58,0.59]\nfor i in num:\n    cm1 = metrics.confusion_matrix(y_train_sm_pred_final.Converted, y_train_sm_pred_final[i] )\n    total1=sum(sum(cm1))\n    accuracy = (cm1[0,0]+cm1[1,1])/total1\n    \n    specificity = cm1[0,0]/(cm1[0,0]+cm1[0,1])\n    sensitivity = cm1[1,1]/(cm1[1,0]+cm1[1,1])\n    cutoff_df.loc[i] =[ i ,accuracy,sensitivity,specificity]\nprint(cutoff_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.881828Z","iopub.execute_input":"2022-08-09T14:29:43.882362Z","iopub.status.idle":"2022-08-09T14:29:43.962894Z","shell.execute_reply.started":"2022-08-09T14:29:43.882320Z","shell.execute_reply":"2022-08-09T14:29:43.961806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plotting accuracy sensitivity and specificity for various probabilities calculated above.\ncutoff_df.plot.line(x='probability', y=['accuracy','sensitivity','specificity'])\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:43.964539Z","iopub.execute_input":"2022-08-09T14:29:43.965194Z","iopub.status.idle":"2022-08-09T14:29:44.189894Z","shell.execute_reply.started":"2022-08-09T14:29:43.965151Z","shell.execute_reply":"2022-08-09T14:29:44.189025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**From the above graph we can conclude, the optimal cutoff point in the probability to define the predicted churn variabe converges at `0.54`**","metadata":{}},{"cell_type":"code","source":"#### From the curve above,we can take 0.54 is the optimum point to take it as a cutoff probability.\n\ny_train_sm_pred_final['final_churn_pred'] = y_train_sm_pred_final.Converted_prob.map( lambda x: 1 if x > 0.53 else 0)\n\ny_train_sm_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.191249Z","iopub.execute_input":"2022-08-09T14:29:44.191682Z","iopub.status.idle":"2022-08-09T14:29:44.224379Z","shell.execute_reply.started":"2022-08-09T14:29:44.191638Z","shell.execute_reply":"2022-08-09T14:29:44.223363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculating the ovearall accuracy again\nprint(\"The overall accuracy of the model now is:\",metrics.accuracy_score(y_train_sm_pred_final.Converted, y_train_sm_pred_final.final_churn_pred))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.225708Z","iopub.execute_input":"2022-08-09T14:29:44.226092Z","iopub.status.idle":"2022-08-09T14:29:44.234155Z","shell.execute_reply.started":"2022-08-09T14:29:44.226064Z","shell.execute_reply":"2022-08-09T14:29:44.233160Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion2 = metrics.confusion_matrix(y_train_sm_pred_final.Converted, y_train_sm_pred_final.final_churn_pred )\nprint(confusion2)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.235491Z","iopub.execute_input":"2022-08-09T14:29:44.236288Z","iopub.status.idle":"2022-08-09T14:29:44.248562Z","shell.execute_reply.started":"2022-08-09T14:29:44.236254Z","shell.execute_reply":"2022-08-09T14:29:44.247546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TP2 = confusion2[1,1] # true positive \nTN2 = confusion2[0,0] # true negatives\nFP2 = confusion2[0,1] # false positives\nFN2 = confusion2[1,0] # false negatives\n\n# Let's see the sensitivity of our logistic regression model\nprint(\"Sensitivity = \",TP2 / float(TP2+FN2))\n\n# Let us calculate specificity\nprint(\"Specificity = \",TN2 / float(TN2+FP2))\n\n# Calculate false postive rate - predicting churn when customer does not have churned\nprint(\"False Positive Rate = \",FP2/ float(TN2+FP2))\n\n# positive predictive value \nprint (\"Precision = \",TP2 / float(TP2+FP2))\n\n# Negative predictive value\nprint (\"True Negative Prediction Rate = \",TN2 / float(TN2 + FN2))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.250214Z","iopub.execute_input":"2022-08-09T14:29:44.250556Z","iopub.status.idle":"2022-08-09T14:29:44.258629Z","shell.execute_reply.started":"2022-08-09T14:29:44.250527Z","shell.execute_reply":"2022-08-09T14:29:44.257333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Precision and recall tradeoff\n","metadata":{}},{"cell_type":"code","source":"p, r, thresholds = precision_recall_curve(y_train_sm_pred_final.Converted, y_train_sm_pred_final.Converted_prob)\n\n# Plotting the curve\nplt.plot(thresholds, p[:-1], \"g-\")\nplt.plot(thresholds, r[:-1], \"r-\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.260474Z","iopub.execute_input":"2022-08-09T14:29:44.261382Z","iopub.status.idle":"2022-08-09T14:29:44.458085Z","shell.execute_reply.started":"2022-08-09T14:29:44.261337Z","shell.execute_reply":"2022-08-09T14:29:44.457332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Making predictions on the test set\n**Transforming and feature selection for test data**","metadata":{}},{"cell_type":"code","source":"# Scaling the test data\nX_test[num_col] = scaler.transform(X_test[num_col])\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.459245Z","iopub.execute_input":"2022-08-09T14:29:44.459693Z","iopub.status.idle":"2022-08-09T14:29:44.518170Z","shell.execute_reply.started":"2022-08-09T14:29:44.459662Z","shell.execute_reply":"2022-08-09T14:29:44.517099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Feature selection\nX_test=X_test[rfe_columns]\nX_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.519656Z","iopub.execute_input":"2022-08-09T14:29:44.520211Z","iopub.status.idle":"2022-08-09T14:29:44.546905Z","shell.execute_reply.started":"2022-08-09T14:29:44.520179Z","shell.execute_reply":"2022-08-09T14:29:44.545881Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Adding constant to the test model.\nX_test_SM = sm.add_constant(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.548447Z","iopub.execute_input":"2022-08-09T14:29:44.548766Z","iopub.status.idle":"2022-08-09T14:29:44.559288Z","shell.execute_reply.started":"2022-08-09T14:29:44.548738Z","shell.execute_reply":"2022-08-09T14:29:44.558265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Predicting the target variable","metadata":{}},{"cell_type":"code","source":"y_test_pred = res.predict(X_test_SM)\nprint(\"\\n The first ten probability value of the prediction are:\\n\",y_test_pred[:10])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.560821Z","iopub.execute_input":"2022-08-09T14:29:44.561267Z","iopub.status.idle":"2022-08-09T14:29:44.579581Z","shell.execute_reply.started":"2022-08-09T14:29:44.561225Z","shell.execute_reply":"2022-08-09T14:29:44.577890Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred = pd.DataFrame(y_test_pred)\ny_pred.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.581785Z","iopub.execute_input":"2022-08-09T14:29:44.582900Z","iopub.status.idle":"2022-08-09T14:29:44.597636Z","shell.execute_reply.started":"2022-08-09T14:29:44.582826Z","shell.execute_reply":"2022-08-09T14:29:44.595906Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred=y_pred.rename(columns = {0:\"Conv_prob\"})","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.599461Z","iopub.execute_input":"2022-08-09T14:29:44.600226Z","iopub.status.idle":"2022-08-09T14:29:44.608025Z","shell.execute_reply.started":"2022-08-09T14:29:44.600179Z","shell.execute_reply":"2022-08-09T14:29:44.606889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_df = pd.DataFrame(y_test)\ny_test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.609982Z","iopub.execute_input":"2022-08-09T14:29:44.611235Z","iopub.status.idle":"2022-08-09T14:29:44.627712Z","shell.execute_reply.started":"2022-08-09T14:29:44.611195Z","shell.execute_reply":"2022-08-09T14:29:44.626432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_final = pd.concat([y_test_df,y_pred],axis=1)\ny_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.629952Z","iopub.execute_input":"2022-08-09T14:29:44.630950Z","iopub.status.idle":"2022-08-09T14:29:44.648478Z","shell.execute_reply.started":"2022-08-09T14:29:44.630908Z","shell.execute_reply":"2022-08-09T14:29:44.647138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_pred_final['test_churn_pred'] = y_pred_final.Conv_prob.map(lambda x: 1 if x>0.54 else 0)\ny_pred_final.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.650362Z","iopub.execute_input":"2022-08-09T14:29:44.651318Z","iopub.status.idle":"2022-08-09T14:29:44.675777Z","shell.execute_reply.started":"2022-08-09T14:29:44.651271Z","shell.execute_reply":"2022-08-09T14:29:44.674430Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Checking the overall accuracy of the predicted set.\nmetrics.accuracy_score(y_pred_final.churn_probability, y_pred_final.test_churn_pred)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.677623Z","iopub.execute_input":"2022-08-09T14:29:44.678994Z","iopub.status.idle":"2022-08-09T14:29:44.690190Z","shell.execute_reply.started":"2022-08-09T14:29:44.678931Z","shell.execute_reply":"2022-08-09T14:29:44.688595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Metrics Evaluation**","metadata":{}},{"cell_type":"code","source":"# Confusion Matrix\nconfusion2_test = metrics.confusion_matrix(y_pred_final.churn_probability, y_pred_final.test_churn_pred)\nprint(\"Confusion Matrix\\n\",confusion2_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.691826Z","iopub.execute_input":"2022-08-09T14:29:44.692354Z","iopub.status.idle":"2022-08-09T14:29:44.699747Z","shell.execute_reply.started":"2022-08-09T14:29:44.692321Z","shell.execute_reply":"2022-08-09T14:29:44.698573Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculating model validation parameters\nTP3 = confusion2_test[1,1] # true positive \nTN3 = confusion2_test[0,0] # true negatives\nFP3 = confusion2_test[0,1] # false positives\nFN3 = confusion2_test[1,0] # false negatives","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.700827Z","iopub.execute_input":"2022-08-09T14:29:44.701549Z","iopub.status.idle":"2022-08-09T14:29:44.708209Z","shell.execute_reply.started":"2022-08-09T14:29:44.701520Z","shell.execute_reply":"2022-08-09T14:29:44.707080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Let's see the sensitivity of our logistic regression model\nprint(\"Sensitivity = \",TP3 / float(TP3+FN3))\n\n# Let us calculate specificity\nprint(\"Specificity = \",TN3 / float(TN3+FP3))\n\n# Calculate false postive rate - predicting churn when customer does not have churned\nprint(\"False Positive Rate = \",FP3/ float(TN3+FP3))\n\n# positive predictive value \nprint (\"Precision = \",TP3 / float(TP3+FP3))\n\n# Negative predictive value\nprint (\"True Negative Prediction Rate = \",TN3 / float(TN3+FN3))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.709393Z","iopub.execute_input":"2022-08-09T14:29:44.710138Z","iopub.status.idle":"2022-08-09T14:29:44.725047Z","shell.execute_reply.started":"2022-08-09T14:29:44.710105Z","shell.execute_reply":"2022-08-09T14:29:44.724281Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Explaining the results","metadata":{}},{"cell_type":"code","source":"print(\"The accuracy of the predicted model is: \",round(metrics.accuracy_score(y_pred_final.churn_probability, y_pred_final.test_churn_pred),2)*100,\"%\")\nprint(\"The sensitivity of the predicted model is: \",round(TP3 / float(TP3+FN3),2)*100,\"%\")\n\nprint(\"\\nAs the model created is based on a sentivity model, i.e. the True positive rate is given more importance as the actual and prediction of churn by a customer\\n\") ","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.726221Z","iopub.execute_input":"2022-08-09T14:29:44.727203Z","iopub.status.idle":"2022-08-09T14:29:44.737260Z","shell.execute_reply.started":"2022-08-09T14:29:44.727174Z","shell.execute_reply":"2022-08-09T14:29:44.736386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ROC curve for the test dataset\n\n# Defining the variables to plot the curve\nfpr, tpr, thresholds = metrics.roc_curve(y_pred_final.churn_probability,y_pred_final.Conv_prob, drop_intermediate = False )\n# Plotting the curve for the obtained metrics\ndraw_roc(y_pred_final.churn_probability,y_pred_final.Conv_prob)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.738298Z","iopub.execute_input":"2022-08-09T14:29:44.739109Z","iopub.status.idle":"2022-08-09T14:29:44.963470Z","shell.execute_reply.started":"2022-08-09T14:29:44.739073Z","shell.execute_reply":"2022-08-09T14:29:44.962746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The AUC score for train dataset is 0.90 and the test dataset is 0.88.\n# This model can be considered as a good model.**","metadata":{}},{"cell_type":"markdown","source":"# PCA","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.8, test_size=0.2, random_state=10)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.964588Z","iopub.execute_input":"2022-08-09T14:29:44.964868Z","iopub.status.idle":"2022-08-09T14:29:44.980270Z","shell.execute_reply.started":"2022-08-09T14:29:44.964829Z","shell.execute_reply":"2022-08-09T14:29:44.978963Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.981663Z","iopub.execute_input":"2022-08-09T14:29:44.981987Z","iopub.status.idle":"2022-08-09T14:29:44.988312Z","shell.execute_reply.started":"2022-08-09T14:29:44.981958Z","shell.execute_reply":"2022-08-09T14:29:44.986999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pca = PCA(random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.989988Z","iopub.execute_input":"2022-08-09T14:29:44.990625Z","iopub.status.idle":"2022-08-09T14:29:44.997295Z","shell.execute_reply.started":"2022-08-09T14:29:44.990582Z","shell.execute_reply":"2022-08-09T14:29:44.996316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pca.fit(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:44.998324Z","iopub.execute_input":"2022-08-09T14:29:44.999139Z","iopub.status.idle":"2022-08-09T14:29:45.073807Z","shell.execute_reply.started":"2022-08-09T14:29:44.999107Z","shell.execute_reply":"2022-08-09T14:29:45.072357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pca.components_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.076173Z","iopub.execute_input":"2022-08-09T14:29:45.077092Z","iopub.status.idle":"2022-08-09T14:29:45.087390Z","shell.execute_reply.started":"2022-08-09T14:29:45.077036Z","shell.execute_reply":"2022-08-09T14:29:45.086156Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Analysing the explained variance ratio","metadata":{}},{"cell_type":"code","source":"pca.explained_variance_ratio_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.088778Z","iopub.execute_input":"2022-08-09T14:29:45.089491Z","iopub.status.idle":"2022-08-09T14:29:45.102831Z","shell.execute_reply.started":"2022-08-09T14:29:45.089449Z","shell.execute_reply":"2022-08-09T14:29:45.101425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"var_cumu = np.cumsum(pca.explained_variance_ratio_)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.104909Z","iopub.execute_input":"2022-08-09T14:29:45.105796Z","iopub.status.idle":"2022-08-09T14:29:45.112050Z","shell.execute_reply.started":"2022-08-09T14:29:45.105738Z","shell.execute_reply":"2022-08-09T14:29:45.110739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=[12,8])\nplt.vlines(x=15, ymax=1, ymin=0, colors=\"r\", linestyles=\"--\")\nplt.hlines(y=0.95, xmax=30, xmin=0, colors=\"g\", linestyles=\"--\")\nplt.plot(var_cumu)\nplt.ylabel(\"Cumulative variance explained\")\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.114084Z","iopub.execute_input":"2022-08-09T14:29:45.114867Z","iopub.status.idle":"2022-08-09T14:29:45.353290Z","shell.execute_reply.started":"2022-08-09T14:29:45.114799Z","shell.execute_reply":"2022-08-09T14:29:45.352220Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"we can use IncrementalPCA for the best result","metadata":{}},{"cell_type":"code","source":"pca_final = IncrementalPCA(n_components=16)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.354614Z","iopub.execute_input":"2022-08-09T14:29:45.354983Z","iopub.status.idle":"2022-08-09T14:29:45.360085Z","shell.execute_reply.started":"2022-08-09T14:29:45.354950Z","shell.execute_reply":"2022-08-09T14:29:45.359001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_pca = pca_final.fit_transform(X_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.361619Z","iopub.execute_input":"2022-08-09T14:29:45.361965Z","iopub.status.idle":"2022-08-09T14:29:45.629353Z","shell.execute_reply.started":"2022-08-09T14:29:45.361933Z","shell.execute_reply":"2022-08-09T14:29:45.627877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_pca.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.631246Z","iopub.execute_input":"2022-08-09T14:29:45.632539Z","iopub.status.idle":"2022-08-09T14:29:45.645179Z","shell.execute_reply.started":"2022-08-09T14:29:45.632480Z","shell.execute_reply":"2022-08-09T14:29:45.643852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corrmat = np.corrcoef(df_train_pca.transpose())","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.647244Z","iopub.execute_input":"2022-08-09T14:29:45.649774Z","iopub.status.idle":"2022-08-09T14:29:45.668205Z","shell.execute_reply.started":"2022-08-09T14:29:45.649714Z","shell.execute_reply":"2022-08-09T14:29:45.666591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"corrmat.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.670370Z","iopub.execute_input":"2022-08-09T14:29:45.672913Z","iopub.status.idle":"2022-08-09T14:29:45.686064Z","shell.execute_reply.started":"2022-08-09T14:29:45.672853Z","shell.execute_reply":"2022-08-09T14:29:45.684615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test_pca = pca_final.transform(X_test)\ndf_test_pca.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.688396Z","iopub.execute_input":"2022-08-09T14:29:45.695071Z","iopub.status.idle":"2022-08-09T14:29:45.721749Z","shell.execute_reply.started":"2022-08-09T14:29:45.694999Z","shell.execute_reply":"2022-08-09T14:29:45.720525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Applying logistic regression on the Principal components","metadata":{}},{"cell_type":"code","source":"learner_pca = LogisticRegression()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.723536Z","iopub.execute_input":"2022-08-09T14:29:45.725711Z","iopub.status.idle":"2022-08-09T14:29:45.733173Z","shell.execute_reply.started":"2022-08-09T14:29:45.725650Z","shell.execute_reply":"2022-08-09T14:29:45.731748Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_pca = learner_pca.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.734826Z","iopub.execute_input":"2022-08-09T14:29:45.736373Z","iopub.status.idle":"2022-08-09T14:29:45.966395Z","shell.execute_reply.started":"2022-08-09T14:29:45.736326Z","shell.execute_reply":"2022-08-09T14:29:45.964888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Making predictions on the test set\n","metadata":{}},{"cell_type":"code","source":"pred_probs_test = model_pca.predict_proba(df_test_pca)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.967947Z","iopub.execute_input":"2022-08-09T14:29:45.968728Z","iopub.status.idle":"2022-08-09T14:29:45.982132Z","shell.execute_reply.started":"2022-08-09T14:29:45.968683Z","shell.execute_reply":"2022-08-09T14:29:45.980472Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"{:2.2}\".format(metrics.roc_auc_score(y_test, pred_probs_test[:,1]))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:45.983785Z","iopub.execute_input":"2022-08-09T14:29:45.984484Z","iopub.status.idle":"2022-08-09T14:29:46.009163Z","shell.execute_reply.started":"2022-08-09T14:29:45.984433Z","shell.execute_reply":"2022-08-09T14:29:46.007665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Confusion matrix, Sensitivity and Specificity\n","metadata":{}},{"cell_type":"code","source":"pred_probs_test1 = model_pca.predict(df_test_pca)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.015979Z","iopub.execute_input":"2022-08-09T14:29:46.019699Z","iopub.status.idle":"2022-08-09T14:29:46.029711Z","shell.execute_reply.started":"2022-08-09T14:29:46.019622Z","shell.execute_reply":"2022-08-09T14:29:46.028326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Confusion matrix\nconfusion = metrics.confusion_matrix(y_test, pred_probs_test1)\nprint(confusion)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.033951Z","iopub.execute_input":"2022-08-09T14:29:46.037397Z","iopub.status.idle":"2022-08-09T14:29:46.053483Z","shell.execute_reply.started":"2022-08-09T14:29:46.037338Z","shell.execute_reply":"2022-08-09T14:29:46.052222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TP = confusion[1,1] # true positive \nTN = confusion[0,0] # true negatives\nFP = confusion[0,1] # false positives\nFN = confusion[1,0] # false negatives","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.058913Z","iopub.execute_input":"2022-08-09T14:29:46.063185Z","iopub.status.idle":"2022-08-09T14:29:46.074908Z","shell.execute_reply.started":"2022-08-09T14:29:46.063130Z","shell.execute_reply":"2022-08-09T14:29:46.073452Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Accuracy\nprint(\"Accuracy:-\",metrics.accuracy_score(y_test, pred_probs_test1))\n\n# Sensitivity\nprint(\"Sensitivity:-\",TP / float(TP+FN))\n\n# Specificity\nprint(\"Specificity:-\", TN / float(TN+FP))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.077420Z","iopub.execute_input":"2022-08-09T14:29:46.080531Z","iopub.status.idle":"2022-08-09T14:29:46.100677Z","shell.execute_reply.started":"2022-08-09T14:29:46.080469Z","shell.execute_reply":"2022-08-09T14:29:46.099200Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Making predictions on the train set","metadata":{}},{"cell_type":"code","source":"pred_probs_train = model_pca.predict_proba(df_train_pca)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.103504Z","iopub.execute_input":"2022-08-09T14:29:46.104529Z","iopub.status.idle":"2022-08-09T14:29:46.116048Z","shell.execute_reply.started":"2022-08-09T14:29:46.104471Z","shell.execute_reply":"2022-08-09T14:29:46.114350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\"{:2.2}\".format(metrics.roc_auc_score(y_train, pred_probs_train[:,1]))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.118778Z","iopub.execute_input":"2022-08-09T14:29:46.119981Z","iopub.status.idle":"2022-08-09T14:29:46.157914Z","shell.execute_reply.started":"2022-08-09T14:29:46.119920Z","shell.execute_reply":"2022-08-09T14:29:46.156462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Confusion matrix, Sensitivity and Specificity\n","metadata":{}},{"cell_type":"code","source":"pred_probs_train1 = model_pca.predict(df_train_pca)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.163118Z","iopub.execute_input":"2022-08-09T14:29:46.167929Z","iopub.status.idle":"2022-08-09T14:29:46.178958Z","shell.execute_reply.started":"2022-08-09T14:29:46.167871Z","shell.execute_reply":"2022-08-09T14:29:46.177259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Confusion matrix\nconfusion = metrics.confusion_matrix(y_train, pred_probs_train1)\nprint(confusion)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.185722Z","iopub.execute_input":"2022-08-09T14:29:46.186967Z","iopub.status.idle":"2022-08-09T14:29:46.209576Z","shell.execute_reply.started":"2022-08-09T14:29:46.186904Z","shell.execute_reply":"2022-08-09T14:29:46.208123Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"TP = confusion[1,1] # true positive \nTN = confusion[0,0] # true negatives\nFP = confusion[0,1] # false positives\nFN = confusion[1,0] # false negatives","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.214877Z","iopub.execute_input":"2022-08-09T14:29:46.215967Z","iopub.status.idle":"2022-08-09T14:29:46.229353Z","shell.execute_reply.started":"2022-08-09T14:29:46.215904Z","shell.execute_reply":"2022-08-09T14:29:46.227917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Accuracy\nprint(\"Accuracy:-\",metrics.accuracy_score(y_train, pred_probs_train1))\n\n# Sensitivity\nprint(\"Sensitivity:-\",TP / float(TP+FN))\n\n# Specificity\nprint(\"Specificity:-\", TN / float(TN+FP))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.232020Z","iopub.execute_input":"2022-08-09T14:29:46.233754Z","iopub.status.idle":"2022-08-09T14:29:46.261186Z","shell.execute_reply.started":"2022-08-09T14:29:46.233699Z","shell.execute_reply":"2022-08-09T14:29:46.254910Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Decision Tree with PCA","metadata":{}},{"cell_type":"code","source":"from sklearn.tree import DecisionTreeClassifier","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.263637Z","iopub.execute_input":"2022-08-09T14:29:46.264761Z","iopub.status.idle":"2022-08-09T14:29:46.273404Z","shell.execute_reply.started":"2022-08-09T14:29:46.264706Z","shell.execute_reply":"2022-08-09T14:29:46.272048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt = DecisionTreeClassifier(random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.275786Z","iopub.execute_input":"2022-08-09T14:29:46.278822Z","iopub.status.idle":"2022-08-09T14:29:46.291206Z","shell.execute_reply.started":"2022-08-09T14:29:46.278765Z","shell.execute_reply":"2022-08-09T14:29:46.289695Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.293613Z","iopub.execute_input":"2022-08-09T14:29:46.294594Z","iopub.status.idle":"2022-08-09T14:29:46.301261Z","shell.execute_reply.started":"2022-08-09T14:29:46.294561Z","shell.execute_reply":"2022-08-09T14:29:46.300317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'max_depth': [2, 3, 5, 10, 20],\n    'min_samples_leaf': [5, 10, 20, 50, 100],\n    'min_samples_split': [50, 150, 50]\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.302592Z","iopub.execute_input":"2022-08-09T14:29:46.303172Z","iopub.status.idle":"2022-08-09T14:29:46.312128Z","shell.execute_reply.started":"2022-08-09T14:29:46.303140Z","shell.execute_reply":"2022-08-09T14:29:46.311312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Instantiate the grid search model\ngrid_search = GridSearchCV(estimator=dt, \n                           param_grid=params, \n                           cv=4, n_jobs=-1, verbose=1, scoring = \"accuracy\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.313414Z","iopub.execute_input":"2022-08-09T14:29:46.313961Z","iopub.status.idle":"2022-08-09T14:29:46.323131Z","shell.execute_reply.started":"2022-08-09T14:29:46.313910Z","shell.execute_reply":"2022-08-09T14:29:46.322249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:29:46.324426Z","iopub.execute_input":"2022-08-09T14:29:46.324993Z","iopub.status.idle":"2022-08-09T14:30:04.937029Z","shell.execute_reply.started":"2022-08-09T14:29:46.324962Z","shell.execute_reply":"2022-08-09T14:30:04.935903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score_df = pd.DataFrame(grid_search.cv_results_)\nscore_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:04.938920Z","iopub.execute_input":"2022-08-09T14:30:04.939270Z","iopub.status.idle":"2022-08-09T14:30:04.966822Z","shell.execute_reply.started":"2022-08-09T14:30:04.939235Z","shell.execute_reply":"2022-08-09T14:30:04.965583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score_df.nlargest(5,\"mean_test_score\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:04.969642Z","iopub.execute_input":"2022-08-09T14:30:04.969958Z","iopub.status.idle":"2022-08-09T14:30:04.995230Z","shell.execute_reply.started":"2022-08-09T14:30:04.969930Z","shell.execute_reply":"2022-08-09T14:30:04.994207Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:04.996764Z","iopub.execute_input":"2022-08-09T14:30:04.997198Z","iopub.status.idle":"2022-08-09T14:30:05.005047Z","shell.execute_reply.started":"2022-08-09T14:30:04.997159Z","shell.execute_reply":"2022-08-09T14:30:05.003897Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_best = DecisionTreeClassifier( random_state = 42,\n                                  max_depth=10, \n                                  min_samples_leaf=20,\n                                  min_samples_split=50)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.006919Z","iopub.execute_input":"2022-08-09T14:30:05.007339Z","iopub.status.idle":"2022-08-09T14:30:05.013721Z","shell.execute_reply.started":"2022-08-09T14:30:05.007298Z","shell.execute_reply":"2022-08-09T14:30:05.012987Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dt_best.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.015693Z","iopub.execute_input":"2022-08-09T14:30:05.016637Z","iopub.status.idle":"2022-08-09T14:30:05.346555Z","shell.execute_reply.started":"2022-08-09T14:30:05.016597Z","shell.execute_reply":"2022-08-09T14:30:05.345377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import confusion_matrix, accuracy_score","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.348109Z","iopub.execute_input":"2022-08-09T14:30:05.348436Z","iopub.status.idle":"2022-08-09T14:30:05.353120Z","shell.execute_reply.started":"2022-08-09T14:30:05.348405Z","shell.execute_reply":"2022-08-09T14:30:05.352096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def evaluate_model(dt_classifier):\n    print(\"Train Accuracy :\", accuracy_score(y_train, dt_classifier.predict(df_train_pca)))\n    print(\"Train Confusion Matrix:\")\n    print(confusion_matrix(y_train, dt_classifier.predict(df_train_pca)))\n    print(\"-\"*50)\n    print(\"Test Accuracy :\", accuracy_score(y_test, dt_classifier.predict(df_test_pca)))\n    print(\"Test Confusion Matrix:\")\n    print(confusion_matrix(y_test, dt_classifier.predict(df_test_pca)))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.354384Z","iopub.execute_input":"2022-08-09T14:30:05.354669Z","iopub.status.idle":"2022-08-09T14:30:05.364675Z","shell.execute_reply.started":"2022-08-09T14:30:05.354641Z","shell.execute_reply":"2022-08-09T14:30:05.364009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"evaluate_model(dt_best)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.365583Z","iopub.execute_input":"2022-08-09T14:30:05.366253Z","iopub.status.idle":"2022-08-09T14:30:05.391619Z","shell.execute_reply.started":"2022-08-09T14:30:05.366222Z","shell.execute_reply":"2022-08-09T14:30:05.390178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##  Random Forest with PCA","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.393281Z","iopub.execute_input":"2022-08-09T14:30:05.393966Z","iopub.status.idle":"2022-08-09T14:30:05.399456Z","shell.execute_reply.started":"2022-08-09T14:30:05.393923Z","shell.execute_reply":"2022-08-09T14:30:05.398015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_features = int(round(np.sqrt(X_train.shape[1])))    # number of variables to consider to split each node\nprint(max_features)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.401157Z","iopub.execute_input":"2022-08-09T14:30:05.402450Z","iopub.status.idle":"2022-08-09T14:30:05.414301Z","shell.execute_reply.started":"2022-08-09T14:30:05.402403Z","shell.execute_reply":"2022-08-09T14:30:05.412773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf = RandomForestClassifier(n_estimators=100, max_depth=4, max_features=7, random_state=100, oob_score=True, verbose=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.416213Z","iopub.execute_input":"2022-08-09T14:30:05.417336Z","iopub.status.idle":"2022-08-09T14:30:05.429316Z","shell.execute_reply.started":"2022-08-09T14:30:05.417293Z","shell.execute_reply":"2022-08-09T14:30:05.427227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:05.434513Z","iopub.execute_input":"2022-08-09T14:30:05.435865Z","iopub.status.idle":"2022-08-09T14:30:09.956513Z","shell.execute_reply.started":"2022-08-09T14:30:05.435782Z","shell.execute_reply":"2022-08-09T14:30:09.955490Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf.oob_score_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:09.957934Z","iopub.execute_input":"2022-08-09T14:30:09.958252Z","iopub.status.idle":"2022-08-09T14:30:09.965121Z","shell.execute_reply.started":"2022-08-09T14:30:09.958222Z","shell.execute_reply":"2022-08-09T14:30:09.964263Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import plot_roc_curve","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:09.966745Z","iopub.execute_input":"2022-08-09T14:30:09.967055Z","iopub.status.idle":"2022-08-09T14:30:09.975136Z","shell.execute_reply.started":"2022-08-09T14:30:09.967027Z","shell.execute_reply":"2022-08-09T14:30:09.974148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_roc_curve(rf, df_train_pca, y_train)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:09.977389Z","iopub.execute_input":"2022-08-09T14:30:09.977785Z","iopub.status.idle":"2022-08-09T14:30:10.326908Z","shell.execute_reply.started":"2022-08-09T14:30:09.977733Z","shell.execute_reply":"2022-08-09T14:30:10.326150Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Hyper-parameter tuning for the Random Forest","metadata":{}},{"cell_type":"code","source":"rf = RandomForestClassifier(random_state=42, n_jobs=-1)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:10.328122Z","iopub.execute_input":"2022-08-09T14:30:10.329046Z","iopub.status.idle":"2022-08-09T14:30:10.333789Z","shell.execute_reply.started":"2022-08-09T14:30:10.329012Z","shell.execute_reply":"2022-08-09T14:30:10.332729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"params = {\n    'max_depth': [2,3,5],\n    'min_samples_leaf': [50,100],\n    'min_samples_split': [ 100, 150, ],\n    'n_estimators': [100, 200 ]\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:10.335317Z","iopub.execute_input":"2022-08-09T14:30:10.336326Z","iopub.status.idle":"2022-08-09T14:30:10.345976Z","shell.execute_reply.started":"2022-08-09T14:30:10.336283Z","shell.execute_reply":"2022-08-09T14:30:10.345001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search = GridSearchCV(estimator=rf,\n                           param_grid=params,\n                           cv = 4,\n                           n_jobs=-1, verbose=1, scoring=\"accuracy\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:10.347146Z","iopub.execute_input":"2022-08-09T14:30:10.347972Z","iopub.status.idle":"2022-08-09T14:30:10.356936Z","shell.execute_reply.started":"2022-08-09T14:30:10.347928Z","shell.execute_reply":"2022-08-09T14:30:10.356037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:30:10.358437Z","iopub.execute_input":"2022-08-09T14:30:10.359053Z","iopub.status.idle":"2022-08-09T14:31:41.190076Z","shell.execute_reply.started":"2022-08-09T14:30:10.359022Z","shell.execute_reply":"2022-08-09T14:31:41.188970Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:31:41.191798Z","iopub.execute_input":"2022-08-09T14:31:41.192473Z","iopub.status.idle":"2022-08-09T14:31:41.198968Z","shell.execute_reply.started":"2022-08-09T14:31:41.192437Z","shell.execute_reply":"2022-08-09T14:31:41.197994Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:31:41.200581Z","iopub.execute_input":"2022-08-09T14:31:41.200906Z","iopub.status.idle":"2022-08-09T14:31:41.211316Z","shell.execute_reply.started":"2022-08-09T14:31:41.200875Z","shell.execute_reply":"2022-08-09T14:31:41.210590Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rfc_model = RandomForestClassifier(bootstrap=True,\n                             max_depth=5,\n                             min_samples_leaf=50, \n                             min_samples_split=100,\n                             n_estimators=200)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:31:41.212147Z","iopub.execute_input":"2022-08-09T14:31:41.212438Z","iopub.status.idle":"2022-08-09T14:31:41.221275Z","shell.execute_reply.started":"2022-08-09T14:31:41.212401Z","shell.execute_reply":"2022-08-09T14:31:41.220435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rfc_model.fit(df_train_pca, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:31:41.222228Z","iopub.execute_input":"2022-08-09T14:31:41.222547Z","iopub.status.idle":"2022-08-09T14:31:47.515499Z","shell.execute_reply.started":"2022-08-09T14:31:41.222519Z","shell.execute_reply":"2022-08-09T14:31:47.514711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"evaluate_model(rfc_model)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:31:47.516705Z","iopub.execute_input":"2022-08-09T14:31:47.517192Z","iopub.status.idle":"2022-08-09T14:31:48.304188Z","shell.execute_reply.started":"2022-08-09T14:31:47.517159Z","shell.execute_reply":"2022-08-09T14:31:48.303429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rfc_model.feature_importances_","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:34.035580Z","iopub.execute_input":"2022-08-09T14:57:34.036069Z","iopub.status.idle":"2022-08-09T14:57:34.070016Z","shell.execute_reply.started":"2022-08-09T14:57:34.036026Z","shell.execute_reply":"2022-08-09T14:57:34.068750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Note:\n\nNote that the best parameters procuded the accuracy of 91% which is not significantly deterred than the accuracy of original random forest, which is pegged around 92%\n\n## Conclusion :\n\nThe best model to predict the churn is observed to be Random Forest based on the accuracy as performance measure.\n\n\nThe incoming calls (with local same operator mobile/other operator mobile/fixed lines, STD or Special) plays a vital role in understanding the possibility of churn. Hence, the operator should focus on incoming calls data and has to provide some kind of special offers to the customers whose incoming calls turning lower.\n\n## Details:\n\n After cleaning the data, we broadly employed three models as mentioned below including some variations within these models in order to arrive at the best model in each of the cases.\n\n### Logistic Regression  :\n\nLogistic Regression with RFE Logistic regression with PCA Random Forest For each of these models, the summary of performance measures are as follows:\n\n#### Logistic Regression\n\n.  Train Accuracy : ~90%\n. Test Accuracy : ~88%\n\n#### Logistic regression with PCA\n\n. Train Accuracy : ~92%\n. Test Accuracy : ~92%\n\n#### Decision Tree with PCA:\n\n. Train Accuracy : ~94%\n. Test Accuracy : ~93%\n\n\n#### Random Forest with PCA:\n. Train Accuracy :~ 92%\n. Test Accuracy :~ 92%","metadata":{}},{"cell_type":"code","source":"churn_test = pd.read_csv(\"../input/telecom-churn-case-study-hackathon-38/test (1).csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:37.738461Z","iopub.execute_input":"2022-08-09T14:57:37.738851Z","iopub.status.idle":"2022-08-09T14:57:38.624105Z","shell.execute_reply.started":"2022-08-09T14:57:37.738806Z","shell.execute_reply":"2022-08-09T14:57:38.622690Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:46.646780Z","iopub.execute_input":"2022-08-09T14:57:46.647978Z","iopub.status.idle":"2022-08-09T14:57:46.793588Z","shell.execute_reply.started":"2022-08-09T14:57:46.647933Z","shell.execute_reply":"2022-08-09T14:57:46.792310Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:51.656813Z","iopub.execute_input":"2022-08-09T14:57:51.659534Z","iopub.status.idle":"2022-08-09T14:57:51.667553Z","shell.execute_reply.started":"2022-08-09T14:57:51.659491Z","shell.execute_reply":"2022-08-09T14:57:51.666413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:53.614809Z","iopub.execute_input":"2022-08-09T14:57:53.615499Z","iopub.status.idle":"2022-08-09T14:57:53.647772Z","shell.execute_reply.started":"2022-08-09T14:57:53.615462Z","shell.execute_reply":"2022-08-09T14:57:53.646455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_id = churn_test['id']","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:57:59.358936Z","iopub.execute_input":"2022-08-09T14:57:59.359349Z","iopub.status.idle":"2022-08-09T14:57:59.365157Z","shell.execute_reply.started":"2022-08-09T14:57:59.359314Z","shell.execute_reply":"2022-08-09T14:57:59.363919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test['tenure'] = (churn_test['aon']/30).round(0)\nchurn_test[\"avg_arpu_6_7\"]= (churn_test['arpu_6']+churn_test['arpu_7'])/2\n\nchurn_test = churn_test[X.columns]","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:01.693262Z","iopub.execute_input":"2022-08-09T14:58:01.693647Z","iopub.status.idle":"2022-08-09T14:58:01.793985Z","shell.execute_reply.started":"2022-08-09T14:58:01.693614Z","shell.execute_reply":"2022-08-09T14:58:01.792440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:04.432031Z","iopub.execute_input":"2022-08-09T14:58:04.432467Z","iopub.status.idle":"2022-08-09T14:58:04.440669Z","shell.execute_reply.started":"2022-08-09T14:58:04.432432Z","shell.execute_reply":"2022-08-09T14:58:04.439090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test_null = churn_test.isnull().sum().sum() / np.product(churn_test.shape) * 100\nchurn_test_null","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:06.488557Z","iopub.execute_input":"2022-08-09T14:58:06.489028Z","iopub.status.idle":"2022-08-09T14:58:06.501372Z","shell.execute_reply.started":"2022-08-09T14:58:06.488992Z","shell.execute_reply":"2022-08-09T14:58:06.500090Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in churn_test.columns:\n    null_col = churn_test[col].isnull().sum() / churn_test.shape[0] * 100\n    print(\"{} : {:.2f}\".format(col,null_col))","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:09.362147Z","iopub.execute_input":"2022-08-09T14:58:09.362607Z","iopub.status.idle":"2022-08-09T14:58:09.388084Z","shell.execute_reply.started":"2022-08-09T14:58:09.362569Z","shell.execute_reply":"2022-08-09T14:58:09.387014Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in churn_test.columns:\n    null_col = churn_test[col].isnull().sum() / churn_test.shape[0] * 100\n    if null_col > 0:\n        churn_test[col] = churn_test[col].fillna(churn_test[col].mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:12.784645Z","iopub.execute_input":"2022-08-09T14:58:12.785647Z","iopub.status.idle":"2022-08-09T14:58:12.837266Z","shell.execute_reply.started":"2022-08-09T14:58:12.785605Z","shell.execute_reply":"2022-08-09T14:58:12.836292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test.isnull().sum().sum()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:15.214137Z","iopub.execute_input":"2022-08-09T14:58:15.215162Z","iopub.status.idle":"2022-08-09T14:58:15.227116Z","shell.execute_reply.started":"2022-08-09T14:58:15.215109Z","shell.execute_reply":"2022-08-09T14:58:15.226104Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test_final = pca_final.transform(churn_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:17.578770Z","iopub.execute_input":"2022-08-09T14:58:17.579937Z","iopub.status.idle":"2022-08-09T14:58:17.609765Z","shell.execute_reply.started":"2022-08-09T14:58:17.579878Z","shell.execute_reply":"2022-08-09T14:58:17.608184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"churn_test_final.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:21.764534Z","iopub.execute_input":"2022-08-09T14:58:21.765351Z","iopub.status.idle":"2022-08-09T14:58:21.771177Z","shell.execute_reply.started":"2022-08-09T14:58:21.765308Z","shell.execute_reply":"2022-08-09T14:58:21.770434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_probalbilty = rfc_model.predict(churn_test_final)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:24.535569Z","iopub.execute_input":"2022-08-09T14:58:24.536470Z","iopub.status.idle":"2022-08-09T14:58:25.006635Z","shell.execute_reply.started":"2022-08-09T14:58:24.536431Z","shell.execute_reply":"2022-08-09T14:58:25.005752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_probalbilty.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:26.658996Z","iopub.execute_input":"2022-08-09T14:58:26.659398Z","iopub.status.idle":"2022-08-09T14:58:26.666696Z","shell.execute_reply.started":"2022-08-09T14:58:26.659362Z","shell.execute_reply":"2022-08-09T14:58:26.665360Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(churn_id)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:28.813124Z","iopub.execute_input":"2022-08-09T14:58:28.813912Z","iopub.status.idle":"2022-08-09T14:58:28.820560Z","shell.execute_reply.started":"2022-08-09T14:58:28.813872Z","shell.execute_reply":"2022-08-09T14:58:28.819512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_prediction = pd.DataFrame({'id':churn_id,'churn_probability':predict_probalbilty})","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:31.385311Z","iopub.execute_input":"2022-08-09T14:58:31.386224Z","iopub.status.idle":"2022-08-09T14:58:31.391727Z","shell.execute_reply.started":"2022-08-09T14:58:31.386183Z","shell.execute_reply":"2022-08-09T14:58:31.390557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_prediction.to_csv('submission.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-09T14:58:36.467392Z","iopub.execute_input":"2022-08-09T14:58:36.467822Z","iopub.status.idle":"2022-08-09T14:58:36.507379Z","shell.execute_reply.started":"2022-08-09T14:58:36.467788Z","shell.execute_reply":"2022-08-09T14:58:36.506176Z"},"trusted":true},"execution_count":null,"outputs":[]}]}