{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### <span style=\"color:#458B74;font-weight: bolder;font-family: cursive;font-size: 24px\"> <center>“Computers are able to see, hear and learn.<br> Welcome to the future”.<br> -Dave Walters- </center> </span> ","metadata":{}},{"cell_type":"markdown","source":" <p style=\"color:#458B74;text-align:center;border-radius:10px 10px;font-weight:bold;border:2px dotted #458B74;font-size:22px\"> ✨Spaceship-Titanic : EDA & classification models ✨<span style='font-size:28px; background-color:blue ;'></span></p>","metadata":{}},{"cell_type":"markdown","source":"<img src= \"https://i.ytimg.com/vi/bY4nfC1knJc/maxresdefault.jpg\" alt =\"Titanic\" style='width: 100%;height:300px'>","metadata":{}},{"cell_type":"markdown","source":"### Table of Contents\n\n* [**Task and probleme statement**](#Chapter1)\n* [**EDA and correlation anlysis**](#Chapter2)\n    * [**Managing Missing Values**](#section_2_1)\n    * [**Univariate Analysis**](#section_2_2) \n        * [**Numerical Data**](#section_2_2_1)\n        * [**Categorical Data**](#section_2_2_2)\n    * [**Bivariate Analysis**](#section_2_3)\n        * [**Numerical-Numerical analysis**](#section_2_3_1)\n        * [**Numerical-Categorical analysis**](#section_2_3_2)\n        * [**Categorical-Categorial analysis**](#section_2_3_3)\n    * [**Managing Outliers**](#section_2_4)\n    * [**Features Seletion using statistical tests**](#section_2_5)  \n        * [**Anova Test**](#section_2_5_1)\n        * [**Chi-square Test**](#section_2_5_2)\n* [**Model Building**](#Chapter3)\n* [**Model choice and submission**](#Chapter4)\n    ","metadata":{}},{"cell_type":"markdown","source":"<span style=\" padding: 5px; border-radius:5px;color:red;font-weight: bolder;font-size: 24px;font-family: cursive\">If you like my work, don't forget to upvote and leave me a comment.This will help me to continue and sharing more topics in the coming days. I hope this work will impress you.</span>","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  1  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Task and probleme statement</span> <a class=\"anchor\" id=\"Chapter1\"></a>","metadata":{}},{"cell_type":"markdown","source":"In this competition our task is to predict whether a passenger was transported to an alternate dimension during the Spaceship Titanic's collision with the spacetime anomaly. To make these predictions, we're given a set of personal records recovered from the ship's damaged computer system.\n## 📁File and Data Field Descriptions\n* <span style=\"color:green;font-weight: bolder\">train.csv</span> - Personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n  * <span style=\"color:blue;font-weight: bolder\">PassengerId</span> - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n  * <span style=\"color:blue;font-weight: bolder\">HomePlanet</span> - The planet the passenger departed from, typically their planet of permanent residence.\n  * <span style=\"color:blue;font-weight: bolder\">Cryosleep</span> - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n  * <span style=\"color:blue;font-weight: bolder\">Cabin</span> - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n  * <span style=\"color:blue;font-weight: bolder\">Destination</span> - The planet the passenger will be debarking to.\n  * <span style=\"color:blue;font-weight: bolder\">Age</span> - The age of the passenger.\n  * <span style=\"color:blue;font-weight: bolder\">VIP</span> - Whether the passenger has paid for special VIP service during the voyage.\n  * <span style=\"color:blue;font-weight: bolder\">RoomService, FoodCourt, ShoppingMall, Spa, VRDeck</span> - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n  * <span style=\"color:blue;font-weight: bolder\">Name</span> - The first and last names of the passenger.\n  * <span style=\"color:blue;font-weight: bolder\">Transported</span> - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.\n* <span style=\"color:green;font-weight: bolder\">test.csv</span> - Personal records for the remaining one-third (~4300) of the passengers, to be used as test data. Your task is to predict the value of Transported for the passengers in this set.\n* <span style=\"color:green;font-weight: bolder\">sample_submission.csv</span> - A submission file in the correct format.\nPassengerId - Id for each passenger in the test set.\nTransported - The target. For each passenger, predict either True or False.","metadata":{}},{"cell_type":"code","source":"# Import Libraries\nimport numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\n\nfrom scipy import stats\nimport scipy.stats as stats\nfrom scipy.stats import chi2_contingency\nfrom scipy.stats import chi2\nimport missingno as msno\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import GridSearchCV\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom xgboost import XGBClassifier\nimport xgboost as xgb\nfrom lightgbm import LGBMClassifier\n\nfrom sklearn import metrics\nfrom sklearn.metrics import confusion_matrix, accuracy_score, classification_report,plot_confusion_matrix\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        \nimport warnings\nwarnings.filterwarnings('always')\nwarnings.filterwarnings('ignore')\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-06-24T10:32:49.403044Z","iopub.execute_input":"2022-06-24T10:32:49.404152Z","iopub.status.idle":"2022-06-24T10:32:52.081541Z","shell.execute_reply.started":"2022-06-24T10:32:49.404095Z","shell.execute_reply":"2022-06-24T10:32:52.080499Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">EDA and Correlation analysis</span> <a class=\"anchor\" id=\"Chapter2\"></a>","metadata":{}},{"cell_type":"code","source":"#Data loading and overview\ndf_train=pd.read_csv('../input/spaceship-titanic/train.csv')\ndf_test=pd.read_csv('../input/spaceship-titanic/test.csv')\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.08402Z","iopub.execute_input":"2022-06-24T10:32:52.084559Z","iopub.status.idle":"2022-06-24T10:32:52.186131Z","shell.execute_reply.started":"2022-06-24T10:32:52.08451Z","shell.execute_reply":"2022-06-24T10:32:52.185143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Print the dimentions of training dataset and test dataset\nprint(f'The dataset df_train contains {df_train.shape[0]} rows and {df_train.shape[1]} columns : {df_train.shape}')\nprint(f'The dataset df_test contains {df_test.shape[0]} rows and {df_test.shape[1]} columns : {df_test.shape}')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.187587Z","iopub.execute_input":"2022-06-24T10:32:52.188215Z","iopub.status.idle":"2022-06-24T10:32:52.194601Z","shell.execute_reply.started":"2022-06-24T10:32:52.188171Z","shell.execute_reply":"2022-06-24T10:32:52.193394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Print summary of training dataset\ndf_train.info()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.195906Z","iopub.execute_input":"2022-06-24T10:32:52.196408Z","iopub.status.idle":"2022-06-24T10:32:52.237012Z","shell.execute_reply.started":"2022-06-24T10:32:52.196367Z","shell.execute_reply":"2022-06-24T10:32:52.235709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the table above we see that:\n* we have quite some categorical columns as well as numerical columns (6 numerial columns and 8 categorical columns ).\n* All columns have missing values exept for 'PassengerId' and 'Transported'. These missing values will be dealt with later in [**Managing Missing values**](#Chapter3) part.","metadata":{}},{"cell_type":"code","source":"#Plot data types\nresult=df_train.dtypes.value_counts(normalize=True)\nplt.figure(figsize=(13,5))\nplt.subplot(1,2,1)\nresult.plot(kind='pie',wedgeprops = { 'linewidth' : 3, 'edgecolor' : 'white' },colors=['blue', 'orange', 'green'], autopct=\"%.0f%%\",explode = (0.05, 0.05, 0.05))\nplt.subplot(1,2,2)\nax = result.plot(kind='bar',figsize=(15,4),width = 0.8,color=['blue', 'orange', 'green'],edgecolor=None)\nplt.xticks(fontsize=14)\nfor spine in plt.gca().spines.values():\n    spine.set_visible(False)\nplt.yticks([])\n# Add this loop to add the annotations\nfor p in ax.patches:\n    width = p.get_width()\n    height = p.get_height()\n    x, y = p.get_xy() \n    ax.annotate(f'{height:.0%}', (x + width/2, y + height*1.02), ha='center')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.240697Z","iopub.execute_input":"2022-06-24T10:32:52.241183Z","iopub.status.idle":"2022-06-24T10:32:52.505971Z","shell.execute_reply.started":"2022-06-24T10:32:52.241141Z","shell.execute_reply":"2022-06-24T10:32:52.504981Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Loking at the visualisations above, we see that **50%** of columns are object, **43%** are float and **7%** are booleans","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2.1  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Managing Missing Values</span> <a class=\"anchor\" id=\"section_2_1\"></a>","metadata":{}},{"cell_type":"code","source":"#Checking for missing values in each column of the training dataset\ndf_train.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.50712Z","iopub.execute_input":"2022-06-24T10:32:52.507458Z","iopub.status.idle":"2022-06-24T10:32:52.519773Z","shell.execute_reply.started":"2022-06-24T10:32:52.507428Z","shell.execute_reply":"2022-06-24T10:32:52.518553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.displot(\n    data=df_train.isna().melt(value_name=\"missing\"),\n    y=\"variable\",\n    hue=\"missing\",\n    multiple=\"fill\",\n    aspect=1.25\n)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:52.522286Z","iopub.execute_input":"2022-06-24T10:32:52.523429Z","iopub.status.idle":"2022-06-24T10:32:53.335994Z","shell.execute_reply.started":"2022-06-24T10:32:52.523378Z","shell.execute_reply":"2022-06-24T10:32:53.334605Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:53.337375Z","iopub.execute_input":"2022-06-24T10:32:53.337814Z","iopub.status.idle":"2022-06-24T10:32:54.270616Z","shell.execute_reply.started":"2022-06-24T10:32:53.337781Z","shell.execute_reply":"2022-06-24T10:32:54.269543Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Split df_train and df_test columns into numerical and categorical \nnum_features=[col for col in df_train.select_dtypes('number')]\ncateg_features=[col for col in df_train.select_dtypes(exclude=['number'])]\ntest_categ_features=[col for col in df_test.select_dtypes(exclude=['number'])]\n\n#replace missing values in each numerical column with the mediane\nfor col in num_features:\n    df_train[col].fillna(df_train[col].median(), inplace=True)\n    df_test[col].fillna(df_test[col].median(), inplace=True)\n    \n#replace missing values in each categorical column with the most frequent value\nfor col in categ_features:\n    df_train[col].fillna(df_train[col].value_counts().index[0], inplace=True)\nfor col in test_categ_features:\n    df_test[col].fillna(df_test[col].value_counts().index[0], inplace=True)   ","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.271989Z","iopub.execute_input":"2022-06-24T10:32:54.27245Z","iopub.status.idle":"2022-06-24T10:32:54.329381Z","shell.execute_reply.started":"2022-06-24T10:32:54.272411Z","shell.execute_reply":"2022-06-24T10:32:54.328494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Checking for missing values in df_train after imputation\ndf_train.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.330438Z","iopub.execute_input":"2022-06-24T10:32:54.330927Z","iopub.status.idle":"2022-06-24T10:32:54.343486Z","shell.execute_reply.started":"2022-06-24T10:32:54.330899Z","shell.execute_reply":"2022-06-24T10:32:54.342519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Checking for missing values in df_test after imputation\ndf_test.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.345087Z","iopub.execute_input":"2022-06-24T10:32:54.346145Z","iopub.status.idle":"2022-06-24T10:32:54.359641Z","shell.execute_reply.started":"2022-06-24T10:32:54.346098Z","shell.execute_reply":"2022-06-24T10:32:54.358785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2.2  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Univariate Analysis</span> <a class=\"anchor\" id=\"Section_2_2\"></a>","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Numerical Data</span> <a class=\"anchor\" id=\"section_2_2_1\"></a>","metadata":{}},{"cell_type":"code","source":"#Print numerical columns\nprint(\"numerical columns are : \",num_features)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.361018Z","iopub.execute_input":"2022-06-24T10:32:54.361637Z","iopub.status.idle":"2022-06-24T10:32:54.374516Z","shell.execute_reply.started":"2022-06-24T10:32:54.361603Z","shell.execute_reply":"2022-06-24T10:32:54.373366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Print summary of numerical columns\ndf_train.describe()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.376024Z","iopub.execute_input":"2022-06-24T10:32:54.376416Z","iopub.status.idle":"2022-06-24T10:32:54.422022Z","shell.execute_reply.started":"2022-06-24T10:32:54.376371Z","shell.execute_reply":"2022-06-24T10:32:54.421044Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From the above informations about numericals features we see that:\n* The average age of passenger is 29 and the oldest passenger has 79 years.\n* The average Amount the passenger has billed for room service is around 225,and the heighest amount is around 14327.\n*  The average Amount the passenger has billed at the food court is around 458, and the heighest amount is around 29813.\n* The average Amount the passenger has billed at the shopping mall is around 174, and the heighest amount is around 23492.\n* The average Amount the passenger has billed at the spa is around 311, and the heighest amount is around 22408.\n* The average Amount the passenger has billed at the VR deck is around 305, and the heighest amount is around 24133.\n* There are somme passengers who hasn't billed any amount.\n\n","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:blue\">How many passengers who have not billed any amount? what's their average age?</span>","metadata":{}},{"cell_type":"code","source":"data_amount=df_train[(df_train['RoomService']==0) & (df_train['Spa']==0) & (df_train['VRDeck']==0) & (df_train['FoodCourt']==0) & (df_train['ShoppingMall']==0)]\nnbr=data_amount.shape[0]\naverage_age=data_amount['Age'].mean().round()\nprint(f'There are {nbr} passengers who have not billed any amount, and their average age is {average_age}')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.424706Z","iopub.execute_input":"2022-06-24T10:32:54.426938Z","iopub.status.idle":"2022-06-24T10:32:54.438062Z","shell.execute_reply.started":"2022-06-24T10:32:54.426894Z","shell.execute_reply":"2022-06-24T10:32:54.436932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> ### <span style=\"color:blue\">What's the most frequent age categorie?</span>","metadata":{}},{"cell_type":"markdown","source":"In this setion we'll devide the age variable into 4 main sections in order to identify the most frequent age categorie.\n* **Children (00-14 years)**\n* **Youth (15-24 years)**\n* **Adults (25-64 years)**\n* **Seniors (65 years and over)**","metadata":{}},{"cell_type":"code","source":"Children_nbr=df_train[df_train['Age']<=14].shape[0]\nYouth_nbr=df_train[(df_train['Age']>=15)&(df_train['Age']<=24)].shape[0]\nAdult_nbr=df_train[(df_train['Age']>=25)&(df_train['Age']<=64)].shape[0]\nSeniors_nbr=df_train[(df_train['Age']>=65)].shape[0]\nprint('Number of children passengers on titanic spaceship is :',Children_nbr )\nprint('Number of young passengers on titanic spaceship is :',Youth_nbr )\nprint('Number of adults passengers on titanic spaceship is :',Adult_nbr )\nprint('Number of seniors passengers on titanic spaceship is :',Seniors_nbr )\n","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.439533Z","iopub.execute_input":"2022-06-24T10:32:54.440013Z","iopub.status.idle":"2022-06-24T10:32:54.45479Z","shell.execute_reply.started":"2022-06-24T10:32:54.439955Z","shell.execute_reply":"2022-06-24T10:32:54.453603Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that the majority of passengers are adult with age between 24 and 64.","metadata":{}},{"cell_type":"code","source":"# creating the bar plot\nsorted_list = [(107,'Seniors'),(1085, 'Children'),(2568, 'Youth'),(4933, 'Adults')]\nfeatures_sorted = []\nimportance_sorted = []\n\nfor i in sorted_list:\n    features_sorted += [i[1]]\n    importance_sorted += [i[0]]\n\nplt.figure(figsize=(10,6))\nplt.title(\"Categories\", fontsize=15)\nplt.xlabel(\"number of passengers\", fontsize=13)\n\nplt.barh(features_sorted, importance_sorted, color=sns.color_palette(\"inferno_r\", 7), edgecolor='green', height=0.4)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.456148Z","iopub.execute_input":"2022-06-24T10:32:54.456731Z","iopub.status.idle":"2022-06-24T10:32:54.64161Z","shell.execute_reply.started":"2022-06-24T10:32:54.456666Z","shell.execute_reply":"2022-06-24T10:32:54.640567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plot distribution of numerical columns\nfig = plt.figure(figsize= (12,4))\nfor i, col in enumerate(num_features):\n    \n    ax=fig.add_subplot( 2, 3, i+1)\n    \n    sns.distplot(df_train[col])\n\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:54.642893Z","iopub.execute_input":"2022-06-24T10:32:54.643453Z","iopub.status.idle":"2022-06-24T10:32:56.108322Z","shell.execute_reply.started":"2022-06-24T10:32:54.643415Z","shell.execute_reply":"2022-06-24T10:32:56.107469Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at the distriutions of each numerical feature, we see that:\n* The distribution of age is bimodale, because it has two peaks \n* The distribution of the variables: RoomServie, FoodCourt, ShoppingMall, Spa, RDeck is skewed, we can confirme that using skew function.","metadata":{}},{"cell_type":"code","source":"# Skew function of Pandas\nskew = df_train[num_features].skew(skipna = True).sort_values(ascending=False)\nskew","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.109388Z","iopub.execute_input":"2022-06-24T10:32:56.109756Z","iopub.status.idle":"2022-06-24T10:32:56.120154Z","shell.execute_reply.started":"2022-06-24T10:32:56.109725Z","shell.execute_reply":"2022-06-24T10:32:56.119301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see all the skewness values are greater than 0, we can understand that our numerical features : RoomServie, FoodCourt, ShoppingMall, Spa, RDeck are right-skewed.\n\nwe can say that these variables have outliers that skew the data. These values will be dealt with later in [**Managing Outliers**](#section_2_4) part.","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Categorical Data</span> <a class=\"anchor\" id=\"section_2_2_2\"></a>","metadata":{}},{"cell_type":"code","source":"df_train['Transported'].replace(to_replace=[False,True],value=['No','Yes'],inplace=True)\ndf_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.121735Z","iopub.execute_input":"2022-06-24T10:32:56.122179Z","iopub.status.idle":"2022-06-24T10:32:56.147542Z","shell.execute_reply.started":"2022-06-24T10:32:56.122136Z","shell.execute_reply":"2022-06-24T10:32:56.146739Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Print summary of categorical columns in training dataset\ndf_train.describe(include='object')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.14948Z","iopub.execute_input":"2022-06-24T10:32:56.149925Z","iopub.status.idle":"2022-06-24T10:32:56.189456Z","shell.execute_reply.started":"2022-06-24T10:32:56.149868Z","shell.execute_reply":"2022-06-24T10:32:56.188372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at the informations above we see that:\n* All categorical columns have missing values, exept the target variable **\"Transported\"**.\n* the vast majority of passengers were traveling to **TRAPPIST-1e**.\n* Most passengers are from **Earth** .\n* Most of passengers have paid for special **VIP** service during the voyage.\n* The variable **cabin** has 6560 categories, therefore, we cannot use it directly for our model.\n","metadata":{}},{"cell_type":"code","source":"#print categories of each categorical column\nfor col in df_train.select_dtypes(exclude=['number']):\n  print(f'{col:-<30},{df_train[col].unique()}')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.191045Z","iopub.execute_input":"2022-06-24T10:32:56.191415Z","iopub.status.idle":"2022-06-24T10:32:56.207161Z","shell.execute_reply.started":"2022-06-24T10:32:56.191385Z","shell.execute_reply":"2022-06-24T10:32:56.206089Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Creating new features for training dataset and test dataset\n\n#Cabin has three values deck/num/side, so we'll create two columns for deck and side\ndf_train[\"Deck\"] = df_train[\"Cabin\"].apply(lambda x: str(x).split(\"/\")[0])\ndf_test[\"Deck\"] = df_test[\"Cabin\"].apply(lambda x: str(x).split(\"/\")[0])\ndf_train[\"side\"] =df_train[\"Cabin\"].apply(lambda x: x.split(\"/\")[2])\ndf_test[\"side\"] = df_test[\"Cabin\"].apply(lambda x: x.split(\"/\")[2])\n\n#Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group, so we'll create two new features GroupId and GroupIdNumber\ndf_train[\"GroupId\"] = df_train[\"PassengerId\"].apply(lambda x: x.split(\"_\")[0])\ndf_test[\"GroupId\"] = df_test[\"PassengerId\"].apply(lambda x: x.split(\"_\")[0])\ndf_train[\"GroupIdNumber\"] =df_train[\"PassengerId\"].apply(lambda x: x.split(\"_\")[1])\ndf_test[\"GroupIdNumber\"] = df_test[\"PassengerId\"].apply(lambda x: x.split(\"_\")[1])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.208348Z","iopub.execute_input":"2022-06-24T10:32:56.208745Z","iopub.status.idle":"2022-06-24T10:32:56.247213Z","shell.execute_reply.started":"2022-06-24T10:32:56.208715Z","shell.execute_reply":"2022-06-24T10:32:56.246372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating new feature InGroup to indicate if a passenger is alone or in group\nGroup_train =df_train[df_train[\"GroupId\"].duplicated()][\"GroupId\"]\nGroup_test =df_test[df_test[\"GroupId\"].duplicated()][\"GroupId\"]\ndf_train[\"InGroup\"] = df_train[\"GroupId\"].apply(lambda x: x in Group_train.values)\ndf_test[\"InGroup\"] = df_test[\"GroupId\"].apply(lambda x: x in Group_test.values)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.248863Z","iopub.execute_input":"2022-06-24T10:32:56.249658Z","iopub.status.idle":"2022-06-24T10:32:56.889796Z","shell.execute_reply.started":"2022-06-24T10:32:56.249612Z","shell.execute_reply":"2022-06-24T10:32:56.888731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Drop 'PassengerId','Cabin','Name','GroupId','GroupIdNumber' from df_train\ndf_train.drop(['PassengerId','Cabin','Name','GroupId','GroupIdNumber'], axis=1, inplace=True)\n#Save PassengerId and Name\nId_test_list = df_test[\"PassengerId\"].tolist()\n#Drop 'PassengerId','Cabin','Name','GroupId','GroupIdNumber' from df_test\ndf_test.drop(['PassengerId','Cabin','Name','GroupId','GroupIdNumber'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.891211Z","iopub.execute_input":"2022-06-24T10:32:56.891792Z","iopub.status.idle":"2022-06-24T10:32:56.903261Z","shell.execute_reply.started":"2022-06-24T10:32:56.891745Z","shell.execute_reply":"2022-06-24T10:32:56.902289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#print categories of each categorical column after removing unnecessary columns\nfor col in df_train.select_dtypes(exclude=['number']):\n  print(f'{col:-<30},{df_train[col].unique()}')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.904499Z","iopub.execute_input":"2022-06-24T10:32:56.904826Z","iopub.status.idle":"2022-06-24T10:32:56.919374Z","shell.execute_reply.started":"2022-06-24T10:32:56.904796Z","shell.execute_reply":"2022-06-24T10:32:56.918249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Print categorical columns in df_train after removing unnecessary columns\nnew_categ_features=[col for col in df_train.select_dtypes(exclude=['number'])]","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.921062Z","iopub.execute_input":"2022-06-24T10:32:56.921538Z","iopub.status.idle":"2022-06-24T10:32:56.930745Z","shell.execute_reply.started":"2022-06-24T10:32:56.921493Z","shell.execute_reply":"2022-06-24T10:32:56.92998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"colors =sns.color_palette(\"inferno_r\", 7)\nfig = plt.figure(figsize= (15,9))\nfor i, col in enumerate(new_categ_features):\n    \n    ax=fig.add_subplot(3, 3, i+1)\n    \n    sns.countplot(x=df_train[col],palette=colors, ax=ax)\n\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:56.931661Z","iopub.execute_input":"2022-06-24T10:32:56.932477Z","iopub.status.idle":"2022-06-24T10:32:57.783871Z","shell.execute_reply.started":"2022-06-24T10:32:56.932431Z","shell.execute_reply":"2022-06-24T10:32:57.782649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Most passengers were traveling on both Deck F, G and on Starboard-side.\n* More than half of passengers choose to travel alone.\n* Few passengers were in CryoSleep.\n* Passengers had an overall even chance of being transported\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"#Plot donut chart for each categorical column\nfig = plt.figure(figsize= (16,10))\nfor i, col in enumerate(new_categ_features):\n    \n    ax=fig.add_subplot( 3, 3, i+1)\n    \n    df_train[col].value_counts().plot.pie(autopct='%.0f%%', pctdistance=0.80, colors=sns.color_palette(\"inferno_r\", 7))\n    # draw circle\n    centre_circle = plt.Circle((0, 0), 0.60, fc='white')\n    fig1 = plt.gcf()\n    # Adding Circle in Pie chart\n    fig1.gca().add_artist(centre_circle)\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:57.785414Z","iopub.execute_input":"2022-06-24T10:32:57.78588Z","iopub.status.idle":"2022-06-24T10:32:58.938329Z","shell.execute_reply.started":"2022-06-24T10:32:57.785838Z","shell.execute_reply":"2022-06-24T10:32:58.937296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2.3  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Bivariate Analysis</span> <a class=\"anchor\" id=\"section_2_3\"></a>","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Numerical-Numerical Analysis</span> <a class=\"anchor\" id=\"section_2_3_1\"></a>","metadata":{}},{"cell_type":"code","source":"#plot the pair plot of numerical features\nsns.pairplot(data = df_train, vars=num_features, diag_kind=\"kde\")","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:32:58.939485Z","iopub.execute_input":"2022-06-24T10:32:58.939936Z","iopub.status.idle":"2022-06-24T10:33:05.194027Z","shell.execute_reply.started":"2022-06-24T10:32:58.939894Z","shell.execute_reply":"2022-06-24T10:33:05.192751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set()\n#define plotting region (6 rows, 6 columns)\nfig, axes = plt.subplots(6, 6,figsize=(18, 10))\nfig.suptitle('Correlation between numerial features', fontsize=24)\nfor i,col1 in enumerate(num_features):\n    for j,col2 in enumerate(num_features):\n        sns.regplot(x=col1,y=col2,data=df_train,color='blue', scatter_kws={\n                    \"color\": \"deepskyblue\"}, line_kws={\"color\": \"red\"}, ax=axes[i,j])\nfig.tight_layout()\nplt.subplots_adjust(top=0.90)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:05.195488Z","iopub.execute_input":"2022-06-24T10:33:05.1959Z","iopub.status.idle":"2022-06-24T10:33:29.877059Z","shell.execute_reply.started":"2022-06-24T10:33:05.195861Z","shell.execute_reply":"2022-06-24T10:33:29.876085Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(16,10))\nsns.heatmap(df_train.corr(),cmap='BuPu',annot=True)\nplt.title ('Correlation HeatMap', fontsize=20)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:29.878637Z","iopub.execute_input":"2022-06-24T10:33:29.879174Z","iopub.status.idle":"2022-06-24T10:33:30.453775Z","shell.execute_reply.started":"2022-06-24T10:33:29.879141Z","shell.execute_reply":"2022-06-24T10:33:30.452627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It seems that numericals features are not strongly correlated.","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Numerical-Categorical Analysis</span> <a class=\"anchor\" id=\"section_2_3_2\"></a>","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize= (16,10))\nfor i, col in enumerate(num_features):\n    \n    ax=fig.add_subplot( 2, 3, i+1)\n\n    df_train.groupby(['Transported'])[col].mean().plot(kind='bar',color=[\"#7AC5CD\",\"#53868B\"])\n    ax.set_ylabel(col)\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:30.455161Z","iopub.execute_input":"2022-06-24T10:33:30.455764Z","iopub.status.idle":"2022-06-24T10:33:31.182324Z","shell.execute_reply.started":"2022-06-24T10:33:30.455718Z","shell.execute_reply":"2022-06-24T10:33:31.181176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at the visualisations above we see that:\n* Majority of passengers, who spent more on Room Service, Spa and VRDeck were not transported.\n* Majority of passengers who spent more on Food Court and Shopping Mall got transported.","metadata":{}},{"cell_type":"code","source":"Transported_df=df_train[df_train['Transported']=='Yes']\nNotTransported_df=df_train[df_train['Transported']=='No']\nfig = plt.figure(figsize= (16,10))\nfor i, col in enumerate(num_features):\n    \n    ax=fig.add_subplot( 3, 2, i+1)\n    \n    sns.distplot(Transported_df[col],label='Transp')\n    sns.distplot(NotTransported_df[col],label='not_Transp')\n    plt.legend()\n    \nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:31.185674Z","iopub.execute_input":"2022-06-24T10:33:31.186064Z","iopub.status.idle":"2022-06-24T10:33:33.918663Z","shell.execute_reply.started":"2022-06-24T10:33:31.186028Z","shell.execute_reply":"2022-06-24T10:33:33.917615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The distribution of each numerical feature for Transported and NoTransported passengers seems different. Therefore, these variables affect chances of being transported.","metadata":{}},{"cell_type":"code","source":"#Plot stucked bar graph\ndf_columns=[col for col in new_categ_features if col!='Transported']\nfor col in df_columns:\n    df_people = df_train.groupby([ \"Transported\",col])[\"RoomService\"]\n    df_people = df_people.sum().reset_index()\n    fig_people = px.bar(df_people, x=col, y=\"RoomService\", color=\"Transported\", barmode=\"stack\", text=\"RoomService\")\n    fig_people.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:33.919987Z","iopub.execute_input":"2022-06-24T10:33:33.920344Z","iopub.status.idle":"2022-06-24T10:33:35.439523Z","shell.execute_reply.started":"2022-06-24T10:33:33.920312Z","shell.execute_reply":"2022-06-24T10:33:35.438572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Categorical-Categorical Analysis</span> <a class=\"anchor\" id=\"section_2_3_3\"></a>","metadata":{}},{"cell_type":"code","source":"for col in new_categ_features:\n    plt.figure(figsize=(12,8))\n    if col!=\"Transported\":\n        sns.countplot(x=col,hue='Transported',data=df_train, palette=colors)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:35.440559Z","iopub.execute_input":"2022-06-24T10:33:35.440863Z","iopub.status.idle":"2022-06-24T10:33:37.049033Z","shell.execute_reply.started":"2022-06-24T10:33:35.440836Z","shell.execute_reply":"2022-06-24T10:33:37.047911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Passengers from earth are less likely to be transported.\n* Passengers in CryoSleep are more likely to be transported.\n* Passengers who are traveling to Trappist-le are less likely to be transported.\n* Being a VIP doesn't seem to significantly affect chances of being transported.\n* passengers on decks F and port-side are more likely to be transported.\n* Passengers who are traveling in group are less likely to be transported.","metadata":{}},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2.4  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Managing Outliers</span> <a class=\"anchor\" id=\"section_2_4\"></a>","metadata":{}},{"cell_type":"code","source":"fig = plt.figure(figsize= (12,4))\nfor i, col in enumerate(num_features):\n    \n    ax=fig.add_subplot( 3, 2, i+1)\n    \n    sns.stripplot(x=df_train[col], ax=ax)\n\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:37.049961Z","iopub.execute_input":"2022-06-24T10:33:37.050305Z","iopub.status.idle":"2022-06-24T10:33:38.216354Z","shell.execute_reply.started":"2022-06-24T10:33:37.050258Z","shell.execute_reply":"2022-06-24T10:33:38.215415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize= (10,4))\nfor i, col in enumerate(num_features):\n    \n    ax=fig.add_subplot( 2, 3, i+1)\n    \n    sns.boxenplot(x=df_train[col],ax=ax)\nfig.tight_layout()  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:38.217402Z","iopub.execute_input":"2022-06-24T10:33:38.217705Z","iopub.status.idle":"2022-06-24T10:33:39.091107Z","shell.execute_reply.started":"2022-06-24T10:33:38.217678Z","shell.execute_reply":"2022-06-24T10:33:39.090093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def outliers_visualisation(var,var_out):\n  f, ax = plt.subplots(nrows=1, ncols=4, figsize=(18, 3))\n  sns.distplot(var, ax=ax[0])\n  sns.boxenplot(var, ax=ax[1])\n  sns.stripplot(var, ax=ax[2])\n  stats.probplot(var, plot=ax[3])\n  plt.suptitle('data with ouliers',fontsize=20)\n  plt.subplots_adjust(left=0.1,\n                    bottom=0.1, \n                    right=0.9, \n                    top=0.8, \n                    wspace=0.4, \n                    hspace=0.4)\n\n  fig.tight_layout()\n  plt.show()\n\n  f, ax = plt.subplots(nrows=1, ncols=4, figsize=(18, 3))\n  sns.distplot(var_out, ax=ax[0])\n  sns.boxenplot(var_out, ax=ax[1])\n  sns.stripplot(var_out, ax=ax[2])\n  stats.probplot(var_out, plot=ax[3])\n  plt.suptitle('data without ouliers',fontsize=20)\n  plt.subplots_adjust(left=0.1,\n                    bottom=0.1, \n                    right=0.9, \n                    top=0.8, \n                    wspace=0.4, \n                    hspace=0.4)\n  fig.tight_layout()\n  plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:39.092599Z","iopub.execute_input":"2022-06-24T10:33:39.093041Z","iopub.status.idle":"2022-06-24T10:33:39.102612Z","shell.execute_reply.started":"2022-06-24T10:33:39.092998Z","shell.execute_reply":"2022-06-24T10:33:39.101763Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Remove some outliers from \"FoodCourt\"\ndf=df_train.copy()\ndf=df[(df['FoodCourt']<20000)]\noutliers_visualisation(df_train['FoodCourt'],df['FoodCourt'])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:39.103924Z","iopub.execute_input":"2022-06-24T10:33:39.104429Z","iopub.status.idle":"2022-06-24T10:33:40.889306Z","shell.execute_reply.started":"2022-06-24T10:33:39.104394Z","shell.execute_reply":"2022-06-24T10:33:40.888273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Remove some outliers from \"RoomService\"\ndf=df[(df['RoomService']<7500)]\noutliers_visualisation(df_train['RoomService'],df['RoomService'])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:40.890771Z","iopub.execute_input":"2022-06-24T10:33:40.891629Z","iopub.status.idle":"2022-06-24T10:33:42.605191Z","shell.execute_reply.started":"2022-06-24T10:33:40.891583Z","shell.execute_reply":"2022-06-24T10:33:42.604213Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Remove some outliers from \"Spa\"\ndf=df[(df['Spa']<15000)]\noutliers_visualisation(df_train['Spa'],df['Spa'])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:42.609682Z","iopub.execute_input":"2022-06-24T10:33:42.610054Z","iopub.status.idle":"2022-06-24T10:33:44.340058Z","shell.execute_reply.started":"2022-06-24T10:33:42.610022Z","shell.execute_reply":"2022-06-24T10:33:44.338762Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Remove some outliers from \"VRDeck\"\ndf=df[(df['VRDeck']<15000)]\noutliers_visualisation(df_train['VRDeck'],df['VRDeck'])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:44.341912Z","iopub.execute_input":"2022-06-24T10:33:44.342528Z","iopub.status.idle":"2022-06-24T10:33:46.132053Z","shell.execute_reply.started":"2022-06-24T10:33:44.342373Z","shell.execute_reply":"2022-06-24T10:33:46.13105Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df=df.reset_index(drop=True)\nprint('shape of df_train befor removing outliers:',df_train.shape)\nprint('shape of df_test after removing outliers:',df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:46.133205Z","iopub.execute_input":"2022-06-24T10:33:46.133557Z","iopub.status.idle":"2022-06-24T10:33:46.141263Z","shell.execute_reply.started":"2022-06-24T10:33:46.133526Z","shell.execute_reply":"2022-06-24T10:33:46.139999Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  2.5  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Feature Seletion using statistical test</span> <a class=\"anchor\" id=\"section_2_5\"></a>","metadata":{}},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Anova Test</span> <a class=\"anchor\" id=\"section_2_5_1\"></a>","metadata":{}},{"cell_type":"code","source":"from scipy import stats\nfor col in num_features:\n  df_anova = df[[col,'Transported']]\n  grouped_anova = df_anova.groupby(['Transported'])\n  f_value, p_value = stats.f_oneway(grouped_anova.get_group('Yes')[col],grouped_anova.get_group('No')[col])\n  result = \"\"\n  if p_value<0.05:\n    result=\"{0} is IMPORTANT for Prediction\".format(col)\n  else:\n    result=\"{0} is NOT an important predictor. (Discard {0} from model)\".format(col)\n  print(result)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:46.142466Z","iopub.execute_input":"2022-06-24T10:33:46.142866Z","iopub.status.idle":"2022-06-24T10:33:46.176864Z","shell.execute_reply.started":"2022-06-24T10:33:46.142836Z","shell.execute_reply":"2022-06-24T10:33:46.175847Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"> ### <span style=\"color:orange;font-weight: bolder\">Chi Square Test</span> <a class=\"anchor\" id=\"section_2_5_2\"></a>","metadata":{}},{"cell_type":"code","source":"# Plot contingency table\ndata_train_categ=df[new_categ_features]\nsns.set(rc={\"figure.figsize\": (10, 7)})\nX = [col for col in new_categ_features if col!='Transported']\nY = ['Transported'] * len(X)\n\n# Parameters for Chi-squared test (5% significance level)\nprob = 0.95\nalpha = 1.0 - prob\n\nfor i, j in zip(X, Y):\n    # Contingency table\n    cont = data_train_categ[[i, j]].pivot_table(\n        index=i, columns=j, aggfunc=len, margins=True, margins_name=\"Total\")\n    tx = cont.loc[:, [\"Total\"]]\n    ty = cont.loc[[\"Total\"], :]\n    n = len(data_train_categ)\n    indep = tx.dot(ty) / n\n    c = cont.fillna(0)  # Replace NaN with 0 in the contingency table\n    measure = (c - indep) ** 2 / indep\n    xi_n = measure.sum().sum()\n    table = measure / xi_n\n\n    # Plot contingency table\n    p = sns.heatmap(table.iloc[:-1, :-1],\n                    annot=c.iloc[:-1, :-1], fmt=\".0f\", cmap=\"Oranges\")\n    p.set_xlabel(j, fontsize=18)\n    p.set_ylabel(i, fontsize=18)\n    p.set_title(f\"\\nχ² test between groups {i} and groups {j}\\n\", size=18)\n    plt.show()\n\n    # Performing Chi-sq test\n    CrosstabResult = pd.crosstab(\n        index=data_train_categ[i], columns=data_train_categ[j])\n    ChiSqResult = chi2_contingency(CrosstabResult)\n    # P-Value is the Probability of H0 being True\n    print(f\"P-Value of the ChiSq Test bewteen {i} and {j} is: {ChiSqResult[1]}\\n\")\n    print('significance=%.3f, p=%.3f' % (alpha, ChiSqResult[1]))\n    if ChiSqResult[1] <= alpha:\n        print('Dependent (reject H0)')\n    else:\n        print('Independent (fail to reject H0)')","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:46.178454Z","iopub.execute_input":"2022-06-24T10:33:46.179193Z","iopub.status.idle":"2022-06-24T10:33:48.731545Z","shell.execute_reply.started":"2022-06-24T10:33:46.179144Z","shell.execute_reply":"2022-06-24T10:33:48.730433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Loking at the result of Chi-square test we see that the variable \"ShoppingMall\" is not correlated with the target variable,so, we'll remove it from both training and test dataset.","metadata":{}},{"cell_type":"code","source":"#Remove \"ShoppingMall\" from both training and test dataset\ndf.drop(['ShoppingMall'], axis=1, inplace=True)\ndf_test.drop(['ShoppingMall'], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.733199Z","iopub.execute_input":"2022-06-24T10:33:48.733864Z","iopub.status.idle":"2022-06-24T10:33:48.743788Z","shell.execute_reply.started":"2022-06-24T10:33:48.733722Z","shell.execute_reply":"2022-06-24T10:33:48.742644Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Model Building</span> <a class=\"anchor\" id=\"Chapter3\"></a>","metadata":{}},{"cell_type":"code","source":"X= pd.get_dummies(df.drop(['Transported'],axis=1),drop_first=True)\ny= df['Transported']","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.74545Z","iopub.execute_input":"2022-06-24T10:33:48.746178Z","iopub.status.idle":"2022-06-24T10:33:48.773093Z","shell.execute_reply.started":"2022-06-24T10:33:48.746137Z","shell.execute_reply":"2022-06-24T10:33:48.772036Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_test=pd.get_dummies(df_test,drop_first=True)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.774661Z","iopub.execute_input":"2022-06-24T10:33:48.774993Z","iopub.status.idle":"2022-06-24T10:33:48.790364Z","shell.execute_reply.started":"2022-06-24T10:33:48.774966Z","shell.execute_reply":"2022-06-24T10:33:48.789074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.792466Z","iopub.execute_input":"2022-06-24T10:33:48.793562Z","iopub.status.idle":"2022-06-24T10:33:48.803715Z","shell.execute_reply.started":"2022-06-24T10:33:48.793509Z","shell.execute_reply":"2022-06-24T10:33:48.802768Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def model_performance(model,model_name,X_train = X_train,y_train = y_train,X_test = X_test,y_test = y_test):\n    \n    model.fit(X_train,y_train)\n    y_test_pred = model.predict(X_test)\n    \n    Training_Accuracy = np.round(model.score(X_train,y_train),3)\n    Testing_Accuracy = np.round(model.score(X_test,y_test),3)\n    cm=confusion_matrix(y_test, y_test_pred)\n    \n    print(\"Model Performance for:\",model_name)\n    #print(\"Best_estimator:\",model.best_params_)\n    print(\"\")\n    print(\"Training Accuracy:\",Training_Accuracy)\n    print(\"Testing Accuracy:\",Testing_Accuracy)\n    print(\"classification_report:\\n\",classification_report(y_test,y_test_pred))\n    print(\"\")\n\n    print(\"confusion_matrix:\\n\",sns.heatmap(cm,annot=True,cmap=\"Blues\",fmt=\"d\",cbar=False, annot_kws={\"size\": 24}))\n    \n    return Training_Accuracy,Testing_Accuracy","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.805501Z","iopub.execute_input":"2022-06-24T10:33:48.806659Z","iopub.status.idle":"2022-06-24T10:33:48.817077Z","shell.execute_reply.started":"2022-06-24T10:33:48.806606Z","shell.execute_reply":"2022-06-24T10:33:48.816129Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"knn_model=KNeighborsClassifier()\nparam_grid= {'n_neighbors':range(1,50), 'metric': ['minkowski','manhattan','euclidean']}\nknn_grid_model = GridSearchCV(knn_model,param_grid,cv=5)\nKNN=model_performance(knn_grid_model,model_name = knn_grid_model)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:33:48.818608Z","iopub.execute_input":"2022-06-24T10:33:48.818999Z","iopub.status.idle":"2022-06-24T10:37:29.620093Z","shell.execute_reply.started":"2022-06-24T10:33:48.818957Z","shell.execute_reply":"2022-06-24T10:37:29.618954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3.1  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Logistic Regression</span> <a class=\"anchor\" id=\"section_3_1\"></a>","metadata":{}},{"cell_type":"code","source":"logreg = LogisticRegression()\nparam_grid_2= {'penalty': ['l1','l2'], 'C': [0.001,0.01,0.1,1,10,100,1000]}\nlogreg_grid_model = GridSearchCV(logreg,param_grid_2,cv=5)\nlogistic_Regression=model_performance(logreg_grid_model,model_name = logreg_grid_model)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:37:29.621757Z","iopub.execute_input":"2022-06-24T10:37:29.622301Z","iopub.status.idle":"2022-06-24T10:37:37.010315Z","shell.execute_reply.started":"2022-06-24T10:37:29.622255Z","shell.execute_reply":"2022-06-24T10:37:37.009441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3.2  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Random Forest Classifier</span> <a class=\"anchor\" id=\"section_3_2\"></a>","metadata":{}},{"cell_type":"code","source":"Rf=RandomForestClassifier(random_state=42)\nparam_grid_3 = {\n     'criterion':['gini','entropy'],\n    'max_depth':[2,3,4,5,20,30],\n    'min_samples_split':[5,20,50],\n    'min_samples_leaf':[15,20,30],\n    'n_estimators': [1,5,10]\n}\nRf_grid_model = GridSearchCV(Rf,param_grid_3,cv=5)\nRandomForest=model_performance(Rf_grid_model,model_name = Rf_grid_model)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:37:37.011456Z","iopub.execute_input":"2022-06-24T10:37:37.012206Z","iopub.status.idle":"2022-06-24T10:38:44.301669Z","shell.execute_reply.started":"2022-06-24T10:37:37.012164Z","shell.execute_reply":"2022-06-24T10:38:44.30019Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3.3  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Decision Tree Classifier</span> <a class=\"anchor\" id=\"section_3_3\"></a>","metadata":{}},{"cell_type":"code","source":"param_grid_4 = {'criterion':['gini','entropy'],'max_depth':np.arange(1,10),'min_samples_split':np.arange(2,10),'min_samples_leaf':np.arange(1,10)}\nDT_grid_model  = GridSearchCV(DecisionTreeClassifier(random_state=42),param_grid = param_grid_4,cv=5)\nDecisionTree=model_performance(DT_grid_model,model_name = DT_grid_model)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:38:44.303706Z","iopub.execute_input":"2022-06-24T10:38:44.304408Z","iopub.status.idle":"2022-06-24T10:41:58.635912Z","shell.execute_reply.started":"2022-06-24T10:38:44.304347Z","shell.execute_reply":"2022-06-24T10:41:58.63478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3.4  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">AdaBoost Classifier</span> <a class=\"anchor\" id=\"section_3_4\"></a>","metadata":{}},{"cell_type":"code","source":"DTC = DecisionTreeClassifier(random_state = 42)\nparam_grid_6 = {'base_estimator__max_depth':[i for i in range(2,11,2)],\n             'base_estimator__min_samples_leaf':[5,10],\n             'n_estimators':[10,50,250,1000],\n             'learning_rate':[0.01,0.1]}\nABC = AdaBoostClassifier(base_estimator = DTC)\ngrid_search_ABC = GridSearchCV(ABC, param_grid=param_grid_6, cv=5)\nAdaboost=model_performance(DTC,model_name = DTC)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:41:58.637177Z","iopub.execute_input":"2022-06-24T10:41:58.638305Z","iopub.status.idle":"2022-06-24T10:41:58.91681Z","shell.execute_reply.started":"2022-06-24T10:41:58.638252Z","shell.execute_reply":"2022-06-24T10:41:58.915753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  3.5  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Gradient Boosting Classifier</span> <a class=\"anchor\" id=\"section_3_5\"></a>","metadata":{}},{"cell_type":"code","source":"param_grid_7 = {\n    'num_leaves': [31, 127],\n    'reg_alpha': [0.1, 0.5],\n    'min_data_in_leaf': [30, 50, 100, 300, 400],\n    'lambda_l1': [0, 1, 1.5],\n    'lambda_l2': [0, 1]\n    }\nLGBM_grid_model  = GridSearchCV(LGBMClassifier(random_state=42),param_grid = param_grid_7,cv=5)\nLGBM=model_performance(LGBM_grid_model,model_name = LGBM_grid_model)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:41:58.918934Z","iopub.execute_input":"2022-06-24T10:41:58.919299Z","iopub.status.idle":"2022-06-24T10:43:52.95228Z","shell.execute_reply.started":"2022-06-24T10:41:58.919267Z","shell.execute_reply":"2022-06-24T10:43:52.951311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">  4  </span> &nbsp; <span style=\"background-color:#DFFF00; padding: 5px; border-radius:5px;color:#DE3163;font-weight: bolder\">Model choice and submission</span> <a class=\"anchor\" id=\"Chapter4\"></a>","metadata":{}},{"cell_type":"code","source":"model_performance = [[\"KNN\",KNN[0],KNN[1]],\n                     [\"Logistic Regression\",logistic_Regression[0],logistic_Regression[1]],\n                     [ \"Random Forest\",RandomForest[0],RandomForest[1]],\n                     [\"Decision Tree\",DecisionTree[0],DecisionTree[1]],\n                     [\"Adaboost\",Adaboost[0],Adaboost[1]],\n                     [\"LGBM\",LGBM[0],LGBM[1]]]","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:43:52.953631Z","iopub.execute_input":"2022-06-24T10:43:52.954789Z","iopub.status.idle":"2022-06-24T10:43:52.962023Z","shell.execute_reply.started":"2022-06-24T10:43:52.95474Z","shell.execute_reply":"2022-06-24T10:43:52.960927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"performance = pd.DataFrame(model_performance,columns = ['Model_Name',\"Train Score\",\"Test Score\"])","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:43:52.963357Z","iopub.execute_input":"2022-06-24T10:43:52.963977Z","iopub.status.idle":"2022-06-24T10:43:52.977864Z","shell.execute_reply.started":"2022-06-24T10:43:52.963929Z","shell.execute_reply":"2022-06-24T10:43:52.97689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"display(performance)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:43:52.980853Z","iopub.execute_input":"2022-06-24T10:43:52.981703Z","iopub.status.idle":"2022-06-24T10:43:53.00008Z","shell.execute_reply.started":"2022-06-24T10:43:52.981665Z","shell.execute_reply":"2022-06-24T10:43:52.999292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.barplot(x='Test Score',\n            y=\"Model_Name\", \n            data=performance, \n            order=performance.sort_values('Test Score',ascending = False).Model_Name)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T10:43:53.001454Z","iopub.execute_input":"2022-06-24T10:43:53.002729Z","iopub.status.idle":"2022-06-24T10:43:53.269575Z","shell.execute_reply.started":"2022-06-24T10:43:53.002676Z","shell.execute_reply":"2022-06-24T10:43:53.268477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"best_model=LGBM_grid_model.best_estimator_\nbest_model","metadata":{"execution":{"iopub.status.busy":"2022-06-24T11:16:16.783119Z","iopub.execute_input":"2022-06-24T11:16:16.783581Z","iopub.status.idle":"2022-06-24T11:16:16.790859Z","shell.execute_reply.started":"2022-06-24T11:16:16.783543Z","shell.execute_reply":"2022-06-24T11:16:16.789983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred=pd.Series(best_model.predict(df_test)).map({'No':False, 'Yes':True})","metadata":{"execution":{"iopub.status.busy":"2022-06-24T11:16:19.347525Z","iopub.execute_input":"2022-06-24T11:16:19.347951Z","iopub.status.idle":"2022-06-24T11:16:19.386718Z","shell.execute_reply.started":"2022-06-24T11:16:19.347914Z","shell.execute_reply":"2022-06-24T11:16:19.385863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({'PassengerId': Id_test_list,\n                       'Transported': pred})\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-06-24T11:16:25.694988Z","iopub.execute_input":"2022-06-24T11:16:25.695435Z","iopub.status.idle":"2022-06-24T11:16:25.707476Z","shell.execute_reply.started":"2022-06-24T11:16:25.695397Z","shell.execute_reply":"2022-06-24T11:16:25.706512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-06-24T11:16:28.972121Z","iopub.execute_input":"2022-06-24T11:16:28.972583Z","iopub.status.idle":"2022-06-24T11:16:28.986748Z","shell.execute_reply.started":"2022-06-24T11:16:28.972544Z","shell.execute_reply":"2022-06-24T11:16:28.985535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<span style=\" padding: 5px; border-radius:5px;color:red;font-weight: bolder;font-size: 24px;font-family: cursive\">If you like my work, don't forget to upvote and leave me a comment.This will help me to continue and sharing more topics in the coming days.</span>","metadata":{}}]}