{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Detailed and Full Solution (Step by Step , > 80% score)\n#### By: Oday Mourad\n##### 13 -  8 - 2022\n\n\nHello kagglers ..\n\nThis notebook designed to be as detailed as possible solution for the Houses pricing problem, I tried to make it typical, clear, tidy and **beginner-friendly**.\n\nIf you find this notebook useful press the **UPVOTE** button, This helps me a lot ^-^.  \n\nI hope you find it helpful.\n\n<img src=\"https://storage.googleapis.com/kaggle-media/competitions/Spaceship%20Titanic/joel-filipe-QwoNAhbmLLo-unsplash.jpg\" width=\"600\"/>\n","metadata":{}},{"cell_type":"markdown","source":"#### Table Of Content: <a class = \"anchor\" id = \"toc\" ></a>\n- [1 - Introduction](#introduction)\n- [2 - Importing](#import)\n- [3 - Descovering the data](#dtd)\n- [4 - Exploratory Data Analysis](#eda)\n    - [target](#eda_target)\n    - [categorical features with target](#cat_with_tar)\n    - [numerical features with target](#num_with_tar)\n    - [correlation between numerical features](#num_features)\n    - [correlation between categorical and numerical features](#cat_and_num)\n- [5 - Data Processing](#dp)\n    - [Filling Missed Values](#fmv)\n    - [Data Engineering](#de)\n    - [Preparing For Trainging](#prfortr)\n- [6 - Modeling](#modeling)\n","metadata":{}},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"introduction\">\n    <div style=\"color:#00ADB5;\n               display:fill;\n               border-radius:5px;\n               background-color:#393E46;\n               font-size:20px;\n               font-family:sans-serif;\n               letter-spacing:0.5px\">\n            <p style=\"padding: 10px;\n                  color:white;\">\n                <b>1 ) Introduction:</b>\n            </p>\n    </div>\n</a>","metadata":{}},{"cell_type":"markdown","source":"The competition is organised by **Kaggle** and is in the GettingStarted Prediction Competition series.\n\nIn this competition, you are supposed to predict predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.\n\nSubmissions are evaluated on **Classification Accuracy.**","metadata":{}},{"cell_type":"markdown","source":"<a class=\"anchor\" id=\"import\">\n    <div style=\"color:#00ADB5;\n               display:fill;\n               border-radius:5px;\n               background-color:#393E46;\n               font-size:20px;\n               font-family:sans-serif;\n               letter-spacing:0.5px\">\n            <p style=\"padding: 10px;\n                  color:white;\">\n                <b>2 ) Importing:</b>\n            </p>\n    </div>\n</a>","metadata":{}},{"cell_type":"code","source":"#=======================================================================================\n# Importing the libaries:\n#=======================================================================================\nimport numpy as np \nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport plotly.express as px\nimport plotly.graph_objects as go\nfrom plotly.subplots import make_subplots\nimport warnings\nwarnings.filterwarnings(\"ignore\")\npd.set_option('display.max_columns', None)\npd.options.display.max_seq_items = 8000\npd.options.display.max_rows = 8000","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-14T14:09:03.039099Z","iopub.execute_input":"2022-08-14T14:09:03.039533Z","iopub.status.idle":"2022-08-14T14:09:03.047671Z","shell.execute_reply.started":"2022-08-14T14:09:03.039497Z","shell.execute_reply":"2022-08-14T14:09:03.046081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#=======================================================================================\n# Importing the data:\n#=======================================================================================\n\ndef read_data():\n    train_data = pd.read_csv(\"/kaggle/input/spaceship-titanic/train.csv\")\n    print(\"Train data imported successfully!!\")\n    print(\"-\"*50)\n    test_data = pd.read_csv(\"/kaggle/input/spaceship-titanic/test.csv\")\n    print(\"Test data imported successfully!!\")\n    return train_data , test_data","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:07.792269Z","iopub.execute_input":"2022-08-14T14:09:07.792726Z","iopub.status.idle":"2022-08-14T14:09:07.799052Z","shell.execute_reply.started":"2022-08-14T14:09:07.792690Z","shell.execute_reply":"2022-08-14T14:09:07.798010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data , test_data = read_data()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:10.556261Z","iopub.execute_input":"2022-08-14T14:09:10.556703Z","iopub.status.idle":"2022-08-14T14:09:10.612588Z","shell.execute_reply.started":"2022-08-14T14:09:10.556670Z","shell.execute_reply":"2022-08-14T14:09:10.611232Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n<a class = \"anchor\"  id = \"dtd\"  >\n    <div style=\"color:#00ADB5;\n               display:fill;\n               border-radius:5px;\n               background-color:#393E46;\n               font-size:20px;\n               font-family:sans-serif;\n               letter-spacing:0.5px\">\n            <p style=\"padding: 10px;\n                  color:white;\">\n                <b> 3 ) Discovering the data:</b>\n            </p>\n    </div>\n</a>\n","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:13.160494Z","iopub.execute_input":"2022-08-14T14:09:13.160910Z","iopub.status.idle":"2022-08-14T14:09:13.184862Z","shell.execute_reply.started":"2022-08-14T14:09:13.160874Z","shell.execute_reply":"2022-08-14T14:09:13.183437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:15.481246Z","iopub.execute_input":"2022-08-14T14:09:15.481666Z","iopub.status.idle":"2022-08-14T14:09:15.506692Z","shell.execute_reply.started":"2022-08-14T14:09:15.481631Z","shell.execute_reply":"2022-08-14T14:09:15.505514Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"That's an interesting PassengerId .. maybe we can use it ..\n","metadata":{}},{"cell_type":"code","source":"print(train_data.columns.values)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:18.360702Z","iopub.execute_input":"2022-08-14T14:09:18.361169Z","iopub.status.idle":"2022-08-14T14:09:18.366810Z","shell.execute_reply.started":"2022-08-14T14:09:18.361127Z","shell.execute_reply":"2022-08-14T14:09:18.365914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info()\nprint(\"-\"*50)\ntest_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:20.643298Z","iopub.execute_input":"2022-08-14T14:09:20.643717Z","iopub.status.idle":"2022-08-14T14:09:20.671740Z","shell.execute_reply.started":"2022-08-14T14:09:20.643684Z","shell.execute_reply":"2022-08-14T14:09:20.670431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Transported feature is object, I will convert it to int for visualization step.","metadata":{}},{"cell_type":"code","source":"train_data[\"Transported\"] = train_data[\"Transported\"].astype(\"int\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:24.020526Z","iopub.execute_input":"2022-08-14T14:09:24.020968Z","iopub.status.idle":"2022-08-14T14:09:24.027034Z","shell.execute_reply.started":"2022-08-14T14:09:24.020927Z","shell.execute_reply":"2022-08-14T14:09:24.026177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Train data shape = \" , train_data.shape)\nprint(\"Test data shape = \" , test_data.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:26.168732Z","iopub.execute_input":"2022-08-14T14:09:26.169842Z","iopub.status.idle":"2022-08-14T14:09:26.175972Z","shell.execute_reply.started":"2022-08-14T14:09:26.169792Z","shell.execute_reply":"2022-08-14T14:09:26.174889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test data is about **50%** of the training data.","metadata":{}},{"cell_type":"code","source":"print(\"Missed Data in train data:\")\nprint(train_data.isnull().sum())\nprint(\"-\" * 50)\nprint(\"Missed Data in test data:\")\ntest_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:28.819319Z","iopub.execute_input":"2022-08-14T14:09:28.819750Z","iopub.status.idle":"2022-08-14T14:09:28.839785Z","shell.execute_reply.started":"2022-08-14T14:09:28.819714Z","shell.execute_reply":"2022-08-14T14:09:28.838288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a lot of missed data .. we are going to process them in the data processing step.","metadata":{}},{"cell_type":"code","source":"train_data.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:32.375432Z","iopub.execute_input":"2022-08-14T14:09:32.375819Z","iopub.status.idle":"2022-08-14T14:09:32.412807Z","shell.execute_reply.started":"2022-08-14T14:09:32.375779Z","shell.execute_reply":"2022-08-14T14:09:32.411932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>insights:</u></b><br>\n \n* <i> There are an approximately <b>equal</b> number of transported passengers and non-transported passengers.</i><br>\n* <i> More than <b>75%</b> of the passengers are under the age of <b>38</b> and there some passengers are over <b>70</b> years old.</i><br>\n* <i> More than <b>50%</b> of the passengers didn't spend any money for RoomService, FoodCourt, ShoppingMall, Spa, VRDeck.  </i><br>\n* <i> here are too high outliers in RoomService, FoodCourt, ShoppingMall, Spa, VRDeck.</i><br>\n</div>","metadata":{}},{"cell_type":"code","source":"train_data.describe(include = [\"O\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:35.676804Z","iopub.execute_input":"2022-08-14T14:09:35.677202Z","iopub.status.idle":"2022-08-14T14:09:35.723228Z","shell.execute_reply.started":"2022-08-14T14:09:35.677170Z","shell.execute_reply":"2022-08-14T14:09:35.722087Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>insights:</u></b><br>\n \n* <i> <b>Earth</b> is the most common HomePlanet.</i><br>\n* <i>Most of the passengers were not put into a cryosleep state.</i><br>\n* <i>There are many passengers with same Cabin (they shared the same cabin). </i><br>\n* <i> Most of the passengers going to <b>TRAPPIST-1e</b>.</i><br>\n* <i>only <b>199</b> passengers are VIP.</i><br>\n</div>","metadata":{}},{"cell_type":"code","source":"# saving the test ids:\nTest_Id = test_data.PassengerId","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:38.964634Z","iopub.execute_input":"2022-08-14T14:09:38.965107Z","iopub.status.idle":"2022-08-14T14:09:38.970961Z","shell.execute_reply.started":"2022-08-14T14:09:38.965066Z","shell.execute_reply":"2022-08-14T14:09:38.969621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### I will do a little of data engineering on **PassengerId** and **Cabin** to use it in EDA Step.","metadata":{}},{"cell_type":"code","source":"combine = [train_data , test_data]\n\n\nfor dataset in combine:\n    \n    # =======================================================================\n    # Extract Passenger Group:\n    # =======================================================================\n\n    dataset[\"PassengerGroup\"] = dataset[\"PassengerId\"].str.split('_' , expand = True)[1].astype(int).astype(str)\n    dataset.drop(columns = [\"PassengerId\"] , inplace = True)\n    # =======================================================================\n    # Extract Cabin num, deck, side:\n    # =======================================================================\n\n    dataset[\"deck\"] = (dataset.Cabin.str.split('/' , expand = True))[0]\n    dataset[\"num\"] = np.nan_to_num(dataset.Cabin.str.split('/', expand = True)[1].astype(float)).astype(int)\n    dataset[\"side\"] = dataset.Cabin.str.split('/', expand = True)[2]\n    dataset.drop(columns = [\"Cabin\"] , inplace = True)\n    \nprint(f\"Available decks are ({train_data.deck.unique().shape[0]} decks): {train_data.deck.unique()}\")\nprint(f\"Available nums are ({train_data.num.unique().shape[0]} nums): {train_data.num.unique()}\")\nprint(f\"Available sides are ({train_data.side.unique().shape[0]} sides): {train_data.side.unique()}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:43.789655Z","iopub.execute_input":"2022-08-14T14:09:43.790040Z","iopub.status.idle":"2022-08-14T14:09:44.057973Z","shell.execute_reply.started":"2022-08-14T14:09:43.790007Z","shell.execute_reply":"2022-08-14T14:09:44.056741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >Back to Table of Contents  ⬆️</a>","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"eda\">\n    <div style=\"color:#00ADB5;\n           display:fill;\n           border-radius:5px;\n           background-color:#393E46;\n           font-size:20px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            <b> 4 ) Exploratory Data Analysis (EDA):</b>\n        </p>\n</div>\n</a>\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"# Helper functions:\n# ====================================================================\ndef survived_bar_plot(feature , ax = None , font_scale = 0.8):\n    sns.set(font_scale=font_scale)  \n    data = train_data[[feature, \"Transported\"]].groupby([feature], as_index=False).mean().sort_values(by='Transported', ascending=False)\n    plot = sns.barplot(data = data , x = feature , y = \"Transported\" ,ci=None , ax = ax )\n    plot.set_title(f\"{feature} Vs Transported\")\n    plot.set(xlabel=None)\n    plot.set(ylabel=None)\n    sns.set(font_scale=font_scale)  \n    plot.bar_label(plot.containers[0],fmt='%.2f')\n# ====================================================================\n\ndef survived_table(feature):\n    return train_data[[feature, \"Transported\"]].groupby([feature], as_index=False).mean().sort_values(by='Transported', ascending=False).style.background_gradient(low=0.75,high=1)\ndef survived_hist_plot(feature , bin_width = 5):\n    plt.figure(figsize = (6,4))\n    sns.histplot(data = train_data , x = feature , hue = \"Transported\",binwidth=bin_width,palette = sns.color_palette([\"yellow\" , \"green\"]) ,multiple = \"stack\" ).set_title(f\"{feature} Vs Transported\")\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:49.849040Z","iopub.execute_input":"2022-08-14T14:09:49.849495Z","iopub.status.idle":"2022-08-14T14:09:49.861960Z","shell.execute_reply.started":"2022-08-14T14:09:49.849458Z","shell.execute_reply":"2022-08-14T14:09:49.860638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class = \"anchor\" id = eda_target >\n    <div style=\"color:black;\n               border-radius:0px;\n               background-color:#00ADB5;\n               font-size:14px;\n               font-family:sans-serif;\n               letter-spacing:0.5px\">\n            <p style=\"padding: 6px;\n                  color:white;\">\n                <b>Target:</b>\n            </p>\n    </div>\n</a>\n","metadata":{}},{"cell_type":"code","source":"# ===================================================================\n# Count of Transported Passengers:\n# ===================================================================\nf,ax=plt.subplots(1,2,figsize=(8,4))\ntrain_data['Transported'].replace({0:\"Not Transported\",1:\"Transported\"}).value_counts().plot.pie(explode=[0,0.1],autopct='%1.1f%%',ax=ax[0],shadow=True)\nax[0].set_ylabel('')\nsns.countplot(x = train_data[\"Transported\"].replace({0:\"Not Transported\",1:\"Transported\"}) , ax = ax[1])\nax[1].set_ylabel('')\nax[1].set_xlabel('')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:51.349574Z","iopub.execute_input":"2022-08-14T14:09:51.350012Z","iopub.status.idle":"2022-08-14T14:09:51.607722Z","shell.execute_reply.started":"2022-08-14T14:09:51.349976Z","shell.execute_reply":"2022-08-14T14:09:51.604942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are approximately equal number of transported and non-transported passengers.","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"cat_with_tar\">\n</a>\n<div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Categorical features with target:</b>\n        </p>\n</div>\n","metadata":{}},{"cell_type":"code","source":"fig , ax = plt.subplots(3,3 , figsize=(18 , 15))\nsurvived_bar_plot(\"HomePlanet\" , ax[0][0])\nsurvived_bar_plot('CryoSleep' , ax[0][1])\nsurvived_bar_plot('Destination' , ax[0][2])\nsurvived_bar_plot('VIP'  , ax[1][0])\nsurvived_bar_plot('PassengerGroup' , ax[1][1] , font_scale=0.7)\nsurvived_bar_plot('deck' , ax[1][2],font_scale=0.7)\nsurvived_bar_plot('side' , ax[2][0])\n","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:09:55.095811Z","iopub.execute_input":"2022-08-14T14:09:55.096238Z","iopub.status.idle":"2022-08-14T14:09:56.360910Z","shell.execute_reply.started":"2022-08-14T14:09:55.096203Z","shell.execute_reply":"2022-08-14T14:09:56.359630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\" style=\"font-size:14px; font-family:verdana; line-height: 1.7em;\">\n    📌 &nbsp;<b><u>insights:</u></b><br>\n \n* <i>Passengers who came from Europa Planet is most common to be transported, then Mars then Earth.</i><br>\n* <i>Passengers going to 55 \"Cancri e\" are most common to be transported.</i><br>\n* <i>The non-VIP passengers are most common to be transported.</i><br>\n* <i>There are varying Transported proportions between different passenger groups.</i><br>\n* <i>As shown in the deck plot .. The highest Transport proportion is in \"B\" and \"C\", And the lowest in \"T\".</i><br>\n* <i>Passengers in the \"S\" side is most common to be transported than the \"P\" side. </i>\n</div>\n\n","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"num_with_tar\">\n    <div style=\"color:black;\n               border-radius:0px;\n               background-color:#00ADB5;\n               font-size:14px;\n               font-family:sans-serif;\n               letter-spacing:0.5px\">\n            <p style=\"padding: 6px;\n                  color:white;\">\n                <b>Numerical features with target:</b>\n            </p>\n    </div>                          \n</a>\n\n","metadata":{}},{"cell_type":"markdown","source":"**1 ) Age:**","metadata":{}},{"cell_type":"markdown","source":"**Note:** This plot is a stack plot.","metadata":{}},{"cell_type":"code","source":"sns.set_style(\"dark\") # to remove the grid.\nsurvived_hist_plot(\"Age\") ","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:05.211818Z","iopub.execute_input":"2022-08-14T14:10:05.212256Z","iopub.status.idle":"2022-08-14T14:10:05.594115Z","shell.execute_reply.started":"2022-08-14T14:10:05.212220Z","shell.execute_reply":"2022-08-14T14:10:05.592917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"children below 10 years old age are most common to be Transported. I am going to make is_child in Data Engineering step.","metadata":{}},{"cell_type":"markdown","source":"**2 ) RoomService, FoodCourt, ShoppingMall, Spa, VrDeck:**","metadata":{}},{"cell_type":"code","source":"plot , ax = plt.subplots(2 , 3, figsize = (18,8))\nsns.boxplot(data = train_data , x = \"Transported\" , y = \"RoomService\" , ax = ax[0][0]).set_title(\"RoomService\")\nsns.boxplot(data = train_data , x = \"Transported\" , y = \"FoodCourt\" , ax = ax[0][1]).set_title(\"FoodCourt\")\nsns.boxplot(data = train_data , x = \"Transported\" , y = \"ShoppingMall\" , ax = ax[0][2]).set_title(\"ShoppingMall\")\nsns.boxplot(data = train_data , x = \"Transported\" , y = \"Spa\" , ax = ax[1][0]).set_title(\"Spa\")\nsns.boxplot(data = train_data , x = \"Transported\" , y = \"VRDeck\" , ax = ax[1][1]).set_title(\"VRDeck\")\nplt.subplots_adjust(wspace=0.4,hspace=0.4)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:11.250252Z","iopub.execute_input":"2022-08-14T14:10:11.251861Z","iopub.status.idle":"2022-08-14T14:10:12.108763Z","shell.execute_reply.started":"2022-08-14T14:10:11.251795Z","shell.execute_reply":"2022-08-14T14:10:12.107598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we saw above, The most values of these features is 0. and There are many outliers.","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"num_features\">\n\n<div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Correlation between numerical features:</b>\n        </p>\n</div>\n</a>\n\n\n","metadata":{}},{"cell_type":"code","source":"sns.set(font_scale=0.8)\nplt.figure(figsize = (8,8))\nsns.heatmap(train_data.corr(),annot=True,fmt='.2f',cmap=\"Blues\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:17.793669Z","iopub.execute_input":"2022-08-14T14:10:17.794113Z","iopub.status.idle":"2022-08-14T14:10:18.538764Z","shell.execute_reply.started":"2022-08-14T14:10:17.794070Z","shell.execute_reply":"2022-08-14T14:10:18.537614Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"insights:\n- luxury features haves some positive correlation with each other.\n- There are negtive correlation between luxury features and target feature.","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"cat_and_num\"><div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Correlation between numerical and categorical features:</b>\n        </p>\n</div></a>\n\n","metadata":{}},{"cell_type":"markdown","source":"**1 ) Age:**","metadata":{}},{"cell_type":"code","source":"plot , ax  = plt.subplots(2,3 , figsize = (16,6))\nsns.boxplot(data = train_data , y = \"Age\" , x = \"HomePlanet\"  , ax = ax[0][0])\nsns.boxplot(data = train_data , y = \"Age\" , x = \"Destination\"  , ax = ax[0][1])\nsns.boxplot(data = train_data , y = \"Age\" , x = \"CryoSleep\"  , ax = ax[0][2])\nsns.boxplot(data = train_data , y = \"Age\" , x = \"VIP\"  , ax = ax[1][0])\nsns.boxplot(data = train_data , y = \"Age\" , x = \"PassengerGroup\"  , ax = ax[1][1])\nsns.boxplot(data = train_data , y = \"Age\" , x = \"VIP\"  , ax = ax[1][0])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:24.798961Z","iopub.execute_input":"2022-08-14T14:10:24.799432Z","iopub.status.idle":"2022-08-14T14:10:25.784420Z","shell.execute_reply.started":"2022-08-14T14:10:24.799394Z","shell.execute_reply":"2022-08-14T14:10:25.783230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"from the plots above we can see that PassengerGroup is good to use for filling Age missed data.","metadata":{}},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >Back to Table of Contents  ⬆️</a>","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"dp\">\n    \n<div style=\"color:#00ADB5;\n           display:fill;\n           border-radius:5px;\n           background-color:#393E46;\n           font-size:20px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            <b> 4 ) Data Processing:</b>\n        </p>\n</div>\n</a>\n\n\n","metadata":{}},{"cell_type":"code","source":"transported = train_data[\"Transported\"]\nall_data = pd.concat([train_data , test_data]).reset_index(drop = True)\nall_data.drop(columns = [\"Transported\"] , inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:32.990228Z","iopub.execute_input":"2022-08-14T14:10:32.990641Z","iopub.status.idle":"2022-08-14T14:10:33.014432Z","shell.execute_reply.started":"2022-08-14T14:10:32.990607Z","shell.execute_reply":"2022-08-14T14:10:33.013195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"fmv\"><div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Filling Missed Values:</b>\n        </p>\n</div></a>\n\n","metadata":{}},{"cell_type":"code","source":"all_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:38.163449Z","iopub.execute_input":"2022-08-14T14:10:38.163839Z","iopub.status.idle":"2022-08-14T14:10:38.179391Z","shell.execute_reply.started":"2022-08-14T14:10:38.163808Z","shell.execute_reply":"2022-08-14T14:10:38.178299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling HomePlanet, CryoSleep, Destination, VIP:\nall_data[\"HomePlanet\"] = all_data[\"HomePlanet\"].fillna(all_data[\"HomePlanet\"].mode()[0]) \nall_data[\"CryoSleep\"] = all_data[\"CryoSleep\"].fillna(all_data[\"CryoSleep\"].mode()[0]) \nall_data[\"Destination\"] = all_data[\"Destination\"].fillna(all_data[\"Destination\"].mode()[0]) \nall_data[\"VIP\"] = all_data[\"VIP\"].fillna(all_data[\"VIP\"].mode()[0]) ","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:48.362826Z","iopub.execute_input":"2022-08-14T14:10:48.363338Z","iopub.status.idle":"2022-08-14T14:10:48.385524Z","shell.execute_reply.started":"2022-08-14T14:10:48.363253Z","shell.execute_reply":"2022-08-14T14:10:48.384471Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling Age Feature by PassengerGroup:\nPassengerGroups = [\"1\" , \"2\" , \"3\" , \"4\" , \"5\" , \"6\" , \"7\" , \"8\"]\nmedian_ages = {}\nfor passengerGroup in PassengerGroups :\n    median_ages[passengerGroup] = all_data.loc[all_data[\"PassengerGroup\"] == passengerGroup , [\"Age\"]].median()\n\nfor index , passenger in all_data.iterrows():\n    if pd.isna(passenger[\"Age\"]):\n        all_data.at[index , \"Age\"] = median_ages[passenger[\"PassengerGroup\"]]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:50.921439Z","iopub.execute_input":"2022-08-14T14:10:50.922529Z","iopub.status.idle":"2022-08-14T14:10:51.652081Z","shell.execute_reply.started":"2022-08-14T14:10:50.922488Z","shell.execute_reply":"2022-08-14T14:10:51.650745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling RoomService, FoodCourt, ShoppingMall, Spa, VRDeck:\nall_data[\"RoomService\"] = all_data[\"RoomService\"].fillna(all_data[\"RoomService\"].mode()[0]) \nall_data[\"FoodCourt\"] = all_data[\"FoodCourt\"].fillna(all_data[\"FoodCourt\"].mode()[0]) \nall_data[\"ShoppingMall\"] = all_data[\"ShoppingMall\"].fillna(all_data[\"ShoppingMall\"].mode()[0]) \nall_data[\"Spa\"] = all_data[\"Spa\"].fillna(all_data[\"Spa\"].mode()[0]) \nall_data[\"VRDeck\"] = all_data[\"VRDeck\"].fillna(all_data[\"VRDeck\"].mode()[0]) ","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:54.345918Z","iopub.execute_input":"2022-08-14T14:10:54.346462Z","iopub.status.idle":"2022-08-14T14:10:54.363657Z","shell.execute_reply.started":"2022-08-14T14:10:54.346411Z","shell.execute_reply":"2022-08-14T14:10:54.362554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling Cabin information:\nall_data.deck = all_data.deck.fillna(all_data.deck.mode()[0])\nall_data.num = all_data.num.fillna(all_data.num.mode()[0])\nall_data.side = all_data.side.fillna(all_data.side.mode()[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:10:57.500356Z","iopub.execute_input":"2022-08-14T14:10:57.500873Z","iopub.status.idle":"2022-08-14T14:10:57.517275Z","shell.execute_reply.started":"2022-08-14T14:10:57.500833Z","shell.execute_reply":"2022-08-14T14:10:57.516250Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Filling Name feature:\nall_data.Name = all_data.Name.fillna(\"None\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:00.339422Z","iopub.execute_input":"2022-08-14T14:11:00.339842Z","iopub.status.idle":"2022-08-14T14:11:00.346899Z","shell.execute_reply.started":"2022-08-14T14:11:00.339809Z","shell.execute_reply":"2022-08-14T14:11:00.345793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:03.059635Z","iopub.execute_input":"2022-08-14T14:11:03.060070Z","iopub.status.idle":"2022-08-14T14:11:03.075107Z","shell.execute_reply.started":"2022-08-14T14:11:03.060032Z","shell.execute_reply":"2022-08-14T14:11:03.073539Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"No More Missed Data !!","metadata":{}},{"cell_type":"markdown","source":"<a class = \"anchor\" id = \"de\"><div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Data Engineering:</b>\n        </p>\n</div>\n    </a>\n","metadata":{}},{"cell_type":"markdown","source":"**1 ) Family Size:**","metadata":{}},{"cell_type":"code","source":"all_data[\"LastName\"] = all_data.Name.str.split(\" \",expand = True)[1]\nlast_name_count = all_data.Name.str.split(\" \",expand = True)[1].value_counts()\nall_data[\"FamilySize\"] = [last_name_count[x] if not pd.isna(x) else None for x in all_data[\"LastName\"]]\nall_data[\"FamilySize\"] = all_data[\"FamilySize\"].fillna(0)\nall_data.drop(columns = [\"LastName\" , \"Name\"] , inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:13.169466Z","iopub.execute_input":"2022-08-14T14:11:13.169860Z","iopub.status.idle":"2022-08-14T14:11:13.301020Z","shell.execute_reply.started":"2022-08-14T14:11:13.169827Z","shell.execute_reply":"2022-08-14T14:11:13.300058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**2 ) MoneySpent:**","metadata":{}},{"cell_type":"code","source":"all_data[\"MoneySpent\"] = all_data[\"RoomService\"] + all_data[\"FoodCourt\"] + all_data[\"ShoppingMall\"] + \\\nall_data[\"Spa\"] + all_data[\"VRDeck\"] ","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:16.592434Z","iopub.execute_input":"2022-08-14T14:11:16.592801Z","iopub.status.idle":"2022-08-14T14:11:16.601587Z","shell.execute_reply.started":"2022-08-14T14:11:16.592771Z","shell.execute_reply":"2022-08-14T14:11:16.600332Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**3 ) Spend Category:**","metadata":{}},{"cell_type":"code","source":"all_data['SpendCategory'] = ''\nall_data.loc[all_data['MoneySpent'].between(0, 1, 'left'), 'SpendCategory'] = 'Zero_Spend'\nall_data.loc[all_data['MoneySpent'].between(1, 800, 'both'), 'SpendCategory'] = 'Under_800'\nall_data.loc[all_data['MoneySpent'].between(800, 1200, 'right'), 'SpendCategory'] = 'Median_1200'\nall_data.loc[all_data['MoneySpent'].between(1200, 2700, 'right'), 'SpendCategory'] = 'Upper_2700'\nall_data.loc[all_data['MoneySpent'].between(2700, 100000, 'right'), 'SpendCategory'] = 'Big_Spender'\nall_data['SpendCategory'] = all_data['SpendCategory'].astype('category')","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:41.451466Z","iopub.execute_input":"2022-08-14T14:11:41.451866Z","iopub.status.idle":"2022-08-14T14:11:41.472895Z","shell.execute_reply.started":"2022-08-14T14:11:41.451832Z","shell.execute_reply":"2022-08-14T14:11:41.471747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**4 ) Any_Spend:**","metadata":{}},{"cell_type":"code","source":"all_data[\"AnySpend\"] = all_data[\"MoneySpent\"] > 0","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:44.329928Z","iopub.execute_input":"2022-08-14T14:11:44.330497Z","iopub.status.idle":"2022-08-14T14:11:44.338024Z","shell.execute_reply.started":"2022-08-14T14:11:44.330446Z","shell.execute_reply":"2022-08-14T14:11:44.336904Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**5 ) Is Child:**","metadata":{}},{"cell_type":"code","source":"all_data[\"IsChild\"] = all_data[\"Age\"] <= 10","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:48.200461Z","iopub.execute_input":"2022-08-14T14:11:48.201044Z","iopub.status.idle":"2022-08-14T14:11:48.208497Z","shell.execute_reply.started":"2022-08-14T14:11:48.200986Z","shell.execute_reply":"2022-08-14T14:11:48.207039Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:11:49.350501Z","iopub.execute_input":"2022-08-14T14:11:49.350886Z","iopub.status.idle":"2022-08-14T14:11:49.381660Z","shell.execute_reply.started":"2022-08-14T14:11:49.350846Z","shell.execute_reply":"2022-08-14T14:11:49.379460Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a class = \"anchor\" id  = \"prfortr\">\n\n<div style=\"color:black;\n           border-radius:0px;\n           background-color:#00ADB5;\n           font-size:14px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 6px;\n              color:white;\">\n            <b>Preparing for Training:</b>\n        </p>\n</div>\n\n\n</a>\n\n\n","metadata":{}},{"cell_type":"code","source":"# =========================================================================\n#  Converting Bool to Int\n# =========================================================================\n\nall_data[\"CryoSleep\"] = all_data[\"CryoSleep\"].astype(int)\nall_data[\"VIP\"] = all_data[\"VIP\"].astype(int)\nall_data[\"IsChild\"] = all_data[\"IsChild\"].astype(int)\nall_data[\"FamilySize\"] = all_data[\"FamilySize\"].astype(int)\nall_data[\"AnySpend\"] = all_data[\"AnySpend\"].astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:12:13.562617Z","iopub.execute_input":"2022-08-14T14:12:13.563007Z","iopub.status.idle":"2022-08-14T14:12:13.574942Z","shell.execute_reply.started":"2022-08-14T14:12:13.562975Z","shell.execute_reply":"2022-08-14T14:12:13.573687Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data = pd.get_dummies(all_data)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:12:35.263625Z","iopub.execute_input":"2022-08-14T14:12:35.264018Z","iopub.status.idle":"2022-08-14T14:12:35.291976Z","shell.execute_reply.started":"2022-08-14T14:12:35.263985Z","shell.execute_reply":"2022-08-14T14:12:35.290998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:14:37.526167Z","iopub.execute_input":"2022-08-14T14:14:37.526595Z","iopub.status.idle":"2022-08-14T14:14:37.559419Z","shell.execute_reply.started":"2022-08-14T14:14:37.526559Z","shell.execute_reply":"2022-08-14T14:14:37.558163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = all_data[:len(train_data)]\ntest_data = all_data[len(train_data):]","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:14:51.143260Z","iopub.execute_input":"2022-08-14T14:14:51.143676Z","iopub.status.idle":"2022-08-14T14:14:51.149968Z","shell.execute_reply.started":"2022-08-14T14:14:51.143643Z","shell.execute_reply":"2022-08-14T14:14:51.148761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a href=\"#toc\" role=\"button\" aria-pressed=\"true\" >Back to Table of Contents  ⬆️</a>","metadata":{}},{"cell_type":"markdown","source":"\n<a class = \"anchor\" id = \"modeling\">\n\n\n<div style=\"color:#00ADB5;\n           display:fill;\n           border-radius:5px;\n           background-color:#393E46;\n           font-size:20px;\n           font-family:sans-serif;\n           letter-spacing:0.5px\">\n        <p style=\"padding: 10px;\n              color:white;\">\n            <b> 5 ) Modeling:</b>\n        </p>\n</div>\n\n\n\n</a>\n\n\n\n\n","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier, AdaBoostClassifier, GradientBoostingClassifier, ExtraTreesClassifier, VotingClassifier\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.model_selection import GridSearchCV, cross_val_score, StratifiedKFold, learning_curve\nfrom xgboost import XGBClassifier","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:14:54.590975Z","iopub.execute_input":"2022-08-14T14:14:54.591441Z","iopub.status.idle":"2022-08-14T14:14:55.101313Z","shell.execute_reply.started":"2022-08-14T14:14:54.591400Z","shell.execute_reply":"2022-08-14T14:14:55.100257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ==================================================================================\n# Preparing Data For Training:\n# ==================================================================================\n\nY_train = transported\nX_train = train_data\nX_test = test_data\nprint(f\"X_train shape is = {X_train.shape}\" )\nprint(f\"Y_train shape is = {Y_train.shape}\" )\nprint(f\"Test shape is = {X_test.shape}\" )","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:14:59.068500Z","iopub.execute_input":"2022-08-14T14:14:59.068893Z","iopub.status.idle":"2022-08-14T14:14:59.075736Z","shell.execute_reply.started":"2022-08-14T14:14:59.068862Z","shell.execute_reply":"2022-08-14T14:14:59.074625Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cross validate model with Kfold stratified cross val\nkfold = StratifiedKFold(n_splits=12)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:15:02.273521Z","iopub.execute_input":"2022-08-14T14:15:02.274625Z","iopub.status.idle":"2022-08-14T14:15:02.281410Z","shell.execute_reply.started":"2022-08-14T14:15:02.274569Z","shell.execute_reply":"2022-08-14T14:15:02.279969Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_learning_curve(estimator, title, X, y, ylim=None, cv=None,\n                        n_jobs=-1, train_sizes=np.linspace(.1, 1.0, 5)):\n    \"\"\"Generate a simple plot of the test and training learning curve\"\"\"\n    plt.figure()\n    plt.title(title)\n    if ylim is not None:\n        plt.ylim(*ylim)\n    plt.xlabel(\"Training examples\")\n    plt.ylabel(\"Score\")\n    train_sizes, train_scores, test_scores = learning_curve(\n        estimator, X, y, cv=cv, n_jobs=n_jobs, train_sizes=train_sizes)\n    train_scores_mean = np.mean(train_scores, axis=1)\n    train_scores_std = np.std(train_scores, axis=1)\n    test_scores_mean = np.mean(test_scores, axis=1)\n    test_scores_std = np.std(test_scores, axis=1)\n    plt.grid()\n\n    plt.fill_between(train_sizes, train_scores_mean - train_scores_std,\n                     train_scores_mean + train_scores_std, alpha=0.1,\n                     color=\"r\")\n    plt.fill_between(train_sizes, test_scores_mean - test_scores_std,\n                     test_scores_mean + test_scores_std, alpha=0.1, color=\"g\")\n    plt.plot(train_sizes, train_scores_mean, 'o-', color=\"r\",\n             label=\"Training score\")\n    plt.plot(train_sizes, test_scores_mean, 'o-', color=\"g\",\n             label=\"Cross-validation score\")\n\n    plt.legend(loc=\"best\")\n    return plt","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:15:05.922543Z","iopub.execute_input":"2022-08-14T14:15:05.923523Z","iopub.status.idle":"2022-08-14T14:15:05.935190Z","shell.execute_reply.started":"2022-08-14T14:15:05.923477Z","shell.execute_reply":"2022-08-14T14:15:05.933421Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Modeling step Test different algorithms \nrandom_state = 2\nclassifiers = []\nclassifiers.append(DecisionTreeClassifier(random_state=random_state))\nclassifiers.append(AdaBoostClassifier(DecisionTreeClassifier(random_state=random_state),random_state=random_state,learning_rate=0.1))\nclassifiers.append(RandomForestClassifier(random_state=random_state))\nclassifiers.append(ExtraTreesClassifier(random_state=random_state))\nclassifiers.append(GradientBoostingClassifier(random_state=random_state))\nclassifiers.append(MLPClassifier(random_state=random_state))\nclassifiers.append(KNeighborsClassifier())\nclassifiers.append(LogisticRegression(random_state = random_state))\nclassifiers.append(LinearDiscriminantAnalysis())\nclassifiers.append(XGBClassifier(random_state = random_state))\ncv_results = []\nfor classifier in classifiers :\n    cv_results.append(cross_val_score(classifier, X_train, y = Y_train, scoring = \"accuracy\", cv = kfold))\n\ncv_means = []\ncv_std = []\nfor cv_result in cv_results:\n    cv_means.append(cv_result.mean())\n    cv_std.append(cv_result.std())\n\ncv_res = pd.DataFrame({\"CrossValMeans\":cv_means,\"CrossValerrors\": cv_std,\"Algorithm\":[\"DecisionTree\",\"AdaBoost\",\n\"RandomForest\",\"ExtraTrees\",\"GradientBoosting\",\"MultipleLayerPerceptron\",\"KNeighboors\",\"LogisticRegression\",\n                                                                                      \"LinearDiscriminantAnalysis\" ,\"XGBoost\"]})\n\ng = sns.barplot(\"CrossValMeans\",\"Algorithm\",data = cv_res, palette=\"Set3\",orient = \"h\",**{'xerr':cv_std})\ng.set_xlabel(\"Mean Accuracy\")\ng = g.set_title(\"Cross validation scores\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:15:12.844878Z","iopub.execute_input":"2022-08-14T14:15:12.845293Z","iopub.status.idle":"2022-08-14T14:17:13.335095Z","shell.execute_reply.started":"2022-08-14T14:15:12.845258Z","shell.execute_reply":"2022-08-14T14:17:13.333929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results = pd.DataFrame({\"Model\" : [\"DecisionTree\",\"AdaBoost\",\n\"RandomForest\",\"ExtraTrees\",\"GradientBoosting\",\"MultipleLayerPerceptron\",\"KNeighboors\",\"LogisticRegression\",\"LinearDiscriminantAnalysis\" ,\"XGBoost\"],\"Score\" : cv_means , \"Std\" : cv_std})\nresults.sort_values(\"Score\" , ascending = False)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T14:17:19.899237Z","iopub.execute_input":"2022-08-14T14:17:19.899698Z","iopub.status.idle":"2022-08-14T14:17:19.916158Z","shell.execute_reply.started":"2022-08-14T14:17:19.899661Z","shell.execute_reply":"2022-08-14T14:17:19.914820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gradient boosting tunning\n\nGBC = GradientBoostingClassifier(random_state=random_state)\n\n# gb_param_grid = {'loss' : [\"deviance\"],\n#               'n_estimators' : [600 , 800],\n#               'learning_rate': [0.01],\n#               'max_depth': [14 , 16],\n#               'min_samples_leaf': [20 , 25],\n#               'max_features': [0.03 , 0.05 ,0.1] \n#               }\n\ngb_param_grid = {\n              'learning_rate': [0.01],\n                \"max_depth\" : [14],\n              'min_samples_leaf': [25],\n              'max_features': [0.05] , \n                \"n_estimators\" : [600]\n              }\n\ngsGBC = GridSearchCV(GBC,param_grid = gb_param_grid, cv=5, scoring=\"accuracy\", verbose = 1)\n\ngsGBC.fit(X_train,Y_train)\n\nGBC_best = gsGBC.best_estimator_\n\nprint(\"The best Model Parameters is :\")\nprint(GBC_best)\nprint(f\"With Cross Validation Score = {gsGBC.best_score_}\")\n","metadata":{"execution":{"iopub.status.busy":"2022-08-14T15:11:33.217899Z","iopub.execute_input":"2022-08-14T15:11:33.218359Z","iopub.status.idle":"2022-08-14T15:11:56.030853Z","shell.execute_reply.started":"2022-08-14T15:11:33.218288Z","shell.execute_reply":"2022-08-14T15:11:56.029533Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curve(GBC_best , \"Gradient Boosting\" , X_train , Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T16:00:50.972548Z","iopub.execute_input":"2022-08-14T16:00:50.972992Z","iopub.status.idle":"2022-08-14T16:01:29.288436Z","shell.execute_reply.started":"2022-08-14T16:00:50.972956Z","shell.execute_reply":"2022-08-14T16:01:29.287262Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# RFC Parameters tunning \nRFC = RandomForestClassifier(random_state=random_state)\n\n## Search grid for optimal parameters\n# rf_param_grid = {\"max_depth\": [16],\n#               \"max_features\": [0.2 ],\n#               \"min_samples_split\": [5],\n#               \"min_samples_leaf\": [15],\n#               \"bootstrap\": [True , False],\n#               \"n_estimators\" :[550 , 600],\n#               \"criterion\": [\"gini\"]}\n\nrf_param_grid = {\"max_depth\": [16],\n              \"max_features\": [0.2],\n              \"min_samples_leaf\": [15],\n              \"min_samples_split\": [5],\n              \"bootstrap\": [False],\n              \"n_estimators\" :[560],\n              \"criterion\": [\"gini\"]}\n\ngsRFC = GridSearchCV(RFC,param_grid = rf_param_grid, cv=5, scoring=\"accuracy\", verbose = 1)\n\ngsRFC.fit(X_train,Y_train)\n\nRFC_best = gsRFC.best_estimator_\n\nprint(\"The best Model Parameters is :\")\nprint(RFC_best)\nprint(f\"With Cross Validation Score = {gsRFC.best_score_}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T16:09:22.593074Z","iopub.execute_input":"2022-08-14T16:09:22.593491Z","iopub.status.idle":"2022-08-14T16:10:05.572331Z","shell.execute_reply.started":"2022-08-14T16:09:22.593458Z","shell.execute_reply":"2022-08-14T16:10:05.570912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curve(RFC_best , \"Random Forest\" , X_train , Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T15:58:55.801868Z","iopub.execute_input":"2022-08-14T15:58:55.802301Z","iopub.status.idle":"2022-08-14T15:59:35.433593Z","shell.execute_reply.started":"2022-08-14T15:58:55.802262Z","shell.execute_reply":"2022-08-14T15:59:35.432131Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#ExtraTrees \nExtC = ExtraTreesClassifier(random_state=random_state)\n\n## Search grid for optimal parameters\nex_param_grid = {\"max_depth\": [18],\n              \"max_features\": [0.5],\n              \"min_samples_split\": [5],\n              \"min_samples_leaf\": [15],\n              \"bootstrap\": [False],\n              \"n_estimators\" :[550],\n              \"criterion\": [\"gini\"],\n                }\n\n# ex_param_grid = {\"max_depth\": [18],\n#               \"max_features\": [10 , ],\n#               \"min_samples_split\": [10 , 8],\n#               \"min_samples_leaf\": [3 , 5],\n#               \"bootstrap\": [False],\n#               \"n_estimators\" :[300],\n#               \"criterion\": [\"gini\"]}\n\ngsExtC = GridSearchCV(ExtC,param_grid = ex_param_grid, cv=5, scoring=\"accuracy\", verbose = 1)\n\ngsExtC.fit(X_train,Y_train)\n\nExtC_best = gsExtC.best_estimator_\n\nprint(\"The best Model Parameters is :\")\nprint(ExtC_best)\nprint(f\"With Cross Validation Score = {gsExtC.best_score_}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T16:24:49.991707Z","iopub.execute_input":"2022-08-14T16:24:49.992104Z","iopub.status.idle":"2022-08-14T16:29:44.682344Z","shell.execute_reply.started":"2022-08-14T16:24:49.992073Z","shell.execute_reply":"2022-08-14T16:29:44.681153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curve(ExtC_best , \"Extra Trees\" , X_train , Y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Logistic Regression: \nLogReg = LogisticRegression(random_state=random_state)\n\n## Search grid for optimal parameters\nlog_reg_param_grid = {\n\"C\":[0.005]   ,\n    \"max_iter\" : [600]\n}\n\ngsLogReg = GridSearchCV(LogReg,param_grid = log_reg_param_grid, cv=5, scoring=\"accuracy\", verbose = 1)\n\ngsLogReg.fit(X_train,Y_train)\n\nLogReg_best = gsLogReg.best_estimator_\n\nprint(\"The best Model Parameters is :\")\nprint(LogReg_best)\nprint(f\"With Cross Validation Score = {gsLogReg.best_score_}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-14T16:35:29.643611Z","iopub.execute_input":"2022-08-14T16:35:29.644038Z","iopub.status.idle":"2022-08-14T16:36:16.458478Z","shell.execute_reply.started":"2022-08-14T16:35:29.644000Z","shell.execute_reply":"2022-08-14T16:36:16.456997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_learning_curve(LogReg_best , \"Logistic Regression\" , X_train , Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T16:36:44.666980Z","iopub.execute_input":"2022-08-14T16:36:44.668208Z","iopub.status.idle":"2022-08-14T16:36:50.319004Z","shell.execute_reply.started":"2022-08-14T16:36:44.668098Z","shell.execute_reply":"2022-08-14T16:36:50.317567Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from catboost import CatBoostClassifier\ncbc = CatBoostClassifier(verbose=0, n_estimators=600)\ncbc.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T17:37:27.935744Z","iopub.execute_input":"2022-08-14T17:37:27.936833Z","iopub.status.idle":"2022-08-14T17:37:30.974320Z","shell.execute_reply.started":"2022-08-14T17:37:27.936788Z","shell.execute_reply":"2022-08-14T17:37:30.973256Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cross_val_score(cbc, X_train, y = Y_train, scoring = \"accuracy\", cv = 5).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-14T17:38:27.297780Z","iopub.execute_input":"2022-08-14T17:38:27.298233Z","iopub.status.idle":"2022-08-14T17:38:40.846066Z","shell.execute_reply.started":"2022-08-14T17:38:27.298198Z","shell.execute_reply":"2022-08-14T17:38:40.845072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ============================================================\n# Train on all Data\n# ============================================================\nGBC_all_data = GBC_best.fit(X_train , Y_train)\nRFC_all_data = RFC_best.fit(X_train , Y_train)\nExtC_all_data = ExtC_best.fit(X_train , Y_train)\nLogReg_all_data = LogReg_best.fit(X_train , Y_train)\ncbc_all_data = cbc.fit(X_train , Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T17:42:19.135768Z","iopub.execute_input":"2022-08-14T17:42:19.136351Z","iopub.status.idle":"2022-08-14T17:42:44.841638Z","shell.execute_reply.started":"2022-08-14T17:42:19.136284Z","shell.execute_reply":"2022-08-14T17:42:44.840139Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"votingC = VotingClassifier(estimators=[('rfc', RFC_best), ('extc', ExtC_best),\n                                       ('gbc',GBC_best) , (\"logreg\" ,LogReg_best ) , (\"catboost\" , cbc)], voting='soft')\n\nvotingC = votingC.fit(X_train, Y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-14T17:42:50.034626Z","iopub.execute_input":"2022-08-14T17:42:50.035097Z","iopub.status.idle":"2022-08-14T17:43:15.754990Z","shell.execute_reply.started":"2022-08-14T17:42:50.035058Z","shell.execute_reply":"2022-08-14T17:43:15.753464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = pd.Series(votingC.predict(X_test).astype(bool), name=\"Transported\")\n\nresults = pd.concat([Test_Id,predictions],axis=1)\n\nresults.to_csv(\"submission.csv\",index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T17:08:41.398209Z","iopub.execute_input":"2022-08-13T17:08:41.398630Z","iopub.status.idle":"2022-08-13T17:08:42.114264Z","shell.execute_reply.started":"2022-08-13T17:08:41.398588Z","shell.execute_reply":"2022-08-13T17:08:42.112551Z"},"trusted":true},"execution_count":null,"outputs":[]}]}