{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# __Spaceship Titanic__","metadata":{}},{"cell_type":"markdown","source":"I am not a expert in data science. This is my first try in this competition after lot of researches. There might be some issues. I hope you like this. These all written by myself and please leave a comment.. ","metadata":{}},{"cell_type":"markdown","source":"## __Workflow Stages__","metadata":{}},{"cell_type":"markdown","source":"1. Problem definition.\n2. Acquire training and testing data.\n3. Wrangle, Prepare, Cleanse the data.\n4. Analyze, identify patterns and explore the data.\n5. Model, Predict and solve the problem.\n6. Visualize, report, and present the problem solving steps and final solution.\n7. Supply or submit the results.","metadata":{}},{"cell_type":"markdown","source":"The workflow indicates general sequence of how each stage may follow the other. However there are use cases with exceptions.","metadata":{}},{"cell_type":"markdown","source":"* We may combine multiple workflow stages. We may analyze by visualizing data.\n* Perform a stage earlier than indicated. We may analyze data before and after wrangling.\n* Perform a stage multiple times in our workflow. Visualize stage may be used multiple times.\n* We may drop a stage altogether","metadata":{}},{"cell_type":"markdown","source":"## __Problem definition__","metadata":{}},{"cell_type":"markdown","source":"Predict which passengers are transported to an alternate dimension","metadata":{}},{"cell_type":"markdown","source":"Welcome to the year 2912, where your data science skills are needed to solve a cosmic mystery. We've received a transmission from four lightyears away and things aren't looking good.\n\nThe Spaceship Titanic was an interstellar passenger liner launched a month ago. With almost 13,000 passengers on board, the vessel set out on its maiden voyage transporting emigrants from our solar system to three newly habitable exoplanets orbiting nearby stars.\n\nWhile rounding Alpha Centauri en route to its first destination—the torrid 55 Cancri E—the unwary Spaceship Titanic collided with a spacetime anomaly hidden within a dust cloud. Sadly, it met a similar fate as its namesake from 1000 years before. Though the ship stayed intact, almost half of the passengers were transported to an alternate dimension!","metadata":{}},{"cell_type":"markdown","source":"* To help rescue crews and retrieve the lost passengers, we have to predict which passengers were transported by the anomaly using records recovered from the spaceship’s damaged computer system.","metadata":{}},{"cell_type":"markdown","source":"## __Workflow goals__","metadata":{}},{"cell_type":"markdown","source":"The datascience solutions workflow solves for seven major goals.","metadata":{}},{"cell_type":"markdown","source":"__Classifing__. We may want to classify or categorize our samples. We may also want to undestand the implications or correlations of different classes with ourr solution goals\n\n__Correlation__. One can approach the problem based on available features within the training dataset. Which features within the dataset contribute significantly to our solution goal? Statistically speaking is there a correlation among a feature and solution goal? As the feature values change does the solution state change as well, and visa-versa? This can be tested both for numerical and categorical features in the given dataset. We may also want to determine correlation among features other than survival for subsequent goals and workflow stages. Correlating certain features may help in creating, completing, or correcting features.\n\n__Converting__. For modelling stage, one needs to prepare the data. Depending on the choice of model algorithm one may require all feature to be converted to numerical equivalent values. So for instance converting text categorical values to numerical values.\n\n__Completing__. Data preparation may also require us to estimate any missing values within a feature. Model algorithm may work best when there are no missing values.\n\n__Correcting__. We may also analyze the given training dataset for errors or possibly innacurate values within features and try to corrent these values or exclude the samples containing the errors. One way to do this is to detect any outliers among our samples or features. We may also completely discard a feature if it is not contributing to the analysis or may significantly skew the result.\n\n__Creating__. Can we create new feature based on an existing feature or a set of features, such that the new feature follows the correlation, conversion, completeness goals.\n\n__Charting__. How to slect the right visualization plots and charts depending on nature of the data and the solution goals.","metadata":{}},{"cell_type":"code","source":"# data analysis and wrangling\nimport pandas as pd\nimport numpy as np\nimport random as rnd\nfrom sklearn.preprocessing import MinMaxScaler\n\n# visualization\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\n# model training\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC, LinearSVC\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import Perceptron\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.ensemble import GradientBoostingClassifier","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:06.626881Z","iopub.execute_input":"2022-08-03T09:37:06.628153Z","iopub.status.idle":"2022-08-03T09:37:08.199591Z","shell.execute_reply.started":"2022-08-03T09:37:06.627987Z","shell.execute_reply":"2022-08-03T09:37:08.198203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Acquire data__","metadata":{}},{"cell_type":"markdown","source":"Import train and test data sets using python pandas library and we create a variable called combine with combining train and test data sets together.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/spaceship-titanic/train.csv')\ntest_df = pd.read_csv('../input/spaceship-titanic/test.csv')\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.201778Z","iopub.execute_input":"2022-08-03T09:37:08.202335Z","iopub.status.idle":"2022-08-03T09:37:08.289405Z","shell.execute_reply.started":"2022-08-03T09:37:08.202294Z","shell.execute_reply":"2022-08-03T09:37:08.288080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Analize data by describing__","metadata":{}},{"cell_type":"markdown","source":"With pandas library we can use it to answer following questions.","metadata":{}},{"cell_type":"markdown","source":"#### __Which features are available in the dataset?__","metadata":{}},{"cell_type":"code","source":"print(test_df.columns.values)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.290844Z","iopub.execute_input":"2022-08-03T09:37:08.291274Z","iopub.status.idle":"2022-08-03T09:37:08.298662Z","shell.execute_reply.started":"2022-08-03T09:37:08.291231Z","shell.execute_reply":"2022-08-03T09:37:08.297788Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Which features are categorical?__","metadata":{}},{"cell_type":"markdown","source":"THese values classify the samples into sets of similar samples. (nominal, ordinal, ratio, interval based)","metadata":{}},{"cell_type":"markdown","source":"* Categorical\n    * HomePlanet\n    * CryoSleep\n    * Destination\n    * VIP\n    * Transported","metadata":{}},{"cell_type":"markdown","source":"#### __Which features are numerical?__","metadata":{}},{"cell_type":"markdown","source":"These values change from sample to sample. (Discrete, Continuous, Timeseries based)","metadata":{}},{"cell_type":"markdown","source":"* Continuous\n    * Age\n    * RoomService\n    * FoodCourt\n    * ShoppingMall\n    * Spa\n    * VRDeck","metadata":{}},{"cell_type":"code","source":"train_df.head(5)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.300961Z","iopub.execute_input":"2022-08-03T09:37:08.301498Z","iopub.status.idle":"2022-08-03T09:37:08.338670Z","shell.execute_reply.started":"2022-08-03T09:37:08.301462Z","shell.execute_reply":"2022-08-03T09:37:08.337747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Which features are mixed types?__","metadata":{}},{"cell_type":"markdown","source":"Numerical, alphanumerical data within same feature.","metadata":{}},{"cell_type":"markdown","source":"* Alphanumerical\n    * Cabin","metadata":{}},{"cell_type":"markdown","source":"#### __Which feature may contain errors of typos?__","metadata":{}},{"cell_type":"markdown","source":"None of feature contain errors still","metadata":{}},{"cell_type":"markdown","source":"#### __Which feature contain blank, null, empty values?__","metadata":{}},{"cell_type":"markdown","source":"* In train data set,\n    * HomePlanet      201\n    * CryoSleep       217\n    * Cabin           199\n    * Destination     182\n    * Age             179\n    * VIP             203\n    * RoomService     181\n    * FoodCourt       183\n    * ShoppingMall    208\n    * Spa             183\n    * VRDeck          188\n    * Name            200\n* In test data set,\n    * HomePlanet       87\n    * CryoSleep        93\n    * Cabin           100\n    * Destination      92\n    * Age              91\n    * VIP              93\n    * RoomService      82\n    * FoodCourt       106\n    * ShoppingMall     98\n    * Spa             101\n    * VRDeck           80\n    * Name             94","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.340310Z","iopub.execute_input":"2022-08-03T09:37:08.340938Z","iopub.status.idle":"2022-08-03T09:37:08.353354Z","shell.execute_reply.started":"2022-08-03T09:37:08.340883Z","shell.execute_reply":"2022-08-03T09:37:08.352502Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.354891Z","iopub.execute_input":"2022-08-03T09:37:08.355501Z","iopub.status.idle":"2022-08-03T09:37:08.371011Z","shell.execute_reply.started":"2022-08-03T09:37:08.355465Z","shell.execute_reply":"2022-08-03T09:37:08.369979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What are the data types for various features.?__","metadata":{}},{"cell_type":"markdown","source":"* 6 features are float\n* 7 features are object(strings)\n* 1 feature is boolian (bool)","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.373277Z","iopub.execute_input":"2022-08-03T09:37:08.374400Z","iopub.status.idle":"2022-08-03T09:37:08.413378Z","shell.execute_reply.started":"2022-08-03T09:37:08.374351Z","shell.execute_reply":"2022-08-03T09:37:08.412182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What is the distribution of numerical feature values across the sample?__","metadata":{}},{"cell_type":"markdown","source":"* This dataset has 8693 samples\n* A very few elders  within age range 60-79","metadata":{}},{"cell_type":"code","source":"train_df.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.415098Z","iopub.execute_input":"2022-08-03T09:37:08.415552Z","iopub.status.idle":"2022-08-03T09:37:08.426330Z","shell.execute_reply.started":"2022-08-03T09:37:08.415517Z","shell.execute_reply":"2022-08-03T09:37:08.424805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.describe()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.428854Z","iopub.execute_input":"2022-08-03T09:37:08.430005Z","iopub.status.idle":"2022-08-03T09:37:08.477326Z","shell.execute_reply.started":"2022-08-03T09:37:08.429921Z","shell.execute_reply":"2022-08-03T09:37:08.475831Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __What is the distribution of categorical features?__","metadata":{}},{"cell_type":"markdown","source":"* Most of Names are unique but only 2 samples not\n* Earth is the Home planet of most passangers\n* 5439 out of 8492, passangers didn't in CryoSleep\n* There are 6560 cabins\n* Most Passangers wasn't VIP\n* Around half of the samples(4378) were transported","metadata":{}},{"cell_type":"code","source":"train_df.describe(include=['O', bool])","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.483142Z","iopub.execute_input":"2022-08-03T09:37:08.483935Z","iopub.status.idle":"2022-08-03T09:37:08.537196Z","shell.execute_reply.started":"2022-08-03T09:37:08.483881Z","shell.execute_reply":"2022-08-03T09:37:08.535657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Assumptions based on data analysis__","metadata":{}},{"cell_type":"markdown","source":"We arrived following assumptions based on data analysis so far.","metadata":{}},{"cell_type":"markdown","source":"* __Correlation__ : We want to know how well does each feature correlated with Transport","metadata":{}},{"cell_type":"markdown","source":"* __Completeting__: We want to complete HomePlanet, CryoSleep, Cabin, Destination, Age, VIP, RoomService, FoodCourt, ShoppingMall, Spa, VRDeck features. I think they are correlated with transport","metadata":{}},{"cell_type":"markdown","source":"* __Creating__:\n    * We should create a value ranges for Age, RoomService, FoodCourt, ShoppingMall, Spa, VRDeck features","metadata":{}},{"cell_type":"markdown","source":"## __Analyze by pivolating features__","metadata":{}},{"cell_type":"markdown","source":"We can analyze our feature correlation by pivolating features against each other. So we can use HomePlanet, CryoSleep, Destination, VIP features against transported feature.","metadata":{}},{"cell_type":"markdown","source":"* __HomePlanet__ We can observe significant correlation among HomePlanet=Europa. So we decide to use HomePlanet feature for our model training.\n\n* __CryoSleep__ We can see cryo sleepers had higher transport rate\n\n* __Destination__ 55 Cancri e had higher transport rate. So we consider destination feature also in the model training.\n\n* __VIP__ Non vip passangers also had higher transport rate.","metadata":{}},{"cell_type":"code","source":"train_df[['HomePlanet','Transported']].groupby(['HomePlanet'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.539518Z","iopub.execute_input":"2022-08-03T09:37:08.539932Z","iopub.status.idle":"2022-08-03T09:37:08.562577Z","shell.execute_reply.started":"2022-08-03T09:37:08.539898Z","shell.execute_reply":"2022-08-03T09:37:08.561177Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['CryoSleep', 'Transported']].groupby(['CryoSleep'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.564625Z","iopub.execute_input":"2022-08-03T09:37:08.565431Z","iopub.status.idle":"2022-08-03T09:37:08.587372Z","shell.execute_reply.started":"2022-08-03T09:37:08.565384Z","shell.execute_reply":"2022-08-03T09:37:08.586133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['Destination', 'Transported']].groupby(['Destination'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.588691Z","iopub.execute_input":"2022-08-03T09:37:08.589075Z","iopub.status.idle":"2022-08-03T09:37:08.607705Z","shell.execute_reply.started":"2022-08-03T09:37:08.589044Z","shell.execute_reply":"2022-08-03T09:37:08.606606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['VIP', 'Transported']].groupby(['VIP'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.609676Z","iopub.execute_input":"2022-08-03T09:37:08.610599Z","iopub.status.idle":"2022-08-03T09:37:08.630623Z","shell.execute_reply.started":"2022-08-03T09:37:08.610546Z","shell.execute_reply":"2022-08-03T09:37:08.629487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Analyze by visualizing data__","metadata":{}},{"cell_type":"markdown","source":"We can continue our assumptions using visualizing data","metadata":{}},{"cell_type":"markdown","source":"#### __Correlation numerical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * Age < 4 had Higher transport rate\n    * Most passangers were in 15-40 Age range\n    * Passangers who belongs to age 20-30 did not transported","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Consider Age feature for model training\n    * Create a age range","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df,col='Transported')\ngrid.map(plt.hist, 'Age', bins=20)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:08.632065Z","iopub.execute_input":"2022-08-03T09:37:08.632499Z","iopub.status.idle":"2022-08-03T09:37:09.156274Z","shell.execute_reply.started":"2022-08-03T09:37:08.632461Z","shell.execute_reply":"2022-08-03T09:37:09.155030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Correlation numerical and categorical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * Passangers who ware HomePlannet Earth had low Transport rate\n    * HomePlannet mars and Earth passengers who belongs to age < 4 mostly Transported\n    * There is a correlation between VIP and Transport\n    \n* __Decisions__\n    * Consider HomePlanet feature for model training\n    * Introduce VIP and Destination Features also for model training","metadata":{}},{"cell_type":"markdown","source":"We can combine multiple features for identify correlation using a single plot. This can be done with numeric and categorical features which have nymeric values","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, col='Transported', row='HomePlanet', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:09.157983Z","iopub.execute_input":"2022-08-03T09:37:09.158688Z","iopub.status.idle":"2022-08-03T09:37:10.574284Z","shell.execute_reply.started":"2022-08-03T09:37:09.158637Z","shell.execute_reply":"2022-08-03T09:37:10.572989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, col='Transported', row='Destination', size=2.2, aspect=1.6)\ngrid.map(plt.hist, 'Age', alpha=.5, bins=20)\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:10.576577Z","iopub.execute_input":"2022-08-03T09:37:10.577441Z","iopub.status.idle":"2022-08-03T09:37:12.250567Z","shell.execute_reply.started":"2022-08-03T09:37:10.577390Z","shell.execute_reply":"2022-08-03T09:37:12.249517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, row='Destination', col='Transported', size=2.2, aspect=1.6)\ngrid.map(sns.barplot, 'VIP', 'Age', alpha=.5, ci=None)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:12.252046Z","iopub.execute_input":"2022-08-03T09:37:12.252380Z","iopub.status.idle":"2022-08-03T09:37:13.201166Z","shell.execute_reply.started":"2022-08-03T09:37:12.252351Z","shell.execute_reply":"2022-08-03T09:37:13.199546Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Correlation categorical features__","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * Mars value is changing one Destination to another\n    * We can see all 3 destination Cryo sleepers were transported\n\n* __Decisions__\n    * Use CryoSleep to our model training","metadata":{}},{"cell_type":"code","source":"grid = sns.FacetGrid(train_df, row='Destination', size=2.2, aspect=1.6)\ngrid.map(sns.pointplot, 'HomePlanet', 'Transported', 'CryoSleep', pallete='deep')\ngrid.add_legend()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:13.203148Z","iopub.execute_input":"2022-08-03T09:37:13.203596Z","iopub.status.idle":"2022-08-03T09:37:14.518140Z","shell.execute_reply.started":"2022-08-03T09:37:13.203551Z","shell.execute_reply":"2022-08-03T09:37:14.517357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Wrangle data__","metadata":{}},{"cell_type":"markdown","source":"We are collected several assumptions and decisions regarding our datasets and solution requirements. So far we did not change any value or any feature. Lets execute our decisions and assumptions for correcting, creating, and completing goals.","metadata":{}},{"cell_type":"markdown","source":"#### __Creating new feature extracting from existing__","metadata":{}},{"cell_type":"markdown","source":"__Cabin__","metadata":{}},{"cell_type":"markdown","source":"Before dropping Cabin feature we want to create a new feature called CabinLetter by extracting the first Capital letter and last letter in the Cabin feature","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * It seems there might be correlation between CabinLetter and Transported rate\n    * CabinLetter B had higher Transported rate","metadata":{}},{"cell_type":"markdown","source":"* __Decisions__\n    * Consider CabinLatter1 and 2 for model training","metadata":{}},{"cell_type":"code","source":"train_df['CabinLetter1'] = train_df['Cabin'].apply(lambda x: str(x).split('/')[0])\ntest_df['CabinLetter1'] = test_df['Cabin'].apply(lambda x: str(x).split('/')[0])\ntrain_df['CabinLetter2'] = train_df['Cabin'].apply(lambda x: str(x).split('/')[-1])\ntest_df['CabinLetter2'] = test_df['Cabin'].apply(lambda x: str(x).split('/')[-1])\n","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.519302Z","iopub.execute_input":"2022-08-03T09:37:14.520315Z","iopub.status.idle":"2022-08-03T09:37:14.545353Z","shell.execute_reply.started":"2022-08-03T09:37:14.520283Z","shell.execute_reply":"2022-08-03T09:37:14.544297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CabinNumber'] = dataset.Cabin.str.extract('([0-9999].)', expand=False)\n    dataset['CabinNumber'] = dataset['CabinNumber'].apply(lambda x: str(x).split('/')[0])\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.546731Z","iopub.execute_input":"2022-08-03T09:37:14.547206Z","iopub.status.idle":"2022-08-03T09:37:14.573916Z","shell.execute_reply.started":"2022-08-03T09:37:14.547161Z","shell.execute_reply":"2022-08-03T09:37:14.572692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.575767Z","iopub.execute_input":"2022-08-03T09:37:14.576585Z","iopub.status.idle":"2022-08-03T09:37:14.600283Z","shell.execute_reply.started":"2022-08-03T09:37:14.576543Z","shell.execute_reply":"2022-08-03T09:37:14.599038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['CabinLetter1', 'Transported']].groupby(['CabinLetter1'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.601843Z","iopub.execute_input":"2022-08-03T09:37:14.602604Z","iopub.status.idle":"2022-08-03T09:37:14.621439Z","shell.execute_reply.started":"2022-08-03T09:37:14.602559Z","shell.execute_reply":"2022-08-03T09:37:14.619962Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['CabinLetter2', 'Transported']].groupby(['CabinLetter2'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.623292Z","iopub.execute_input":"2022-08-03T09:37:14.623720Z","iopub.status.idle":"2022-08-03T09:37:14.639515Z","shell.execute_reply.started":"2022-08-03T09:37:14.623679Z","shell.execute_reply":"2022-08-03T09:37:14.638507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Name__","metadata":{}},{"cell_type":"markdown","source":"Is there any correlation with name feature?","metadata":{}},{"cell_type":"markdown","source":"* __Observations__\n    * We can see their is a correlation between Family name and Transported rate\n    * Some family name categories had higher transported rate while some not had\n* __Decisions__\n    * We want to group family names and convert their names with count\n    * Use family name feature for model training","metadata":{}},{"cell_type":"markdown","source":"By analyzing the name feature we can see their last name is similar most times. So we can extract that name as a family name. ","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilyName'] = dataset['Name'].apply(lambda x: str(x).split(' ')[-1])\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.641426Z","iopub.execute_input":"2022-08-03T09:37:14.641774Z","iopub.status.idle":"2022-08-03T09:37:14.659286Z","shell.execute_reply.started":"2022-08-03T09:37:14.641745Z","shell.execute_reply":"2022-08-03T09:37:14.658364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df[['FamilyName', 'Transported']].groupby(['FamilyName'], as_index=False).mean()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.662797Z","iopub.execute_input":"2022-08-03T09:37:14.663159Z","iopub.status.idle":"2022-08-03T09:37:14.692738Z","shell.execute_reply.started":"2022-08-03T09:37:14.663114Z","shell.execute_reply":"2022-08-03T09:37:14.691808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets drop Cabin and Name features. It's no need now","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset.drop(['Cabin', 'Name'], axis=1, inplace=True)\ncombine = [train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.694255Z","iopub.execute_input":"2022-08-03T09:37:14.694570Z","iopub.status.idle":"2022-08-03T09:37:14.704838Z","shell.execute_reply.started":"2022-08-03T09:37:14.694537Z","shell.execute_reply":"2022-08-03T09:37:14.704095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Completing a categorical feature__","metadata":{}},{"cell_type":"markdown","source":"__Cabin__","metadata":{}},{"cell_type":"markdown","source":"But there is a problem with Cabbin letters and Cabin number features. Because there are some missing values. So we want to fill these with most frequently values.","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.715639Z","iopub.execute_input":"2022-08-03T09:37:14.716789Z","iopub.status.idle":"2022-08-03T09:37:14.732387Z","shell.execute_reply.started":"2022-08-03T09:37:14.716746Z","shell.execute_reply":"2022-08-03T09:37:14.731560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This shows there isn't any null values in CabinLetter features. So let's change their values manually","metadata":{}},{"cell_type":"code","source":"freq_cb1 = train_df.CabinLetter1.dropna().mode()[0]\nfreq_cb2 = train_df.CabinLetter2.dropna().mode()[0]\nfreq_cb1, freq_cb2","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.733560Z","iopub.execute_input":"2022-08-03T09:37:14.734611Z","iopub.status.idle":"2022-08-03T09:37:14.745802Z","shell.execute_reply.started":"2022-08-03T09:37:14.734577Z","shell.execute_reply":"2022-08-03T09:37:14.744737Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CabinLetter1'] = dataset.CabinLetter1.replace('nan', 'Unknown')\n    dataset['CabinLetter2'] = dataset.CabinLetter2.replace('nan', 'Unknown')\n    dataset['CabinNumber'] = dataset['CabinNumber'].replace('nan', 'Unknown')\ncombine = [train_df,test_df]\n","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.747500Z","iopub.execute_input":"2022-08-03T09:37:14.747918Z","iopub.status.idle":"2022-08-03T09:37:14.759533Z","shell.execute_reply.started":"2022-08-03T09:37:14.747878Z","shell.execute_reply":"2022-08-03T09:37:14.758492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.CabinLetter1.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.761444Z","iopub.execute_input":"2022-08-03T09:37:14.761862Z","iopub.status.idle":"2022-08-03T09:37:14.776631Z","shell.execute_reply.started":"2022-08-03T09:37:14.761822Z","shell.execute_reply":"2022-08-03T09:37:14.775184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.CabinLetter2.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.778767Z","iopub.execute_input":"2022-08-03T09:37:14.779237Z","iopub.status.idle":"2022-08-03T09:37:14.789101Z","shell.execute_reply.started":"2022-08-03T09:37:14.779197Z","shell.execute_reply":"2022-08-03T09:37:14.788162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Completing and Converting a  categotical feature__","metadata":{}},{"cell_type":"markdown","source":"__FamilyName__","metadata":{}},{"cell_type":"markdown","source":"Now we want to complete FamilyName feature and take a count of that category values. Then create a new feature called FamilyMembers","metadata":{}},{"cell_type":"markdown","source":"Replace nan values with 'Unknown'","metadata":{}},{"cell_type":"code","source":"train_df['FamilyName'] = train_df['FamilyName'].replace('nan', 'Unknown')\ntest_df['FamilyName'] = test_df['FamilyName'].replace('nan', 'Unknown')\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.790736Z","iopub.execute_input":"2022-08-03T09:37:14.791438Z","iopub.status.idle":"2022-08-03T09:37:14.799408Z","shell.execute_reply.started":"2022-08-03T09:37:14.791397Z","shell.execute_reply":"2022-08-03T09:37:14.798617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Create a dictionary for familynames","metadata":{}},{"cell_type":"code","source":"family_name_dic = train_df['FamilyName'].value_counts().to_dict()\nfamily_name_dic['Unknown'] = 0","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.800731Z","iopub.execute_input":"2022-08-03T09:37:14.801683Z","iopub.status.idle":"2022-08-03T09:37:14.811978Z","shell.execute_reply.started":"2022-08-03T09:37:14.801643Z","shell.execute_reply":"2022-08-03T09:37:14.811137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Map the dictionary to datasets, fill empty vaues, drop the family name column","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['FamilyMember'] = dataset['FamilyName']\n    dataset['FamilyMember'] = dataset['FamilyMember'].map(family_name_dic)\n    dataset['FamilyMember'] = dataset['FamilyMember'].fillna(0).astype(int)\n    dataset.drop(['FamilyName'], axis=1, inplace=True)\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.813601Z","iopub.execute_input":"2022-08-03T09:37:14.814443Z","iopub.status.idle":"2022-08-03T09:37:14.832291Z","shell.execute_reply.started":"2022-08-03T09:37:14.814401Z","shell.execute_reply":"2022-08-03T09:37:14.831414Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Complleting Categorical features__","metadata":{}},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.833848Z","iopub.execute_input":"2022-08-03T09:37:14.834584Z","iopub.status.idle":"2022-08-03T09:37:14.849228Z","shell.execute_reply.started":"2022-08-03T09:37:14.834541Z","shell.execute_reply":"2022-08-03T09:37:14.848031Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__HomePlanet__","metadata":{}},{"cell_type":"markdown","source":"Lets fill empty values in HomePlanet feature with unknown","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['HomePlanet'] = dataset['HomePlanet'].fillna('Unknown')\ncombine = [train_df,test_df]\ntrain_df[['HomePlanet', 'Transported']].groupby(['HomePlanet'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.851422Z","iopub.execute_input":"2022-08-03T09:37:14.851810Z","iopub.status.idle":"2022-08-03T09:37:14.871331Z","shell.execute_reply.started":"2022-08-03T09:37:14.851771Z","shell.execute_reply":"2022-08-03T09:37:14.870277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__CryoSleep__","metadata":{}},{"cell_type":"markdown","source":"Most passangers were in  without cryo sleeping. So lets fill empty cryosleep values with false","metadata":{}},{"cell_type":"code","source":"freq_sleep = train_df.CryoSleep.dropna().mode()[0]\nfreq_sleep","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.872867Z","iopub.execute_input":"2022-08-03T09:37:14.873969Z","iopub.status.idle":"2022-08-03T09:37:14.883856Z","shell.execute_reply.started":"2022-08-03T09:37:14.873936Z","shell.execute_reply":"2022-08-03T09:37:14.882672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CryoSleep'] = dataset['CryoSleep'].fillna(freq_sleep)\ncombine = [train_df,test_df]\ntrain_df[['CryoSleep', 'Transported']].groupby(['CryoSleep'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.885526Z","iopub.execute_input":"2022-08-03T09:37:14.886347Z","iopub.status.idle":"2022-08-03T09:37:14.906072Z","shell.execute_reply.started":"2022-08-03T09:37:14.886305Z","shell.execute_reply":"2022-08-03T09:37:14.905353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Destination__","metadata":{}},{"cell_type":"markdown","source":"Also fill empty destination values with unknown","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Destination'] = dataset['Destination'].fillna('Unknown')\ncombine = [train_df,test_df]\ntrain_df[['Destination', 'Transported']].groupby(['Destination'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.907094Z","iopub.execute_input":"2022-08-03T09:37:14.907767Z","iopub.status.idle":"2022-08-03T09:37:14.924852Z","shell.execute_reply.started":"2022-08-03T09:37:14.907733Z","shell.execute_reply":"2022-08-03T09:37:14.923813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__VIP__","metadata":{}},{"cell_type":"markdown","source":"Fill vip empty values with most frequency choice","metadata":{}},{"cell_type":"code","source":"freq_vip = train_df['VIP'].dropna().mode()[0]\nfreq_vip","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.926106Z","iopub.execute_input":"2022-08-03T09:37:14.926515Z","iopub.status.idle":"2022-08-03T09:37:14.935685Z","shell.execute_reply.started":"2022-08-03T09:37:14.926485Z","shell.execute_reply":"2022-08-03T09:37:14.934549Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['VIP'] = dataset['VIP'].fillna(freq_vip)\ncombine = [train_df,test_df]\ntrain_df[['VIP', 'Transported']].groupby(['VIP'], as_index=False).mean().sort_values(by='Transported', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.937014Z","iopub.execute_input":"2022-08-03T09:37:14.938173Z","iopub.status.idle":"2022-08-03T09:37:14.957221Z","shell.execute_reply.started":"2022-08-03T09:37:14.938111Z","shell.execute_reply":"2022-08-03T09:37:14.955942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Complete Numerical features__","metadata":{}},{"cell_type":"markdown","source":"Now we should fill null empty values in numerical features. There are several ways doing that","metadata":{}},{"cell_type":"markdown","source":"1. Generate random value between mean and standard deviation.\n2. Guessing missing values by using other correlated features.(Gender, Pclass)\n3. Combining 1 and 2 methods and use random value between mean and std base on set of Pclass and Gender combinations","metadata":{}},{"cell_type":"markdown","source":"Method 1 introduce a random noice for our dataset. But for method 2 and 3 still I can't find a good correlation between numerical features. So I fill these numeric empty values using their median","metadata":{}},{"cell_type":"markdown","source":"__Instead of filling one by one I'll fill all with their median and some with zero__","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['Age'] = dataset['Age'].fillna(dataset.Age.median())\n    dataset['RoomService'] = dataset['RoomService'].fillna(0)\n    dataset['FoodCourt'] = dataset['FoodCourt'].fillna(0)\n    dataset['ShoppingMall'] = dataset['ShoppingMall'].fillna(0)\n    dataset['Spa'] = dataset['Spa'].fillna(0)\n    dataset['VRDeck'] = dataset['VRDeck'].fillna(0)\ncombine =[train_df,test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.958395Z","iopub.execute_input":"2022-08-03T09:37:14.959507Z","iopub.status.idle":"2022-08-03T09:37:14.972044Z","shell.execute_reply.started":"2022-08-03T09:37:14.959475Z","shell.execute_reply":"2022-08-03T09:37:14.971171Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.973388Z","iopub.execute_input":"2022-08-03T09:37:14.974380Z","iopub.status.idle":"2022-08-03T09:37:14.986384Z","shell.execute_reply.started":"2022-08-03T09:37:14.974346Z","shell.execute_reply":"2022-08-03T09:37:14.985372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### __Converting Categorical values to numeric__","metadata":{}},{"cell_type":"markdown","source":"When machine learning we can only train numerical values. Therefore we should convert our categotical features to numeric. ","metadata":{}},{"cell_type":"markdown","source":"__HomePlanet__","metadata":{}},{"cell_type":"code","source":"train_df.HomePlanet.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.987801Z","iopub.execute_input":"2022-08-03T09:37:14.988382Z","iopub.status.idle":"2022-08-03T09:37:14.996217Z","shell.execute_reply.started":"2022-08-03T09:37:14.988351Z","shell.execute_reply":"2022-08-03T09:37:14.994998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_hp = {'Unknown':0, 'Europa':1, 'Earth':2, 'Mars':3}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:14.997592Z","iopub.execute_input":"2022-08-03T09:37:14.997899Z","iopub.status.idle":"2022-08-03T09:37:15.004973Z","shell.execute_reply.started":"2022-08-03T09:37:14.997871Z","shell.execute_reply":"2022-08-03T09:37:15.003867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__CryoSleep__","metadata":{}},{"cell_type":"code","source":"train_df.CryoSleep.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.006764Z","iopub.execute_input":"2022-08-03T09:37:15.007153Z","iopub.status.idle":"2022-08-03T09:37:15.018506Z","shell.execute_reply.started":"2022-08-03T09:37:15.007061Z","shell.execute_reply":"2022-08-03T09:37:15.017606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_cs = {'False':0, 'True':1}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.020310Z","iopub.execute_input":"2022-08-03T09:37:15.020813Z","iopub.status.idle":"2022-08-03T09:37:15.027667Z","shell.execute_reply.started":"2022-08-03T09:37:15.020780Z","shell.execute_reply":"2022-08-03T09:37:15.026594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Destination__","metadata":{}},{"cell_type":"code","source":"train_df.Destination.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.029522Z","iopub.execute_input":"2022-08-03T09:37:15.030318Z","iopub.status.idle":"2022-08-03T09:37:15.042441Z","shell.execute_reply.started":"2022-08-03T09:37:15.030286Z","shell.execute_reply":"2022-08-03T09:37:15.041489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_d = {'Unknown':0, 'TRAPPIST-1e':1, 'PSO J318.5-22':2, '55 Cancri e':3}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.044108Z","iopub.execute_input":"2022-08-03T09:37:15.044784Z","iopub.status.idle":"2022-08-03T09:37:15.051098Z","shell.execute_reply.started":"2022-08-03T09:37:15.044752Z","shell.execute_reply":"2022-08-03T09:37:15.049979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__VIP__","metadata":{}},{"cell_type":"code","source":"train_df.VIP.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.052236Z","iopub.execute_input":"2022-08-03T09:37:15.052919Z","iopub.status.idle":"2022-08-03T09:37:15.063007Z","shell.execute_reply.started":"2022-08-03T09:37:15.052870Z","shell.execute_reply":"2022-08-03T09:37:15.062265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_v = {'False':0, 'True':1}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.064138Z","iopub.execute_input":"2022-08-03T09:37:15.064841Z","iopub.status.idle":"2022-08-03T09:37:15.072342Z","shell.execute_reply.started":"2022-08-03T09:37:15.064810Z","shell.execute_reply":"2022-08-03T09:37:15.071504Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Transported__","metadata":{}},{"cell_type":"code","source":"train_df.Transported.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.073617Z","iopub.execute_input":"2022-08-03T09:37:15.074186Z","iopub.status.idle":"2022-08-03T09:37:15.086736Z","shell.execute_reply.started":"2022-08-03T09:37:15.074140Z","shell.execute_reply":"2022-08-03T09:37:15.085673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_t = {'False':0, 'True':1}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.087971Z","iopub.execute_input":"2022-08-03T09:37:15.088388Z","iopub.status.idle":"2022-08-03T09:37:15.094797Z","shell.execute_reply.started":"2022-08-03T09:37:15.088358Z","shell.execute_reply":"2022-08-03T09:37:15.093940Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__CabinLetter1__","metadata":{}},{"cell_type":"code","source":"train_df.CabinLetter1.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.096385Z","iopub.execute_input":"2022-08-03T09:37:15.097102Z","iopub.status.idle":"2022-08-03T09:37:15.106729Z","shell.execute_reply.started":"2022-08-03T09:37:15.097054Z","shell.execute_reply":"2022-08-03T09:37:15.105931Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_cb1 = {'Unknown':0, 'B':1, 'F':2, 'A':3, 'G':4, 'E':5, 'D':6, 'C':7, 'T':8}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.108408Z","iopub.execute_input":"2022-08-03T09:37:15.109214Z","iopub.status.idle":"2022-08-03T09:37:15.115180Z","shell.execute_reply.started":"2022-08-03T09:37:15.109172Z","shell.execute_reply":"2022-08-03T09:37:15.114001Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__CabinLetter2__","metadata":{}},{"cell_type":"code","source":"train_df.CabinLetter2.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.116509Z","iopub.execute_input":"2022-08-03T09:37:15.117010Z","iopub.status.idle":"2022-08-03T09:37:15.128917Z","shell.execute_reply.started":"2022-08-03T09:37:15.116981Z","shell.execute_reply":"2022-08-03T09:37:15.127711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mapping_cb2 = {'Unknown':0, 'P':1, 'S':2}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.130582Z","iopub.execute_input":"2022-08-03T09:37:15.131204Z","iopub.status.idle":"2022-08-03T09:37:15.137330Z","shell.execute_reply.started":"2022-08-03T09:37:15.131162Z","shell.execute_reply":"2022-08-03T09:37:15.136270Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"__Cabin Number__","metadata":{}},{"cell_type":"code","source":"train_df.CabinNumber.unique()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.138716Z","iopub.execute_input":"2022-08-03T09:37:15.139439Z","iopub.status.idle":"2022-08-03T09:37:15.150973Z","shell.execute_reply.started":"2022-08-03T09:37:15.139399Z","shell.execute_reply":"2022-08-03T09:37:15.149932Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset['CabinNumber'] = dataset['CabinNumber'].replace('Unknown', 100)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.152527Z","iopub.execute_input":"2022-08-03T09:37:15.153837Z","iopub.status.idle":"2022-08-03T09:37:15.162580Z","shell.execute_reply.started":"2022-08-03T09:37:15.153804Z","shell.execute_reply":"2022-08-03T09:37:15.161715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.164776Z","iopub.execute_input":"2022-08-03T09:37:15.165829Z","iopub.status.idle":"2022-08-03T09:37:15.189171Z","shell.execute_reply.started":"2022-08-03T09:37:15.165793Z","shell.execute_reply":"2022-08-03T09:37:15.188203Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can convert categorical features to ordinal with these mappings","metadata":{}},{"cell_type":"code","source":"for dataset in combine:\n    dataset['HomePlanet'] = dataset['HomePlanet'].map(mapping_hp)\n    dataset['CryoSleep'] = dataset['CryoSleep'].astype(str).map(mapping_cs)\n    dataset['Destination'] = dataset['Destination'].map(mapping_d)\n    dataset['VIP'] = dataset['VIP'].astype(str).map(mapping_v)\n    dataset['CabinLetter1'] = dataset['CabinLetter1'].map(mapping_cb1)\n    dataset['CabinLetter2'] = dataset['CabinLetter2'].map(mapping_cb2)\n    \ntrain_df['Transported'] = train_df['Transported'].astype(str).map(mapping_t)\ncombine = [train_df, test_df]","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.190566Z","iopub.execute_input":"2022-08-03T09:37:15.191615Z","iopub.status.idle":"2022-08-03T09:37:15.235888Z","shell.execute_reply.started":"2022-08-03T09:37:15.191580Z","shell.execute_reply":"2022-08-03T09:37:15.234649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is the final look of our data sets","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.237616Z","iopub.execute_input":"2022-08-03T09:37:15.237983Z","iopub.status.idle":"2022-08-03T09:37:15.258063Z","shell.execute_reply.started":"2022-08-03T09:37:15.237952Z","shell.execute_reply":"2022-08-03T09:37:15.257206Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.259258Z","iopub.execute_input":"2022-08-03T09:37:15.260235Z","iopub.status.idle":"2022-08-03T09:37:15.283388Z","shell.execute_reply.started":"2022-08-03T09:37:15.260186Z","shell.execute_reply":"2022-08-03T09:37:15.282165Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.285159Z","iopub.execute_input":"2022-08-03T09:37:15.285859Z","iopub.status.idle":"2022-08-03T09:37:15.311977Z","shell.execute_reply.started":"2022-08-03T09:37:15.285813Z","shell.execute_reply":"2022-08-03T09:37:15.310597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Scalling__","metadata":{}},{"cell_type":"markdown","source":"There  are huge range in numerical features. It can be a problem when model training. So we should scale our data","metadata":{}},{"cell_type":"code","source":"scaller = MinMaxScaler()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.313830Z","iopub.execute_input":"2022-08-03T09:37:15.314442Z","iopub.status.idle":"2022-08-03T09:37:15.321991Z","shell.execute_reply.started":"2022-08-03T09:37:15.314356Z","shell.execute_reply":"2022-08-03T09:37:15.320957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_features = ['HomePlanet','CryoSleep','Destination','Age','VIP','RoomService','FoodCourt','ShoppingMall','Spa','VRDeck','CabinLetter1','CabinLetter2','CabinNumber','FamilyMember']","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.324001Z","iopub.execute_input":"2022-08-03T09:37:15.324887Z","iopub.status.idle":"2022-08-03T09:37:15.332618Z","shell.execute_reply.started":"2022-08-03T09:37:15.324841Z","shell.execute_reply":"2022-08-03T09:37:15.331677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for dataset in combine:\n    dataset_scalled = scaller.fit_transform(dataset[num_features])\n    dataset[num_features] = dataset_scalled\n","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.334677Z","iopub.execute_input":"2022-08-03T09:37:15.335360Z","iopub.status.idle":"2022-08-03T09:37:15.400195Z","shell.execute_reply.started":"2022-08-03T09:37:15.335323Z","shell.execute_reply":"2022-08-03T09:37:15.399067Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.402404Z","iopub.execute_input":"2022-08-03T09:37:15.402885Z","iopub.status.idle":"2022-08-03T09:37:15.430867Z","shell.execute_reply.started":"2022-08-03T09:37:15.402841Z","shell.execute_reply":"2022-08-03T09:37:15.429671Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.432638Z","iopub.execute_input":"2022-08-03T09:37:15.432993Z","iopub.status.idle":"2022-08-03T09:37:15.458741Z","shell.execute_reply.started":"2022-08-03T09:37:15.432964Z","shell.execute_reply":"2022-08-03T09:37:15.457604Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we can use this data to train our model","metadata":{}},{"cell_type":"markdown","source":"## __Model train, predict and slove__","metadata":{}},{"cell_type":"markdown","source":"Now our dataset is looking good and we have to train the model. Then we can use the model to slove the problem solution. There are many model algorithms to use. But our problem is classification and reggression problem in supervised learning. So we can use these models.","metadata":{}},{"cell_type":"markdown","source":"* Logistic Reggression\n* KNN\n* Support vector mask\n* Naive bayes classifier\n* Decision tree\n* Random forest\n* Preception\n* Artificial neural network\n* RVM","metadata":{}},{"cell_type":"markdown","source":"Now we should categorize our data into train and test data","metadata":{}},{"cell_type":"code","source":"X_train = train_df.drop(['PassengerId', 'Transported'], axis=1)\ny_train = train_df['Transported']\nX_test = test_df.drop(['PassengerId'], axis=1)\nX_train.shape, y_train.shape, X_test.shape\n","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.460077Z","iopub.execute_input":"2022-08-03T09:37:15.461028Z","iopub.status.idle":"2022-08-03T09:37:15.474074Z","shell.execute_reply.started":"2022-08-03T09:37:15.460994Z","shell.execute_reply":"2022-08-03T09:37:15.472794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets check each model algorithm one by one. And check the score","metadata":{}},{"cell_type":"code","source":"#  Lodistic regression\nlogreg = LogisticRegression()\nlogreg.fit(X_train, y_train)\nprint(round(logreg.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.476420Z","iopub.execute_input":"2022-08-03T09:37:15.477343Z","iopub.status.idle":"2022-08-03T09:37:15.605097Z","shell.execute_reply.started":"2022-08-03T09:37:15.477288Z","shell.execute_reply":"2022-08-03T09:37:15.603374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# K Nearest Neighbors\nknn = KNeighborsClassifier(n_neighbors=7)\nknn.fit(X_train, y_train)\nprint(round(knn.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:15.613030Z","iopub.execute_input":"2022-08-03T09:37:15.614212Z","iopub.status.idle":"2022-08-03T09:37:16.404133Z","shell.execute_reply.started":"2022-08-03T09:37:15.614145Z","shell.execute_reply":"2022-08-03T09:37:16.402244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Support vector machines\nsvc = SVC()\nsvc.fit(X_train, y_train)\nprint(round(svc.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:16.406629Z","iopub.execute_input":"2022-08-03T09:37:16.407087Z","iopub.status.idle":"2022-08-03T09:37:22.157597Z","shell.execute_reply.started":"2022-08-03T09:37:16.407048Z","shell.execute_reply":"2022-08-03T09:37:22.156313Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gaussian Naive bayes\ngaussian = GaussianNB()\ngaussian.fit(X_train, y_train)\nprint(round(gaussian.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.159168Z","iopub.execute_input":"2022-08-03T09:37:22.159571Z","iopub.status.idle":"2022-08-03T09:37:22.181534Z","shell.execute_reply.started":"2022-08-03T09:37:22.159536Z","shell.execute_reply":"2022-08-03T09:37:22.180598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Perceptron\nperceptron = Perceptron()\nperceptron.fit(X_train, y_train)\nprint(round(perceptron.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.182971Z","iopub.execute_input":"2022-08-03T09:37:22.183633Z","iopub.status.idle":"2022-08-03T09:37:22.211445Z","shell.execute_reply.started":"2022-08-03T09:37:22.183597Z","shell.execute_reply":"2022-08-03T09:37:22.209915Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Linear SVC\nlinearsvc = LinearSVC()\nlinearsvc.fit(X_train, y_train)\nprint(round(linearsvc.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.213359Z","iopub.execute_input":"2022-08-03T09:37:22.214222Z","iopub.status.idle":"2022-08-03T09:37:22.372694Z","shell.execute_reply.started":"2022-08-03T09:37:22.214171Z","shell.execute_reply":"2022-08-03T09:37:22.371106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Stochastic Gradient Descent\nsgd = SGDClassifier()\nsgd.fit(X_train, y_train)\nprint(round(sgd.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.374779Z","iopub.execute_input":"2022-08-03T09:37:22.375661Z","iopub.status.idle":"2022-08-03T09:37:22.466611Z","shell.execute_reply.started":"2022-08-03T09:37:22.375605Z","shell.execute_reply":"2022-08-03T09:37:22.465094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Decision Tree\ndecisiontr = DecisionTreeClassifier()\ndecisiontr.fit(X_train, y_train)\nprint(round(decisiontr.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.469138Z","iopub.execute_input":"2022-08-03T09:37:22.470230Z","iopub.status.idle":"2022-08-03T09:37:22.588616Z","shell.execute_reply.started":"2022-08-03T09:37:22.470168Z","shell.execute_reply":"2022-08-03T09:37:22.587437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Random forest\nrndforest = RandomForestClassifier()\nrndforest.fit(X_train, y_train)\nprint(round(rndforest.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:22.589976Z","iopub.execute_input":"2022-08-03T09:37:22.590946Z","iopub.status.idle":"2022-08-03T09:37:23.941723Z","shell.execute_reply.started":"2022-08-03T09:37:22.590901Z","shell.execute_reply":"2022-08-03T09:37:23.940632Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Gradient boost\ngbc = GradientBoostingClassifier()\ngbc.fit(X_train, y_train)\nprint(round(gbc.score(X_train, y_train)*100,2),'%')","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:23.943442Z","iopub.execute_input":"2022-08-03T09:37:23.943892Z","iopub.status.idle":"2022-08-03T09:37:25.150111Z","shell.execute_reply.started":"2022-08-03T09:37:23.943860Z","shell.execute_reply":"2022-08-03T09:37:25.148862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Model tuning__","metadata":{}},{"cell_type":"markdown","source":"Lets try these different models with different parameters and find the best algorithm for our solution","metadata":{}},{"cell_type":"code","source":"# define a dictionary for models and their parameters\nmodel_param = {\n    'svc': {\n        'model': SVC(gamma='auto'),\n        'params' : {\n            'C': [1,10,20],\n            'kernel': ['rbf','linear']\n        }  \n    },\n    'random_forest': {\n        'model': RandomForestClassifier(),\n        'params' : {\n            'n_estimators': [1,5,10]\n        }\n    },\n    'logistic_regression' : {\n        'model': LogisticRegression(solver='liblinear',multi_class='auto'),\n        'params': {\n            'C': [1,5,10]\n        }\n    },\n    'gaussian' :{\n        'model' : GaussianNB(),\n        'params' : {\n            \n        }\n    },\n    'knn' : {\n        'model' : KNeighborsClassifier(),\n        'params' : {\n            'n_neighbors' : [1,3,5,7,9]\n        }\n    },\n    'tree' : {\n        'model' : DecisionTreeClassifier(),\n        'params' : {\n            'criterion': ['gini','entropy'],\n        }\n    },\n    'perceptron' : {\n        'model' : Perceptron(),\n        'params' : {\n            'penalty' : ['l2','l1','elasticnet']\n        }\n    },\n    'linearsvc' : {\n        'model' : LinearSVC(),\n        'params' : {\n                 \n        }\n    },\n    'sgd' : {\n        'model' : SGDClassifier(),\n        'params' : {\n\n        }\n    },\n    'gbc' : {\n        'model' : GradientBoostingClassifier(),\n        'params' : {\n            'n_estimators': [1,5,10,50,100,500]\n        }\n    }\n}","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:25.160836Z","iopub.execute_input":"2022-08-03T09:37:25.161436Z","iopub.status.idle":"2022-08-03T09:37:25.172698Z","shell.execute_reply.started":"2022-08-03T09:37:25.161396Z","shell.execute_reply":"2022-08-03T09:37:25.171387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#find the best model\nscores = []\nfor model_name, mp in model_param.items():\n  clf = GridSearchCV(mp['model'],mp['params'],cv=3,return_train_score=False)\n  clf.fit(X_train, y_train)\n  scores.append({\n      'model' : model_name,\n      'best_score' : clf.best_score_,\n      'best_params' : clf.best_params_\n  })\ndf = pd.DataFrame(scores,columns=['model','best_score','best_params'])\ndf","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:37:25.174640Z","iopub.execute_input":"2022-08-03T09:37:25.175199Z","iopub.status.idle":"2022-08-03T09:38:31.869839Z","shell.execute_reply.started":"2022-08-03T09:37:25.175152Z","shell.execute_reply":"2022-08-03T09:38:31.868713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## __Model training and predict__","metadata":{}},{"cell_type":"markdown","source":"So according to above chart we can see SVC is the best scored model with these parameters. So let's use that to train our model.","metadata":{}},{"cell_type":"code","source":"model = GradientBoostingClassifier(n_estimators=50)\nmodel.fit(X_train, y_train)\ny_pred = model.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:38:57.628765Z","iopub.execute_input":"2022-08-03T09:38:57.629279Z","iopub.status.idle":"2022-08-03T09:38:58.251030Z","shell.execute_reply.started":"2022-08-03T09:38:57.629242Z","shell.execute_reply":"2022-08-03T09:38:58.249754Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### __Submission__","metadata":{}},{"cell_type":"markdown","source":"Lets do our submission","metadata":{}},{"cell_type":"markdown","source":"Convert y_pred result to boolian","metadata":{}},{"cell_type":"code","source":"result = y_pred.astype(bool)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:39:08.106517Z","iopub.execute_input":"2022-08-03T09:39:08.107021Z","iopub.status.idle":"2022-08-03T09:39:08.114037Z","shell.execute_reply.started":"2022-08-03T09:39:08.106984Z","shell.execute_reply":"2022-08-03T09:39:08.112597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'PassengerId':test_df['PassengerId'],\n    'Transported':result\n})\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:39:10.135739Z","iopub.execute_input":"2022-08-03T09:39:10.136258Z","iopub.status.idle":"2022-08-03T09:39:10.151311Z","shell.execute_reply.started":"2022-08-03T09:39:10.136223Z","shell.execute_reply":"2022-08-03T09:39:10.150134Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-03T09:39:12.457511Z","iopub.execute_input":"2022-08-03T09:39:12.458002Z","iopub.status.idle":"2022-08-03T09:39:12.470516Z","shell.execute_reply.started":"2022-08-03T09:39:12.457966Z","shell.execute_reply":"2022-08-03T09:39:12.469515Z"},"trusted":true},"execution_count":null,"outputs":[]}]}