{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n# Work Plan for the Project\n\nI have divided the work into 2 notebooks. I think it is cleaner and shorter to keep notebooks seperated\n\n1. Data Preprocessing, Feature Engineering, EDA & Feature Importance Analysis (This Notebook)\n    - Data Understanding\n    - Preprocess null values (Imputation)\n    - Create New Features (Categorical Feature Interactions, Binning, Percentiles, etc.)  \n\n\n\n2. EDA & Feature Importance\n\n    - Exploratory data analysis for every feature\n    - Correlation Matrix \n    - Feature Importance Analysis using Cramer V Stat\nhttps://www.kaggle.com/emreuzel/eda-and-feature-importance/edit\n\n\n\n3. ML Models & Feature Selection & Hyperparameter Optimization \n\n    - Applying ML models\n    - Applying feature selection techniques\n    - Hyperparameter Optimization\n\n*I would be glad if you look into my work and give feedback by notes and upvotes. Thanks!*","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.preprocessing import OneHotEncoder\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport plotly.express as px\nfrom warnings import simplefilter\nimport scipy.stats as ss\nimport os \n\n\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-04T09:41:40.870953Z","iopub.execute_input":"2022-08-04T09:41:40.871253Z","iopub.status.idle":"2022-08-04T09:41:40.884138Z","shell.execute_reply.started":"2022-08-04T09:41:40.871222Z","shell.execute_reply":"2022-08-04T09:41:40.883291Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = pd.read_csv('/kaggle/input/titanic/train.csv')\ndf_test = pd.read_csv('/kaggle/input/titanic/test.csv')\ngender_submission = pd.read_csv('/kaggle/input/titanic/gender_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:41.140494Z","iopub.execute_input":"2022-08-04T09:41:41.141007Z","iopub.status.idle":"2022-08-04T09:41:41.165934Z","shell.execute_reply.started":"2022-08-04T09:41:41.140951Z","shell.execute_reply":"2022-08-04T09:41:41.165305Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# A Look Over the Dataset","metadata":{}},{"cell_type":"code","source":"df_train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:41.223922Z","iopub.execute_input":"2022-08-04T09:41:41.224433Z","iopub.status.idle":"2022-08-04T09:41:41.241220Z","shell.execute_reply.started":"2022-08-04T09:41:41.224377Z","shell.execute_reply":"2022-08-04T09:41:41.240314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Info about train dataset \\n')\nprint( df_train.info(), '\\n \\n')\n\nprint(\"Description of the train dataset \\n\")\nprint(df_train.describe(), '\\n \\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:41.290224Z","iopub.execute_input":"2022-08-04T09:41:41.291151Z","iopub.status.idle":"2022-08-04T09:41:41.455016Z","shell.execute_reply.started":"2022-08-04T09:41:41.291105Z","shell.execute_reply":"2022-08-04T09:41:41.454328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Unique values of the features","metadata":{}},{"cell_type":"code","source":"print('Train Set Nunique values')\nfor col in df_train.columns:\n    print(col + ' nunique:',df_train[col].nunique())\n\nprint('\\nTest Set Nunique values')\nfor col in df_test.columns:\n    print(col + ' nunique:',df_test[col].nunique())\n","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:41.456176Z","iopub.execute_input":"2022-08-04T09:41:41.457001Z","iopub.status.idle":"2022-08-04T09:41:41.476183Z","shell.execute_reply.started":"2022-08-04T09:41:41.456966Z","shell.execute_reply":"2022-08-04T09:41:41.475320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Nan Values & Imputation","metadata":{}},{"cell_type":"code","source":"print(\"Train data's nan value distribution \\n\",df_train.isna().sum(), '\\n')\nprint(\"Test data's nan value distribution \\n\", df_test.isna().sum(), '\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:41.700667Z","iopub.execute_input":"2022-08-04T09:41:41.700950Z","iopub.status.idle":"2022-08-04T09:41:41.711419Z","shell.execute_reply.started":"2022-08-04T09:41:41.700921Z","shell.execute_reply":"2022-08-04T09:41:41.710772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Filling Embarked Feature\n\nFill the 2 embarked value as the closest values according to their pclass","metadata":{}},{"cell_type":"code","source":"df_train.Embarked.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:42.040859Z","iopub.execute_input":"2022-08-04T09:41:42.041125Z","iopub.status.idle":"2022-08-04T09:41:42.048866Z","shell.execute_reply.started":"2022-08-04T09:41:42.041096Z","shell.execute_reply":"2022-08-04T09:41:42.048084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(df_train[df_train.Embarked == 'Q'].Pclass.value_counts())\nprint(df_train[df_train.Embarked == 'C'].Pclass.value_counts())\nprint(df_train[df_train.Embarked == 'S'].Pclass.value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:42.181030Z","iopub.execute_input":"2022-08-04T09:41:42.181306Z","iopub.status.idle":"2022-08-04T09:41:42.193326Z","shell.execute_reply.started":"2022-08-04T09:41:42.181274Z","shell.execute_reply":"2022-08-04T09:41:42.192607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.at[61, 'Embarked'] = 'C'\ndf_train.at[829, 'Embarked'] = 'C'\ndf_test.at[152, 'Fare'] = 13.5","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:42.480616Z","iopub.execute_input":"2022-08-04T09:41:42.480917Z","iopub.status.idle":"2022-08-04T09:41:42.486195Z","shell.execute_reply.started":"2022-08-04T09:41:42.480884Z","shell.execute_reply":"2022-08-04T09:41:42.485505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Getting title of the person from the name","metadata":{}},{"cell_type":"code","source":"df_train['title'] = 'empty'\ndf_test['title'] = 'empty'\n\nfor i in range(len(df_train)):\n    the_title = df_train['Name'][i][df_train['Name'][i].find(',') +2: df_train['Name'][i].find('.')]\n    df_train.at[i, 'title'] = the_title\n\nfor i in range(len(df_test)):\n    the_title = df_test['Name'][i][df_test['Name'][i].find(',') +2: df_test['Name'][i].find('.')]\n    df_test.at[i, 'title'] = the_title\n\n#eleminate the titles that are not available at test set\nthe_others = set(df_train.title.unique()) - set(df_test.title.unique())\n\nfor i in range(len(df_train)):\n    if df_train.at[i, 'title'] in the_others:\n        df_train.at[i, 'title'] = 'the_others'","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:42.940830Z","iopub.execute_input":"2022-08-04T09:41:42.941104Z","iopub.status.idle":"2022-08-04T09:41:42.999125Z","shell.execute_reply.started":"2022-08-04T09:41:42.941074Z","shell.execute_reply":"2022-08-04T09:41:42.998512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"title_dict = {'male': 'Mr', 'female': 'Mrs'}\nthe_other_titles = df_train[df_train.title == 'the_others'].Sex.map(title_dict)\n\ndf_train.loc[df_train[df_train.title == 'Ms'].index, 'title'] = 'Mrs'\ndf_train.loc[df_train[df_train.title == 'Col'].index, 'title'] = 'Mr'\ndf_train.loc[df_train[df_train.title == 'Dr'].index, 'title'] = 'Mr'\ndf_train.loc[df_train[df_train.title == 'Rev'].index, 'title'] = 'Mr'\ndf_train.loc[df_train[df_train.title == 'the_others'].index, 'title'] = the_other_titles\n\ndf_test.loc[df_test[df_test.title == 'Ms'].index, 'title'] = 'Mrs'\ndf_test.loc[df_test[df_test.title == 'Col'].index, 'title'] = 'Mr'\ndf_test.loc[df_test[df_test.title == 'Dr'].index, 'title'] = 'Mr'\ndf_test.loc[df_test[df_test.title == 'Rev'].index, 'title'] = 'Mr'\ndf_test.loc[df_test[df_test.title == 'Dona'].index, 'title'] = 'Mrs'","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:43.140692Z","iopub.execute_input":"2022-08-04T09:41:43.140998Z","iopub.status.idle":"2022-08-04T09:41:43.163319Z","shell.execute_reply.started":"2022-08-04T09:41:43.140965Z","shell.execute_reply":"2022-08-04T09:41:43.162670Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_data = pd.concat([df_train, df_test])\nall_data.reset_index(drop = True, inplace = True)\nall_data","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:43.347249Z","iopub.execute_input":"2022-08-04T09:41:43.347977Z","iopub.status.idle":"2022-08-04T09:41:43.377088Z","shell.execute_reply.started":"2022-08-04T09:41:43.347937Z","shell.execute_reply":"2022-08-04T09:41:43.376238Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Filling Age\n\nWe fill the age by title accordingly. \n","metadata":{}},{"cell_type":"code","source":"title_age = all_data.groupby(by = 'title').median()['Age']","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:43.760610Z","iopub.execute_input":"2022-08-04T09:41:43.761456Z","iopub.status.idle":"2022-08-04T09:41:43.771170Z","shell.execute_reply.started":"2022-08-04T09:41:43.761408Z","shell.execute_reply":"2022-08-04T09:41:43.770402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def age_imputer_by_title(df_train):\n    for ind in df_train[df_train.Age.isna()].index:\n        the_title = df_train.at[ind, 'title']\n        the_val = title_age[the_title]\n        df_train.at[ind, 'Age'] = the_val\n    return df_train\ndf_train = age_imputer_by_title(df_train)\ndf_test = age_imputer_by_title(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:43.954268Z","iopub.execute_input":"2022-08-04T09:41:43.955168Z","iopub.status.idle":"2022-08-04T09:41:43.970222Z","shell.execute_reply.started":"2022-08-04T09:41:43.955125Z","shell.execute_reply":"2022-08-04T09:41:43.969584Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Cabin Imputation\n\nWe can see that Cabin feature has high cardinality and lots of missing values. We will try to save the feature with imputation.\n\n 1. Change the cabin types mroe generalized version. For instance: C23, C27 -> C ; F23, F3 -> F\n 2. Impute according to fare types. I have witnessed that Cabin types have an hierarchy. For example, F cabin has the people with lowest Fares (Between 0- 15). While Cabin B and C has the people with highest fares ","metadata":{}},{"cell_type":"code","source":"print(df_train.Cabin.value_counts(), '\\n')\nprint(df_test.Cabin.value_counts())","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:44.380456Z","iopub.execute_input":"2022-08-04T09:41:44.381352Z","iopub.status.idle":"2022-08-04T09:41:44.391465Z","shell.execute_reply.started":"2022-08-04T09:41:44.381302Z","shell.execute_reply":"2022-08-04T09:41:44.390702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train['new_cabin'] = np.nan\ndf_test['new_cabin'] = np.nan\ndef new_cabin_creator(df_train):\n    for ind in range(len(df_train)):\n\n        if 'C' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind,'new_cabin'] = 'C'\n\n        elif 'B' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind, 'new_cabin'] = 'B'\n\n        elif 'F' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind,'new_cabin'] = 'F'\n\n        elif 'D' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind,'new_cabin'] = 'D'\n\n        elif 'E' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind,'new_cabin'] = 'E'\n\n        elif 'A' in str(df_train.loc[ind, 'Cabin']):\n            df_train.loc[ind,'new_cabin'] = 'A'\n\n        else:\n            df_train.loc[ind, 'new_cabin'] = np.nan\n            \n    return df_train\n\ndf_train = new_cabin_creator(df_train)\ndf_test = new_cabin_creator(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:44.550807Z","iopub.execute_input":"2022-08-04T09:41:44.551400Z","iopub.status.idle":"2022-08-04T09:41:45.033207Z","shell.execute_reply.started":"2022-08-04T09:41:44.551344Z","shell.execute_reply":"2022-08-04T09:41:45.032523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\nplt.figure(figsize = (20, 6))\nplt.xticks()\nsns.barplot(data = df_train, x= 'new_cabin', y = 'Survived')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:45.034535Z","iopub.execute_input":"2022-08-04T09:41:45.034777Z","iopub.status.idle":"2022-08-04T09:41:45.449165Z","shell.execute_reply.started":"2022-08-04T09:41:45.034746Z","shell.execute_reply":"2022-08-04T09:41:45.448294Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.new_cabin.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:45.450555Z","iopub.execute_input":"2022-08-04T09:41:45.450803Z","iopub.status.idle":"2022-08-04T09:41:45.457951Z","shell.execute_reply.started":"2022-08-04T09:41:45.450766Z","shell.execute_reply":"2022-08-04T09:41:45.457051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Cardinality is reduced. We have 6 main categories for the cabin feature now. It is time to impute other data points in the cabin feature.","metadata":{}},{"cell_type":"code","source":"def cabin_imputer(df_train):\n    \n\n    cabin_nan = df_train[df_train.new_cabin.isna() == True]\n    cabin_not_nan = df_train[df_train.new_cabin.isna() == False]\n\n    the_dict = {}\n    for i in cabin_nan.Ticket.unique():\n        if i in cabin_not_nan.Ticket.unique():\n            the_dict[i] = cabin_not_nan[cabin_not_nan.Ticket == i]['new_cabin'].values[0]\n\n    for i in the_dict:\n        the_val = the_dict[i]    \n        the_index = cabin_nan[cabin_nan.Ticket == i].index\n        cabin_nan.loc[the_index, 'new_cabin'] = the_val\n\n\n    f_class_index =cabin_nan[cabin_nan.Fare <20].index\n    df_train.loc[f_class_index, 'new_cabin'] = 'F' \n\n    b_class_index =cabin_nan[cabin_nan.Fare >100].index\n    df_train.loc[b_class_index, 'new_cabin'] = 'B' \n    \n    e_class_index =cabin_nan[(cabin_nan.Fare >20) & (cabin_nan.Fare <50)].index\n    df_train.loc[e_class_index, 'new_cabin'] = 'E' \n    \n    d_class_index =cabin_nan[(cabin_nan.Fare >50) & (cabin_nan.Fare <100)].index\n    df_train.loc[d_class_index, 'new_cabin'] = 'D' \n    \n    return df_train","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:45.459382Z","iopub.execute_input":"2022-08-04T09:41:45.459622Z","iopub.status.idle":"2022-08-04T09:41:45.470000Z","shell.execute_reply.started":"2022-08-04T09:41:45.459591Z","shell.execute_reply":"2022-08-04T09:41:45.469152Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.drop(columns = ['Cabin'], inplace = True)\ndf_test.drop(columns = ['Cabin'], inplace = True)\n\ndf_train = cabin_imputer(df_train)\ndf_test = cabin_imputer(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:45.520896Z","iopub.execute_input":"2022-08-04T09:41:45.521202Z","iopub.status.idle":"2022-08-04T09:41:45.628282Z","shell.execute_reply.started":"2022-08-04T09:41:45.521160Z","shell.execute_reply":"2022-08-04T09:41:45.627512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ticket: An Insight","metadata":{}},{"cell_type":"markdown","source":"#### At the first look: Ticket column seem meaningless. Because it has high cardinality and seems doesn't convey any critical information. However, it can be seen that with the same ticket ID; all the people have same survival status \n\n#### Example","metadata":{}},{"cell_type":"code","source":"#We identified with the same \ndf_train[df_train.Ticket == '347088']","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:46.080657Z","iopub.execute_input":"2022-08-04T09:41:46.081315Z","iopub.status.idle":"2022-08-04T09:41:46.099924Z","shell.execute_reply.started":"2022-08-04T09:41:46.081274Z","shell.execute_reply":"2022-08-04T09:41:46.099272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train[df_train.Ticket == 'CA. 2343']","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:46.261409Z","iopub.execute_input":"2022-08-04T09:41:46.261874Z","iopub.status.idle":"2022-08-04T09:41:46.282148Z","shell.execute_reply.started":"2022-08-04T09:41:46.261834Z","shell.execute_reply":"2022-08-04T09:41:46.281594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### We can see above there are people with the same ticket in the both train and test set.","metadata":{}},{"cell_type":"code","source":"the_list = []\n\nfor i in set(df_train.Ticket):\n    for j in set(df_test.Ticket):\n        if i == j:\n            the_list.append(i)\n            #print(i)\n            \n            \nfor i in the_list:\n    the_value = df_train[df_train.Ticket == i].Survived.mean()\n    the_length = len(df_train[df_train.Ticket == i].Survived)\n    the_index = df_test[df_test.Ticket == i].index\n    df_test.loc[the_index, 'survived_ticket'] = the_value\n    df_test.loc[the_index, 'number_of_supportage'] = the_length","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:46.673702Z","iopub.execute_input":"2022-08-04T09:41:46.674116Z","iopub.status.idle":"2022-08-04T09:41:47.040107Z","shell.execute_reply.started":"2022-08-04T09:41:46.674083Z","shell.execute_reply":"2022-08-04T09:41:47.039431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### *Conclusion: We conclude that with the same ticket it is probably to have same survived situation. That's why we created two new features for final prediction.*\n#### *If we know a ticket in the train set also exist in the test set, probably have the same survival status. The same features will be created in the modelling phase in order to prevent leaking*","metadata":{}},{"cell_type":"markdown","source":"### Bining the Fare ","metadata":{}},{"cell_type":"code","source":"def fare_binning(df_train):\n    df_train['fare_bin'] = np.nan\n    \n    lower_than_20= df_train[df_train['Fare']<=20].index\n    twenty_seventy= df_train[(df_train['Fare'] >20) & (df_train['Fare']<=70)].index\n    higher_than_seventy= df_train[df_train['Fare'] >70].index\n\n    df_train.loc[lower_than_20, 'fare_bin'] = '10'\n    df_train.loc[twenty_seventy, 'fare_bin'] = '50'\n    df_train.loc[higher_than_seventy, 'fare_bin'] = '100'\n    \n    return df_train\n\ndf_train = fare_binning(df_train)\ndf_test = fare_binning(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:47.260451Z","iopub.execute_input":"2022-08-04T09:41:47.260873Z","iopub.status.idle":"2022-08-04T09:41:47.276855Z","shell.execute_reply.started":"2022-08-04T09:41:47.260841Z","shell.execute_reply":"2022-08-04T09:41:47.276107Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Logging the Fare","metadata":{}},{"cell_type":"code","source":"df_train['Fare_logged'] = np.log1p(df_train.Age)\ndf_test['Fare_logged'] = np.log1p(df_test.Age)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:47.590904Z","iopub.execute_input":"2022-08-04T09:41:47.591458Z","iopub.status.idle":"2022-08-04T09:41:47.597685Z","shell.execute_reply.started":"2022-08-04T09:41:47.591418Z","shell.execute_reply":"2022-08-04T09:41:47.596655Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Binning the age by 15 range bins","metadata":{}},{"cell_type":"code","source":"bins = list(np.arange(0,80, 20))\nlabels = bins[:-1]\ndf_train['Age_binned'] = pd.cut(df_train['Age'], bins = bins, labels = labels, include_lowest = True)\ndf_test['Age_binned'] = pd.cut(df_test['Age'], bins = bins, labels = labels, include_lowest = True)\n\ndf_train.Age_binned.fillna(40, inplace = True)\ndf_test.Age_binned.fillna(40, inplace = True)\n\n#df_train.Age_binned.fillna(75, inplace = True)\n#df_test.Age_binned.fillna(75, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:48.011209Z","iopub.execute_input":"2022-08-04T09:41:48.011514Z","iopub.status.idle":"2022-08-04T09:41:48.025588Z","shell.execute_reply.started":"2022-08-04T09:41:48.011474Z","shell.execute_reply":"2022-08-04T09:41:48.024789Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Logging the Age","metadata":{}},{"cell_type":"code","source":"df_train['Age_logged'] = np.log1p(df_train.Age)\ndf_test['Age_logged'] = np.log1p(df_test.Age)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:48.380289Z","iopub.execute_input":"2022-08-04T09:41:48.380642Z","iopub.status.idle":"2022-08-04T09:41:48.387769Z","shell.execute_reply.started":"2022-08-04T09:41:48.380606Z","shell.execute_reply":"2022-08-04T09:41:48.386633Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Two-way Feature Interactions","metadata":{}},{"cell_type":"code","source":"cat_cols = ['Sex', 'Pclass', 'Embarked',  'fare_bin',  'Age_binned' ]\n\ndef two_way_interactions(df_train):\n    count = 0\n    for i in range(len(cat_cols)):\n        if i == len(cat_cols) -1 :\n            break\n\n        for j in range(i+1, len(cat_cols)):\n            stabilized_feature = cat_cols[i]\n            #print(i , j)\n            if j!= i:\n                other_feature = cat_cols[j]\n                #print(stabilized_feature, other_feature)\n                df_train[stabilized_feature + '__' + other_feature] = df_train[stabilized_feature].astype('str') + '__' + df_train[other_feature].astype('str')\n    return df_train\n\ndf_train = two_way_interactions(df_train)\ndf_test = two_way_interactions(df_test)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:48.820423Z","iopub.execute_input":"2022-08-04T09:41:48.820712Z","iopub.status.idle":"2022-08-04T09:41:48.858521Z","shell.execute_reply.started":"2022-08-04T09:41:48.820683Z","shell.execute_reply":"2022-08-04T09:41:48.857691Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', None)\ndf_train[df_train.Age > 70]","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:49.000134Z","iopub.execute_input":"2022-08-04T09:41:49.000827Z","iopub.status.idle":"2022-08-04T09:41:49.029703Z","shell.execute_reply.started":"2022-08-04T09:41:49.000789Z","shell.execute_reply":"2022-08-04T09:41:49.029015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Percentile Features","metadata":{}},{"cell_type":"code","source":"numeric_cols = ['Age', 'Fare']\nfor col in df_train[numeric_cols].rank(pct = True).columns:\n    the_col_train = df_train[numeric_cols].rank(pct = True)[col]\n    the_col_test = df_test[numeric_cols].rank(pct = True)[col]\n    \n    df_train['percentile_' + col ] = the_col_train\n    df_test['percentile_' + col ] = the_col_test","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:49.420198Z","iopub.execute_input":"2022-08-04T09:41:49.420758Z","iopub.status.idle":"2022-08-04T09:41:49.435141Z","shell.execute_reply.started":"2022-08-04T09:41:49.420715Z","shell.execute_reply":"2022-08-04T09:41:49.434395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Two Way Interaction Features","metadata":{}},{"cell_type":"code","source":"df_train['Pclass__title']= df_train['Pclass'].astype('str') + '_' + df_train['title'].astype('str')\ndf_train['title__fare_bin']= df_train['Pclass'].astype('str') + '_' + df_train['fare_bin'].astype('str')\ndf_train['Pclass__Age_binned']= df_train['Pclass'].astype('str') + '_' + df_train['Age_binned'].astype('str')\n\ndf_test['Pclass__Age_binned']= df_test['Pclass'].astype('str') + '_' + df_test['Age_binned'].astype('str')\ndf_test['Pclass__title']= df_test['Pclass'].astype('str') + '_' + df_test['title'].astype('str')\ndf_test['title__fare_bin']= df_test['Pclass'].astype('str') + '_' + df_test['fare_bin'].astype('str')\n\ndf_train['family_size'] = df_train.SibSp + df_train.Parch\ndf_test['family_size'] = df_test.SibSp + df_test.Parch\n\nalone_fam_index = df_train[df_train.family_size.where((df_train.family_size == 0)).isna() == False].index\nnot_alone_fam_index = df_train[df_train.family_size.where((df_train.family_size != 0)).isna() == False].index\n\ndf_train.loc[alone_fam_index, 'is_alone'] = 1\ndf_train.loc[not_alone_fam_index, 'is_alone'] = 0\n\n\nalone_fam_index = df_test[df_test.family_size.where((df_test.family_size == 0)).isna() == False].index\nnot_alone_fam_index = df_test[df_test.family_size.where((df_test.family_size != 0)).isna() == False].index\n\ndf_test.loc[alone_fam_index, 'is_alone'] = 1\ndf_test.loc[not_alone_fam_index, 'is_alone'] = 0\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:50.041122Z","iopub.execute_input":"2022-08-04T09:41:50.041604Z","iopub.status.idle":"2022-08-04T09:41:50.078006Z","shell.execute_reply.started":"2022-08-04T09:41:50.041565Z","shell.execute_reply":"2022-08-04T09:41:50.077335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"###  Categorization of Family Size Feature","metadata":{}},{"cell_type":"code","source":"def family_size_category(df_train):\n    small_fam_index = df_train[df_train.family_size.where(df_train.family_size < 2).isna() == False].index\n    df_train.loc[small_fam_index, 'family_size_category'] = 'small'\n\n    medium_fam_index = df_train[df_train.family_size.where((df_train.family_size >= 2) & (df_train.family_size < 5)).isna() == False].index\n    df_train.loc[medium_fam_index, 'family_size_category'] = 'medium'\n\n    large_fam_index = df_train[df_train.family_size.where((df_train.family_size >= 5)).isna() == False].index\n    df_train.loc[large_fam_index, 'family_size_category'] = 'large'\n    \n    return df_train\n\ndf_train = family_size_category(df_train)\ndf_test = family_size_category(df_test)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:50.391560Z","iopub.execute_input":"2022-08-04T09:41:50.392411Z","iopub.status.idle":"2022-08-04T09:41:50.418749Z","shell.execute_reply.started":"2022-08-04T09:41:50.392329Z","shell.execute_reply":"2022-08-04T09:41:50.418078Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Creating the count of people with similar ticket as a feature","metadata":{}},{"cell_type":"code","source":"df_all = pd.concat([df_train, df_test], axis = 0)\n\nticket_count = df_all.groupby(by = 'Ticket').count()['PassengerId']\nticket_count = ticket_count.reset_index()\nticket_count.rename(columns = {'PassengerId': 'Ticket_count'}, inplace = True)\n\ndf_train = df_train.merge(ticket_count)\ndf_test = df_test.merge(ticket_count)","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:50.970394Z","iopub.execute_input":"2022-08-04T09:41:50.971164Z","iopub.status.idle":"2022-08-04T09:41:51.016171Z","shell.execute_reply.started":"2022-08-04T09:41:50.971126Z","shell.execute_reply":"2022-08-04T09:41:51.015566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train.to_csv('df_train.csv')\ndf_test.to_csv('df_test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-04T09:41:51.340324Z","iopub.execute_input":"2022-08-04T09:41:51.340682Z","iopub.status.idle":"2022-08-04T09:41:51.372220Z","shell.execute_reply.started":"2022-08-04T09:41:51.340638Z","shell.execute_reply":"2022-08-04T09:41:51.371602Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are now ready to apply EDA and Feature Importance techniques. Please use links above in order to look at the","metadata":{}}]}