{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Introduction \n\nIn this notebook, I have tried to illustrate the best way (in this case) to feature engineering and handle missing values. I have shown some handy tricks one can use to better their score in Kaggle competitions.\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# Imports ","metadata":{}},{"cell_type":"code","source":"\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\npd.set_option('max_rows',100)\npd.set_option('max_columns',100)\nimport seaborn as sns\nimport matplotlib.pyplot as  plt\nimport scipy\nfrom sklearn.preprocessing import StandardScaler\n\n# from pycaret.classification import setup, compare_models\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-13T07:36:51.677363Z","iopub.execute_input":"2022-07-13T07:36:51.677891Z","iopub.status.idle":"2022-07-13T07:36:51.690915Z","shell.execute_reply.started":"2022-07-13T07:36:51.677853Z","shell.execute_reply":"2022-07-13T07:36:51.689081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"# combining train and test \n\nWe have been given train and test data by the competition, the standard way to do this is taking the train data and spliting it for test or using KFold. However in this setting of competition we can use the extra bit of information provided to us by the test set to impute and scale the data. This will boost the accuracy","metadata":{}},{"cell_type":"code","source":"\ndf_train = pd.read_csv('../input/spaceship-titanic/train.csv')\ndf_test = pd.read_csv('../input/spaceship-titanic/test.csv')\nsample_sub = pd.read_csv('../input/spaceship-titanic/sample_submission.csv')\ntarget = df_train['Transported']\n\ndf_train.drop(['Transported'], axis= 1, inplace = True)\ncdata = pd.concat([df_train, df_test],axis = 0,ignore_index=True)\ncdata.drop(['Name'], axis= 1, inplace = True)\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-13T07:36:51.738996Z","iopub.execute_input":"2022-07-13T07:36:51.740551Z","iopub.status.idle":"2022-07-13T07:36:51.825619Z","shell.execute_reply.started":"2022-07-13T07:36:51.740480Z","shell.execute_reply":"2022-07-13T07:36:51.824361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Split to Create more Features\nWe can split the seeming unimportant string type feature to give our model a better chance to understand what is going on. However, we need to filter these features and check their impact.","metadata":{}},{"cell_type":"code","source":"cdata[['pass_grp','pass_no']]= cdata['PassengerId'].str.split('_', n = -1, expand = True)\ncdata.drop('PassengerId', axis = 1, inplace = True)\ncdata[['deck','num','side']]= cdata['Cabin'].str.split('/', n = -1, expand = True)\ncdata.drop('Cabin', axis = 1, inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T07:36:51.864665Z","iopub.execute_input":"2022-07-13T07:36:51.865370Z","iopub.status.idle":"2022-07-13T07:36:51.945142Z","shell.execute_reply.started":"2022-07-13T07:36:51.865323Z","shell.execute_reply":"2022-07-13T07:36:51.943730Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\n# Handling missing values and Adding new feature\n\nwith trial and error i found that replacing 'RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck' missing values with zero works best. I guess is these missing values could be there by design. num and pass grp dont add much value to the overall performance and thus can be dropped.\n","metadata":{}},{"cell_type":"code","source":"\ncdata['Total_cost'] = pd.Series(np.zeros(cdata.shape[0]))\nfor feat in ['RoomService', 'FoodCourt', 'ShoppingMall', 'Spa', 'VRDeck']:\n    cdata[feat].fillna(cdata[feat].mean(), inplace = True)\n    cdata['Total_cost'] += cdata[feat]\n#     cdata3.drop(feat, axis = 1,inplace = True )\ncdata.drop(['num', 'pass_grp'],axis = 1, inplace = True)\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-07-13T08:09:11.052268Z","iopub.execute_input":"2022-07-13T08:09:11.052803Z","iopub.status.idle":"2022-07-13T08:09:11.075911Z","shell.execute_reply.started":"2022-07-13T08:09:11.052759Z","shell.execute_reply":"2022-07-13T08:09:11.074583Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple imputing ","metadata":{}},{"cell_type":"code","source":"categorical_features = cdata.select_dtypes(object)\nnum_features = cdata.select_dtypes(np.number)\nfor feat in categorical_features:\n    cdata[feat].fillna(cdata[feat].mode()[0], inplace= True)\n\ncdata['Age'].fillna(int(cdata['Age'].mean()), inplace = True)","metadata":{"execution":{"iopub.status.busy":"2022-07-13T08:09:14.453426Z","iopub.execute_input":"2022-07-13T08:09:14.453903Z","iopub.status.idle":"2022-07-13T08:09:14.509208Z","shell.execute_reply.started":"2022-07-13T08:09:14.453866Z","shell.execute_reply":"2022-07-13T08:09:14.507789Z"},"trusted":true},"execution_count":null,"outputs":[]}]}