{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler, Normalizer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.cluster import KMeans\nfrom sklearn.model_selection import cross_val_score\nimport scipy.stats as stats\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-05T16:50:47.037865Z","iopub.execute_input":"2022-08-05T16:50:47.038467Z","iopub.status.idle":"2022-08-05T16:50:47.765090Z","shell.execute_reply.started":"2022-08-05T16:50:47.038373Z","shell.execute_reply":"2022-08-05T16:50:47.763046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Note Book Resume\nThe problem is classify into 2 value or Binary classification. To deal with the problem, there are several steps on this notebook.\n1. Divided features into categorical dan numerical.\n2. Deal with Null value using simple imputer.\n3. Create additional features using feature engineering.\n4. Check for best machine learning model.\n5. Tune chosen machine learning to increase model performance.\n\nIn the end, the model shows quite good result. Although I have a feeling that the result still can be improved.","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv(\"/kaggle/input/spaceship-titanic/train.csv\")\ntest_df = pd.read_csv(\"/kaggle/input/spaceship-titanic/test.csv\")\n\nsubm = pd.read_csv(\"/kaggle/input/spaceship-titanic/sample_submission.csv\")\nsubm.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.766718Z","iopub.execute_input":"2022-08-05T16:50:47.767090Z","iopub.status.idle":"2022-08-05T16:50:47.841322Z","shell.execute_reply.started":"2022-08-05T16:50:47.767056Z","shell.execute_reply":"2022-08-05T16:50:47.839873Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Examine Data","metadata":{}},{"cell_type":"code","source":"train_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.843036Z","iopub.execute_input":"2022-08-05T16:50:47.843451Z","iopub.status.idle":"2022-08-05T16:50:47.869833Z","shell.execute_reply.started":"2022-08-05T16:50:47.843416Z","shell.execute_reply":"2022-08-05T16:50:47.868231Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Features Examination\nCheck features will be used in model. Group features into nominal and categorical features. ","metadata":{}},{"cell_type":"code","source":"train_df.info()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.875049Z","iopub.execute_input":"2022-08-05T16:50:47.875929Z","iopub.status.idle":"2022-08-05T16:50:47.899662Z","shell.execute_reply.started":"2022-08-05T16:50:47.875874Z","shell.execute_reply":"2022-08-05T16:50:47.898264Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. \"PassengerId\" and \"Name\" are unique for each data. \"Name\" will be dropped while \"PassengerId\" is kept to create submission.\n2. \"Transported\" is target feature.\n3. Nominal Variable: \"Age\", \"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\", \"VRDeck\".\n4. Categorical Variable: \"HomePlanet\", \"CryoSleep\", \"Cabin\", \"Destination\", \"VIP\"","metadata":{}},{"cell_type":"markdown","source":"## Separate Nominal and Categorical Features\nNominal and categorical features need to be separated because they need different treatment. For example, in data imputing nominal feature uses mean or median strategy, while categorical feature use most_frequent.","metadata":{}},{"cell_type":"code","source":"cat_cols = [\"HomePlanet\", \"CryoSleep\", \"Cabin\", \"Destination\", \"VIP\"]\nnom_cols = [\"Age\", \"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\", \"VRDeck\"]\n\npass_id = test_df[\"PassengerId\"].values\ntarget = train_df[\"Transported\"].values\n\ntrain_df.drop(columns=[\"Name\", \"PassengerId\", \"Transported\"], inplace=True)\ntest_df.drop(columns=[\"Name\", \"PassengerId\"], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.901332Z","iopub.execute_input":"2022-08-05T16:50:47.901732Z","iopub.status.idle":"2022-08-05T16:50:47.913392Z","shell.execute_reply.started":"2022-08-05T16:50:47.901698Z","shell.execute_reply":"2022-08-05T16:50:47.911654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Count for Null Values\nFeatures with too much null value should be dropped and the others will be imputed.","metadata":{}},{"cell_type":"code","source":"columns = train_df.columns\nn = len(train_df)\nnull_values = []\n\nfor i in range(len(columns)):\n    perc = np.sum(train_df[columns[i]].isnull())/n\n    null_values.append([columns[i], round(perc*100)])\n\n# Show Null\nnull_values = pd.DataFrame(null_values, columns=[\"Column\", \"Perc Null\"])\nnull_values","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.915618Z","iopub.execute_input":"2022-08-05T16:50:47.916255Z","iopub.status.idle":"2022-08-05T16:50:47.945650Z","shell.execute_reply.started":"2022-08-05T16:50:47.916190Z","shell.execute_reply":"2022-08-05T16:50:47.944168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Majority features have minimum null. They will be fixed using simple imputer.","metadata":{}},{"cell_type":"markdown","source":"## Deal with Null Data Using Simple Imputer\nNominal data will be imputed using \"Median\" and categorical data using \"most_frequent\".","metadata":{}},{"cell_type":"code","source":"# Nominal Features Imputing\nimputer = SimpleImputer(missing_values=np.nan, strategy=\"median\")\n\ntrain_nom = imputer.fit_transform(train_df[nom_cols])\ntrain_nom = pd.DataFrame(train_nom, columns=nom_cols)\n\ntest_nom = imputer.fit_transform(test_df[nom_cols])\ntest_nom = pd.DataFrame(test_nom, columns=nom_cols)\n\n# Categorical Features Imputing\nimputer = SimpleImputer(missing_values=np.nan, strategy=\"most_frequent\")\n\ntrain_cat = imputer.fit_transform(train_df[cat_cols])\ntrain_cat = pd.DataFrame(train_cat, columns=cat_cols)\n\ntest_cat = imputer.fit_transform(test_df[cat_cols])\ntest_cat = pd.DataFrame(test_cat, columns=cat_cols)\n\n# Convert VIP Feature\nconv = [\"VIP\", \"CryoSleep\"]\nfor col in conv:\n    train_cat[col] = train_cat[col].apply(lambda x: str(x))\n    test_cat[col] = test_cat[col].apply(lambda x: str(x))","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:47.947369Z","iopub.execute_input":"2022-08-05T16:50:47.948269Z","iopub.status.idle":"2022-08-05T16:50:48.009836Z","shell.execute_reply.started":"2022-08-05T16:50:47.948225Z","shell.execute_reply":"2022-08-05T16:50:48.008470Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert Cabin Data\n\"Cabin\" feature is presented by some kind of codes with 3 part. The feature have more than 6000 unique value, which is impossible to be used as categorical feature. So, I decide to split the feature into 3 different feature, \"Cab_1\", \"Cab_2\", and \"Cab_3\". \"Cab_1\" and \"Cab_3\" are put into categorical variable, and \"Cab_2\" is nominal since it is consisted of number.","metadata":{}},{"cell_type":"code","source":"# Convert Train Data\ncab = train_cat[\"Cabin\"].apply(lambda x: x.split(\"/\"))\n\n# Append and Drop Features\ntrain_cat[\"Cab_1\"] = cab.apply(lambda x: x[0])\ntrain_cat[\"Cab_3\"] = cab.apply(lambda x: x[2])\ntrain_nom[\"Cab_2\"] = cab.apply(lambda x: float(x[1]))\ntrain_cat.drop(columns=[\"Cabin\"], inplace=True)\n\n# Convert Test Data\ncab = test_cat[\"Cabin\"].apply(lambda x: x.split(\"/\"))\n\n# Categorical Data\ntest_cat[\"Cab_1\"] = cab.apply(lambda x: x[0])\ntest_cat[\"Cab_3\"] = cab.apply(lambda x: x[2])\ntest_nom[\"Cab_2\"] = cab.apply(lambda x: float(x[1]))\ntest_cat.drop(columns=[\"Cabin\"], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:48.011430Z","iopub.execute_input":"2022-08-05T16:50:48.012585Z","iopub.status.idle":"2022-08-05T16:50:48.054807Z","shell.execute_reply.started":"2022-08-05T16:50:48.012529Z","shell.execute_reply":"2022-08-05T16:50:48.053253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Nominal Data Distribution\nData distribution is checked for several reason.\n1. Check for skewed data.\n2. Check correlation between feature and target","metadata":{}},{"cell_type":"code","source":"# cols = train_nom.columns\nfig, axs = plt.subplots(2, 4)\nfig.set_size_inches(20, 8)\n\ncols = train_nom.columns\nfor i in range(2):\n    for j in range(4):\n        idx = i*4+j\n        if idx < len(cols):\n            temp = pd.DataFrame()\n            temp[cols[idx]] = train_nom[cols[idx]]\n            temp[\"Transported\"] = target\n            _ = sns.histplot(x=cols[idx], data=temp, hue=\"Transported\", element=\"step\", \n                             bins=12, ax=axs[i, j])\n        else:\n            axs[i, j].axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:48.056719Z","iopub.execute_input":"2022-08-05T16:50:48.057227Z","iopub.status.idle":"2022-08-05T16:50:49.461583Z","shell.execute_reply.started":"2022-08-05T16:50:48.057175Z","shell.execute_reply":"2022-08-05T16:50:49.459807Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"1. Althoug skewed, \"Age\" and \"Cab_2\" seems have data distributed into all the value.\n2. But \"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\" seems have data concentrated on certain value, near 0 value.\n3. \"FoodCourt\" and \"ShoppingMall\" dont show different distribution of Transported data. It seems that the features does not correlated with the target.\n4. Examine \"RoomService\" and \"FoodCourt\" to deal with their problem.","metadata":{}},{"cell_type":"markdown","source":"### Deal with skewed Feature","metadata":{}},{"cell_type":"code","source":"cols = [\"RoomService\", \"FoodCourt\", \"ShoppingMall\", \"Spa\", \"VRDeck\", \"Cab_2\"]\nfor col in cols:\n    train_nom[col] = (train_nom[col])**(1/2)\n    test_nom[col] = (test_nom[col])**(1/2)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:49.463265Z","iopub.execute_input":"2022-08-05T16:50:49.463663Z","iopub.status.idle":"2022-08-05T16:50:49.476936Z","shell.execute_reply.started":"2022-08-05T16:50:49.463629Z","shell.execute_reply":"2022-08-05T16:50:49.475438Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_nom[\"Age\"] = np.log10(train_nom[\"Age\"]+2)\ntest_nom[\"Age\"] = np.log10(test_nom[\"Age\"]+2)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:49.479185Z","iopub.execute_input":"2022-08-05T16:50:49.479559Z","iopub.status.idle":"2022-08-05T16:50:49.488729Z","shell.execute_reply.started":"2022-08-05T16:50:49.479526Z","shell.execute_reply":"2022-08-05T16:50:49.487435Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Create Categorical Data\nData with dominant 0 values will be added to categorical data. 0 and bigger than 0 as 1.","metadata":{}},{"cell_type":"markdown","source":"# Categorical Features Distribution\nCheck categorical feature correlation with target. The strategy is to see wheter each category inside the features have different Transported count. Features with similar Transported count are assumed as no connection with target.","metadata":{}},{"cell_type":"code","source":"cols = train_cat.columns\nprint(cols)\nfig, axs = plt.subplots(2, 4)\nfig.set_size_inches(18, 8)\n\nfor i in range(2):\n    for j in range(4):\n        idx = i*4+j\n        \n        if idx < len(cols):\n            col = cols[idx]\n            if col != \"Cabin\":        \n                temp = pd.DataFrame()\n                temp[col] = train_cat[col]\n                temp[\"Transported\"] = target\n\n                _=sns.histplot(x=col, data=temp, hue=\"Transported\", ax=axs[i, j])\n        else:\n            axs[i, j].axis(\"off\")","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:49.490214Z","iopub.execute_input":"2022-08-05T16:50:49.490628Z","iopub.status.idle":"2022-08-05T16:50:50.940552Z","shell.execute_reply.started":"2022-08-05T16:50:49.490555Z","shell.execute_reply":"2022-08-05T16:50:50.938491Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only \"VIP\" feature have similar Transported count for each category, it means that the features inside \"VIP\" did not affect Transported. But let's check numerically using chi squre test.","metadata":{}},{"cell_type":"markdown","source":"## Check Correlation using Chi Square Test\nLet's confirm categorical value using numerical technique, chi square test.","metadata":{}},{"cell_type":"code","source":"# Calculate Chi Square - for later use\ndef chi_square(cat_1, cat_2):\n    ctab = pd.crosstab(cat_1, cat_2, margins=True, margins_name=\"Total\")\n\n    row, col = ctab.shape\n    chi = 0\n    n = len(train_df)\n\n    for i in range(row-1):\n        for j in range(col-1):\n            E = ctab.iloc[i, -1]*ctab.iloc[-1, j]/ctab.iloc[-1, -1]\n            chi += (ctab.iloc[i, j] - E)**2 / E\n\n    p = 1 - stats.chi2.cdf(chi, (row-1)*(col-1))\n    return p\n\n# Calculate for All Category Data\ncorr = []\ncols = train_cat.columns\n\nfor cat in cols:\n    p = chi_square(train_cat[cat], target)\n    corr.append([cat, p])\n    \ncorr = pd.DataFrame(corr, columns=[\"Feature\", \"pValue\"])\ncorr","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:50.945678Z","iopub.execute_input":"2022-08-05T16:50:50.946055Z","iopub.status.idle":"2022-08-05T16:50:51.191764Z","shell.execute_reply.started":"2022-08-05T16:50:50.946023Z","shell.execute_reply":"2022-08-05T16:50:51.190189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"All of the categorical features have <0.05 value, so they are related to the target.","metadata":{}},{"cell_type":"markdown","source":"# Add Feature by Grouping\nCreate additional features by group passenger with similar attributes using clustering method, and add as categorical data.","metadata":{}},{"cell_type":"code","source":"# Prepare Data\nscaler = Normalizer()\n\n# Train Data\na = scaler.fit_transform(train_nom)\nb = pd.get_dummies(train_cat, drop_first=True)\nX_train = np.concatenate((a, b), axis=1)\n\nkmean = KMeans(n_clusters=2)\nkmean.fit(X_train)\npred = kmean.predict(X_train)\n\ntrain_cat[\"Group\"] = pred\n\n# Test Data\na = scaler.fit_transform(test_nom)\nb = pd.get_dummies(test_cat, drop_first=True)\nX = np.concatenate((a, b), axis=1)\ntest_cat[\"Group\"] = kmean.predict(X)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:51.193292Z","iopub.execute_input":"2022-08-05T16:50:51.193652Z","iopub.status.idle":"2022-08-05T16:50:52.413484Z","shell.execute_reply.started":"2022-08-05T16:50:51.193619Z","shell.execute_reply":"2022-08-05T16:50:52.411846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare Data\nThe data was separated into nominal and categorical data. So, they will be prepared before put on machine learning model. \n1. Scale nominal data. After some trial, normalizer show better result.\n2. Encode Categorical features. Since categorical features are not many, they will be hot encoded. Usage of drop_first shows better result.\n3. Concatenate those data into 1 variable","metadata":{}},{"cell_type":"code","source":"scaler = Normalizer()\n\n# Train Data\na = scaler.fit_transform(train_nom)\nb = pd.get_dummies(train_cat, drop_first=True)\nX_train = np.concatenate((a, b), axis=1)\n\n# Test Data\na = scaler.fit_transform(test_nom)\nb = pd.get_dummies(test_cat, drop_first=True)\nX_test = np.concatenate((a, b), axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:52.416750Z","iopub.execute_input":"2022-08-05T16:50:52.417199Z","iopub.status.idle":"2022-08-05T16:50:52.461742Z","shell.execute_reply.started":"2022-08-05T16:50:52.417158Z","shell.execute_reply":"2022-08-05T16:50:52.460064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Classification Using SVM\nafter trying for several classification model, I decided to use SVM-SVC which provide best result. And after using hyper parameter tuning, I use kernel=\"poly\" for parameter. Cross validation is utilized to check model performance.","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn import svm\n\nmodel = svm.SVC(kernel=\"poly\", shrinking=False)\nmodel.fit(X_train, target)\n\nscores = cross_val_score(model, X_train, target, cv=5)\nprint(np.average(scores))","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:50:52.464813Z","iopub.execute_input":"2022-08-05T16:50:52.465154Z","iopub.status.idle":"2022-08-05T16:51:04.930131Z","shell.execute_reply.started":"2022-08-05T16:50:52.465123Z","shell.execute_reply":"2022-08-05T16:51:04.928779Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"the model shows quite good result, so it ready to make submission prediction.","metadata":{}},{"cell_type":"markdown","source":"# Create Submission","metadata":{}},{"cell_type":"code","source":"# Create Submission\npred = model.predict(X_test)\n\nsubm = pd.DataFrame()\nsubm[\"PassengerId\"] = pass_id\nsubm[\"Transported\"] = pred.astype(str)\n\nsubm.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T16:51:04.931717Z","iopub.execute_input":"2022-08-05T16:51:04.932073Z","iopub.status.idle":"2022-08-05T16:51:05.643007Z","shell.execute_reply.started":"2022-08-05T16:51:04.932040Z","shell.execute_reply":"2022-08-05T16:51:05.641990Z"},"trusted":true},"execution_count":null,"outputs":[]}]}