{"cells":[{"metadata":{"_uuid":"dc9be7c2ae5385e0c39f254a08cda97c1c9f1a81"},"cell_type":"markdown","source":"**1. Introduction: (#1)**\n\n**2. Loading Data and Explanation of Features: (#2)**\n\n**3. Exploratory Data Analysis (EDA): (#3)**\n\n**4. Feature engineering: (#4)**\n\n**5. Modeling: (#5)**\n\n**6. Conclusion: (#6)**"},{"metadata":{"_uuid":"e65d107a7a7af2276a8456fbf493a32a0c072c5a"},"cell_type":"markdown","source":"<a id=\"1\"></a> \n**1. Introduction**\n\nHello everyone!  This is my first competition kernel. I choosed Titanic dataset which is a famous case and also has good features to work with.\n\nThe datasets contains:\n\n* survival==>\t    Survival==>\t                                                       0 = No, 1 = Yes\n* pclass==>\t        Ticket class==>\t                                                 1 = 1st, 2 = 2nd, 3 = 3rd\n* sex==>\t         Sex\t\n* Age==>\t         Age in years\t\n* sibsp==>           # of siblings / spouses aboard the Titanic\t\n* parch==>\t        # of parents / children aboard the Titanic\t\n* ticket==>\t          Ticket number\t\n* fare==>\t          Passenger fare\t\n* cabin==>\t        Cabin number\t\n* embarked==>  Port of Embarkation==>\t                                  C = Cherbourg, Q = Queenstown, S = Southampton"},{"metadata":{"_uuid":"d7a8fb46469fb3386506cb088fa746bedac69b34"},"cell_type":"markdown","source":"<a id=\"2\"></a> \n**2. Loading Data and Explanation of Features**"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"#data analysis libraries \nimport numpy as np\nimport pandas as pd\nimport random\n\n#visualization libraries\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n%matplotlib inline\n\n#ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\nimport os\nprint(os.listdir(\"../input\"))","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"data_train=pd.read_csv(\"../input/train.csv\") # train data\ndata_test=pd.read_csv(\"../input/test.csv\") # test data\n\nprint(\"Train info:\\n\")\ndata_train.info()\nprint(\"-\"*40)\nprint(\"Test info:\\n\")\ndata_test.info()\ndata_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ab10260c7b71aaf39504364a54d164e45739a411"},"cell_type":"markdown","source":"Datas consist different dtypes and some of the columns have NaN values.\n\n* **Numerical Features**: Age (Continuous), Fare (Continuous), SibSp (Discrete), Parch (Discrete)\n* **Categorical Features**: Survived, Sex, Embarked, Pclass\n* **Alphanumeric Features**: Ticket, Cabin"},{"metadata":{"trusted":true,"_uuid":"83e97dae39ad58e53aeae65ea704af05873bc198"},"cell_type":"code","source":"print('Train columns with null values:\\n',data_train.isnull().sum()) # sum. of null values\nprint(\"-\"*40)\nprint('Test columns with null values:\\n',data_test.isnull().sum())\ndata_train.describe(include = 'all')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4ed4da585a6c1b693bfed13bc7135e2d21d27151"},"cell_type":"markdown","source":"Cabin feature has lots of Nan value also Age feature has a good amount of Nan values. Fare and Embark Nan values may be changed easily."},{"metadata":{"_uuid":"a0b0bc209e68bd9b6ca5bf905cd264fd43365a4b"},"cell_type":"markdown","source":"<a id=\"3\"></a> \n**3. Exploratory Data Analysis (EDA)**\n\nLets display each feature to understand better."},{"metadata":{"trusted":true,"_uuid":"f66d5d364ff9b667cfbb495a4bed00d572e68267"},"cell_type":"code","source":"print(data_train['Survived'].value_counts())\n\nsns.set()\nf,ax=plt.subplots(1,2,figsize=(12,5))\ndata_train['Survived'].value_counts().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,ax=ax[1])\nax[1].set_title('Rate of the Survived',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"614bc08710c636311499794614286981a0b3457f"},"cell_type":"code","source":"print(data_train.groupby(['Pclass','Survived'])['Survived'].count())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train.groupby(['Pclass','Survived'])['Survived'].count().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived vs Dead by Pclass',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,hue='Pclass',ax=ax[1])\nax[1].set_title('Survived vs Dead by Pclass',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Pclass'])['Survived'].mean().index,y=data_train.groupby(['Pclass'])['Survived'].mean().values,ax=ax[2])\nax[2].set_title('Rate of the Survived by Pclass',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"aca5f3e5763c761ed6017ae1fffce8f00056c457"},"cell_type":"code","source":"print(data_train.groupby(['Sex','Survived'])['Survived'].count())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train.groupby(['Sex','Survived'])['Survived'].count().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived vs Dead by Sex',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,hue='Sex',ax=ax[1])\nax[1].set_title('Survived vs Dead by Sex',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Sex'])['Survived'].mean().index,y=data_train.groupby(['Sex'])['Survived'].mean().values,ax=ax[2])\nax[2].set_title('Rate of the Survived by Sex',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e8590bdff3135cae94d7d2f6812dabd19cf0239a"},"cell_type":"code","source":"print('Age of the oldest passanger:',data_train['Age'].max(),'years old')\nprint('Age of the youngest passanger:',data_train['Age'].min(),'years old')\nprint('Average Age on the ship:',data_train['Age'].mean(),'years old')\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train[data_train['Survived']==0].Age.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='b')\nax[0].set_title('Dead',color = 'r',fontsize=15)\n\ndata_train[data_train['Survived']==1].Age.plot.hist(ax=ax[1],color='orange',bins=20,edgecolor='black')\nax[1].set_title('Survived',color = 'r',fontsize=15)\n\nsns.kdeplot(data_train[\"Age\"][(data_train[\"Survived\"] == 0) & (data_train[\"Age\"].notnull())], color=\"Red\", shade = True,ax=ax[2])\nsns.kdeplot(data_train[\"Age\"][(data_train[\"Survived\"] == 1) & (data_train[\"Age\"].notnull())], color=\"Blue\", shade= True,ax=ax[2])\nax[2].set_title('Rate of Survived vs Dead',color = 'r',fontsize=15)\nax[2].legend(['Dead','Survived'])\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f66412d6fedd7c806cb8c5100bf81cd036099df"},"cell_type":"code","source":"print(data_train.groupby(['SibSp','Survived'])['Survived'].count())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train[data_train['Survived']==0].SibSp.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='b')\nax[0].set_title('Dead',color = 'r',fontsize=15)\n\ndata_train[data_train['Survived']==1].SibSp.plot.hist(ax=ax[1],color='orange',bins=20,edgecolor='black')\nax[1].set_title('Survived',color = 'r',fontsize=15)\n\nsns.factorplot('SibSp','Survived',data=data_train,ax=ax[2])\nax[2].set_title('Rate of the Survived',color = 'r',fontsize=15)\nplt.close(2)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fdcb577fc8135797964820e60a140c39435e5695"},"cell_type":"code","source":"print(data_train.groupby(['Parch','Survived'])['Survived'].count())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train[data_train['Survived']==0].Parch.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='b')\nax[0].set_title('Dead',color = 'r',fontsize=15)\n\ndata_train[data_train['Survived']==1].Parch.plot.hist(ax=ax[1],color='orange',bins=20,edgecolor='black')\nax[1].set_title('Survived',color = 'r',fontsize=15)\n\nsns.factorplot('Parch','Survived',data=data_train,ax=ax[2])\nax[2].set_title('Rate of the Survived',color = 'r',fontsize=15)\nplt.close(2)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6eebb010d4ba43a957b956e4cb8729171df38a40"},"cell_type":"code","source":"data_train['Ticket'].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"32d47487b14e716ec7d8030add203041f59a919e"},"cell_type":"markdown","source":"The most of the Tickets are different from each other. Too bad there are much values to visualise."},{"metadata":{"trusted":true,"_uuid":"27ea276650b305cfa0f7a89dae5d66186a36f27a"},"cell_type":"code","source":"print('The highest fare was:',data_train['Fare'].max())\nprint('The lowest fare was:',data_train['Fare'].min())\nprint('The avarage was:',data_train['Fare'].mean())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train[data_train['Survived']==0].Fare.plot.hist(ax=ax[0],bins=20,edgecolor='black',color='b')\nax[0].set_title('Dead',color = 'r',fontsize=15)\n\ndata_train[data_train['Survived']==1].Fare.plot.hist(ax=ax[1],color='orange',bins=20,edgecolor='black')\nax[1].set_title('Survived',color = 'r',fontsize=15)\n\nsns.kdeplot(data_train[\"Fare\"][(data_train[\"Survived\"] == 0) & (data_train[\"Age\"].notnull())], color=\"Red\", shade = True,ax=ax[2])\nsns.kdeplot(data_train[\"Fare\"][(data_train[\"Survived\"] == 1) & (data_train[\"Age\"].notnull())], color=\"Blue\", shade= True,ax=ax[2])\nax[2].set_title('Rate of the Survived',color = 'r',fontsize=15)\nax[2].legend(['Dead','Survived'])\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a83d1fc65830f157d8ae74a57566ea89d970acd4"},"cell_type":"code","source":"data_train['Cabin'].describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"278d2f1555de5d602a29bfcb550114bc7b29212b"},"cell_type":"markdown","source":"Although having little information about Cabins, they may be categorized somehow."},{"metadata":{"trusted":true,"_uuid":"07648eddbb5c19474ae7093fb8893b3ad2a5c725"},"cell_type":"code","source":"print(data_train.groupby(['Embarked','Survived'])['Survived'].count())\n\nf,ax=plt.subplots(1,3,figsize=(20,5))\n\ndata_train.groupby(['Embarked','Survived'])['Survived'].count().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived vs Dead by Embarked',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,hue='Embarked',ax=ax[1])\nax[1].set_title('Survived vs Dead by Embarked',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Embarked'])['Survived'].mean().index,y=data_train.groupby(['Embarked'])['Survived'].mean().values,ax=ax[2])\nax[2].set_title('Rate of the Survived by Embarked',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9fbb60ceeefe2a168a1d5dac5129a9f376dd246f"},"cell_type":"markdown","source":"**Observations:**\n\n* Being higher the sosyoeconomic ranking, are more likely to survive.\n* Females are more likely to survive.\n* Children are way more  likely to survive.\n* Having more than 4 family members Decreases survival rate.\n* Passengers having expensive fare are more likely to survive.\n* Passengers from Cherbourg are more likely to survive."},{"metadata":{"_uuid":"08cefbffeb027d366fff07d25ac3865dffdd1bae"},"cell_type":"markdown","source":"<a id=\"4\"></a> \n**4. Feature engineering**\n\nNow its time for predict NaN values from the  features and create new category for useful information. Firstly titles will be stripped from names."},{"metadata":{"trusted":true,"_uuid":"9896ffa1f1db6fd7183719d36d551689cf2d7a51"},"cell_type":"code","source":"Title_train=[i.split(\",\")[1].split(\".\")[0].strip() for i in data_train[\"Name\"]] # split names from , to .\ndata_train[\"Title\"] = pd.Series(Title_train)\n\nTitle_test=[i.split(\",\")[1].split(\".\")[0].strip() for i in data_test[\"Name\"]]\ndata_test[\"Title\"] = pd.Series(Title_test)\n\nTitle=pd.concat([data_train[['Title','Sex']],data_test[['Title','Sex']]],axis=0) \n\nprint(Title.groupby(['Title','Sex'])['Title'].count())\n\nplt.figure(figsize=(15,5))\n\nsns.barplot(x=Title[\"Title\"].value_counts().index,y=Title[\"Title\"].value_counts().values)\nplt.xticks(rotation=45)\nplt.title('Titles',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"54f3b1ac98c926481ef5988275e298a24a0533cb"},"cell_type":"markdown","source":"As you can see above many titles are very low count so they can be placed to more common title."},{"metadata":{"trusted":true,"_uuid":"8b06b05c74014afa5e4b7886e7a49c1bcdf1057c"},"cell_type":"code","source":"data_train['Title'].replace(['Lady', 'the Countess','Countess','Capt', 'Col','Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'],\n                            'Rare',inplace=True)\ndata_train['Title'].replace(['Mlle','Mme','Ms'], 'Miss',inplace=True)\n\ndata_test['Title'].replace(['Lady', 'the Countess','Countess','Capt', 'Col','Don', 'Dr', 'Major', 'Rev', 'Sir', 'Jonkheer', 'Dona'],\n                           'Rare',inplace=True)\ndata_test['Title'].replace(['Mlle','Mme','Ms'], 'Miss',inplace=True)\n\ndata_train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2263258f51b663cb5fe8779e963f834f2dc53e93"},"cell_type":"markdown","source":"Titles are useful as a new feateure which helps to define passengers have NaN Age value."},{"metadata":{"trusted":true,"_uuid":"94aa405b35abac8353af10a5ba7e8933b3dce716"},"cell_type":"code","source":"Age=pd.concat([data_train[['Title','Age']],data_test[['Title','Age']]],axis=0)\n\nprint(Age.groupby('Title')['Age'].mean())\n\ndata_train.loc[(data_train['Age'].isnull())&(data_train['Title']=='Master'),'Age'] = Age[Age['Title']=='Master'].Age.mean()\ndata_train.loc[(data_train['Age'].isnull())&(data_train['Title']=='Miss'),'Age'] = Age[Age['Title']=='Miss'].Age.mean()\ndata_train.loc[(data_train['Age'].isnull())&(data_train['Title']=='Mr'),'Age'] = Age[Age['Title']=='Mr'].Age.mean()\ndata_train.loc[(data_train['Age'].isnull())&(data_train['Title']=='Mrs'),'Age'] = Age[Age['Title']=='Mrs'].Age.mean()\ndata_train.loc[(data_train['Age'].isnull())&(data_train['Title']=='Rare'),'Age'] = Age[Age['Title']=='Rare'].Age.mean()\n\ndata_test.loc[(data_test['Age'].isnull())&(data_test['Title']=='Master'),'Age'] = Age[Age['Title']=='Master'].Age.mean()\ndata_test.loc[(data_test['Age'].isnull())&(data_test['Title']=='Miss'),'Age'] = Age[Age['Title']=='Miss'].Age.mean()\ndata_test.loc[(data_test['Age'].isnull())&(data_test['Title']=='Mr'),'Age'] = Age[Age['Title']=='Mr'].Age.mean()\ndata_test.loc[(data_test['Age'].isnull())&(data_test['Title']=='Mrs'),'Age'] = Age[Age['Title']=='Mrs'].Age.mean()\ndata_test.loc[(data_test['Age'].isnull())&(data_test['Title']=='Rare'),'Age'] = Age[Age['Title']=='Rare'].Age.mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"73898a0a2005711455e3279d3f28bca746366b06"},"cell_type":"code","source":"f,ax=plt.subplots(1,3,figsize=(23,6))\n\ndata_train.groupby(['Title','Survived'])['Survived'].count().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived vs Dead by Title',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,hue='Title',ax=ax[1])\nax[1].set_title('Survived vs Dead by Title',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Title'])['Survived'].mean().index,y=data_train.groupby(['Title'])['Survived'].mean().values,ax=ax[2])\nax[2].set_title('Rate of the Survived by Title',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eb4053829165e3c4db3f41e56d309054924debf5"},"cell_type":"code","source":"data_test[data_test['Fare'].isnull()]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4527a1e817e901b4e6874ca180ee024685e4ac23"},"cell_type":"markdown","source":"Only one passenger has a NaN Fare value. I am going to fill it with the  passengers' fare whose have common features. As common features I choosed Pclass, Embarked, Parch, Sex, SibSp and Title to be close as possible."},{"metadata":{"trusted":true,"_uuid":"47ed6637a996175e109d6f38af3ee44631180ee5"},"cell_type":"code","source":"Fare=pd.concat([data_train[['Fare','Pclass','Embarked','Parch','Sex','SibSp','Title']],\n                data_test[['Fare','Pclass','Embarked','Parch','Sex','SibSp','Title']]],axis=0)\n\ndata_test['Fare'].fillna(Fare[(Fare[\"Pclass\"]==3) & (Fare[\"Embarked\"]=='S') & (Fare[\"SibSp\"]==0) & \n           (Fare[\"Parch\"]==0) & (Fare[\"Sex\"]=='male') & (Fare[\"Title\"]=='Mr')].Fare.median(),inplace=True)\n\ndata_test.iloc[152]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8120342b28a6ddb9aedd98dd1cce003fc4aefca5"},"cell_type":"markdown","source":"Secondly I am going to categorize Cabins by initials than distinguish them by Pclass, Embarked and Fare."},{"metadata":{"trusted":true,"_uuid":"16f750192b8ccfd54f5ffed3f849cdee873a1e5c"},"cell_type":"code","source":"data_train['Cabin'] = data_train['Cabin'].str[0] # add initial value to same location\n\ndata_test['Cabin'] = data_test['Cabin'].str[0]\n\nCabin=pd.concat([data_train[['Cabin','Embarked','Pclass','Fare']],data_test[['Cabin','Embarked','Pclass','Fare']]],axis=0)\n\nCabin.groupby(['Pclass','Embarked','Cabin'])['Fare'].max()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"fe4bfb6a871a26fb83d64fda1c86ce2b1d560410"},"cell_type":"markdown","source":"I am going to apply NaN Cabin values from information which is above,. This process may be bulky but it helps to fill the NaN values with the logical way."},{"metadata":{"trusted":true,"_uuid":"4b42e0861b9b9a5174f62aa4a3d3f52b83bd576f"},"cell_type":"code","source":"data_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='C')&(data_train.Fare<=56.9292),'Cabin']='A'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='C')&(data_train.Fare>56.9292)&(data_train.Fare<=113.2750),'Cabin']='D'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='C')&(data_train.Fare>113.2750)&(data_train.Fare<=134.5000),'Cabin']='E'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='C')&(data_train.Fare>134.5000)&(data_train.Fare<=227.5250),'Cabin']='C'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='C')&(data_train.Fare>227.5250),'Cabin']='B'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare<=35.5000),'Cabin']='T'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare>35.5000)&(data_train.Fare<=77.9583),'Cabin']='D'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare>77.9583)&(data_train.Fare<=79.6500),'Cabin']='E'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare>79.6500)&(data_train.Fare<=81.8583),'Cabin']='A'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare>81.8583)&(data_train.Fare<=211.3375),'Cabin']='B'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&\n               (data_train.Embarked=='S')&(data_train.Fare>211.3375),'Cabin']='C'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==1)&(data_train.Embarked=='Q'),'Cabin']='C'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==2)&\n               (data_train.Embarked=='S')&(data_train.Fare<=13.0000),'Cabin']=random.sample(['D','E'],1)\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==2)&\n               (data_train.Embarked=='S')&(data_train.Fare>13.0000),'Cabin']='F'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==2)&(data_train.Embarked=='C'),'Cabin']='D'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==2)&(data_train.Embarked=='Q'),'Cabin']='E'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==3)&\n               (data_train.Embarked=='S')&(data_train.Fare<=7.6500),'Cabin']='F'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==3)&\n               (data_train.Embarked=='S')&(data_train.Fare>7.6500)&(data_train.Fare<=12.4750),'Cabin']='E'\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==3)&\n               (data_train.Embarked=='S')&(data_train.Fare>12.4750),'Cabin']='G'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==3)&(data_train.Embarked=='C'),'Cabin']='F'\n\ndata_train.loc[(data_train.Cabin.isnull())&(data_train.Pclass==3)&(data_train.Embarked=='Q'),'Cabin']='F'\n\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='C')&(data_test.Fare<=56.9292),'Cabin']='A'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='C')&(data_test.Fare>56.9292)&(data_test.Fare<=113.2750),'Cabin']='D'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='C')&(data_test.Fare>113.2750)&(data_test.Fare<=134.5000),'Cabin']='E'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='C')&(data_test.Fare>134.5000)&(data_test.Fare<=227.5250),'Cabin']='C'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='C')&(data_test.Fare>227.5250),'Cabin']='B'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare<=35.5000),'Cabin']='T'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare>35.5000)&(data_test.Fare<=77.9583),'Cabin']='D'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare>77.9583)&(data_test.Fare<=79.6500),'Cabin']='E'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare>79.6500)&(data_test.Fare<=81.8583),'Cabin']='A'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare>81.8583)&(data_test.Fare<=211.3375),'Cabin']='B'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&\n               (data_test.Embarked=='S')&(data_test.Fare>211.3375),'Cabin']='C'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==1)&(data_test.Embarked=='Q'),'Cabin']='C'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==2)&\n               (data_test.Embarked=='S')&(data_test.Fare<=13.0000),'Cabin']=random.sample(['D','E'],1)\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==2)&\n               (data_test.Embarked=='S')&(data_test.Fare>13.0000),'Cabin']='F'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==2)&(data_test.Embarked=='C'),'Cabin']='D'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==2)&(data_test.Embarked=='Q'),'Cabin']='E'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==3)&\n               (data_test.Embarked=='S')&(data_test.Fare<=7.6500),'Cabin']='F'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==3)&\n               (data_test.Embarked=='S')&(data_test.Fare>7.6500)&(data_test.Fare<=12.4750),'Cabin']='E'\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==3)&\n               (data_test.Embarked=='S')&(data_test.Fare>12.4750),'Cabin']='G'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==3)&(data_test.Embarked=='C'),'Cabin']='F'\n\ndata_test.loc[(data_test.Cabin.isnull())&(data_test.Pclass==3)&(data_test.Embarked=='Q'),'Cabin']='F'\n\nprint(data_test.Cabin.isnull().any(),'\\n')\n\nprint(data_train.Cabin.isnull().any())","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"28ef41cba23a03d8e7e5f30eb0c598b62b7c2961"},"cell_type":"code","source":"f,ax=plt.subplots(1,3,figsize=(23,6))\n\ndata_train.groupby(['Cabin','Survived'])['Survived'].count().plot.pie(autopct='%1.1f%%',ax=ax[0])\nax[0].set_title('Survived vs Dead by Cabin',color = 'r',fontsize=15)\nax[0].set_ylabel('')\n\nsns.countplot('Survived',data=data_train,hue='Cabin',ax=ax[1])\nax[1].set_title('Survived vs Dead by Cabin',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Cabin'])['Survived'].mean().index,\n            y=data_train.groupby(['Cabin'])['Survived'].mean().values,ax=ax[2])\nax[2].set_title('Rate of the Survived by Cabin',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d25296dc123393842d0cbe34eae1b8188241ce13"},"cell_type":"markdown","source":"Looks like Cabins B and C have more survival rate than others."},{"metadata":{"_uuid":"73f02d6c983f07f47f4b7051ce380f3de99f430a"},"cell_type":"markdown","source":"Thirdly its time for NaN Embarked values. "},{"metadata":{"trusted":true,"_uuid":"9aa1c49ad89ef8f57c463d8781577838bf049ae3"},"cell_type":"code","source":"Cabin=pd.concat([data_train[['Cabin','Embarked','Pclass','Fare']],data_test[['Cabin','Embarked','Pclass','Fare']]],axis=0)\n\ndata_train[data_train['Embarked'].isnull()]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"319dbda8398730147f877273156fcfca7c3dbb43"},"cell_type":"markdown","source":"These two passengers share most of the features so their Embarked can be predicted from passengers that share same features."},{"metadata":{"trusted":true,"_uuid":"0b9907ae12fd610ba94a1dff2eb11ae774d7fa33"},"cell_type":"code","source":"f,ax=plt.subplots(1,2,figsize=(12,5))\n\nsns.countplot('Survived',\n              data=data_train.loc[(data_train.Pclass==1)&(data_train.Cabin=='B')&(data_train.Embarked.notnull())&(data_train.Sex=='female')],\n              hue='Embarked',\n              ax=ax[0])\nax[0].set_title('Survived',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.loc[(data_train.Embarked=='C')|(data_train.Embarked=='S')].groupby(['Embarked'])['Survived'].mean().index,\n            y=data_train.loc[(data_train.Embarked=='C')|(data_train.Embarked=='S')].groupby(['Embarked'])['Survived'].mean().values,\n            ax=ax[1])\nax[1].set_title('Rate of the Survived by Embarked',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"cb419d23e1f69370f95b706c88bd4fc451a15a20"},"cell_type":"markdown","source":"Two Embarked locations are displayed above for these two passenger. C has more survival rate than S. So I am going to Choose C for their Embarked value because these two passengers survived."},{"metadata":{"trusted":true,"_uuid":"1eb4cb4d20dab8cc121d0cad910317f417ac48a1"},"cell_type":"code","source":"data_train['Embarked'].fillna('C',inplace=True)\n\ndata_train.iloc[[61,829]]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ccc0047cce86c7747fb2c55f13e891b5cd9851ee"},"cell_type":"markdown","source":"I completed filling the NaN values. Lets verify it. "},{"metadata":{"trusted":true,"_uuid":"665d8d86edc82d87a433ae7aeebda60b66a2afe4"},"cell_type":"code","source":"print('Train columns with null values:\\n',data_train.isnull().sum())\nprint(\"-\"*40)\nprint('Test columns with null values:\\n',data_test.isnull().sum())","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3766603ef6efc077daa531d169e9b573e20001ef"},"cell_type":"markdown","source":"This time I am going to replace object values to numerical values for modelling and categorize some features for handling much easier."},{"metadata":{"trusted":true,"_uuid":"055317bd104137f5aa4f18ec6694fe1375b9c3b2"},"cell_type":"code","source":"data_train['Sex'].replace(['male','female'],[0,1],inplace=True)\ndata_train['Embarked'].replace(['C','Q','S'],[0,1,2],inplace=True)\ndata_train['Title'].replace(['Master','Miss','Mr','Mrs','Rare'],[0,1,2,3,4],inplace=True)\ndata_train['Cabin'].replace(['A','B','C','D','E','F','G','T'],[0,1,2,3,4,5,6,7],inplace=True)\n\ndata_train.loc[data_train['Age']<=16,'Age']=0\ndata_train.loc[(data_train['Age']>16)&(data_train['Age']<=32),'Age']=1\ndata_train.loc[(data_train['Age']>32)&(data_train['Age']<=48),'Age']=2\ndata_train.loc[(data_train['Age']>48)&(data_train['Age']<=64),'Age']=3\ndata_train.loc[data_train['Age']>64,'Age']=4\n\ndata_test['Sex'].replace(['male','female'],[0,1],inplace=True)\ndata_test['Embarked'].replace(['C','Q','S'],[0,1,2],inplace=True)\ndata_test['Title'].replace(['Master','Miss','Mr','Mrs','Rare'],[0,1,2,3,4],inplace=True)\ndata_test['Cabin'].replace(['A','B','C','D','E','F','G','T'],[0,1,2,3,4,5,6,7],inplace=True)\n\ndata_test.loc[data_test['Age']<=16,'Age']=0\ndata_test.loc[(data_test['Age']>16)&(data_test['Age']<=32),'Age']=1\ndata_test.loc[(data_test['Age']>32)&(data_test['Age']<=48),'Age']=2\ndata_test.loc[(data_test['Age']>48)&(data_test['Age']<=64),'Age']=3\ndata_test.loc[data_test['Age']>64,'Age']=4","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d2f3b6ce45eaea7da803544e666c50c9be9da2a6"},"cell_type":"markdown","source":"I seperated Age values to 5 equal range."},{"metadata":{"trusted":true,"_uuid":"8ea6b6b5dfe85688cb4c8d9bdca3d28d6286c882"},"cell_type":"code","source":"data_train['Family_Size']=0\ndata_train['Family_Size']=data_train['Parch']+data_train['SibSp']\n\ndata_test['Family_Size']=0\ndata_test['Family_Size']=data_test['Parch']+data_test['SibSp']","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3cdd720709a583bac4b44df3a0780be6c942b610"},"cell_type":"markdown","source":"Also added new feature as family which keeps total family count."},{"metadata":{"trusted":true,"_uuid":"1cda94867643135f16b335e3ebc30ccf8e80f837"},"cell_type":"code","source":"sns.factorplot('Family_Size','Survived',data=data_train)\nplt.title('Family_Size',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f983d4a01ac694b1e44bfc20a53a890cbe068924"},"cell_type":"markdown","source":"From the plot above, it show family member's count. Having 0 member means that passenger is single."},{"metadata":{"trusted":true,"_uuid":"5a45592c448cae95301b49cda50a4395c43b0ec2"},"cell_type":"code","source":"data_train.loc[(data_train['Family_Size']>0)&(data_train['Family_Size']<4),'Family_Size']=1\ndata_train.loc[(data_train['Family_Size']>=4),'Family_Size']=2\n\ndata_test.loc[(data_test['Family_Size']>0)&(data_test['Family_Size']<4),'Family_Size']=1\ndata_test.loc[(data_test['Family_Size']>=4),'Family_Size']=2\n\nf,ax=plt.subplots(1,2,figsize=(12,5))\n\nsns.countplot('Family_Size',data=data_train,hue='Survived',ax=ax[0])\nax[0].set_title('Survived',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Family_Size'])['Survived'].mean().index,\n            y=data_train.groupby(['Family_Size'])['Survived'].mean().values,\n            ax=ax[1])\nax[1].set_title('Rate of the Survived by Family_Size',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1ff310d8895d8f897e2413cc9ba82feed29f0f92"},"cell_type":"markdown","source":"In addition I seperated family size to 3 groups. First one is for singles, second one is for small family and third one is for large family. Because from the previous plot there is a obvious distinctions between having small and large families. "},{"metadata":{"trusted":true,"_uuid":"005ac167c86573e719f7ef99678e38d453618746"},"cell_type":"code","source":"Fare.groupby(['Pclass'])['Fare'].median()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1acc7bc84c48b267f7b5643987ab0af4140561f8"},"cell_type":"markdown","source":"I categorize the fare into 4 grups under the Pclass."},{"metadata":{"trusted":true,"_uuid":"88a773f7425bced09d16dbeef396f103ee26ef6b"},"cell_type":"code","source":"data_train.loc[data_train['Fare']<=8.0500,'Fare']=0\ndata_train.loc[(data_train['Fare']>8.0500)&(data_train['Fare']<=15.0458),'Fare']=1\ndata_train.loc[(data_train['Fare']>15.0458)&(data_train['Fare']<=60.0000),'Fare']=2\ndata_train.loc[data_train['Fare']>60.0000,'Fare']=3\n\ndata_test.loc[data_test['Fare']<=8.0500,'Fare']=0\ndata_test.loc[(data_test['Fare']>8.0500)&(data_test['Fare']<=15.0458),'Fare']=1\ndata_test.loc[(data_test['Fare']>15.0458)&(data_test['Fare']<=60.0000),'Fare']=2\ndata_test.loc[data_test['Fare']>60.0000,'Fare']=3\n\nf,ax=plt.subplots(1,2,figsize=(12,5))\n\nsns.countplot('Fare',data=data_train,hue='Survived',ax=ax[0])\nax[0].set_title('Survived',color = 'r',fontsize=15)\n\nsns.barplot(x=data_train.groupby(['Fare'])['Survived'].mean().index,\n            y=data_train.groupby(['Fare'])['Survived'].mean().values,\n            ax=ax[1])\nax[1].set_title('Rate of the Survived by Fare',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a1fe37193abe4ce252eeb9f1f482418bee99f8e"},"cell_type":"markdown","source":"Finally I dropped the features that  is not needed for modeling. After that I visualised the correlation between the features."},{"metadata":{"trusted":true,"_uuid":"e6b307ec01036fb71b37910b3750687d46e07a3e"},"cell_type":"code","source":"data_train.drop(['Name','Ticket','PassengerId'],axis=1,inplace=True)\ndata_test.drop(['Name','Ticket','PassengerId'],axis=1,inplace=True)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f40a2883277e18aa1cd29879e4706a8eb0c46f9a"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(12, 12))\ncmap = sns.diverging_palette(220, 10, as_cmap=True)\nsns.heatmap(data_train.corr(), cmap=cmap, vmax=.3, center=0,square=True,annot=True, linewidths=.5, cbar_kws={\"shrink\": .5})\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c255d70a72d6e3aacfa425d6bf74d6ab48f7ada0"},"cell_type":"markdown","source":"For speaking survival, sex correlated the most. Pclass correlated with Cabin and Fare, Family size correlated with Sibsp and Parch size highly."},{"metadata":{"trusted":true,"_uuid":"8db98168b5870a2d0cb4276e70e37a01c65489d7"},"cell_type":"code","source":"sns.set_context(\"notebook\", font_scale=1.5, rc={\"lines.linewidth\": 2.5})\nsns.pairplot(data_train,diag_kind=\"kde\",hue=\"Survived\")\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"efe2174827dcf2032b79b325ddf500d70909c0ac"},"cell_type":"markdown","source":"<a id=\"5\"></a> \n**5. Modeling**\n\nIn this section I am going to apply each model to train data for find the best survival predict. Before this I spiltted the train data to %80 train and % 20 test size. "},{"metadata":{"trusted":true,"_uuid":"d51460ec84efd66a61aeb8d1baa9f2e9d14dd999"},"cell_type":"code","source":"y=data_train['Survived']\n\nx=data_train.drop(['Survived'],axis=1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"72ac4872a2a5a0a811c8331a2ecd03c4239320e3"},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nx_train, x_test, y_train, y_test = train_test_split(x,y,test_size=0.2,random_state=1)\n\nprint(\"x train: \",x_train.shape)\nprint(\"x test: \",x_test.shape)\nprint(\"y train: \",y_train.shape)\nprint(\"y test: \",y_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"b3f99ea412e3f1c3cf681790a7047a6c81aa6fcd"},"cell_type":"markdown","source":"The models which I am going to use:\n\n1) K-Nearest Neighbours\n\n2) Support Vector Machines\n\n3) Naive Bayes\n\n4) Decision Tree\n\n5) Random Forest\n\n6) Logistic Regression\n\nAlso Confusion matrix and Roc curve will be plotted for better understanding."},{"metadata":{"trusted":true,"_uuid":"41c29b83cdc1766ddb0cfc5354ea7bd60b5c20e6"},"cell_type":"code","source":"from sklearn import metrics #accuracy measure\nfrom sklearn.model_selection import KFold #for K-fold cross validation\nfrom sklearn.model_selection import cross_val_score #score evaluation\nfrom sklearn.model_selection import cross_val_predict #prediction\nfrom sklearn.model_selection import GridSearchCV # # Grid Search Cross Validation\nfrom sklearn.metrics import confusion_matrix #for confusion matrix\nfrom sklearn.metrics import roc_curve # ROC Curve with logistic regression","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c5e2627233b8e9688a953a5376055a28f8ae3e3d"},"cell_type":"code","source":"kfold = KFold(n_splits=10, random_state=22) # k=10, split the data into 10 equal parts\nmean=[] # List for mean of CV scores\naccuracy=[] # List for CV score\nstd=[] # List for CV std\n\n# Main function for models\ndef model(algorithm,x_train_,y_train_,x_test_,y_test_): \n    algorithm.fit(x_train_,y_train_)\n    predicts=algorithm.predict(x_test_)\n    prediction=pd.DataFrame(predicts)\n    prob=algorithm.predict_proba(x_test_)[:,1]\n    cross_val=cross_val_score(algorithm,x_train_,y_train_,cv=kfold)\n    \n    # Appending results to Lists \n    mean.append(cross_val.mean())\n    std.append(cross_val.std())\n    accuracy.append(cross_val)\n    \n    # Printing results  \n    print(('{}'.format(algorithm)).split(\"(\")[0].strip(),'\\n') \n    print(\"CV std :\",cross_val.std(),\"\\n\")\n    print(\"CV scores:\",cross_val,\"\\n\")\n    print(\"CV mean:\",cross_val.mean())\n    \n    # Plot for conf. matrix and roc curve\n    fpr, tpr, thresholds = roc_curve(y_test_, prob)\n    \n    f,ax=plt.subplots(1,2,figsize=(11,4))\n    \n    # Plot ROC curve\n    plt.plot([0, 1], [0, 1], 'k--')\n    plt.plot(fpr, tpr)\n    plt.xlabel('False Positive Rate')\n    plt.ylabel('True Positive Rate')\n    plt.title('ROC')\n    # Plot for Confusion Matrix\n    y_pred = cross_val_predict(algorithm,x,y,cv=10)\n    sns.heatmap(confusion_matrix(y,y_pred),ax=ax[0],annot=True,fmt='2.0f')\n    ax[0].set_title(('Confusion Matrix for {}'.format(algorithm)).split(\"(\")[0].strip())\n    \n    plt.subplots_adjust(wspace=0.3)\n    plt.close(0)\n    plt.show()\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"7c0b76c65e3d7a31f1ac580b9acf7a2f6a5a18fc"},"cell_type":"code","source":"# K-Nearest Neighbours\n\nfrom sklearn.neighbors import KNeighborsClassifier\n\ngrids = {'n_neighbors': np.arange(1,50)}\n\ngrid = GridSearchCV(estimator=KNeighborsClassifier(), param_grid=grids, cv=kfold) # Grid Search for best param.\ngrid.fit(x_train, y_train)\n\n# Print hyperparameter\nprint(\"Tuned hyperparameter k: {}\".format(grid.best_params_),'\\n') \nprint(\"Best score: {}\".format(grid.best_score_))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"69231b93c902f0f9e4a82e75c07cd3d01c02b35a"},"cell_type":"code","source":"knn = KNeighborsClassifier(n_neighbors = grid.best_estimator_.n_neighbors)\n\nmodel(knn,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"dbae60041fc45e9b43342ffc4c99f9d76eb059c1"},"cell_type":"code","source":"# Support Vector Machines\n\nfrom sklearn import svm \n\nCs = [0.001, 0.01, 0.1, 1, 10]\ngammas = [0.001, 0.01, 0.1, 1]\ngrids = {'C': Cs, 'gamma' : gammas}\n\ngrid = GridSearchCV(estimator=svm.SVC(kernel='linear'), param_grid=grids, cv=kfold) # Grid Search for best param.\ngrid.fit(x_train, y_train)\n\n# Print hyperparameter\nprint(\"Tuned hyperparameter k: {}\".format(grid.best_params_),'\\n') \nprint(\"Best score: {}\".format(grid.best_score_))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d83fb0d213fadb0f1e4605908fc4ab39d3e563b6"},"cell_type":"code","source":"svm = svm.SVC(kernel='linear',C=grid.best_estimator_.C,gamma=grid.best_estimator_.gamma,probability=True)\n\nmodel(svm,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"75d1fb057993e8c2d4a1f53d32eccc0b5ec6a584"},"cell_type":"code","source":"# Naive Bayes\n\nfrom sklearn.naive_bayes import GaussianNB \nnb = GaussianNB()\nmodel(nb,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d32c09e91a70172aab128dbe51e390a38a17d37a"},"cell_type":"code","source":"# Decision Tree\n\nfrom sklearn.tree import DecisionTreeClassifier \n\ngrids={'min_samples_split' : range(10,500,20),'max_depth': range(1,20,2)}\n\ngrid = GridSearchCV(estimator=DecisionTreeClassifier(), param_grid=grids, cv=kfold) # Grid Search for best param.\ngrid.fit(x_train, y_train)\n\n# Print hyperparameter\nprint(\"Tuned hyperparameter k: {}\".format(grid.best_params_),'\\n') \nprint(\"Best score: {}\".format(grid.best_score_))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"81d384fc951d13b7e6b5418cdfe35c387d5708a6"},"cell_type":"code","source":"dtc = DecisionTreeClassifier(min_samples_split=grid.best_estimator_.min_samples_split, max_depth=grid.best_estimator_.max_depth)\n\nmodel(dtc,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"12eb454c8646f2edcfa2994430de3df3b4f16b6f","scrolled":false},"cell_type":"code","source":"# Random Forest\n\nfrom sklearn.ensemble import RandomForestClassifier \n\ngrids={'n_estimators':range(100,500,100)}\n\ngrid = GridSearchCV(estimator=RandomForestClassifier(), param_grid=grids, cv=kfold) # Grid Search for best param.\ngrid.fit(x_train, y_train)\n\n# Print hyperparameter\nprint(\"Tuned hyperparameter k: {}\".format(grid.best_params_),'\\n') \nprint(\"Best score: {}\".format(grid.best_score_))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f7a07b8afbadd8060e1586892944379e1b3d34df"},"cell_type":"code","source":"rf = RandomForestClassifier(n_estimators=grid.best_estimator_.n_estimators)\n\nmodel(rf,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8bbf6daa26053a70b0f56cf58a2a5c9e37e3c14d"},"cell_type":"code","source":"# Logistic Regression\n\nfrom sklearn.linear_model import LogisticRegression \n\ngrids = {'C': np.logspace(-3, 3, 7), 'penalty': ['l1', 'l2']}\n\ngrid = GridSearchCV(estimator=LogisticRegression(), param_grid=grids, cv=kfold) # l1 lasso l2 ridge\ngrid.fit(x_train, y_train)\n\n# Print hyperparameter\nprint(\"Tuned hyperparameter k: {}\".format(grid.best_params_),'\\n') \nprint(\"Best score: {}\".format(grid.best_score_))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d405055cfe50c550f30970b096a78c791affc444"},"cell_type":"code","source":"lr = LogisticRegression(C=grid.best_estimator_.C,penalty=grid.best_estimator_.penalty)\n\nmodel(lr,x_train,y_train,x_test,y_test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3e446dd02e1a73faf29fc0a3b4cff7b6e3ac0eef"},"cell_type":"markdown","source":"Lets see the results."},{"metadata":{"trusted":true,"_uuid":"421ba73900c50c761d4566b1aaac50f1258b8358"},"cell_type":"code","source":"classifiers=['KNN','Svm','Naive Bayes','Decision Tree','Random Forest','Logistic Regression']\n\nmodels=pd.DataFrame({'CV mean':mean,'Std':std},index=classifiers)       \nprint(models)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bfcb1b4feceb7907fb5ad5f72f127f507a781fd5"},"cell_type":"code","source":"f, ax = plt.subplots(figsize=(16, 7))\n\nsns.boxplot(x=models.index, y=accuracy)\nplt.xticks(rotation=45)\nplt.title('Models',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"037d91363666dee7bb9db6c84ae2aa0281fe7d5b"},"cell_type":"markdown","source":"Decision Tree has the best result. Lets see what coefficients effects how much."},{"metadata":{"trusted":true,"_uuid":"13ae2a3c5998d3c0829715c02f0747b55f3781f3"},"cell_type":"code","source":"coefficients=pd.DataFrame({'Features':data_test.columns,'Coefficients':dtc.feature_importances_})       \n\nplt.figure(figsize=(15,5))\nsns.barplot(x=coefficients['Features'],y=coefficients['Coefficients'])\nplt.xticks(rotation=45)\nplt.title('Titles',color = 'r',fontsize=15)\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"668e9443029412e1aca0b5827189f4159e0f2664"},"cell_type":"markdown","source":"Finally it comes to creating a submission file. I decided to use Decision Tree model for the submission."},{"metadata":{"trusted":true,"_uuid":"e982b145d30bcf048e055ec92dce2c1be322e4dc"},"cell_type":"code","source":"submission = pd.DataFrame({\"PassengerId\": pd.read_csv(\"../input/test.csv\")[\"PassengerId\"],\"Survived\": dtc.predict(data_test)})\n\nsubmission.to_csv('titanic.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ba77c42245a1f1eba70e80835a30a2fc931aba91"},"cell_type":"markdown","source":"<a id=\"6\"></a> \n**6. Conclusion**\n\nAs I wrote before this dataset is my first competition dataest. I enjoyed very much working with and also learned much studying how to do it. If you come this far, I thank you very much.\n\nIf you like it, thank you for you upvotes.\nIf you have any question, I will happy to hear it\n\nAlso look for https://www.kaggle.com/kanncaa1/machine-learning-tutorial-for-beginners for Machine Learning"}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}