{"cells":[{"metadata":{"_uuid":"727588938d2c9601ed9746247eefd85e3526966b"},"cell_type":"markdown","source":"# Your first ML Project\nThis notebook has been made to help you to go trough, in a really simple way, a machine learning project. In a few minutes you will be able to run your first ML algorithm :-) ."},{"metadata":{"trusted":true,"_uuid":"4833e686c942dbc91932ca7e64bcdf59b16f5c98"},"cell_type":"code","source":"from IPython.display import Image\ndisplay(Image('../input/machine-learning-everywhere-memes/machine-learning-everywhere.jpg', width=500, unconfined=True))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bb7e2107f010ce015d84a17ea2a5da27c1f6e6bf"},"cell_type":"markdown","source":"### Import all libraries required\n* pandas is used for data manipulation (including reading data)\n* numpy is used for mathematacial functions\n* matplotlib and seaborn are for Data Visualisation"},{"metadata":{"_uuid":"51c32aa0cdbd042ad1a3db6728e5e43902892fb0","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d4366aa8680ece31e5a9cbc46c52234ebb535c05"},"cell_type":"markdown","source":"### Read data\nPandas allow us to read data directly from the file. Here we have a csv file :"},{"metadata":{"_uuid":"fc762dd1bab237f1bd82f3ba9672a3c8125a8512","trusted":false},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bc11793777c36d0e8185fbace24c67e6f5bb7144"},"cell_type":"markdown","source":"## Do you understand what we have to do ?\nHere we have to determine if a passenger will survive or not the Titanic tragedy. As you can notice below, the column Survived can only be two numbers : 1 = Survived , 0 = Died. So it seems to be a classification problem. We will have to keep in mind this information to adapt our exploration and modeling strategy.\nIn this case, it is a supervised problem, as we have a column of the label that we want to predict."},{"metadata":{"_uuid":"e47b0766f6cc8d94484938a2388e2cbf7d31219c","trusted":false},"cell_type":"code","source":"train.Survived.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2b525c3892e17bff1bb135cdfdffe111faf063b5"},"cell_type":"markdown","source":"Have a preview of the training dataframe. The training dataframe is made to be able to learn from it, as a development environement."},{"metadata":{"_uuid":"47dac9592fa8f212baabd98ad1771433ab8bd700","trusted":false},"cell_type":"code","source":"train.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"be8853c7114567258e3895a6a8af0455a4fd2c83"},"cell_type":"markdown","source":"Have the preview of the testing dataframe. The testing dataframe is the dataframe that we will have to use to make our prediction in the production environement."},{"metadata":{"_uuid":"ef341af902ed814761273de3142b9c571cb4f05d","trusted":false},"cell_type":"code","source":"test.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f9fa9255dfe15f6a1dfb67587e1c54abdcc4fd8"},"cell_type":"markdown","source":"You can notice that the column \"Survived' is not present in the test dataframe. It's because it is our \"label\"/\"target\" column that we will need to predict. :-)"},{"metadata":{"_uuid":"14dfaec80b5920accf4cdde4159f5a9ad69fb915"},"cell_type":"markdown","source":"## STEP 1) Data Exploration\nThe first step of any ML project is to understand your data. Let's do it ! :-)"},{"metadata":{"_uuid":"2c7c4dbae7ad967a66ee585a116d65597861ee14","trusted":false},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3c8476137806dd6d5c67f8234e7ed1a49b39acf3"},"cell_type":"markdown","source":"Knowing the type of your data is important. Some algorithms heavily depend on the type and you may obtain an Error if you use an unadapted type.\nHere you can notice that PassengerId, Survived, Pclass, Age, SibSp, Parch, Fare are numerical columns. And Name, Sex, Ticket,Cabin, Embarked are object columns (in this case, that means strings)."},{"metadata":{"_uuid":"de9d95cda2a1b3b63cb6d8d7aac3b0ac7520a7ef"},"cell_type":"markdown","source":"### Trying to find why a person would survive, or not\nDo you remember that it is a classification problem ? We will try to find what is different between the two categories of person using the training dataframe."},{"metadata":{"_uuid":"3b6ca3ab7a869629da71ef305028a9549f102600"},"cell_type":"markdown","source":"## Fill empty values\nNULL values ? Empty values?\nOne of the first thing to do is to check for empty cells, that can be a common source of error when your run an algorithm !"},{"metadata":{"_uuid":"d0b3d138445b0716108c43f08fb27a21563ef252","trusted":false},"cell_type":"code","source":"train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"14b53524abacd48a85cbf277cab20dfc5bbda945"},"cell_type":"markdown","source":"Fortunately we made this check ! We got 177 empty values for Age, 687 for Cabin and 2 for Embarked ! You will notice along your data science journey that every dataset has to be cleaned before use ;) let's deal with those issues."},{"metadata":{"_uuid":"361725a1c0c04202d3082e24138671303d98b801"},"cell_type":"markdown","source":"#### Fill Age\n\nFill the Age with the mean value to make it not empty. We may do a more sophisticated filling, as for example taking the mean for each Pclass and fill the value based on this feature... you can try to do it !"},{"metadata":{"_uuid":"949fc10a0b4f8f827dce4e935c17b5de592b62cd","trusted":false},"cell_type":"code","source":"train.loc[train.Age.isna(), 'Age'] = train[~train.Age.isna()].Age.mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f7f40d146a807081631aad135831307568e79f2"},"cell_type":"markdown","source":"#### Fill Cabin\n\nCabin being empty... Perhaps they do not have a Cabin number ? Let make this assumption !\n"},{"metadata":{"_uuid":"a4f992fa1f549c72c56dfbdefb0e7765f476da9d","trusted":false},"cell_type":"code","source":"train.loc[train.Cabin.isna(),'Cabin'] = \"No Cabin\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7726cda2a517b8af6efe71c99f4d27a905d93bc9"},"cell_type":"markdown","source":"The Embarked feature has only two values empty, it may not really be smart to fill those value with \"empty\" values that will create a new class and make your future model more complex. \n\nLet's use the more representative port here, which is S (Southampton)."},{"metadata":{"_uuid":"628676513ea27d76085c0544819746b7eb1fc650","trusted":false},"cell_type":"code","source":"print(train.Embarked.value_counts())\ntrain.loc[train.Embarked.isna(),'Embarked'] = \"S\"","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f4c39a51bf31ea583a17e095a618bde29ef3ec30"},"cell_type":"markdown","source":"Let's check if all nan value have been filled. Indeed, everything seems okay now :"},{"metadata":{"_uuid":"74ac04554aa751d7c309975ca92125379ebab370","trusted":false},"cell_type":"code","source":"train.isna().sum()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e26421409da8f7f2b13da8f74f2f82c4090e3752"},"cell_type":"markdown","source":"## Outliers\nWe now need to check for outliers. Outliers are values that are very different compared to what we could expect from all other values.\nHere we will try to verify some assumptions. The first one we could make is that older people have more money (and thus a better Class).\n\nWe will use a boxplot. What is a boxplot ? Check this link : https://towardsdatascience.com/understanding-boxplots-5e2df7bcbd51"},{"metadata":{"_uuid":"79050ac7125cb2661c28bd8b7c688e598b94bbf9","trusted":false},"cell_type":"code","source":"fig,axes = plt.subplots(1, 2,figsize=(25,8))\nprint(axes)\n\nsns.boxplot(x='Pclass',y='Age',data=train, palette='viridis',ax=axes[0])\n\n# We now need to check for outliers (values that seem irregular compared to the others)\nsns.boxplot(x='Pclass',y='Fare',data=train, palette='viridis',ax=axes[1])\n\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"42989cf7d81a1e58e546a818339d2aa073565131"},"cell_type":"markdown","source":"It seems that only a few tickets have been really expensive compared to the others ! Let's take a closer look at those."},{"metadata":{"_uuid":"a656d5342d2f812c95c6ad6bdd0e814410470405","trusted":false},"cell_type":"code","source":"train.loc[train.Fare > 200]","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"3f5f4e1ee028cea769ab3da0afe6770beb4cf2c5"},"cell_type":"markdown","source":"As we can notice : some have several Cabins, but some not, and they get only one but really expensive Cabin.\n\nLine 258 : Ward Anna seems to have paid 512$ ! As we know in the Titanic some Cabin can be really expensive compared to others - perhaps it was the best Cabin ever made !\n\nLine 679 & Line 737 : We can also notice that Mr Thomas Drake Martinez and Mr Gustance had the same Fare, with the same number Ticket, perhaps they are traveling together ?\n\nYou can try to change those outliers Fares, and manage to do some nice modifications ! The idea is to have an intuition about the veracity of the values, and implement it."},{"metadata":{"_uuid":"eeb33a67aa33b7dd47c57121f98d7d48eea544d3"},"cell_type":"markdown","source":"## Correlation matrix\nA common way to find relationships between variables is to plot a correlation matrix. \nHere you can already see that the Survived feature is higly dependent on the Ticket Class (Pclass), interesting, isn't it ?"},{"metadata":{"_uuid":"07f3ee9194f4ebc9f703a00871026774a9c087c4","trusted":false},"cell_type":"code","source":"\nnumerical_column = ['int64','float64'] #select only numerical features to find correlation\nplt.figure(figsize=(10,10))\nsns.heatmap(\n    train.select_dtypes(include=numerical_column).corr(),\n    cmap=plt.cm.RdBu,\n    vmax=1.0,\n    linewidths=0.1,\n    linecolor='white',\n    square=True,\n    annot=True\n)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e9c5dd765504e1f203695c2d491216ae6e9d089c"},"cell_type":"markdown","source":"## Pairplot / scatter matrix\nAnother comomon thing to do in a Classification is to plot a pair plot, which is a figure that allows you to see the distribution of each data compared to others, in different colors regarding your label column (here Survived). You can already notice the different distribution of Pclass.\n\nIt seems that the cheapeast class has a lowest chance of survival... It seems also that people with no parent/child aboard the titanic (regarding Parch) has the highest change to survive... hum... let's keep this in mind this for later."},{"metadata":{"_uuid":"717dcbf770f7bd75efb4a83e4c4c73c3330c9796","trusted":false},"cell_type":"code","source":"plt.figure(figsize=(10,10))\nsns.pairplot(train.select_dtypes(include=numerical_column), hue = 'Survived')","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1b22cd264867b937369be57bd0ba4430cc32baf"},"cell_type":"markdown","source":"## Few other checks\nWe can notice, regarding the Age distrubition that children have been (most of the time) saved !\n'Save women and children in first' seems at last to really be true... but wait ! We don't have Sex ditrbution here\nas it's an object column, and pairplot do not support \"object\" columns. Let's change that."},{"metadata":{"_uuid":"061587b109d96e686bed13edbc4246e10908bea3","trusted":false},"cell_type":"code","source":"train.Sex.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f1c9b96297d60605ebb8c30a15da44f8f6045c31","trusted":false},"cell_type":"code","source":"# thanks to the open-source world, we do not need to waste that much time here ! LabelEncore from sckit-learn allow us\n# to convert this text categorical data into numbers ! :O\n\nfrom sklearn.preprocessing import LabelEncoder\nlabelencoder = LabelEncoder()\ntrain['Sex'] = labelencoder.fit_transform(train['Sex'])\ntrain.Sex.value_counts()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"03527dd5d9b6e69d87d247f042af9360daec0b34","trusted":false},"cell_type":"code","source":"#Let's now print the Age distrubtion regarding who survived\n\npalette ={1:\"g\", 0:\"r\"}\nsns.countplot(x='Sex',data=train,hue=\"Survived\", palette=palette)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6328d0022a2ff18611cfcdea5291a7f8754cb57c"},"cell_type":"markdown","source":"Indeed, our assumption was correct. We can notice that indeed women have a better chance of surviving during this tragedy."},{"metadata":{"_uuid":"4d81be923ea628babb2bca32bb2d2548e0b2e39d"},"cell_type":"markdown","source":"## STEP 2) Features Engineering\nThis step is made to be able to get a maximum of information from the data, by, for example, creating new columns or changing something in a column as we just already did with the \"Sex\" column.\nI recommend to create a features_engineering function to have a common modification between your \"training\" and \"testing\", through just one function :-).\n\nI will be really simple here to give you the opportunity to create and add some new features and test it in your model. Do not forget to add here all features modification you might already have done during the previous step.\n### !! Be sure to never include your target/label column in a features engineering process or an error will occur during the feature engineering process of your testing dataframe !!"},{"metadata":{"_uuid":"b16d4ec629593b5fe12b6993420cac21c830eeb1","trusted":false},"cell_type":"code","source":"def features_engineering(df):\n    df.loc[df.Age.isna(), 'Age'] = df[~df.Age.isna()].Age.mean()\n    df.loc[df.Cabin.isna(),'Cabin'] = \"No Cabin\"\n    df.loc[df.Embarked.isna(),'Embarked'] = \"S\"\n    df['persons_abroad_size'] = (df['Parch']+df['SibSp']).astype(int)\n    df['alone'] = np.where(df['Parch']==0,1,0)\n    df['Embarked'] = df['Embarked'].map( {'S': 0, 'C': 1, 'Q': 2} ).astype(int)\n    df['Sex'] = df['Sex'].map( {'male': 1, 'female': 2} ).astype(int)\n    df['log_fare'] = df['Fare'].apply(np.log)\n    df['Room'] = (df['Cabin']\n                    .str.slice(1,5).str.extract('([0-9]+)', expand=False)\n                    .fillna(0)\n                    .astype(int))\n    df['RoomBand'] = 0\n    df.loc[(df.Room > 0) & (df.Room <= 20), 'RoomBand'] = 1\n    df.loc[(df.Room > 20) & (df.Room <= 40), 'RoomBand'] = 2\n    df.loc[(df.Room > 40) & (df.Room <= 80), 'RoomBand'] = 3\n    df.loc[df.Room > 80, 'RoomBand'] = 4\n    df_id = df.PassengerId\n    df = df.drop('PassengerId', axis=1)\n    return df,df_id","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4766da7075d2bc8f6cb85855452cbe6ba0954e80","trusted":false},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv')\ntest = pd.read_csv('../input/test.csv')\ntrain,train_id = features_engineering(train)\ntest,test_id = features_engineering(test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"528a3539d175e2a38443ae5e13d4ae7eb7b7de93","trusted":false},"cell_type":"code","source":"train.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"a258a42e891e2914bca6d1ef81e2a9ae92960b59"},"cell_type":"markdown","source":"## STEP 3) Model and save your prediction\nHere we will construct our model. We will use all numerical columns and the XGboost algorithm. XGB is one of the most famous algorithm you may see one the Kaggle platform by its capacity to automatize a large number of process."},{"metadata":{"_uuid":"12aef7c07f7a5700a482ce468fdef997f1dc9d31","trusted":false},"cell_type":"code","source":"import xgboost as xgb\nfrom sklearn import model_selection\nX_train = train.drop('Survived',axis=1).select_dtypes(include=['int32','int64','float64'])\ny_train = train['Survived']\nX_test = test.select_dtypes(include=['int32','int64','float64'])\n\nxg_boost = xgb.XGBClassifier(base_score=0.5, booster='gbtree', colsample_bylevel=1,\n       colsample_bytree=0.65, gamma=2, learning_rate=0.3, max_delta_step=1,\n       max_depth=4, min_child_weight=2, missing=None, n_estimators=280,\n       n_jobs=1, nthread=None, objective='binary:logistic', random_state=0,\n       reg_alpha=0, reg_lambda=1, scale_pos_weight=1, seed=None,\n       silent=True, subsample=1)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"efa3a24038836a0144d1c6ec6c6e01026d8131eb"},"cell_type":"markdown","source":"Train your model with your trained dataset :"},{"metadata":{"_uuid":"80f364110c60c6bfe2dceb4725121ada9bcf3b66","trusted":false},"cell_type":"code","source":"xg_boost.fit(X_train, y_train)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"74fd6240c1e17968890d6992398a15cb2942ba23"},"cell_type":"markdown","source":"Here we can do an evaluation of our model :"},{"metadata":{"_uuid":"60bb0559a2a78d710509ae91d4c2b8dc25a53e25","trusted":false},"cell_type":"code","source":"print(xg_boost.score(X_train, y_train))\n\nscores = model_selection.cross_val_score(xg_boost, X_train, y_train, cv=5, scoring='accuracy')\nprint(scores)\nprint(\"Kfold on XGBClassifier: %0.4f (+/- %0.4f)\" % (scores.mean(), scores.std()))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"54d7f63d7fb38ee7f474e04a62584f5d650a8588","trusted":false},"cell_type":"code","source":"xgb.plot_importance(xg_boost)\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8ba9458f73bff459e9c98fde66b269b3977bc99c"},"cell_type":"markdown","source":"82% of accuracy during a cross validation is a correct score for the first shot in a binary classification. Try to improve this ! :) \nNow let's predict our testing value."},{"metadata":{"_uuid":"1ccc1dd8f7c98243a2fdf25c61caefda5717f6e9","trusted":false},"cell_type":"code","source":"Y_pred = xg_boost.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"19215db8e7370d128e8e13c6bc2da37bb6549a03","trusted":false},"cell_type":"code","source":"submission = pd.DataFrame({\n    \"PassengerId\": test_id, \n    \"Survived\": Y_pred \n})\nsubmission.head(10)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ca30e522ec9cf6d4b957e4d07cbfad2d55cad59"},"cell_type":"markdown","source":"Your file with the prediction with the testing dataframe is ready to be saved !  Congratulations you just made your first machine learning project !"},{"metadata":{"_uuid":"8644b1f977fb065e6587440df7b32a3126eecf4b","trusted":false},"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}