{"cells":[{"metadata":{"_uuid":"bab13e4f8aee192d2c8ae4789670f6deeeeea407"},"cell_type":"markdown","source":"<h1 style=\"font-family:courier;color:#04abed;font-size:40px\">Hello Data Science, Hello Kaggle _</h1>\n<h3 style=\"font-family:courier;\">Today we're gonna discover Data Science World, so take a deep breath since we're gonna go deeper than the Titanic ... </h3>"},{"metadata":{"_uuid":"3c3e76ebfa738b36b6b4ddfaa194aec1676f2be5"},"cell_type":"markdown","source":"<h2 style=\"font-size:20px;color:red;\">Let’s Understand Why We Need Data Science ?</h2>"},{"metadata":{"_uuid":"8c66db5b8328d6edb5bbbb3ead1c5b3845d71417"},"cell_type":"markdown","source":"<img src=\"https://d1jnx9ba8s6j9r.cloudfront.net/blog/wp-content/uploads/2017/01/Flow-of-unstructured-data.png\" >\nTraditionally, the data that we had was mostly structured and small in size, which could be analyzed by using the simple BI tools. Unlike data in the traditional systems which was mostly structured, today most of the data is unstructured or semi-structured. Let’s have a look at the data trends in the image given above which shows that by 2020, more than 80 % of the data will be unstructured.\nThis data is generated from different sources like financial logs, text files, multimedia forms, sensors, and instruments. Simple BI tools are not capable of processing this huge volume and variety of data. This is why we need more complex and advanced analytical tools and algorithms for processing, analyzing and drawing meaningful insights out of it."},{"metadata":{"_uuid":"252f0eeb2ed5e0ef9d45897c3f7da42df27693d2"},"cell_type":"markdown","source":"<h1 style=\"font-family:courier;color:#04abed;font-size:25px;\">What is Data Science ?</h1>\nData science is an interdisciplinary field that uses scientific methods, processes, algorithms and systems to extract knowledge and insights from data in various forms, both structured and unstructured, similar to data mining. ~Wikipedia\n\nData Science is a field where we apply ‘science’ to available ‘data’ in order to get the ‘patterns’ or ‘insights’ which can help a business to optimize operations or improvise decisions.\n\n<img src=\"https://cdn-images-1.medium.com/max/800/1*ewxqYVXny5jyDQlo-xMuLA.png\"/>"},{"metadata":{"_uuid":"4472f3076a5c6bf3e00af833a0ead6e9083cdd47"},"cell_type":"markdown","source":"<h1 style=\"font-family:courier;color:#04abed;font-size:25px;\">Why Data Science is important?</h1>\nEvery phenomenon has a reason behind its occurrence. So has data science. Therefore it would be interesting to know the emerging trends that give data science the utmost importance to keep pace with the changing scenario:\n\n<img src=\"https://cdn.intellipaat.com/blog/wp-content/uploads/2016/11/Why-do-we-need-Data-Science.jpg\">\nEvery business has data but its business value depends on how much they know about the data they have.\nData Science has gained importance in recent times because it can help businesses to increase business value of its available data which in turn can help them to take competitive advantage against their competitors.\nIt can help us to know our customers better, it can help us to optimize our processes, it can help us to take better decisions. Because of data science, data has become strategic asset.\nIn the following chart, you can have a look at the business use cases where data science is being used in the industry.\n\n<img src=\"https://d1jnx9ba8s6j9r.cloudfront.net/blog/wp-content/uploads/2017/01/Data-Science-use-cases.png\">"},{"metadata":{"_uuid":"9896631605b471415c243aaffa35d9ed5fb05245"},"cell_type":"markdown","source":"<h1 style=\"font-family:courier;color:#04abed;font-size:25px;\">How to do Data Science?</h1>\nA typical data science process looks like this, which can be modified for specific use case:"},{"metadata":{"_uuid":"7604061748ec701b33b252a8a643d256e7e2e353"},"cell_type":"markdown","source":"### Understand the business or the Problem\nLearn about the issue at ground, ask the right questions which is at the center of what a Data Scientist does and forms the foundation for the later stages of the Data Scientist’s role. Define the problem and convert it into a concrete framework which can then be worked upon.\n### Collect & explore the data\nAs the name implies the Data Scientist has to collect enough data in order to make sense of the problem at hand and get a better grip of the issue with respect to the time, money and resources needed to make the process successful.\n### Prepare & process the data\nData can rarely be used in its original form. It needs to be processed and various methods exist to convert it into a usable format. This is an essential part of every Data Scientist’s job routine and this consumes a major chunk of his time and resources.\n### Explore the Data\nAfter the data has been processed and converted into a form that can then be used for the later stages, you need to explore it further so as to get the characteristics of the data and find out more about the obvious trends, correlation and the not so obvious hidden relationships and more.\n### Build & validate the models (Analyze the Data)\nThis is where the magic happens. The data scientist deploys the various arsenals in his repository like machine learning, statistics and probability, linear and logistic regression, time-series analysis and more in order to make sense of the data. At the end of this step the Data Scientist would be able to gain valuable business insights like predictions, business process optimization, finding new ways of doing the same old things among other things.\n### Communicate the Results then Deploy & monitor the performance\nAt the end of the entire process there is a need to communicate the findings to the right stake-holders in order to get the groundwork done for the action to be taken and deployment of the decisions that are taken."},{"metadata":{"_uuid":"bbb128ecf0e34b1e0a26b6d4915557aeffcc49c0"},"cell_type":"markdown","source":"<img src=\"https://cdn-images-1.medium.com/max/800/0*tkonfeb6Z9ubAllz.png\">"},{"metadata":{"_uuid":"3957f2873dbb9c4e51e35d623097df857ff83f70"},"cell_type":"markdown","source":"<h1 style=\"font-family:courier;color:#04abed;font-size:30px;\">Case study: Titanic Disaster</h1>\n<img src=\"http://image.noelshack.com/fichiers/2019/05/1/1548690862-0.jpg\">"},{"metadata":{"_uuid":"72b8fa1f54736de052a8619a28dc6a75b52a782e"},"cell_type":"markdown","source":"## loading packages and libraries"},{"metadata":{"trusted":true,"_uuid":"47f06ce612261fe3796942e4c86f341767ad17b0"},"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nimport os\nprint(os.listdir('../input'))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"c088818c9e422e7a362a39960a2b918934293c60"},"cell_type":"markdown","source":"## Exploratory Data Analysis"},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"train_data= pd.read_csv('../input/train.csv')\nprint(train_data.columns)\ntrain_data.describe()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6fc788f89e3e055fc6cb010e177d50c5ee80e339"},"cell_type":"markdown","source":"<h1 style=\"color:red;font-size:medium\">where is Name, Sex, .... !!!!!!</h1>"},{"metadata":{"trusted":true,"_uuid":"ba355e96055560587cc2f9e18fb9a2b3e7012a50"},"cell_type":"code","source":"train_data.info()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"f3351d244cac7e59de68904cd44fd6e11875b21e"},"cell_type":"markdown","source":"## The relation between the age and the survival odds"},{"metadata":{"trusted":true,"_uuid":"8f179b02ab55f763db0add83bc0514066b2cdb96"},"cell_type":"code","source":"train_data['Died'] = 1-train_data['Survived']\ntrain_data.groupby('Sex').agg('sum')[['Survived','Died']].plot(kind='bar', stacked=True, figsize=(10,6));","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"1941d35d2c0e7b61904437606032c8359087ed3e"},"cell_type":"markdown","source":"Chart above says that more male passengers are died compared to females"},{"metadata":{"_uuid":"4155b70e78ee386fa248011a22dfe6ff7e63d067"},"cell_type":"markdown","source":"## Getting more insights"},{"metadata":{"trusted":true,"_uuid":"5d9cef2a00d9d2d0beeeea8f1fdc69231c31fd30"},"cell_type":"code","source":"fig, axis = plt.subplots(1,2,figsize=(15,8))\nsns.barplot(x=\"Embarked\", y=\"Survived\", hue=\"Sex\", ax=axis[(0)], data=train_data);\nsns.barplot(x=\"Pclass\", y=\"Survived\", hue=\"Sex\", ax=axis[(1)], data=train_data);\nplt.figure(figsize=(10,5))\nsns.barplot(x=\"Parch\", y=\"Survived\", hue=\"Sex\", data=train_data);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"c69c589be1f54b5e7ace0bef21b6266b0853739d"},"cell_type":"code","source":"plt.figure(figsize=(10,6))\nsns.violinplot(x='Sex',y='Age',hue='Survived', data=train_data, split=True);","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"55283c2d66dc4e46e7df357960228cde8d3ad720"},"cell_type":"markdown","source":"Above, we can see men who are between 20 to 40 are survived more compared to older aged men. But for the women odds don't depend on their age."},{"metadata":{"trusted":true,"_uuid":"b12ed00905f3cfae6dac40d5ae9bd02562f1e964"},"cell_type":"code","source":"plt.figure(figsize=(15,10))\nplt.hist([train_data[train_data['Survived'] == 1]['Fare'], train_data[train_data['Died'] == 1]['Fare']],\n        stacked=True, color=['g','r'], bins=70, label = ['Survived','Died'])\nplt.xlabel('Fare')\nplt.ylabel('Number of passnegers')\nplt.legend()\nplt.grid()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"728f549a1dbc3eb5b57ab85aa456facbe8907bce"},"cell_type":"markdown","source":"we can see that the passengers with cheaper ticket fares are more likely to die"},{"metadata":{"trusted":true,"_uuid":"d1da84a7ea57d4ab9cc2cd8c935d759e38609229"},"cell_type":"code","source":"plt.figure(figsize=(25,10))\nax=plt.subplot()\n\nax.scatter(train_data[train_data['Survived'] == 1]['Age'], train_data[train_data['Survived'] == 1]['Fare'],\n          c='green', s=train_data[train_data['Survived'] == 1]['Fare'])\nax.scatter(train_data[train_data['Died'] == 1]['Age'], train_data[train_data['Died'] == 1]['Fare'],\n          c='red', s=train_data[train_data['Died'] == 1]['Fare']);\nplt.xlabel('Age')\nplt.ylabel('Fare');","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"dec74d12178b3f5e76016362610dd72a18f1f3c5"},"cell_type":"markdown","source":"Small green dots between x=0 & x=10 : Children who were survived.\n\nSmall red dots between x=10 & x=45: Adults who died (from a lower classes).\n\nLarge green dots between x=20 & x=45 : Adults with larger ticket fares who are survived."},{"metadata":{"trusted":true,"_uuid":"723ba3230d97734530bc502079a1a2fa3e70f1e5"},"cell_type":"code","source":"ax=plt.subplot()\nax.set_ylabel('Average fare')\ntrain_data.groupby('Pclass').mean()['Fare'].plot(kind='bar', ax=ax, figsize=(10,6) );\n#the line above is the same as :\n#train_data.groupby('Pclass').agg({'Fare':'mean'}).plot(kind='bar', ax=ax)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6b4f26782e9a7b363967c0b3e18430929bdf511d"},"cell_type":"markdown","source":"## Feature Engineering"},{"metadata":{"_uuid":"fcb014edcfccd1616108d45cae642095fe0b0516"},"cell_type":"markdown","source":"<h1 style=\"color:red;font-size:20px;font-family:courier;\">First rule !!!! _COMBINE THE TRAIN AND THE TEST SET TO APPLY THE SAME CHANGES ON BOTH.</h1>"},{"metadata":{"trusted":true,"_uuid":"98fb1119d28ef20cdd84cc86ce4a80e1bd21f6c9"},"cell_type":"code","source":"X_train = train_data.drop(['Survived','Died'],axis=1)\ny_train = train_data['Survived']\nX_test = pd.read_csv('../input/test.csv')\ndf_combined = X_train.append(X_test)\nprint(X_train.shape[1])\nprint(df_combined.shape[1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"4e95f315796dd3df9b66e0baa7948f279bf57986"},"cell_type":"markdown","source":"Processing Family column :"},{"metadata":{"trusted":true,"_uuid":"ff6deb89781b302e46e7f7f49eef4bd0759d1a16"},"cell_type":"code","source":"def process_family(df):\n    # introducing a new feature : the size of families (including the passenger)\n    df['FamilySize'] = df['Parch'] + df['SibSp'] + 1\n    \n    # introducing other features based on the family size\n    df['Singleton'] = df['FamilySize'].map(lambda s: 1 if s == 1 else 0)\n    df['SmallFamily'] = df['FamilySize'].map(lambda s: 1 if 2 <= s <= 4 else 0)\n    df['LargeFamily'] = df['FamilySize'].map(lambda s: 1 if 5 <= s else 0)    \n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"6c6579484e23617d1f6bdf70fb4ac2099bcce0c4"},"cell_type":"code","source":"df_combined = process_family(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d3cd216f042ca3c10650fd3878b69b0e8a5d4ca"},"cell_type":"markdown","source":"Processing Embarked :"},{"metadata":{"trusted":true,"_uuid":"b06eef1e5cd2b66b89a3591e23b235494da27c72"},"cell_type":"code","source":"print(df_combined.Embarked.describe())\ndf_combined.loc[df_combined.Embarked.isna()]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e6ee57a4655be4ade523736f94440c2952a16107"},"cell_type":"code","source":"def process_embarked(df):\n    # two missing embarked values - filling them with the most frequent one in the train  set(S)\n    df.Embarked.fillna('S', inplace=True)\n    # dummy encoding \n    df_dummies = pd.get_dummies(df['Embarked'], prefix='Embarked')\n    df = pd.concat([df, df_dummies], axis=1)\n    df.drop('Embarked', axis=1, inplace=True)\n#     status('embarked')\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a722f4c488ef34f644c3364064603dd6be5e726f"},"cell_type":"code","source":"df_combined = process_embarked(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ce34e459adcca7189f576c3b9376c4567985fb22"},"cell_type":"markdown","source":"Processing Cabin :"},{"metadata":{"trusted":true,"_uuid":"a1486386bfd96c16be4ddcfbb2a8fc42b43e01ca"},"cell_type":"code","source":"def process_cabin(df):\n    # replacing missing cabins with U (for Uknown)\n    df.Cabin.fillna('U', inplace=True)\n    \n    # mapping each Cabin value with the cabin letter\n    df['Cabin'] = df['Cabin'].apply(lambda x: x[0])\n    \n    # dummy encoding ...\n    cabin_dummies = pd.get_dummies(df['Cabin'], prefix='Cabin')    \n    df = pd.concat([df, cabin_dummies], axis=1)\n\n    df.drop('Cabin', axis=1, inplace=True)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"253a3b8e0ec039e18f7d26824e827fd9df1c8be9"},"cell_type":"code","source":"df_combined = process_cabin(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b720de420d288fc5e0c13bccc4d7c2ee1380560b"},"cell_type":"code","source":"titles = set()\nfor name in df_combined['Name']:\n    titles.add(name.split(',')[1].split('.')[0].strip())\ntitles","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fc6b251c6927e6a3342ee77e1c5300d651659815"},"cell_type":"code","source":"Title_Dictionary = {\n     'Capt':'Officier',\n     'Col':'Officier',\n     'Don':'Royalty',\n     'Dona':'Royalty',\n     'Dr':'Officier',\n     'Jonkheer':'Royalty',\n     'Lady':'Royalty',\n     'Major':'Officier',\n     'Master':'Master',\n     'Miss':'Miss',\n     'Mlle':'Miss',\n     'Mme':'Mrs',\n     'Mr':'Mr',\n     'Mrs':'Mrs',\n     'Ms':'Mrs',\n     'Rev':'Officier',\n     'Sir':'Royalty',\n     'the Countess':'Royalty'   \n}\ndef passenger_title(df):\n    df['Title'] = df['Name'].apply(lambda x:x.split(',')[1].split('.')[0].strip())\n    df['Title'] = df['Title'].apply( lambda x : Title_Dictionary[x])\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"444c692a1a971510737432261a3e1f2a5ab5f567"},"cell_type":"code","source":"df_combined = passenger_title(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2cea6a45afca49bfabdc831296a65181a3e8401f"},"cell_type":"code","source":"grouped_train = df_combined.groupby(['Sex','Pclass','Title'])\ngrouped_median_train = grouped_train.median()\ngrouped_median_train = grouped_median_train.reset_index()[['Sex', 'Pclass', 'Title', 'Age']]\ngrouped_median_train.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1e57afcbd51469ac50cf61620e31b6752df2fe93"},"cell_type":"code","source":"df_combined.groupby(['Sex','Pclass','Title']).agg({'Age':'median'}).reset_index().head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0bf145cbd83e5530b0f107d3a1eca4d6d5a8052a"},"cell_type":"markdown","source":"Adding the value of age for missing values based on the group"},{"metadata":{"trusted":true,"_uuid":"463c19625cfff72dbda9bc1d44d4cbe6606ccb25"},"cell_type":"code","source":"def fill_age(row):\n    condition = (\n        (grouped_median_train['Sex'] == row['Sex']) & \n        (grouped_median_train['Title'] == row['Title']) & \n        (grouped_median_train['Pclass'] == row['Pclass'])\n    ) \n    return grouped_median_train[condition]['Age'].values[0]\n\ndef process_age(df):\n    # a function that fills the missing values of the Age variable\n    df['Age'] = df.apply(lambda row: fill_age(row) if np.isnan(row['Age']) else row['Age'], axis=1)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d8deef58511bd1e85f78ca2b64f0fe27faa8e49f"},"cell_type":"code","source":"df_combined = process_age(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"2d4027bff750653a102fd2911fd8276937724653"},"cell_type":"markdown","source":"Processing Name :"},{"metadata":{"trusted":true,"_uuid":"e7e52f97e5f1ff90b30366488600cd85d485f964"},"cell_type":"code","source":"def process_name(df):\n    #removing the name column since we have the title column\n    df.drop('Name', axis=1, inplace=True)\n    \n    #dummification Title column\n    titles_dummies = pd.get_dummies(df['Title'], prefix='Title')\n    df = pd.concat([df, titles_dummies], axis=1)\n    \n    #removing the title column since we have its dummies\n    df.drop('Title', axis=1, inplace=True)\n    return df","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9e880ea112e2d99418c50cc9a2c5f454e9a174e8"},"cell_type":"code","source":"df_combined = process_name(df_combined)\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"7a1b9badb49421e6bb8fcfb8860cf0584452e625"},"cell_type":"markdown","source":"Processing Sex :"},{"metadata":{"trusted":true,"_uuid":"9b857cb167554ad3237bbcdb26755c95c6fe0c78"},"cell_type":"code","source":"from sklearn.preprocessing import LabelEncoder\nenc = LabelEncoder()\ndf_combined['Sex'] = enc.fit_transform(df_combined['Sex'])\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"6ce9b7cfa2631b7d0ec4c7ce204c43cc0711fbe7"},"cell_type":"markdown","source":"Processing ticket :"},{"metadata":{"trusted":true,"_uuid":"0e32be69002d28d348a01904ba6b3fb4451f9141"},"cell_type":"code","source":"enc = LabelEncoder()\ndf_combined['Ticket'] = enc.fit_transform(df_combined['Ticket'])\ndf_combined.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d274872a3470207c736f188751d6caf68d7634ae"},"cell_type":"markdown","source":"## Building and training the model"},{"metadata":{"trusted":true,"_uuid":"bcd9f0a7fe945c0aa5a0791e1b545c55d1809fac"},"cell_type":"code","source":"train_data.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4473a94cacd4a63b2c6d82ef265bc5616fa4e550"},"cell_type":"code","source":"X_train = df_combined[:891]\nX_test = df_combined[891:]\nprint(X_train.shape)\nprint(X_test.shape)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"0a5209a797829852bcf7570edab713dc24d5e636"},"cell_type":"markdown","source":"### Splitting the train set into train and dev set"},{"metadata":{"trusted":true,"_uuid":"d22481e5d3d7cd163ff9e3a2b867288d1da6ebea"},"cell_type":"code","source":"def split_vals(a,n): return a[:n], a[n:]\nvalid_count =60\nn_trn = len(X_train)-valid_count\nX_train1, X_valid1 = split_vals(X_train, n_trn)\ny_train1, y_valid1 = split_vals(y_train, n_trn)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"9d8be36289a2f00306d285161e5b1557ed8bad86"},"cell_type":"code","source":"X_train1.shape,y_train1.shape,X_valid1.shape,y_valid1.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b9769d27393be428b45d54603cd5772a3672d54b"},"cell_type":"code","source":"from sklearn import metrics\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import GridSearchCV\n\nrfc = RandomForestClassifier(n_estimators=180,\n                             min_samples_leaf=3,\n                             max_features=0.5,\n                             n_jobs=-1)\nrfc.fit(X_train1,y_train1)\nrfc.score(X_train1,y_train1)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"85f513f2a79ad97d69fa2a871c2f47e4470de05a"},"cell_type":"code","source":"y_predict=rfc.predict(X_valid1)\nfrom sklearn.metrics import accuracy_score\naccuracy_score(y_valid1,y_predict)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d32edc76a4305dae2cac196398290374e5a5414d"},"cell_type":"code","source":"from sklearn.metrics import classification_report, confusion_matrix\nprint(classification_report(y_valid1,y_predict))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9ea509c221237792e39939dd98057d6b6c8ad1a2"},"cell_type":"markdown","source":"# confusion matrix \n<img src=\"https://www.researchgate.net/profile/D_Soeffker/publication/328860691/figure/fig2/AS:691499431907328@1541877720208/Confusion-matrix-and-related-performance-measures.png\">"},{"metadata":{"trusted":true,"_uuid":"1d475509c53b86527c2e33cbb111811873beec4d"},"cell_type":"code","source":"print(confusion_matrix(y_valid1,y_predict))","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bb2f8a862cfba9aa0fbb7950fc4febe99d31341a"},"cell_type":"markdown","source":"## Feature importance"},{"metadata":{"trusted":true,"_uuid":"8919cbf625730ded4ecd683ff6923519b7fac24e"},"cell_type":"code","source":"!pip install git+https://github.com/fastai/fastai@2e1ccb58121dc648751e2109fc0fbf6925aa8887\n!apt update && apt install -y libsm6 libxext6","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"218c10ba81aeefe4e0002f865080447358ebd724"},"cell_type":"code","source":"from fastai.imports import *\nfrom fastai.structured import *\nfrom pandas_summary import DataFrameSummary\nfrom sklearn.ensemble import RandomForestRegressor, RandomForestClassifier,GradientBoostingClassifier\nfrom IPython.display import display\nfrom sklearn import metrics\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import GridSearchCV\nimport seaborn as sns\nimport pylab as plot\n#Feature importance\nfi = rf_feat_importance(rfc, X_train1); fi[:10]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"fd8a3627ba233cd02991cdee27e33f296150c122"},"cell_type":"code","source":"def plot_fi(fi): return fi.plot('cols', 'imp', 'barh', figsize=(12,7), legend=False)\nplot_fi(fi[:30]);","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f7957c61def0d10745fd403d7eefe1b6d7fb7006"},"cell_type":"code","source":"# Keeping only the variables which are significant for the model(>0.01)\nto_keep = fi[fi.imp>0.01].cols; len(to_keep)\nto_keep","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"058e4e278181ad1bbd141f1aad0473a3c58e25f6"},"cell_type":"code","source":"#Now training the model on the entire data with only the important features.\nX_train = X_train[to_keep]\nX_train","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"5b0b9f7a0e4f6dcb4e943c9a06a71f0ea6558b8c"},"cell_type":"code","source":"rfc = RandomForestClassifier(n_estimators=200,min_samples_leaf=3,max_features=0.5,n_jobs=-1)\nrfc.fit(X_train,y_train)\nrfc.score(X_train,y_train)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"24fc28e963dca6fc3de0825f26b4244af3dcb75d"},"cell_type":"code","source":"X_test = X_test[to_keep]\nX_test.isna().sum()\nX_test.Fare.fillna(200, inplace=True)\noutput=rfc.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"f23de3f9a8d01f00da2e2b4e6e41c1a654786a4a"},"cell_type":"code","source":"output.size","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"a8ef4ad0883c380cd86af2af0c740a534c298597"},"cell_type":"code","source":"data_test = pd.read_csv('../input/test.csv')\ndf_output = pd.DataFrame()\ndf_output['PassengerId'] = data_test['PassengerId']\ndf_output['Survived'] = output\ndf_output[['PassengerId','Survived']].to_csv('submission2', index=False)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"78d95744d80d900e1594cff225ea12db49828392"},"cell_type":"code","source":"df_output.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b46705efe6ef9de0dd461ba243bb3832cce9a60f"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}