{"cells":[{"metadata":{"_uuid":"c15ce893164b257bb9510b0bf3d2c5869529b40b"},"cell_type":"markdown","source":"## Titanc Solution using LSTM network##\nThis is my first approach to solve the titanic Kaggle problem using LSTMs, of course improvements can be made.\nThe main way the algorithm works is to use all the features as a time-series data. \nI am open to comments and possible corrections on my code. \n"},{"metadata":{"_uuid":"ab02333d47719dc80798c014f1593352f12b5fa5"},"cell_type":"markdown","source":"### Frameworks\n- pandas\n- numpy\n- seaborn \n- keras\n- scikit-learn"},{"metadata":{"trusted":true,"_uuid":"4b133ce82a55931ed8eec6eb15565c6763039905"},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom keras.layers import Dense, LSTM, Activation, Dropout\nfrom keras.models import Sequential\nfrom keras.optimizers import Adam","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"333cbf44f6664c04d330edfd4e3fee42576c5251"},"cell_type":"code","source":"train = pd.read_csv('../input/train.csv', index_col = [\"PassengerId\"])\ntest = pd.read_csv('../input/test.csv', index_col = [\"PassengerId\"])\ncombination = [train,test]\n\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9b3b6a5c2b429c8ce4a5285024a427b272d6b256"},"cell_type":"markdown","source":"The data at our disposal is numerical (age, Pclass, etc.), alphabetical (name, sex & embarked) and alphanumerical (tickets). Let's get some further insight on our data:"},{"metadata":{"trusted":true,"_uuid":"72cafb4d0179d7a6dda799e57b27ce03dba34dc4"},"cell_type":"code","source":"train.describe(), test.describe()\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"22a717e86bd6b1902207d3c39d565af14edf49c1"},"cell_type":"markdown","source":"From the data above we get information regarding the various means and standard deviation of each data, furthermore we can see where we have missing values. For example in the training data set there are 714 values for the passengers' ages, however, we know that the total number of passengers is 891. \nThis information will be useful later on. For now I will try to find which features will have a heavier influence on the Survival and delete those featurues that will have a lower influence on the output."},{"metadata":{"trusted":true,"_uuid":"dbe1be422662e3b59828b6d6a8bda1c8e72a0e0d"},"cell_type":"code","source":"train[[\"Survived\",\"Pclass\"]].groupby(\"Pclass\").mean()\n\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b1b0ad6c409855e3675a2298ba10fb88e9f6aef4"},"cell_type":"code","source":"train[[\"Survived\", \"Sex\"]].groupby(\"Sex\").mean()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"e82de4377213667126f45b23f69f2d45983d4eed"},"cell_type":"code","source":"train['groups']=pd.cut(train.Age,[0,10,20,30,40,50,60,70,80])\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d9612a9253754bca98105931eee1f327b588d76d"},"cell_type":"code","source":"train[[\"Survived\", \"groups\"]].groupby(\"groups\").mean()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"01e535d7746d7ca7084ea07283d58bb4381e8967"},"cell_type":"markdown","source":"In the latter I have tried to divide the passengers by age group, by doing so we see how age has a big impact on the survival of the passenger, kids between 0-10 yrs have a higher survival rate. "},{"metadata":{"_uuid":"d7e7423098ee39e4da80309b49ee891408cac2c9"},"cell_type":"markdown","source":"### Plotting ###\nTo have a better view of the data I have decided to plot it. "},{"metadata":{"trusted":true,"_uuid":"16a2adc02748978cd359c2d0db17dbdba85e23be"},"cell_type":"code","source":"sns.barplot(train.groups, train.Survived)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2b7be18a16f7efca587dd53735a1f804fd4df326"},"cell_type":"code","source":"sns.barplot(train.Sex, train.Survived)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"22dc119859a56e780016f7b5ac374007721c9744"},"cell_type":"code","source":"sns.barplot(train.Pclass, train.Survived)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"4f01007eb3f8563703a1d9b6c96b370d8196079e"},"cell_type":"code","source":"sns.barplot(train.Pclass, train.Survived, hue=train.Sex)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"d4ec5864a723a9d54cff027993b326f1f97abeab"},"cell_type":"code","source":"train.describe(include = [\"O\"]), test.describe(include = [\"O\"])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"9d236d43156a5c48d24984b8e1edbcfafc5cbb6d"},"cell_type":"markdown","source":"For this version I have decided to delete the data regarding names, cabin and tickets. Cabin has many N/A values, names may not be directly related to the survival rate, however, their title could be relevant for future evaluations. In this case the name values are dropped, but in the future it could be interesting keeping the title of each passenger. "},{"metadata":{"trusted":true,"_uuid":"4744da64f98c3fb5205dc819313b09bf67767162"},"cell_type":"code","source":"train=train.drop([\"Name\", \"Ticket\",\"Cabin\", \"groups\"], axis=1)\ntest=test.drop([\"Name\", \"Ticket\",\"Cabin\"], axis=1)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"38e734217336a2aafd89ad06d13835d821384336"},"cell_type":"markdown","source":"### Converting Data and filling missing data\nSex are either male or females, hence we will convert this data to 1s and 0s respectively. Same thing can be applied to the embarked feature [0,1,2]. "},{"metadata":{"trusted":true,"_uuid":"df0eb8a6de36e8b1c86586dbee6f8b7b1c7dc226"},"cell_type":"code","source":"male_female = {\"male\":1,\n              \"female\":0}\n\ntrain[\"Sex\"]=train[\"Sex\"].map(male_female)\ntest[\"Sex\"]=test[\"Sex\"].map(male_female)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8d7f4b23c5612f8d880e89abda6d0737dfb4d52f"},"cell_type":"code","source":"embar = {\"C\":2,\n        \"S\":1,\n        \"Q\": 0}\ntrain[\"Embarked\"]=train[\"Embarked\"].map(embar)\ntest[\"Embarked\"]=test[\"Embarked\"].map(embar)\ntrain.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b4ba728f8d881e2636e302e5b5f7cad50594c7d4"},"cell_type":"code","source":"\ntrain[\"Age\"]=train[\"Age\"].fillna(value=np.mean(train[\"Age\"]))\ntest[\"Age\"]=test[\"Age\"].fillna(value=np.mean(train[\"Age\"]))\ntest[\"Fare\"]=test[\"Fare\"].fillna(value=np.mean(train[\"Fare\"]))\ntrain[\"Embarked\"]=train[\"Embarked\"].fillna(value=round(np.mean(train[\"Embarked\"])))\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"e2cc5d952508365ae292e02e577612d4f5356225"},"cell_type":"markdown","source":"## Model ##\nExtracting data"},{"metadata":{"trusted":true,"_uuid":"af76a14c553384893006fae0638d64db0e15db15"},"cell_type":"code","source":"train_y = train[\"Survived\"].iloc[:].values\ntrain_x = train.drop([\"Survived\"], axis = 1).iloc[:,:].values\n\ntrain_x = train_x.reshape(train_x.shape[0],-1,1)\ntrain_x.shape","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"67e50b28d182ac7af63fd0216773cadcef1b6771"},"cell_type":"markdown","source":"LSTM parameters"},{"metadata":{"trusted":true,"_uuid":"ffd73ff8096c812001c994a4fa6c641168236a32"},"cell_type":"code","source":"batch_size = 11\nepoch = 20\nhidden_units = 256 ","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"ec8cd7f3654898ad0373039f67e9717c814744b6"},"cell_type":"markdown","source":"LSTM architecture"},{"metadata":{"trusted":true,"_uuid":"b879ded9f59ea27966261a8237198fca3d155c4e"},"cell_type":"code","source":"model = Sequential()\nmodel.add(LSTM(hidden_units, input_shape=train_x.shape[1:],batch_size=batch_size))\nmodel.add(Activation('sigmoid'))\nmodel.add(Dense(1))\nmodel.compile(optimizer='Adam', loss = 'mean_squared_error',metrics = ['accuracy'] )\nmodel.fit(train_x,train_y, batch_size=batch_size, epochs=epoch, verbose = 1)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"8972805a8bf435dccb9cc1ba99af418ee09ee7a6"},"cell_type":"code","source":"out = pd.read_csv('../input/gender_submission.csv', index_col = [\"PassengerId\"])\ny_test = out.iloc[:].values","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"b2dd54f03622d9d042f9158886fa10acadee866b"},"cell_type":"code","source":"test_x = test.iloc[:,:].values\ntest_x = test_x.reshape(test_x.shape[0],-1,1)\nscores = model.evaluate(test_x, y_test, batch_size=batch_size)\npredictions = model.predict(test_x, batch_size = batch_size)\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"eba765c1b200bd965aee4406a49bd6020d2c3455"},"cell_type":"code","source":"\nprint('LSTM test accuracy:', scores[1])","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"247296fbb7d3cfa3a3458003580851b9ae836a78"},"cell_type":"markdown","source":""}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.7.1"}},"nbformat":4,"nbformat_minor":1}