{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder, LabelEncoder, MinMaxScaler, LabelBinarizer\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.linear_model import LogisticRegression, RidgeClassifierCV\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.svm import LinearSVC, SVC\n\nfrom mlxtend.classifier import StackingClassifier\nfrom catboost import CatBoostClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\n\nimport plotly.express as px\nfrom matplotlib import pyplot as plt\nimport scikitplot as skplt\nimport missingno as msno\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport os\nimport re\n\npd.set_option('display.max_rows', 500)\npd.set_option('display.max_columns', 500)\npd.set_option('display.width', 1000)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-size:250%;color:#292F55\"> Titanic - Machine Learning from Disaster</h1>\n\n*Top 3% Titanic solution - For this last version best model score is 0.81100*.\n![](https://images.fineartamerica.com/images/artworkimages/mediumlarge/1/2-rms-titanic-ship-plans-jose-elias-sofia-pereira.jpg)","metadata":{}},{"cell_type":"code","source":"cm = [\"#273176\",\"#3B61A3\",\"#76A4AC\",\"#BFD4B2\",\"#DAD8A1\"]\ngradient = [\"#292F55\",\"#273176\",\"#223A92\",\"#3B61A3\",\"#76A4AC\",\"#BFD4B2\",\"#DAD8A1\",\"#C7B679\",\"#957447\"]\nsns.palplot(gradient)","metadata":{"_uuid":"33d7c81c-db2c-4b04-bda7-5bdc7844222d","_cell_guid":"81ca7ea7-f034-4b79-accd-b19737dcbc9e","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a style=\"font-size:200%;color:#292F55\">Table Of Content\n* [<a style=\"font-size:150%;color:#292F55\">1. EDA & Feature Engeneering](#1_bullet)\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 1.1 Passengers location analisys](#1.1_bullet) - Survival for different Deck / Cabin numbers\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 1.2 Groups and family bonds analisys](#1.2_bullet) - Ticket numbers analisys, Names / Surnames\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 1.3 Personal features analisys](#1.3_bullet) - Age / Status analisys\n* [<a style=\"font-size:150%;color:#292F55\">2. Data preparation](#2_bullet)\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 2.1 Filling None values](#2.1_bullet)\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 2.2 Encoding features and droping unnecessary](#2.2_bullet)\n* [<a style=\"font-size:150%;color:#292F55\">3. Model development](#3_bullet)\n    * [<a style=\"font-size:130%;color:#3B61A3\"> 2.1 Catboost baseline](#2.1_bullet)\n    \n","metadata":{}},{"cell_type":"markdown","source":"# <a class=\"anchor\" id=\"1_bullet\" style=\"color:#292F55\"> 1. Exploratory Data Analysis (EDA)","metadata":{"_uuid":"5202c7e5-6b8d-4c54-ad88-9c79678d822d","_cell_guid":"33b08eac-de20-43ad-b320-490dcc25b860","trusted":true}},{"cell_type":"code","source":"path = \"/kaggle/input/titanic/\"\ndf_tr = pd.read_csv(f\"{path}train.csv\").set_index(\"PassengerId\", drop=True)\ndf_ts = pd.read_csv(f\"{path}test.csv\").set_index(\"PassengerId\", drop=True)\ndf = pd.concat([df_tr, df_ts], axis=0)\ndf.head(10).style.background_gradient(cmap='Blues')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.bar(df, figsize=(30,2), color=gradient)","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a class=\"anchor\" id=\"1.1_bullet\" style=\"color:#3B61A3\"> 1.1 Passengers location analisys \n\nI assume, that the location of passengers at the moment of the disaster may affect on the surviving rate.\n\nFor this case we will try to analyze such location, referring to the time of the disaster - 23:40. It means that some passengers may already have been in their cabins, some of them was hanging on some restaurants and \"bars\" (linked to there Pclass and maby cabin location) etc.\n\nAnyway, we do not know their location at that moment, but we can create some feature to describe it somehow to help our model.\n\nOffcourse, we have smth like 80% None values in \"Cabin\" feature - but still some useful information can be extracted + we need to be careful with filling the gaps.","metadata":{"_uuid":"1516af28-5838-42af-b0a8-2489625b6417","_cell_guid":"aa3f011e-6a39-41e9-93a2-e07ab39dab27","trusted":true}},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"1.1.1_bullet\" style=\"color:#3B61A3\"> 1.1.1 Survival for different Deck numbers\nDeck descriptors are coded inside of the Cabin numbers\nWe will extract them as a new feature.\nUnknowns we will mark as \"N/A\"","metadata":{"_uuid":"80181582-6313-4bae-91fd-0efa3c8394c0","_cell_guid":"393adf89-c2ba-4367-85ea-c1fe6abed04e","trusted":true}},{"cell_type":"code","source":"df[\"Deck\"] = df[\"Cabin\"].str[:1]\ndf[\"Deck\"] = df[\"Deck\"].replace(np.nan,\"N/A\")\nprint(\"All Deck descriptors:\")\nprint(set(df[\"Deck\"].values))","metadata":{"_uuid":"a4934617-4f93-4a35-b062-e330380db5ef","_cell_guid":"2362473f-904d-4edd-ab44-2150bc1fee0d","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\ndfplt = dfplt[dfplt[\"Deck\"]!=\"N/A\"]\nfig = px.histogram(dfplt, x=\"Deck\",color=\"Survived\",\n                   color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"1d46a8c7-ef89-43d9-8714-0adc452183b6","_cell_guid":"e4cf78ba-be4d-4973-96a5-36485073d648","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"What we can see from such graphs:\n*     Only one PassengerId have been on the Deck T - and he is in train set, so we can drop the \"T\" value\n*     There ara some dependancies between surviving rate and the Deck number:\n\n|deck|died|survived|ratio|\n|:--:|:--:|:------:|:---:|\n|C   |24  |35      |0.4  |\n|G   |2   |2       |0.5  |\n|D   |8   |25      |0.24 |\n|A   |8   |7       |0.53 |\n|B   |12  |35      |0.26 |\n|F   |5   |8       |0.38 |\n|E   |8   |24      |0.25 |\n        \nIn this case - we can put such new features into the model and encode them.","metadata":{"_uuid":"ee62d350-4215-431e-a2bd-aa6a1d6952cf","_cell_guid":"c45aee6d-53bc-4f78-acfc-64cb4f5ee631","trusted":true}},{"cell_type":"code","source":"df.loc[df[\"Deck\"]=='T',\"Deck\"] = 'N/A'","metadata":{"_uuid":"76f56537-1e67-4029-8b5a-4524dcdb3430","_cell_guid":"556a5031-69bb-42e7-8f5c-f2f221bcefad","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"1.1.2_bullet\" style=\"color:#3B61A3\"> 1.1.2 Cabin numbers analisys\nSome of the passangers have shared several cabins - thay may have the same cabin number or leave in 3 of 4 cabins at the same time (families). But actually not only families or reletives shared cabins, but some collegues, friends etc. In this case i will try to create a features, which will describe the location of passangers, linked to there Cabin number.\n\nSome of the cabins has ambiquouse values - i will hardcode them - giving maximum values to multiple cabins \nfor a passanger. Or just rename them is several other ways, according to the Deck plans from https://www.encyclopedia-titanica.org/titanic-deckplans/","metadata":{"_uuid":"7b11b92e-1041-4633-898b-4e36fe2a288a","_cell_guid":"d52f81b3-2441-428f-be6e-0647f2db270e","trusted":true}},{"cell_type":"code","source":"replaces = {'B51 B53 B55': 'B55', 'B52 B54 B56': 'B56', 'B57 B59 B63 B66': 'B66', 'B58 B60': 'B60', \n            'B82 B84': 'B84', 'B96 B98': 'B98', 'C22 C26': 'C26', 'C23 C25 C27': 'C27', 'C55 C57': 'C57',\n            'C62 C64': 'C64', 'D10 D12': 'D12', 'E39 E41': 'E41', 'F E46': 'E46', 'F E57': 'E57',\n            'F E69': 'E69', 'F G63': 'G63', 'F G73': 'G73', 'F': None, 'D': None, ' ': None, 'T': None, np.nan: None}\ndf[\"Cabin\"] = df[\"Cabin\"].replace(replaces)\ndf[\"Cabin\"] = df.fillna(np.nan)[\"Cabin\"].str[1:].astype(float)","metadata":{"_uuid":"c859d9e8-19a9-4993-bd35-2a34069eb08c","_cell_guid":"50fd1102-27a0-467d-a58c-6196ef5902c5","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.histogram(dfplt, x=\"Cabin\",color=\"Survived\", height=300,\n                   color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"cb10a404-c348-41ec-a746-820e6862b2d2","_cell_guid":"f2c6157a-ed4b-4329-a553-1825e61365bb","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that there is no accurate dependency between Cabin number and the Survival in general - maby it can work on connection with some other features. Some extra information we can take from the Cabin number - is the Side of the ship (left or right) ,according to the even or odd number of it. The side of the ship can effect much on the serviving rate.","metadata":{"_uuid":"cb9bb37c-8b54-4f05-b05f-76f6ad3f98ae","_cell_guid":"34c4b6e8-8ce3-488b-a7ab-b9d2362de7a3","trusted":true}},{"cell_type":"code","source":"df[\"Side\"] = df[\"Cabin\"]\ndf.loc[df[\"Side\"]!=0,\"Side\"] = (df[\"Cabin\"][df[\"Cabin\"]!=0]%2-0.5)*2\n\ns = df[df[\"Side\"]==1]\nprint(f'Survived for side 1\\t {len(s[s[\"Survived\"]==1])/len(s)}')\ns = df[df[\"Side\"]==-1]\nprint(f'Survived for side -1\\t {len(s[s[\"Survived\"]==1])/len(s)}')","metadata":{"_uuid":"0986fef2-98be-4abc-b622-45eac9255f9e","_cell_guid":"7637a5db-d796-4aa7-8630-57fb2dd2d927","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see - passangers survived more often, living on the 1, then on the side -1\n\nTo fill the missing values for \"Cabin\", we will put median values for all passangers in specific Decks.\nBefore this, we will devide number of the Cabin by two - because the \"side\" location has been already taken from the data.\nIf Deck number is unknown - we will just mark Cabin as -1 - mwaning \"Unknown\"","metadata":{"_uuid":"6773376d-31ac-47df-aaba-14e5de786bbe","_cell_guid":"18a17eef-9d17-4e23-8c68-fe83161575df","trusted":true}},{"cell_type":"code","source":"for i in set(df[\"Deck\"].values):\n    v = df[df[\"Deck\"]==i][\"Cabin\"]//2\n    df.loc[df[\"Deck\"]==i, \"Cabin\"]= v\n    df.loc[(df[\"Deck\"]==i) & (df[\"Cabin\"]==0),\"Cabin\"] = np.median(v)\n    \ndf.loc[df[\"Cabin\"].isna(),\"Cabin\"]=-1\ndf[\"Cabin\"] = df[\"Cabin\"].astype(int)","metadata":{"_uuid":"6a23462d-d394-49f7-9d2b-308b42752688","_cell_guid":"1608909e-1b08-4dcc-8887-7a6e58f8de09","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we gonna plot some developed \"coordinates\" of the passangers, to find some patterns in survival.","metadata":{}},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt = dfplt[~dfplt[\"Survived\"].isna()]\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.scatter_3d(dfplt, x=\"Cabin\", y=\"Side\", z= \"Deck\", color=\"Survived\",\n                    color_discrete_sequence=cm, size_max=6, width=1000, height=1000)\nfig.show()","metadata":{"_uuid":"d504012f-6958-4bf6-b57b-02e88a2203bf","_cell_guid":"72bab3ec-7486-472e-bbc1-f38a6662b26a","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see some patterns in the data:\n  - For example, all of the passangers on the \"1\" side of the Deck D survived.\n  - For only some of decks the \"closer\" location to zero may cause the better survival (maby those passangers was closer to the ladders)","metadata":{"_uuid":"d864b5a9-5662-4769-8621-2b8b8ea67fb4","_cell_guid":"5dd0b6ae-5a63-4a89-94d4-101b16c2d2dd","trusted":true}},{"cell_type":"code","source":"df[\"Side\"] = df[\"Side\"].fillna(0)\nmsno.bar(df, figsize=(30,2), color=cm)","metadata":{"_uuid":"80cf7463-85d1-4825-a74f-7e8834b129c7","_cell_guid":"6aab6563-0a93-4b0b-ba41-e01e25e8e207","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the end - we have to fillna for \"Side\" feature with 0 - refering to unknown side of the ship","metadata":{"_uuid":"a1f5dae3-4c10-4b48-b855-eba55375e38d","_cell_guid":"ece36e69-9f7a-4833-a314-b65136952865","trusted":true}},{"cell_type":"markdown","source":"## <a class=\"anchor\" id=\"1.2_bullet\" style=\"color:#3B61A3\">  1.2 Groups and family bonds analisys\nSome groups of passangers, should be connected via ticket numbers, surnames, classP and etc. We will try to put the data into the shape in which our model will detect such connections\n### <a class=\"anchor\" id=\"1.2.1_bullet\" style=\"color:#3B61A3\">  1.2.1 Ticket numbers analisys","metadata":{"_uuid":"705b1795-a65c-414f-9ef8-00421d77fc65","_cell_guid":"76e5b855-6e3d-4f71-89c9-deb45fb139a9","trusted":true}},{"cell_type":"markdown","source":"**\"LINE\" tickets:**\nAll Tickets number contains some numbers, except \"LINE\" tickets www.encyclopedia-titanica.org\nPhiladelphia's westbound voyage was cancelled, and several shipmates forced to travel aboard Titanic as passengers (Some of them have LINE ticket number):\n - August Johnson (Johnson, Mr. Alfred) - **LINE** - 370160\n - William Cahoone Jr. Johnson - **LINE** - 370160\n - William Henry Törnquist - **LINE** - 370160\n - Andrew John Shannon (Lionel Leonard) - **LINE** - 370160\n\nWe will not analize them much, in this part, but we should take it into account","metadata":{"_uuid":"3b2759b6-d0f9-46b5-a7bf-145f54263bb2","_cell_guid":"1c5d3d5c-41e1-40ea-94a0-23bcc73118f1","trusted":true}},{"cell_type":"code","source":"lin_rep = lambda x: x.replace({'LINE':\"370160\"})\ndf = lin_rep(df)","metadata":{"_uuid":"f6a6048d-3c22-496a-8b91-5cd79038d292","_cell_guid":"3ba1057d-8e48-4710-a7d5-27dc68401ec2","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some of the tickets contains several prefixes refering for some specific sale policies - maby it may be usefull for the model (we will save it into Ticket_p) but at the same time wew will clear our ticket number from umbiques information (It may be some hidden connections in the ticket numbers).","metadata":{"_uuid":"df429238-96a0-4ef8-99c4-c54f44c2e1af","_cell_guid":"ab4cc549-f2f6-49b0-af82-6cf8980ba87e","trusted":true}},{"cell_type":"code","source":"prefixes = []\nnums, prefs = [],[]\nfor i in df[\"Ticket\"].values:   \n    if not i.isdigit():\n        nums.append(int(re.search('.* {1}([0-9]+)', i).groups()[0]))\n        prefix = re.search('(.*)( {1})[0-9]+', i).groups()[0]\n        prefs.append(prefix.replace(\".\",\"\").replace(\" \", \"\").replace(\"/\",\"\")) # Needed to put in one group such prefixes as \"A/5\", \"A/5.\", \"A.5\" etc.\n    else:\n        nums.append(int(i))\n        prefs.append(\"\")\ndf[\"Ticket\"] = nums\ndf[\"Ticket_p\"] = prefs","metadata":{"_uuid":"85ea3798-a81e-4d32-813d-dec4bcb1ccd2","_cell_guid":"5ee04e59-0850-4c50-9f00-477c8af92da7","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\nfig = px.scatter(dfplt.astype(str), x=\"Ticket_p\", y=\"Name\", color= \"Survived\",\n                 color_discrete_sequence=cm, size_max=6,width=1200, height=500)\nfig.show()","metadata":{"_uuid":"3cc290b8-125a-4fdf-85d9-af4c41e4763c","_cell_guid":"8c19c9cd-9854-4b93-8593-d4933bae9cb7","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that some of prefixes appear only one time or only in train dataset. Lets clean them in order not to confuse our model.\n\nWe wil do it manualy - couse it is not much off them, and we gonna check them, if maby there is some sintax mistakes.","metadata":{"_uuid":"ad42f9d2-7e05-4afc-9328-c386e07676f8","_cell_guid":"5605a53e-1c3b-4db9-a24f-01b800899d83","trusted":true}},{"cell_type":"code","source":"drop = [\"SP\", \"SOP\", \"Fa\", \"SCOW\", \"PPP\", \"AS\", \"CASOTON\", \"SWPP\", \"SCAHBasle\", \"SCA3\", \"STONOQ\", \"AQ4\", \"A2\", \"LP\", \"AQ3\", \"\"]\ndf = df.replace(drop, 'N/A')","metadata":{"_uuid":"b871f7ff-2dae-4647-9dd4-7b142fa05862","_cell_guid":"6b09e9eb-69bd-47c1-887c-eacf1a83d473","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt= dfplt[dfplt[\"Ticket_p\"] != \"N/A\"]\nfig = px.scatter(dfplt.astype(str), x=\"Ticket_p\", y=\"Name\", color= \"Survived\",\n                 color_discrete_sequence=cm, size_max=6,width=1200, height=500)\nfig.show()","metadata":{"_uuid":"eb33907f-10a0-48d5-99c2-601c1da8df89","_cell_guid":"0eface82-f518-48fd-9d3a-dfc4c23fc1e0","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We do not know much about the meaning of this prefixes, but there are some dependencies:\n* No passangers survived with the prefix A4\n* Most of the passangers with WC prefix - died","metadata":{"_uuid":"be55324a-3aa8-4841-9978-48fc628b23a9","_cell_guid":"5cbca22d-6f4d-4f38-9778-0b14791017ee","trusted":true}},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"1.2.2_bullet\" style=\"color:#3B61A3\">  1.2.2 Names, surnames and status feature\nTo get more information about the family bonds - we will extract surnames instead of Names, and put them into features list/","metadata":{"_uuid":"2da48d5e-ad1c-4ffe-b23c-b7b0be0b790b","_cell_guid":"72410d2c-549c-4292-9f35-86071687a154","trusted":true}},{"cell_type":"code","source":"df[[\"Surname\",\"Name\"]] = [i.split(\",\") for i in df[\"Name\"].values]","metadata":{"_uuid":"b93b4921-953f-4dfd-abdc-c85d4b11edc8","_cell_guid":"9ce8b816-3fa3-4484-b55e-20095f1fbedc","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will get all same-surname groups into one list, to put some feature for it\nAll others will be marked with \"Other\".(*)","metadata":{"_uuid":"4f217c15-c310-4f4b-b14b-28b72e10a3b2","_cell_guid":"ccf0507b-fcb6-439b-8baa-28322fc225a3","trusted":true}},{"cell_type":"code","source":"a = df.groupby(\"Surname\")[\"Surname\"].count()\nfam_list = a[a>1].index.values\ndf.loc[~df[\"Surname\"].isin(fam_list),\"Surname\"] = \"Other\"","metadata":{"_uuid":"044429f1-47cc-4469-80e1-c1df409c8c59","_cell_guid":"432c7092-bf65-48a6-a86f-10a73cf32199","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Lets visualize Surviving rates for families:","metadata":{"_uuid":"d1680b73-8f98-4d20-8fc0-2c1b5f8eb232","_cell_guid":"dc80b257-2f42-4aef-97d9-4fb3d7e9bde2","trusted":true}},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\ndfplt = dfplt[dfplt[\"Surname\"]!=\"Other\"]\nfig = px.histogram(dfplt, x=\"Surname\",color=\"Survived\",color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"85cf66c6-ae8a-40cb-ba5c-003be9174d9c","_cell_guid":"1b227f4e-bb79-44e9-a03a-ea3b2aa8c359","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The same thing as in (*) we gonna do for those Surnames who do not exist in test set.\nAgain, i gonna do it manually, cause i dont wanna work with automatization here =)\n\nBut just before this we will put some extra feature which mean \"Has any namesakes\" to determin if the person in som relatio-like group. It may look the same as SibSp, but guess, that som connections between people is hidden.","metadata":{"_uuid":"599c20bc-f55a-427f-9308-b7bdaf09e0ed","_cell_guid":"b2b3f804-bd8a-462b-b836-ae79a053b517","trusted":true}},{"cell_type":"code","source":"df[\"Namesakes\"] = 1\ndf.loc[df[\"Surname\"]==\"Other\",\"Namesakes\"] = 0","metadata":{"_uuid":"f1ffbe22-b595-490b-845f-c9b51ce56e4a","_cell_guid":"5099e84e-1dfe-445b-9f13-60cc2ab4fd00","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"not_imp_s = [\"Braund\",\"Allen\",\"Moran\",\"Meyer\",\"Holverson\",\"Turpin\",\"Arnold-Franchi\",\"Panula\",\"Harris\",\"Skoog\",\"Kantor\",\"Petroff\",\"Gustafsson\",\"Zabour\",\n             \"Jussila\",\"Attalah\",\"Baxter\",\"Hickman\",\"Nasser\",\"Futrelle\",\"Navratil\",\"Calic\",\"Bourke\",\"Strom\",\"Backstrom\",\"Ali\",\"Jacobsohn\",\"Larsson\",\n             \"Carter\",\"Lobb\",\"Taussig\",\"Johnson\",\"Abelson\",\"Hart\",\"Graham\",\"Pears\",\"Barbara\",\"O'Brien\",\"Hakkarainen\",\"Van Impe\",\"Flynn\",\"Silvey\",\"Hagland\",\n             \"Morley\",\"Renouf\",\"Stanley\",\"Penasco y Castellana\",\"Webber\",\"Coleff\",\"Yasbeck\",\"Collyer\",\"Thorneycroft\",\"Jensen\",\"Newell\",\"Saad\",\"Thayer\",\"Hoyt\",\n             \"Andrews\",\"Lam\",\"Harper\",\"Nicola-Yarred\",\"Doling\",\"Hamalainen\",\"Beckwith\",\"Mellinger\",\"Bishop\",\"Hippach\",\"Richards\",\"Baclini\",\"Goldenberg\",\n             \"Beane\",\"Duff Gordon\",\"Tylor\",\"Dick\",\"Chambers\",\"Moor\",\"Snyder\", \"Howard\", \"Jefferys\", \"Franklin\",\"Abelseth\",\"Straus\",\"Khalil\",\"Dyker\",\"Stengel\",\n             \"Foley\",\"Buckley\",\"Zakarian\",\"Peacock\",\"Mahon\",\"Clark\",\"Pokrnic\",\"Ware\",\"Gibson\",\"Taylor\"]\ndf = df.replace(not_imp_s,'Other')","metadata":{"_uuid":"c9c96df2-6dff-4a2f-a396-2d374b5bb8d9","_cell_guid":"5e90c3ce-8afe-443c-b2eb-749cec997fe6","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[(df[\"Surname\"]==\"Other\") & (df[\"Namesakes\"]==True)].head(10).style.background_gradient(cmap=\"Blues\")","metadata":{"_uuid":"db011d17-10e8-4c2f-8677-2310e03b6736","_cell_guid":"e1605265-49c8-445a-a3e2-0d00d37eabdb","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"s = df[df[\"Namesakes\"]==0]\nprint(f'Have no Namesakes \\t {len(s[s[\"Survived\"]==1])/len(s)}')\ns = df[df[\"Namesakes\"]==1]\nprint(f'Have Namesakes \\t\\t {len(s[s[\"Survived\"]==1])/len(s)}')","metadata":{"_uuid":"b1c23264-1eb5-4168-93a9-f6dcf1563d54","_cell_guid":"2b5e1e3b-ed52-4a4a-bf86-f5ce57d3fb79","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\ndfplt = dfplt[dfplt[\"Surname\"]!=\"Other\"]\nfig = px.histogram(dfplt, x=\"Surname\",color=\"Survived\",color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"964bbf2b-235c-4f78-9579-c1e8ea7b2260","_cell_guid":"d658375a-bd32-40b9-aee5-f0623f0a6a10","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will delete all the Families, for which number of Survived equel to not Survived (Again ill do it manually)","metadata":{"_uuid":"fa6ddf7f-c74b-4403-8218-7e8dd21ddd3d","_cell_guid":"4e0d2517-3859-4f61-9060-77ecc217fe5d","trusted":true}},{"cell_type":"code","source":"drop = [\"Abbott\",\"Keane\",\"Minahan\",\"Crosby\",\"Hocking\",\"Dean\",\"Mallet\",\"\"]\ndf = df.replace(drop,'Other')","metadata":{"_uuid":"81621009-f3bf-4fbe-81f5-2717297a8631","_cell_guid":"2a21adf1-e65e-4271-9939-619a79e17730","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a class=\"anchor\" id=\"1.3_bullet\" style=\"color:#3B61A3\">  1.3 Personal features analisys\n### <a class=\"anchor\" id=\"1.3.1_bullet\" style=\"color:#3B61A3\">  1.3.1 Title features\nPassangers with different Titles may hav different survival rate (may be connected to so some position on groups, status of inner features)","metadata":{}},{"cell_type":"code","source":"df[\"Title\"] = pd.DataFrame(df[\"Name\"].str.strip().str.split(\".\").tolist()).set_index(df.index).iloc[:,0]\ndf[\"Title\"] = df[\"Title\"].fillna(\"Others\")","metadata":{"_uuid":"289a2842-0ca4-4c61-b1c6-2ffc03ae3507","_cell_guid":"55027e26-39a2-47b4-9dce-6ffa89a392bf","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.histogram(dfplt, x=\"Title\",color=\"Survived\",color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"a3cbbc5e-334b-4b58-82f7-34bb325c59e7","_cell_guid":"87d67c2b-7080-473c-8a5f-c2f7f3e69f39","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rename = {\"Miss\":\"Ms\",\n          \"Mrs\": \"Mme\",\n          \"Others\": [\"Don\",\"Rev\",\"Dr\",\"Lady\",\"Sir\",\"Mlle\",\"Col\",\"the Countess\",\"Mme\",\"Major\",\"Capt\",\"Jonkheer\",\"Dona\"]}\nfor k in rename:\n    df[\"Title\"] = df[\"Title\"].replace(rename[k],k)","metadata":{"_uuid":"a11c1314-ae6d-4f87-8579-9adc5b970a80","_cell_guid":"002387bd-b51d-4155-87fd-1a3bb3c239b3","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I desided to delete values with low frequancy - for our model not to be messed up.","metadata":{}},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.histogram(dfplt, x=\"Title\",color=\"Survived\",color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"7a1f23f4-64e2-4362-a16a-0609f414a7c5","_cell_guid":"d731197e-ba2d-47f5-a7b3-b8c9312571ac","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"1.3.2_bullet\" style=\"color:#3B61A3\">  1.3.2 Parch and Age\nThis features may have crucial effect on the target. Parants was saving there children befor them, and many other connections may be inside of the just two attributes","metadata":{"_uuid":"42ae9525-34b5-4777-b88a-a17cd0c1e462","_cell_guid":"5fba8bba-9bce-4eb2-9b87-888323796a97","trusted":true}},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt = dfplt[~dfplt[\"Survived\"].isna()]\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.scatter(dfplt, x=\"Age\", y=\"Parch\", color = \"Survived\", size_max=6\n                 ,color_discrete_sequence=cm,width=1200, height=500)\nfig.show()","metadata":{"_uuid":"d1d27d28-5021-44fc-84c2-581495bb1beb","_cell_guid":"53a9ecd8-aa07-44dd-a240-bc03ec665cfa","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[(df[\"Age\"]==5) & (df[\"Parch\"]==0)]\n# 5 y.o child traveling by hereself?","metadata":{"_uuid":"27b4d576-987b-4901-92d7-22400f6ab6c4","_cell_guid":"44932bfa-147e-497d-8773-f13e885ca0db","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the end i desided not to fix it, cause if takes some extra information about the data, and it is kind of cheating.\nBut i will leave it there just in case someone will need it","metadata":{"_uuid":"e5ded82d-413c-4dd5-81a3-a4fa536f7827","_cell_guid":"60687615-4a79-49ab-a189-d809bee004cd","trusted":true}},{"cell_type":"code","source":"#df.loc[df[\"Name\"]==\"Emanuel, Miss. Virginia Ethel\",\"Parch\"]=1\n#df.loc[df[\"Name\"]==\"Dowdell, Miss. Elizabeth\",\"Parch\"]=1\n\n#df.loc[df[\"Name\"]==\"Albimona, Mr. Nassef Cassem\",\"Parch\"]=1\n#df.loc[df[\"Name\"]==\"Hassan, Mr. Houssein G N\",\"Parch\"]=1\n\n#df.loc[df[\"Name\"]=='Watt, Mrs. James (Elizabeth \"Bessie\" Inglis Milne)',\"Parch\"]=1\n#df.loc[df[\"Name\"]==\"Watt, Miss. Bertha J\",\"Parch\"]=1","metadata":{"_uuid":"e501740c-4e71-4c05-8dd1-01bb50192d20","_cell_guid":"4c51069f-a32b-4261-8b34-1e60928b443e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We will manually set two more features:\n* \"Kid\" feature - for less then 18 y.o passangers - there survival rate is higher \n* 'Alone' feature - for those who was travalling alone (>18 e.o, Parch & SibSp ==0). \n* 'Old' feature - elder people died more often. I took 60 y.o as a treashold, after checking several and chooseing the one","metadata":{"_uuid":"1d7c1d4a-087f-41ef-a288-b868321b2058","_cell_guid":"ad92ae12-0519-43c1-8bb6-60a8d4d9d527","trusted":true}},{"cell_type":"code","source":"df[\"Kid\"]=0\ndf.loc[(df[\"Age\"]<18),\"Kid\"]=1\nprint(f'Kids survived koeff:\\t{len(df[(df[\"Kid\"]==1) & (df[\"Survived\"]==0)])/len(df[(df[\"Kid\"]==1) & (df[\"Survived\"]==1)])}')\nprint(f'Others survived koeff:\\t{len(df[(df[\"Kid\"]==0) & (df[\"Survived\"]==0)])/len(df[(df[\"Kid\"]==0) & (df[\"Survived\"]==1)])}')","metadata":{"_uuid":"a705d64f-278c-4f64-a801-5e1e4d67d096","_cell_guid":"b628a26c-c84f-4b85-9400-802b7b3ac67b","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"Old\"]=0\ndf.loc[(df[\"Age\"]>60),\"Old\"]=1\nprint(f'Elder survived koeff:\\t{len(df[(df[\"Old\"]==1) & (df[\"Survived\"]==0)])/len(df[(df[\"Old\"]==1) & (df[\"Survived\"]==1)])}')\nprint(f'Others survived koeff:\\t{len(df[(df[\"Old\"]==0) & (df[\"Survived\"]==0)])/len(df[(df[\"Old\"]==0) & (df[\"Survived\"]==1)])}')","metadata":{"_uuid":"23f6734b-1b39-430a-a1ea-a072e99e766d","_cell_guid":"aba991d8-5840-4979-947b-4a69cd5c6d3e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"Alone\"] = 0\ndf.loc[(df[\"Parch\"]==0) & (df[\"SibSp\"]==0),\"Alone\"]=1\nprint(f'Alone survived koeff:\\t\\t{len(df[(df[\"Alone\"]==1) & (df[\"Survived\"]==0)])/len(df[(df[\"Alone\"]==1) & (df[\"Survived\"]==1)])}')\nprint(f'Not Alone survived koeff:\\t{len(df[(df[\"Alone\"]==0) & (df[\"Survived\"]==0)])/len(df[(df[\"Alone\"]==0) & (df[\"Survived\"]==1)])}')","metadata":{"_uuid":"7393379c-82b8-4571-a0cb-11e9b9465854","_cell_guid":"019d185a-cd22-4f6d-b2a9-8d02ebe1aab7","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <a class=\"anchor\" id=\"2_bullet\" style=\"color:#292F55\">  2. Data preparation\n## <a class=\"anchor\" id=\"2.1_bullet\" style=\"color:#3B61A3\">  2.1 Filling None values\n### <a class=\"anchor\" id=\"2.1.1_bullet\" style=\"color:#3B61A3\">  2.1.1 Filling \"Age\" None values","metadata":{"_uuid":"292ee101-7b25-43fc-aa7e-efa0838fef62","_cell_guid":"61ba1684-1b40-49a4-8916-c8bb91991463","trusted":true}},{"cell_type":"code","source":"msno.bar(df, figsize=(30,2), color=gradient)","metadata":{"_uuid":"e09567d5-cdbb-442a-b4d8-28a3bca74978","_cell_guid":"e720f1db-4944-4b80-be88-62b00bee0915","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the \"Age\" feature we just can replace None values with some mean value, but we gonna do it in a bit more complex way.\nWe will locate some groups, based on Sex and Pclass - lets check the difference in ages in several groups.","metadata":{"_uuid":"a4a5c0c9-3cdf-4649-8a27-64042204a124","_cell_guid":"38f70eff-cb08-49ed-bb91-b380cd46dd38","trusted":true}},{"cell_type":"code","source":"print(df[(df[\"Pclass\"]==1) & (df[\"Sex\"]==\"female\")][\"Age\"].median())\nprint(df[(df[\"Pclass\"]==1) & (df[\"Sex\"]==\"male\")][\"Age\"].median())\nprint(df[(df[\"Pclass\"]==2) & (df[\"Sex\"]==\"female\")][\"Age\"].median())\nprint(df[(df[\"Pclass\"]==2) & (df[\"Sex\"]==\"male\")][\"Age\"].median())\nprint(df[(df[\"Pclass\"]==3) & (df[\"Sex\"]==\"female\")][\"Age\"].median())\nprint(df[(df[\"Pclass\"]==3) & (df[\"Sex\"]==\"male\")][\"Age\"].median())","metadata":{"_uuid":"f11ad35f-030d-45c3-aa46-0f3dfd229cdf","_cell_guid":"e9680737-ca11-4bff-ae68-54d951d554a6","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from itertools import *\nl1, l2 = [1,2,3], [\"female\",\"male\"]\nfor c,s in product(l1,l2):\n    msk = (df[\"Pclass\"]==c) & (df[\"Sex\"]==s)\n    df.loc[msk,\"Age\"] = df[msk][\"Age\"].fillna(df[msk][\"Age\"].median())","metadata":{"_uuid":"8db6ef6b-2a55-4e50-8b41-ba2b4dd36cf7","_cell_guid":"0fb987c6-ab87-4d50-9593-20c82aa28c7b","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.scatter(dfplt, x=\"Age\", y=\"Name\", color = \"Survived\", size_max=6,\n                 width=1200, height=500,color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"ab1d1834-1350-48cf-bec5-d2c6308e1819","_cell_guid":"437f28ed-2a0e-4ca9-8c16-a73c1390ea78","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"2.1.2_bullet\" style=\"color:#3B61A3\">  2.1.2 Filling \"Fare\" feature\nI will manually fill None value with the mean value for passanger class.\nAfter that we can replace Fare with a rank, because i assume that the value of difference is not important.","metadata":{"_uuid":"e2120264-33ea-48ec-8983-e52d8dad2b19","_cell_guid":"a3c3236b-6766-4263-a71b-09be078c95a7","trusted":true}},{"cell_type":"code","source":"print(df.loc[1044])\ndf.loc[1044,\"Fare\"] = df[df[\"Pclass\"]==3][\"Fare\"].mean()","metadata":{"_uuid":"63d789be-4e77-4225-a6c1-751d26ad39af","_cell_guid":"9ffb1dfc-ac94-4ad5-9224-e8c5dd6a86a3","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df[\"Fare\"] = df[\"Fare\"].rank(method='max')","metadata":{"_uuid":"98052483-20f1-41ca-a8c6-e13119d08402","_cell_guid":"b0c72175-2ec3-4fc7-ae66-e0ed14d86b8f","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dfplt = df.copy(deep=True)\ndfplt[\"Survived\"] = dfplt[\"Survived\"].astype(str)\nfig = px.scatter(dfplt, x=\"Fare\", y=\"Name\", color = \"Survived\", size_max=6,\n                 width=1200, height=500,color_discrete_sequence=cm)\nfig.show()","metadata":{"_uuid":"167ed852-4705-4a7a-b505-1dd50d743280","_cell_guid":"3c07ea5f-fbc9-4625-ae56-772f7dd27875","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"2.1.3_bullet\" style=\"color:#3B61A3\">  2.1.3 Filling \"Embarked\" feature\nWe have two passangers, whose \"Embarked\" information is unknown","metadata":{"_uuid":"a992670c-47e3-4373-ab2d-9ee53f058c27","_cell_guid":"01496d42-f552-48b5-a9de-13ce8bbb66bb","trusted":true}},{"cell_type":"code","source":"df[df[\"Embarked\"].isna()]","metadata":{"_uuid":"ca5beac8-adc1-4aba-a41e-2100cc1f8d79","_cell_guid":"35740500-4e16-4769-9537-f7d220b74e0a","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"According to https://www.encyclopedia-titanica.org/titanic-survivor/amelia-icard.html we can fill nan values with \"S\"","metadata":{"_uuid":"abc79687-4bc2-49e5-a25b-e57ce87d7ffc","_cell_guid":"1d7f9d52-7e69-426f-8260-63e904a9dbd8","trusted":true}},{"cell_type":"code","source":"df.loc[df[\"Embarked\"].isna(),\"Embarked\"] = \"S\"","metadata":{"_uuid":"170ff031-6e80-47d7-9523-984a5206df0d","_cell_guid":"23149575-ae80-4e4d-82dd-22586c09848e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a class=\"anchor\" id=\"2.2_bullet\" style=\"color:#3B61A3\">  2.2 Encoding features and droping unnecessary","metadata":{"_uuid":"4cf6f7fe-dead-4266-943b-b19396a86ff0","_cell_guid":"db49fdb9-d9f3-42f7-861a-84910dea0f90","trusted":true}},{"cell_type":"markdown","source":"We can try to apply OneHot encoder for all columns to see how spars will be the matrix","metadata":{"_uuid":"fa48b2ba-9d88-4a7e-90d9-b1e21b4f640b","_cell_guid":"dfd1141a-c893-4153-ab1e-c1fd15d565c8","trusted":true}},{"cell_type":"code","source":"onehot_df = pd.DataFrame(index=df.index)\n\nfor c in [\"Pclass\",\"Sex\",\"Embarked\",\"Deck\",\"Ticket_p\",\"Surname\",\"Title\"]:\n    encoded = OneHotEncoder().fit_transform(df[c].to_numpy().reshape(-1,1)).toarray()\n    columns = [f\"{c}_{i}\" for i in range(encoded.shape[1])]\n    _df =pd.DataFrame(data=encoded, columns=columns, index=df.index)\n    onehot_df = pd.concat([_df,onehot_df], axis=1)\n    \nonehot_df = pd.concat([onehot_df,df[[\"Survived\",\"Age\",\"SibSp\",\"Parch\",\"Fare\",\"Cabin\",\"Namesakes\",\"Kid\",\"Alone\",\"Side\"]]], axis=1)\n\nfor c in [\"Age\",\"Fare\",\"Cabin\",\"SibSp\",\"Parch\"]:\n    onehot_df[c] = MinMaxScaler().fit_transform(onehot_df[c].to_numpy().reshape(-1,1))","metadata":{"_uuid":"195eeae4-3849-4c78-986f-1fb067302623","_cell_guid":"741ab149-49a4-4e42-90ee-7e9446c5ad4d","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"onehot_df.head(10).style.background_gradient(cmap=\"Blues\")","metadata":{"_uuid":"d64e0ab9-8ec0-42f9-b10d-0ad92384a2f6","_cell_guid":"21dec68b-ecda-42fb-860e-93bcf2d12e26","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Resulting matrix looks like too spars, because of Surnames features, which should provide us some groups of survivals.\nNow we gonna try to use it for the baseline. In the next versions, i'll try some other configurations","metadata":{"_uuid":"e82c090e-cfe3-45be-9955-4cd44eda79de","_cell_guid":"1d0d9573-65ff-4c3f-ad8f-003a86d9332a","trusted":true}},{"cell_type":"markdown","source":"# <a class=\"anchor\" id=\"3_bullet\" style=\"color:#292F55\">  3. Model development","metadata":{"_uuid":"cac35a7f-7e66-4bf1-9bc3-a9323df4e029","_cell_guid":"f371a517-365b-4bd6-901b-216b8b7ddc57","trusted":true}},{"cell_type":"code","source":"df_train = onehot_df.copy(deep=True)\nmask = df_train[\"Survived\"].isna()\ntrain, deploy = df_train[~mask], df_train[mask]\ndeploy = deploy.drop(\"Survived\", axis=1)\ntrain.loc[:,\"Survived\"] = train.loc[:,\"Survived\"].astype(bool)\nx_train, y_train = train.drop(\"Survived\", axis=1), train[\"Survived\"].astype(int)","metadata":{"_uuid":"372d7b8a-4901-4d3d-ad75-58b7ff0c5ca9","_cell_guid":"cae6f771-7cc4-4a45-9def-a0af1a3e9fb7","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <a class=\"anchor\" id=\"3.1_bullet\" style=\"color:#3B61A3\">  3.1 Models baselines","metadata":{"_uuid":"32d4dcbc-9174-4d5c-8ac3-c218846ade32","_cell_guid":"7ec678c2-9cdb-4916-bf99-7ee97ef30d02","trusted":true}},{"cell_type":"markdown","source":"We will define the method, which fits the given model and puts all the neccesery data into different structures to use them further:\n\n* deploy_acc  - model accuracy for deploy dataset\n* train_acc   - model accuracy after the validation\n* models_dict - dictionary of the models\n","metadata":{}},{"cell_type":"code","source":"deploy_acc, train_acc, models_dict= {},{},{}\n\ndef baseline(name, model, verbose=True):\n    models_dict[name] = model\n    models_dict[name].fit(x_train,y_train)\n    y_train_hat = models_dict[name].predict(x_train)\n    train_acc[name] = accuracy_score(y_train,y_train_hat)\n    if verbose:\n        skplt.metrics.plot_confusion_matrix(y_train, y_train_hat, normalize=True, figsize=(5,5))\n    submition = pd.DataFrame(models_dict[name].predict(deploy), index= deploy.index,columns = [\"Survived\"]).astype(int)\n    submition.to_csv(f'{name}.csv')","metadata":{"_uuid":"e3a1d109-d844-4dfe-afd9-d846ecd8b996","_cell_guid":"d8a87f4f-f972-464a-9d28-c482878d48be","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I've tried several Gridsearches for CVs with Kfold. And after a bunch of deploys, i chose those models.\nI chose models which gave me the best deploy accuracy. This models gonna be not \"Best model for Survival prediction\" , but \"The best model to predict the test set\", which is not the exect solution for the task.","metadata":{"_uuid":"596e662e-0691-49c7-bbd6-e9c6bec6c28e","_cell_guid":"803c1c5c-f248-4f4a-8ba8-7ac0a210a7d6","trusted":true}},{"cell_type":"code","source":"params = {\"penalty\":\"l2\",\"solver\": \"liblinear\",\"C\":0.2,}\nname, model = \"lr_baseline\", LogisticRegression(**params)\nbaseline(name, model, verbose=False)\n\nname, model = \"svm_baseline\", SVC(**{'C': 5, 'degree': 2, 'gamma': 0.1, 'kernel': 'poly'})\nbaseline(name, model, verbose=False)\n\nparams = {\"eta\":0.1,\"gamma\":0,\"max_depth\":6,\"lambda\":0.1,\"alpha\":10}\nname, model = \"xg_baseline\", XGBClassifier(**params)\nbaseline(name, model, verbose=False)\n\nparams = {\"rsm\":0.1, \"learning_rate\":0.005,\"iterations\":500,\"l2_leaf_reg\":5,\"verbose\":False}\nname, model = \"cb_baseline\", CatBoostClassifier(**params)\nbaseline(name, model, verbose=False)","metadata":{"_uuid":"53c36f6c-fbf3-4fe9-bd3a-54e6d02961ba","_cell_guid":"61e0bb7c-72ac-4372-a7e4-64ace9513b52","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = plt.figure(figsize=(35,10))\ncount=1\nfor k in models_dict:\n    ax = fig.add_subplot(1,len(models_dict),count)\n    count+=1\n    skplt.metrics.plot_confusion_matrix(y_train, models_dict[k].predict(x_train), normalize=True, figsize=(5,5),ax=ax, cmap=\"Blues\")\n    ax.set_title(k)\nplt.show()","metadata":{"_uuid":"5a56d86c-253d-40f2-9e2f-8bb85e1e044e","_cell_guid":"4350c146-dfbd-47ff-a463-6f868515222a","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"deploy_acc[\"lr_baseline\"]  = 0.79904\ndeploy_acc[\"svm_baseline\"] = 0.80382\ndeploy_acc[\"xg_baseline\"]  = 0.78947\ndeploy_acc[\"cb_baseline\"]  = 0.79425 \nprint(\"Accuracy on the deployment set:\")\nfor k in deploy_acc:\n    print(f\"{k}\\t:\\t{deploy_acc[k]}\")","metadata":{"_uuid":"a9da422e-d5bb-4ab1-abb2-5a33dfd8f7da","_cell_guid":"df1dd3c2-e8d6-4d05-930a-edd816dc6b9e","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Accuracy on the training set:\")\nfor k in train_acc:\n    print(f\"{k}\\t:\\t{train_acc[k]}\")","metadata":{"_uuid":"65abf3ed-70f8-4214-b727-a36f96f6eb41","_cell_guid":"aa6b9012-0d91-4e79-9db5-d038ab503fd2","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <a class=\"anchor\" id=\"3.2_bullet\" style=\"color:#3B61A3\"> 3.2 Ensemble models","metadata":{"_uuid":"b10c79cf-8ce9-4ac9-a178-dd02af01ecc4","_cell_guid":"9ac22ede-a999-4e1a-9bd7-1b1a02bd6549","trusted":true}},{"cell_type":"code","source":"name, model = \"ensemble\", StackingClassifier(classifiers=(models_dict[\"svm_baseline\"],models_dict[\"lr_baseline\"],\n                                                                    models_dict[\"xg_baseline\"], models_dict[\"cb_baseline\"]),\n                                               meta_classifier=LogisticRegression(**{\"penalty\":\"l2\",\"solver\": \"liblinear\",\"C\":0.2,}),\n                                               use_features_in_secondary=True)\nbaseline(name, model, verbose=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"deploy_acc[\"ensemble\"]  = 0.80622","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see, that all 5 models has different sepations on the data (even on the trains set)\nNow we gonna combine this models into peculiar ensemble.\nWe gonna put som waights, according to there deploy results.","metadata":{"_uuid":"20cda916-1671-4cd8-8318-c8c31e5f9620","_cell_guid":"44599150-a833-4e1f-84a2-a2ae3e8746e7","trusted":true}},{"cell_type":"code","source":"ens_train, ens_deploy = {}, {}\nfor k in models_dict:\n    ens_train[k] = models_dict[k].predict(x_train) * deploy_acc[k]\n    ens_deploy[k] = models_dict[k].predict(deploy) * deploy_acc[k]\nx_train = pd.concat([pd.DataFrame(ens_train, index=x_train.index),x_train], axis=1)        \ndeploy = pd.concat([pd.DataFrame(ens_deploy, index=deploy.index),deploy], axis=1)","metadata":{"_uuid":"2b1f1ed5-35ba-41a8-8435-58dbee39beca","_cell_guid":"098c489a-2ac5-4b43-b459-6fcc5032dc58","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = SVC()\nmodel.fit(x_train,y_train)\nsubmition = pd.DataFrame({\"PassengerId\":deploy.index,\"Survived\":model.predict(deploy)}).astype(int)","metadata":{"_uuid":"7da65375-1506-49d1-a4a7-ec6d5baed188","_cell_guid":"6c11288c-b221-4f04-809f-27482fdafda8","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submition.to_csv(\"mixture.csv\",index=False) #0.81100","metadata":{"_uuid":"4828e092-33fb-418b-8771-5abba83a2a35","_cell_guid":"fb680e71-e00e-4f85-b763-ea12d6450190","collapsed":false,"jupyter":{"outputs_hidden":false},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![Upvote!](https://img.shields.io/badge/Upvote-If%20you%20like%20my%20work-07b3c8?style=for-the-badge&logo=kaggle)","metadata":{}}]}