{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center>Spaceship Titanic</center>\n### <center>RandomForestClassifier</center>","metadata":{}},{"cell_type":"markdown","source":"## Оглавление\n - [Импорт библиотек](#Импорт-библиотек)\n - [Описание данных](#File-and-Data-Field-Descriptions(Описание-данных))\n - [Загрузка и предпросмотр данных](#Загрузка-и-предпросмотр-данных)\n - [Функции для обработки данных](#Функции-для-обработки-данных)\n     - [Функция преобразования/разбиения столбцов с данными на категории с численными значениями](#Функция-преобразования/разбиения-столбцов-с-данными-на-категории-с-численными-значениями)\n     - [Функция разбиения текстовых данных по символу](#Функция-разбиения-текстовых-данных-по-символу)\n     - [Функция конечного преобразования данных](#Функция-конечного-преобразования-данных)\n - [Обработка данных(перевод в численные значения)](#Обработка-данных(перевод-в-численные-значения))\n     - [Переход к численному виду данных](#Переход-к-численному-виду-данных)\n     - [Проверка корреляции данных](#Проверка-корреляции-данных)\n - [Пробное обучение](#Пробное-обучение)\n - [Подбор атрибутов и параметров модели](#Подбор-атрибутов-и-параметров-модели)\n     - [Важность атрибутов для RF](#Важность-атрибутов-для-RF)\n     - [Подбор параметров для модели](#Подбор-параметров-для-модели)\n - [Финальное обучение](#Финальное-обучение) ","metadata":{}},{"cell_type":"markdown","source":"### Импорт библиотек\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\n\nimport missingno as msno\n\n\nfrom sklearn import model_selection\n\n\nfrom sklearn.preprocessing import LabelEncoder \nfrom sklearn.preprocessing import OneHotEncoder \nfrom sklearn import preprocessing\n\n\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import classification_report\n\n\nfrom sklearn.model_selection import RandomizedSearchCV\nfrom sklearn.model_selection import GridSearchCV\n\n\nfrom sklearn.ensemble import RandomForestClassifier\n\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:42:29.086229Z","iopub.execute_input":"2022-07-14T10:42:29.086711Z","iopub.status.idle":"2022-07-14T10:42:29.095331Z","shell.execute_reply.started":"2022-07-14T10:42:29.086678Z","shell.execute_reply":"2022-07-14T10:42:29.093693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install missingno","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### File and Data Field Descriptions(Описание данных)\n[Обратно к оглавлению](#Оглавление)\n\n   __train.csv__ - Personal records for about two-thirds (~8700) of the passengers, to be used as training data.\n1. __PassengerId__ - A unique Id for each passenger. Each Id takes the form gggg_pp where gggg indicates a group the passenger is travelling with and pp is their number within the group. People in a group are often family members, but not always.\n2. __HomePlanet__ - The planet the passenger departed from, typically their planet of permanent residence.\n3. __CryoSleep__ - Indicates whether the passenger elected to be put into suspended animation for the duration of the voyage. Passengers in cryosleep are confined to their cabins.\n4. __Cabin__ - The cabin number where the passenger is staying. Takes the form deck/num/side, where side can be either P for Port or S for Starboard.\n5. __Destination__ - The planet the passenger will be debarking to.\n6. __Age__ - The age of the passenger.\n7. __VIP__ - Whether the passenger has paid for special VIP service during the voyage.\n8. __RoomService, FoodCourt, ShoppingMall, Spa, VRDeck__ - Amount the passenger has billed at each of the Spaceship Titanic's many luxury amenities.\n9. __Name__ - The first and last names of the passenger.\n10. __Transported__ - Whether the passenger was transported to another dimension. This is the target, the column you are trying to predict.\n\n\n   __test.csv__ - Personal records for the remaining one-third (~4300) of the passengers, to be used as test data. Your task is to predict the value of Transported for the passengers in this set.\n   \n   \n   __sample_submission.csv__ - A submission file in the correct format.\nPassengerId - Id for each passenger in the test set.\nTransported - The target. For each passenger, predict either True or False.","metadata":{}},{"cell_type":"markdown","source":"### Загрузка и предпросмотр данных\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"result_data = pd.read_csv('../input/spaceship-titanic/sample_submission.csv')\ntest_data = pd.read_csv('../input/spaceship-titanic/test.csv')\nmain_data = pd.read_csv('../input/spaceship-titanic/train.csv')\n\nmain_data","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:42:31.261543Z","iopub.execute_input":"2022-07-14T10:42:31.262443Z","iopub.status.idle":"2022-07-14T10:42:31.337288Z","shell.execute_reply.started":"2022-07-14T10:42:31.262395Z","shell.execute_reply":"2022-07-14T10:42:31.336214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"main_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:42:33.888131Z","iopub.execute_input":"2022-07-14T10:42:33.889420Z","iopub.status.idle":"2022-07-14T10:42:33.908392Z","shell.execute_reply.started":"2022-07-14T10:42:33.889375Z","shell.execute_reply":"2022-07-14T10:42:33.907073Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.matrix(main_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:42:35.433998Z","iopub.execute_input":"2022-07-14T10:42:35.435177Z","iopub.status.idle":"2022-07-14T10:42:36.037084Z","shell.execute_reply.started":"2022-07-14T10:42:35.435125Z","shell.execute_reply":"2022-07-14T10:42:36.035429Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"msno.matrix(test_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:42:36.258553Z","iopub.execute_input":"2022-07-14T10:42:36.259715Z","iopub.status.idle":"2022-07-14T10:42:36.800420Z","shell.execute_reply.started":"2022-07-14T10:42:36.259658Z","shell.execute_reply":"2022-07-14T10:42:36.798997Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Функции для обработки данных\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"markdown","source":"#### Функция преобразования/разбиения столбцов с данными на категории с численными значениями\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"def category_data(data_in, columns = None, how = 'one_column', destroy = False):\n    \"\"\"Разбивает текстовые(с типом object или category) столбцы на категории с численными значениями\n        - data_in - входные данные pd.DataFrame\n        - columns - столбцы которые нужно перевести в категории(если не указан, то будут проверены все столбцы)\n        - how - тип разбиения на категории \n                    'bool' - создает n столбцов(категорий) на основе данного со значениями 1 или 0,\n                             где n - число уникальных значений в столбце,\n                             каждому значению соответствует свой столбец,\n                             значение 1 - соответстует тому что в данной строчке изначального столбца было данное значение\n                    'one column' - ранжирование значения в одном столбце по возрастанию(1,2,3,4,5,6,7...),\n                                   одинаковые значения буд равны.\n                    'parity' - меняет все численные значения в столбце на 0 - если число нечетное и 1 - если четное.\n        - destroy - удалять ли изначальный столбец(актуально только для 'bool')\"\"\"\n    def sort_value(data, column):\n        \"\"\"Сортирует pd.DataFrame(data) по колонке(column) и возвращает весь pd.DataFrame\"\"\"\n        if data[column].dtype == 'category' or data[column].dtype == 'object':\n            data[column] = data[column].astype(object)\n            data = data.sort_values(column)\n            data[column] = data[column].astype('category')\n        return data\n    def one_column(data, column):\n        \"\"\"Разбивает на категории в рамках одного столбца\"\"\"\n        if data[column].dtype == 'category' or data[column].dtype == 'object':\n            data = sort_value(data, column)\n\n            data[column] = pd.Categorical(data[column])\n            data[column] = data[column].cat.codes\n\n            print(f\"Столбец {column} теперь переведен в категории там {len(data[column].value_counts())} категорий\")\n        return data\n    def bool_column(data, column):\n        \"\"\"Разбивает на логические категории/столбцы со значения 0 или 1 в ячейках.\"\"\"\n        \n        array_encoded = np.array(data[column])\n        array_encoded = array_encoded.reshape(len(array_encoded), 1)\n        \n        onehot_encoder = OneHotEncoder(sparse=False)\n        onehot_encoded = onehot_encoder.fit_transform(array_encoded)\n        \n        new_columns = []\n        for c in onehot_encoder.get_feature_names_out():\n            new_columns.append(column + ' ' + c[3:])\n\n        new_columns = pd.DataFrame(onehot_encoded, columns = new_columns)\n        \n        data = data.join(new_columns)\n                                               \n        if destroy:\n            del data[column]\n        print(f\"Столбец {column} разбит на {len(onehot_encoder.get_feature_names())} логических категорий/столбцов\")\n        return data\n    def parity(data, column):\n        \"\"\"Разбивает числовой столбец на два значения: четные числа и нечетные числа\"\"\"\n        if data[column].dtype == 'Int64' or data[column].dtype == 'Float64':\n            data[f'{column} parity'] = data[column].apply(lambda x:x % 2)\n            if destroy == True:\n                data = data.drop(column, axis = 1)\n        else:\n            print(f'Типа данных в ячейках {column} не Int64 или Float64')\n            return data\n        return data\n    \n    def main(data, column, how):\n        if how == 'one_column':\n            data = one_column(data, column)\n        elif how == 'bool':\n            data = bool_column(data, column)\n        elif how == 'parity':\n            data = parity(data, column)    \n        else:\n            print('Параметр how указан не верно, возможные значения: \\n    -bool\\n    -one_column')\n            return data\n        return data\n    data_out = data_in.copy()\n    \n    if columns == None or columns == 'all':\n        for col in data_in.columns:\n            data_out = main(data_out, col, how)\n    else:\n        for col in columns:\n            data_out = main(data_out, col, how)\n    return data_out","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:46:52.982206Z","iopub.execute_input":"2022-07-14T10:46:52.982649Z","iopub.status.idle":"2022-07-14T10:46:53.000435Z","shell.execute_reply.started":"2022-07-14T10:46:52.982610Z","shell.execute_reply":"2022-07-14T10:46:52.998939Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Функция разбиения текстовых данных по символу\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"def split_series(data, name_series, split_simbol, destroy = False):\n    end_data = data.copy()\n    for column in name_series:\n        max_len = max(data[column].str.split(split_simbol).str.len())\n        print(f'Максимальная длина разбиения столбца {column} -', max_len)\n        \n        columns_names = {}\n        \n        for i in range(int(max_len)):\n            columns_names[int(i)] = f'{column}_{i}'\n        end_data = end_data.join(data[column].str.split(split_simbol, expand=True))\n        end_data = end_data.rename(columns=columns_names)\n        \n        if destroy:\n            del end_data[column]\n    return end_data","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:46:54.032266Z","iopub.execute_input":"2022-07-14T10:46:54.033008Z","iopub.status.idle":"2022-07-14T10:46:54.039933Z","shell.execute_reply.started":"2022-07-14T10:46:54.032964Z","shell.execute_reply":"2022-07-14T10:46:54.039003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Функция конечного преобразования данных\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"def data_transform(data):\n    \n    data_out = data.copy()\n    \n    data_out = split_series(data_out, ['Cabin'], '/')\n    \n    #data_out = split_series(data_out, ['Name'], ' ')\n    #data_out['Name_1'] = data_out['Name_1'].apply(lambda x: str(x)[:1])\n    #data_out['Name_0'] = data_out['Name_0'].apply(lambda x: str(x)[:1])\n    #data_out = category_data(data_out, ['Name_0'], 'bool', destroy = False)\n    #data_out = category_data(data_out, ['Name_1'], 'bool', destroy = False)\n    #data_out = data_out.replace({'Name_1':dict(data_out.Name_1.value_counts()), 'Name_0':dict(data_out.Name_0.value_counts())})\n    #data_out = data_out.drop(['Name_0', 'Name_1'], axis = 1)\n\n    data_out = category_data(data_out, ['Cabin_0'], 'bool', destroy = True)\n    data_out['Cabin_1'] = data_out['Cabin_1'].fillna('0')\n    data_out['Cabin_1'] = data_out['Cabin_1'].astype('int64')\n    \n    data_out = category_data(data_out, ['Cabin_1'], 'parity')\n    data_out = category_data(data_out, ['Cabin_2'], 'bool', destroy = True)\n    \n    data_out = data_out.drop(['PassengerId', 'Name', 'Cabin'], axis = 1)\n    data_out = category_data(data_out, \n                             ['HomePlanet','CryoSleep','Destination','VIP'], \n                             'bool', \n                             destroy = True)\n    \n    # Заполняем пустые значения на среднее значение по столбцу\n    for column in data_out.columns:\n        data_out = data_out.replace({column:{-1: np.NaN}})\n        data_out[column] = data_out[column].fillna(int(data_out[column].mean()))\n        \n    data_out = data_out.drop(['CryoSleep False'], axis = 1)\n    \n    \n\n    #data_array = data_out.values\n    #min_max_scaler = preprocessing.MinMaxScaler()\n    #x_scaled = min_max_scaler.fit_transform(data_array)\n    #data_out = pd.DataFrame(x_scaled, columns = data_out.columns)\n    \n    \n    return data_out","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:40.393866Z","iopub.execute_input":"2022-07-14T10:53:40.394385Z","iopub.status.idle":"2022-07-14T10:53:40.405171Z","shell.execute_reply.started":"2022-07-14T10:53:40.394342Z","shell.execute_reply":"2022-07-14T10:53:40.403998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Обработка данных(перевод в численные значения)\n[Обратно к оглавлению](#Оглавление)\n","metadata":{}},{"cell_type":"markdown","source":"#### Переход к численному виду данных\n[Обратно к оглавлению](#Оглавление)\n","metadata":{}},{"cell_type":"code","source":"train_data = data_transform(main_data)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:41.359450Z","iopub.execute_input":"2022-07-14T10:53:41.359913Z","iopub.status.idle":"2022-07-14T10:53:41.518716Z","shell.execute_reply.started":"2022-07-14T10:53:41.359876Z","shell.execute_reply":"2022-07-14T10:53:41.517155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:42.433331Z","iopub.execute_input":"2022-07-14T10:53:42.433779Z","iopub.status.idle":"2022-07-14T10:53:42.475544Z","shell.execute_reply.started":"2022-07-14T10:53:42.433742Z","shell.execute_reply":"2022-07-14T10:53:42.474248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.columns","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:43.931020Z","iopub.execute_input":"2022-07-14T10:53:43.931474Z","iopub.status.idle":"2022-07-14T10:53:43.940647Z","shell.execute_reply.started":"2022-07-14T10:53:43.931439Z","shell.execute_reply":"2022-07-14T10:53:43.939149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Проверка корреляции данных\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"f, ax = plt.subplots(figsize=(18, 18))\ncorr = np.round_(train_data.corr(), decimals=2)\nsns.heatmap(corr,annot=False,cmap='RdYlGn',linewidths=0.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:45.710243Z","iopub.execute_input":"2022-07-14T10:53:45.711000Z","iopub.status.idle":"2022-07-14T10:53:46.610330Z","shell.execute_reply.started":"2022-07-14T10:53:45.710958Z","shell.execute_reply":"2022-07-14T10:53:46.609180Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data = train_data.drop(['Cabin_2 nan'], axis = 1)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:46.612120Z","iopub.execute_input":"2022-07-14T10:53:46.612733Z","iopub.status.idle":"2022-07-14T10:53:46.619914Z","shell.execute_reply.started":"2022-07-14T10:53:46.612694Z","shell.execute_reply":"2022-07-14T10:53:46.618627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f, ax = plt.subplots(figsize=(18, 18))\ncorr = np.round_(train_data.corr(), decimals=2)\nsns.heatmap(corr,annot=False,cmap='RdYlGn',linewidths=0.2)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:49.189224Z","iopub.execute_input":"2022-07-14T10:53:49.189722Z","iopub.status.idle":"2022-07-14T10:53:50.065550Z","shell.execute_reply.started":"2022-07-14T10:53:49.189672Z","shell.execute_reply":"2022-07-14T10:53:50.064137Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:50.067513Z","iopub.execute_input":"2022-07-14T10:53:50.068169Z","iopub.status.idle":"2022-07-14T10:53:50.086779Z","shell.execute_reply.started":"2022-07-14T10:53:50.068125Z","shell.execute_reply":"2022-07-14T10:53:50.085588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = train_data['Transported']\nx = train_data.drop('Transported', axis = 1)\n\ntrain_x, test_x, train_y, test_y = model_selection.train_test_split(x, y, test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:50.088555Z","iopub.execute_input":"2022-07-14T10:53:50.089020Z","iopub.status.idle":"2022-07-14T10:53:50.101807Z","shell.execute_reply.started":"2022-07-14T10:53:50.088974Z","shell.execute_reply":"2022-07-14T10:53:50.100551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Пробное обучение\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"model=RandomForestClassifier(n_estimators=100, max_depth = 10)\n\n# обучаем модель\nmodel.fit(train_x,train_y)\n\n\nmodel_pred = model.predict(test_x)\n\nprint(accuracy_score(test_y,model_pred))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:51.132684Z","iopub.execute_input":"2022-07-14T10:53:51.133565Z","iopub.status.idle":"2022-07-14T10:53:52.088722Z","shell.execute_reply.started":"2022-07-14T10:53:51.133496Z","shell.execute_reply":"2022-07-14T10:53:52.087422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_pred = model.predict(train_x)\n\nprint(accuracy_score(train_y,model_pred))","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:52.091137Z","iopub.execute_input":"2022-07-14T10:53:52.092046Z","iopub.status.idle":"2022-07-14T10:53:52.192274Z","shell.execute_reply.started":"2022-07-14T10:53:52.091996Z","shell.execute_reply":"2022-07-14T10:53:52.191138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Подбор атрибутов и параметров модели\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"markdown","source":"#### Важность атрибутов для RF\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"headers = list(x.columns.values)\n\nfeature_imp = pd.Series(model.feature_importances_,index=headers).sort_values(ascending=False)\n\nf, ax = plt.subplots(figsize=(16,16))\nsns.barplot(x=feature_imp, y=feature_imp.index)\n\nplt.xlabel('Важность атрибутов')\nplt.ylabel('Атрибуты')\nplt.title(\"Наиболее важные атрибуты\")\nplt.legend()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:52.493226Z","iopub.execute_input":"2022-07-14T10:53:52.493998Z","iopub.status.idle":"2022-07-14T10:53:52.966128Z","shell.execute_reply.started":"2022-07-14T10:53:52.493958Z","shell.execute_reply":"2022-07-14T10:53:52.964606Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feature_imp[feature_imp > 0.005]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:52.968513Z","iopub.execute_input":"2022-07-14T10:53:52.969258Z","iopub.status.idle":"2022-07-14T10:53:52.981233Z","shell.execute_reply.started":"2022-07-14T10:53:52.969194Z","shell.execute_reply":"2022-07-14T10:53:52.979752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Подбор параметров для модели\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"%%time\nn_estimators = [int(x) for x in np.linspace(start = 100, stop = 1000, num = 10)]\nmax_features = ['log2', 'sqrt']\nmax_depth = [int(x) for x in np.linspace(start = 1, stop = 40, num = 20)]\nmin_samples_split = [int(x) for x in np.linspace(start = 2, stop = 50, num = 10)]\nmin_samples_leaf = [int(x) for x in np.linspace(start = 2, stop = 50, num = 10)]\nbootstrap = [True, False]\nparam_dist = {'n_estimators': n_estimators,\n              'max_features': max_features,\n              'max_depth': max_depth,\n              'min_samples_split': min_samples_split,\n              'min_samples_leaf': min_samples_leaf,\n              'bootstrap': bootstrap}\nrs = RandomizedSearchCV(model, \n                        param_dist, \n                        n_iter = 100, \n                        cv = 3, \n                        verbose = 1, \n                        n_jobs=-1, \n                        random_state=0)\nrs.fit(train_x,train_y)\nrs.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:53:57.290010Z","iopub.execute_input":"2022-07-14T10:53:57.290427Z","iopub.status.idle":"2022-07-14T10:57:45.211207Z","shell.execute_reply.started":"2022-07-14T10:53:57.290394Z","shell.execute_reply":"2022-07-14T10:57:45.209783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rs_df = pd.DataFrame(rs.cv_results_).sort_values('rank_test_score').reset_index(drop=True)\nrs_df = rs_df.drop([\n            'mean_fit_time', \n            'std_fit_time', \n            'mean_score_time',\n            'std_score_time', \n            'params', \n            'split0_test_score', \n            'split1_test_score', \n            'split2_test_score', \n            'std_test_score'],\n            axis=1)\nrs_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:57:45.214183Z","iopub.execute_input":"2022-07-14T10:57:45.214607Z","iopub.status.idle":"2022-07-14T10:57:45.243295Z","shell.execute_reply.started":"2022-07-14T10:57:45.214566Z","shell.execute_reply":"2022-07-14T10:57:45.241905Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=3, nrows=2)\nsns.set(style=\"whitegrid\", color_codes=True, font_scale = 2)\nfig.set_size_inches(30,30)\nsns.barplot(x='param_n_estimators', y='mean_test_score', data=rs_df, ax=axs[0,0], color='lightgrey')\naxs[0,0].set_ylim([.725,.825])\naxs[0,0].set_title(label = 'n_estimators', size=30, weight='bold')\nsns.barplot(x='param_min_samples_split', y='mean_test_score', data=rs_df, ax=axs[0,1], color='coral')\naxs[0,1].set_ylim([.725,.825])\naxs[0,1].set_title(label = 'min_samples_split', size=30, weight='bold')\nsns.barplot(x='param_min_samples_leaf', y='mean_test_score', data=rs_df, ax=axs[0,2], color='lightgreen')\naxs[0,2].set_ylim([.725,.825])\naxs[0,2].set_title(label = 'min_samples_leaf', size=30, weight='bold')\nsns.barplot(x='param_max_features', y='mean_test_score', data=rs_df, ax=axs[1,0], color='wheat')\naxs[1,0].set_ylim([.725,.825])\naxs[1,0].set_title(label = 'max_features', size=30, weight='bold')\nsns.barplot(x='param_max_depth', y='mean_test_score', data=rs_df, ax=axs[1,1], color='lightpink')\naxs[1,1].set_ylim([.725,.825])\naxs[1,1].set_title(label = 'max_depth', size=30, weight='bold')\nsns.barplot(x='param_bootstrap',y='mean_test_score', data=rs_df, ax=axs[1,2], color='skyblue')\naxs[1,2].set_ylim([.725,.825])\naxs[1,2].set_title(label = 'bootstrap', size=30, weight='bold')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:58:34.602481Z","iopub.execute_input":"2022-07-14T10:58:34.602974Z","iopub.status.idle":"2022-07-14T10:58:37.024140Z","shell.execute_reply.started":"2022-07-14T10:58:34.602935Z","shell.execute_reply":"2022-07-14T10:58:37.022825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nn_estimators = [600]\nmax_depth = [19]\nmin_samples_split = [2,7,12,17,22,28,34]\nmin_samples_leaf = [1,4,7,10,12,14]\nbootstrap = [False]\nparam_grid = {'n_estimators': n_estimators,\n              'max_depth': max_depth,\n              'min_samples_split': min_samples_split,\n              'min_samples_leaf': min_samples_leaf,\n              'bootstrap': bootstrap}\ngs = GridSearchCV(model, param_grid, cv = 3, verbose = 1, n_jobs=-1)\ngs.fit(train_x, train_y)\nrfc_3 = gs.best_estimator_\ngs.best_params_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T10:58:52.384675Z","iopub.execute_input":"2022-07-14T10:58:52.385140Z","iopub.status.idle":"2022-07-14T11:01:31.782136Z","shell.execute_reply.started":"2022-07-14T10:58:52.385106Z","shell.execute_reply":"2022-07-14T11:01:31.781370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gs_df = pd.DataFrame(gs.cv_results_).sort_values('rank_test_score').reset_index(drop=True)\ngs_df = gs_df.drop([\n            'mean_fit_time', \n            'std_fit_time', \n            'mean_score_time',\n            'std_score_time', \n            'params', \n            'split0_test_score', \n            'split1_test_score', \n            'split2_test_score', \n            'std_test_score','param_bootstrap','param_n_estimators'],\n            axis=1)\ngs_df.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:31.784504Z","iopub.execute_input":"2022-07-14T11:01:31.784999Z","iopub.status.idle":"2022-07-14T11:01:31.806106Z","shell.execute_reply.started":"2022-07-14T11:01:31.784953Z","shell.execute_reply":"2022-07-14T11:01:31.804630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig, axs = plt.subplots(ncols=2, nrows=1)\nsns.set(style=\"whitegrid\", color_codes=True, font_scale = 2)\nfig.set_size_inches(30,10)\n\nsns.barplot(x='param_min_samples_split', y='mean_test_score', data=gs_df, ax=axs[0], color='coral')\naxs[0].set_ylim([.725,.825])\naxs[0].set_title(label = 'min_samples_split', size=30, weight='bold')\nsns.barplot(x='param_min_samples_leaf', y='mean_test_score', data=gs_df, ax=axs[1], color='lightgreen')\naxs[1].set_ylim([.725,.825])\naxs[1].set_title(label = 'min_samples_leaf', size=30, weight='bold')\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:31.809724Z","iopub.execute_input":"2022-07-14T11:01:31.810107Z","iopub.status.idle":"2022-07-14T11:01:32.497445Z","shell.execute_reply.started":"2022-07-14T11:01:31.810072Z","shell.execute_reply":"2022-07-14T11:01:32.496426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model=RandomForestClassifier(bootstrap= False,\n                             max_depth= 19,\n                             min_samples_leaf= 4,\n                             min_samples_split= 2,\n                             n_estimators= 600)\n\n# обучаем модель\nmodel.fit(train_x,train_y)\n\n\nmodel_pred = model.predict(test_x)\n\nprint(accuracy_score(test_y,model_pred))\n","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:32.499935Z","iopub.execute_input":"2022-07-14T11:01:32.500865Z","iopub.status.idle":"2022-07-14T11:01:37.650072Z","shell.execute_reply.started":"2022-07-14T11:01:32.500815Z","shell.execute_reply":"2022-07-14T11:01:37.648867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Финальное обучение\n[Обратно к оглавлению](#Оглавление)","metadata":{}},{"cell_type":"code","source":"y = train_data['Transported']\nx = train_data.drop('Transported', axis = 1)\nx = x[feature_imp[feature_imp > 0.005].index]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:37.651553Z","iopub.execute_input":"2022-07-14T11:01:37.652003Z","iopub.status.idle":"2022-07-14T11:01:37.663648Z","shell.execute_reply.started":"2022-07-14T11:01:37.651972Z","shell.execute_reply":"2022-07-14T11:01:37.662645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_data = test_data.join(result_data.Transported)\npredict_data = data_transform(predict_data)\npredict_data = predict_data[feature_imp[feature_imp > 0.005].index]\n","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:37.664987Z","iopub.execute_input":"2022-07-14T11:01:37.666158Z","iopub.status.idle":"2022-07-14T11:01:37.795582Z","shell.execute_reply.started":"2022-07-14T11:01:37.666102Z","shell.execute_reply":"2022-07-14T11:01:37.794158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_data.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:37.799208Z","iopub.execute_input":"2022-07-14T11:01:37.799632Z","iopub.status.idle":"2022-07-14T11:01:37.815185Z","shell.execute_reply.started":"2022-07-14T11:01:37.799596Z","shell.execute_reply":"2022-07-14T11:01:37.814192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model=RandomForestClassifier(bootstrap= False,\n                             max_depth= 19,\n                             min_samples_leaf= 4,\n                             min_samples_split= 2,\n                             n_estimators= 600)\n# обучаем модель\nmodel.fit(x,y)\n\n\nmodel_pred = model.predict(predict_data)\n","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:37.816470Z","iopub.execute_input":"2022-07-14T11:01:37.817557Z","iopub.status.idle":"2022-07-14T11:01:45.434648Z","shell.execute_reply.started":"2022-07-14T11:01:37.817487Z","shell.execute_reply":"2022-07-14T11:01:45.433512Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_p = pd.Series(model_pred == 1, name = 'Transported')\nmodel_p.value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:45.435941Z","iopub.execute_input":"2022-07-14T11:01:45.436276Z","iopub.status.idle":"2022-07-14T11:01:45.446124Z","shell.execute_reply.started":"2022-07-14T11:01:45.436247Z","shell.execute_reply":"2022-07-14T11:01:45.445057Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame(\n    {'PassengerId':test_data[\"PassengerId\"] ,\n     'Transported': model_p},columns=['PassengerId', 'Transported'])\n\nsubmission.to_csv(\"./submission.csv\", index=False)\nsubmission.head(20)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T11:01:45.478590Z","iopub.execute_input":"2022-07-14T11:01:45.478908Z","iopub.status.idle":"2022-07-14T11:01:45.499575Z","shell.execute_reply.started":"2022-07-14T11:01:45.478881Z","shell.execute_reply":"2022-07-14T11:01:45.498299Z"},"trusted":true},"execution_count":null,"outputs":[]}]}