{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <center> Dota 2 winner prediction\n\n<img src='https://habrastorage.org/webt/ua/vn/pq/uavnpqfoih4zwwznvxubu33ispy.jpeg'>","metadata":{"_uuid":"540935bf571b28452d474fd9191d8e1e20d87005"}},{"cell_type":"markdown","source":"## Data description\n\nWe have the following files:\n\n- `sample_submission.csv`: example of a submission file\n- `train_matches.jsonl`, `test_matches.jsonl`: full \"raw\" training data \n- `train_features.csv`, `test_features.csv`: features created by organizers\n- `train_targets.csv`: results of training games (including the winner)","metadata":{"_uuid":"4339c93b125e5ec10e2fb6e4d5d00892570a2911"}},{"cell_type":"markdown","source":"## Features created by organizers\n\nThese are basic features which include simple players' statistics. Scroll to the end to see how to build these features from raw json files.","metadata":{"_uuid":"2df6f18c884cc8dfa75b584d7f71f7d0e89db897"}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current sessio","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:36.293390Z","iopub.execute_input":"2022-07-14T12:29:36.293845Z","iopub.status.idle":"2022-07-14T12:29:36.301477Z","shell.execute_reply.started":"2022-07-14T12:29:36.293775Z","shell.execute_reply":"2022-07-14T12:29:36.300506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd\n\nPATH_TO_DATA = '../input/'\n\ndf_train_features = pd.read_csv('../input/train_features.csv', index_col='match_id_hash')\ndf_train_targets = pd.read_csv('../input/train_targets.csv', index_col='match_id_hash')","metadata":{"_uuid":"14467d4ccb21360c1d699d0275b87f234f30c961","execution":{"iopub.status.busy":"2022-07-14T12:29:36.303374Z","iopub.execute_input":"2022-07-14T12:29:36.304159Z","iopub.status.idle":"2022-07-14T12:29:38.074668Z","shell.execute_reply.started":"2022-07-14T12:29:36.304091Z","shell.execute_reply":"2022-07-14T12:29:38.073674Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# ## exercise, read test dataframe\n# df_test_features = \n# df_test_targets = ","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:38.075913Z","iopub.execute_input":"2022-07-14T12:29:38.076387Z","iopub.status.idle":"2022-07-14T12:29:38.079682Z","shell.execute_reply.started":"2022-07-14T12:29:38.076216Z","shell.execute_reply":"2022-07-14T12:29:38.078982Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have ~ 40k games, each described by `match_id_hash` (game id) and 245 features. Also `game_time` is given - time (in secs) when the game was over. ","metadata":{"_uuid":"e805d116c02fb349617562accd0b784e199f30dd"}},{"cell_type":"code","source":"df_train_features.shape","metadata":{"_uuid":"3b93c28dfd711f49e4f3a2374175f979547f4b67","execution":{"iopub.status.busy":"2022-07-14T12:29:38.080855Z","iopub.execute_input":"2022-07-14T12:29:38.081104Z","iopub.status.idle":"2022-07-14T12:29:38.094461Z","shell.execute_reply.started":"2022-07-14T12:29:38.081054Z","shell.execute_reply":"2022-07-14T12:29:38.093462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_features.head()","metadata":{"_uuid":"c5de043331cca15e917bc20e4bb8e6bed75ddcc9","execution":{"iopub.status.busy":"2022-07-14T12:29:38.095763Z","iopub.execute_input":"2022-07-14T12:29:38.096255Z","iopub.status.idle":"2022-07-14T12:29:38.236104Z","shell.execute_reply.started":"2022-07-14T12:29:38.096200Z","shell.execute_reply":"2022-07-14T12:29:38.235366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are interested in the `radiant_win` column in `train_targets.csv`. All these features are not known during the game (they come \"from future\" as compared to `game_time`), so we have these features only for training data. ","metadata":{"_uuid":"2ab261af9db00f17af31a68c950912c03ab40ae0"}},{"cell_type":"code","source":"df_train_targets.head()","metadata":{"_uuid":"bd4f390eb21611fb9ce7dd2379923238ba6c692b","execution":{"iopub.status.busy":"2022-07-14T12:29:38.237248Z","iopub.execute_input":"2022-07-14T12:29:38.237468Z","iopub.status.idle":"2022-07-14T12:29:38.257194Z","shell.execute_reply.started":"2022-07-14T12:29:38.237427Z","shell.execute_reply":"2022-07-14T12:29:38.256223Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_targets['radiant_win'] = df_train_targets['radiant_win'].astype(int)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:38.259285Z","iopub.execute_input":"2022-07-14T12:29:38.259708Z","iopub.status.idle":"2022-07-14T12:29:38.264617Z","shell.execute_reply.started":"2022-07-14T12:29:38.259490Z","shell.execute_reply":"2022-07-14T12:29:38.263764Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Training and evaluating a model","metadata":{"_uuid":"c25fbce112c73f64c0fd544c853084e5734de11d"}},{"cell_type":"markdown","source":"#### Let's construct a feature matrix `X` and a target vector `y`","metadata":{"_uuid":"c679d807c3609bf57320c331fefe93bffc72d948"}},{"cell_type":"code","source":"X = df_train_features.values\ny = df_train_targets['radiant_win'].values","metadata":{"_uuid":"37ca1c0b98702d304f90d462379fa62f03571ebc","execution":{"iopub.status.busy":"2022-07-14T12:29:38.265797Z","iopub.execute_input":"2022-07-14T12:29:38.266056Z","iopub.status.idle":"2022-07-14T12:29:38.337938Z","shell.execute_reply.started":"2022-07-14T12:29:38.266000Z","shell.execute_reply":"2022-07-14T12:29:38.337080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Perform  a train/test split (a simple validation scheme)","metadata":{"_uuid":"135d1642aa2fb6865f874b0f64689b5cb6037491"}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nX_train, X_val, y_train, y_val = train_test_split(X, y, \n                                                  test_size=0.2, \n                                                  random_state=17)","metadata":{"_uuid":"58485053cb79d517d709ceaa3e6e31d1d0b8e0f2","execution":{"iopub.status.busy":"2022-07-14T12:29:38.340861Z","iopub.execute_input":"2022-07-14T12:29:38.341226Z","iopub.status.idle":"2022-07-14T12:29:38.558606Z","shell.execute_reply.started":"2022-07-14T12:29:38.341154Z","shell.execute_reply":"2022-07-14T12:29:38.557774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Train the Random Forest model","metadata":{"_uuid":"1062cc9f81b5ef07b03385994fd3d6bbe2ade8bd"}},{"cell_type":"markdown","source":"<img src='https://www.baeldung.com/wp-content/uploads/sites/4/2022/03/decision_tree1.jpg'>\n\nhttps://www.youtube.com/watch?v=cIbj0WuK41w","metadata":{}},{"cell_type":"markdown","source":"Most important hyperparameters of Random Forest:\n\n- n_estimators = n of trees\n- max_features = max number of features considered for splitting a node\n- max_depth = max number of levels in each decision tree\n- min_samples_split = min number of data points placed in a node before the node is split\n- min_samples_leaf = min number of data points allowed in a leaf node\n- bootstrap = method for sampling data points (with or without replacement)","metadata":{}},{"cell_type":"code","source":"%%time\nfrom sklearn.ensemble import RandomForestClassifier\nmodel = RandomForestClassifier(\n                               n_estimators=100, \n                               max_features=5,\n                               max_depth=5,\n                               min_samples_split=10,\n                               min_samples_leaf=10,\n                               n_jobs=-1, \n                               random_state=17\n                              )\nmodel.fit(X_train, y_train)","metadata":{"_uuid":"98687be065f843aa0dd201cb9d20b797584c38f2","execution":{"iopub.status.busy":"2022-07-14T12:29:38.560066Z","iopub.execute_input":"2022-07-14T12:29:38.560344Z","iopub.status.idle":"2022-07-14T12:29:39.950963Z","shell.execute_reply.started":"2022-07-14T12:29:38.560292Z","shell.execute_reply":"2022-07-14T12:29:39.950167Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Make predictions for the holdout set\n\nWe need to predict probabilities of class 1 - that Radiant wins, thus we need index 1 in the matrix returned by the `predict_proba` method.","metadata":{"_uuid":"8b1799aff8d7493557d9bede120ea67af00b14f6"}},{"cell_type":"code","source":"y_pred = model.predict_proba(X_val)[:, 1]","metadata":{"_uuid":"bc1b7df6bf6bbd7910a947de14e3b7d8e2890e1a","execution":{"iopub.status.busy":"2022-07-14T12:29:39.951884Z","iopub.execute_input":"2022-07-14T12:29:39.952094Z","iopub.status.idle":"2022-07-14T12:29:40.061329Z","shell.execute_reply.started":"2022-07-14T12:29:39.952061Z","shell.execute_reply":"2022-07-14T12:29:40.060480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look:","metadata":{"_uuid":"54947aa18a6e9ab0209fc9134085235bd6a6d810"}},{"cell_type":"code","source":"y_pred","metadata":{"_uuid":"ce5d9759ac0bda3920ed526e2d61a7369a3dca1e","execution":{"iopub.status.busy":"2022-07-14T12:29:40.062534Z","iopub.execute_input":"2022-07-14T12:29:40.062785Z","iopub.status.idle":"2022-07-14T12:29:40.068361Z","shell.execute_reply.started":"2022-07-14T12:29:40.062738Z","shell.execute_reply":"2022-07-14T12:29:40.067659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Let's evaluate prediction quality with the holdout set","metadata":{"_uuid":"0195d46c9fc6fc2ab1ab263ab560ddcbfd4bfea7"}},{"cell_type":"code","source":"from sklearn.metrics import accuracy_score, precision_score, recall_score, f1_score, confusion_matrix, roc_auc_score","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.069631Z","iopub.execute_input":"2022-07-14T12:29:40.069888Z","iopub.status.idle":"2022-07-14T12:29:40.079100Z","shell.execute_reply.started":"2022-07-14T12:29:40.069838Z","shell.execute_reply":"2022-07-14T12:29:40.078326Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Out if curiosiry, we can calculate accuracy of a classifier which predicts class 1 if predicted probability is higher than 50%. ","metadata":{"_uuid":"0741b9f8ef727db4f0049af76c672c7d7ab94aa0"}},{"cell_type":"code","source":"valid_accuracy = accuracy_score(y_val, y_pred > 0.5)\nprint('Validation accuracy of P>0.5 classifier:', valid_accuracy)","metadata":{"_uuid":"fb163a80ec9fc92bc2c6356a9193a5356596a42a","execution":{"iopub.status.busy":"2022-07-14T12:29:40.080545Z","iopub.execute_input":"2022-07-14T12:29:40.080809Z","iopub.status.idle":"2022-07-14T12:29:40.092308Z","shell.execute_reply.started":"2022-07-14T12:29:40.080759Z","shell.execute_reply":"2022-07-14T12:29:40.091475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"confusion_matrix(y_val, y_pred > 0.5)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.093217Z","iopub.execute_input":"2022-07-14T12:29:40.093429Z","iopub.status.idle":"2022-07-14T12:29:40.142449Z","shell.execute_reply.started":"2022-07-14T12:29:40.093396Z","shell.execute_reply":"2022-07-14T12:29:40.141426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"precision_score(y_val, y_pred > 0.5)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.143948Z","iopub.execute_input":"2022-07-14T12:29:40.144174Z","iopub.status.idle":"2022-07-14T12:29:40.153989Z","shell.execute_reply.started":"2022-07-14T12:29:40.144140Z","shell.execute_reply":"2022-07-14T12:29:40.153103Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"recall_score(y_val, y_pred > 0.5)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.155407Z","iopub.execute_input":"2022-07-14T12:29:40.155663Z","iopub.status.idle":"2022-07-14T12:29:40.164033Z","shell.execute_reply.started":"2022-07-14T12:29:40.155620Z","shell.execute_reply":"2022-07-14T12:29:40.163173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f1_score(y_val, y_pred > 0.5)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.165369Z","iopub.execute_input":"2022-07-14T12:29:40.165634Z","iopub.status.idle":"2022-07-14T12:29:40.176518Z","shell.execute_reply.started":"2022-07-14T12:29:40.165585Z","shell.execute_reply":"2022-07-14T12:29:40.175517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A confusion matrix is a tool designed to help us understand a little better how well our classifier is performing. An *accuracy score*, like that returned by kaggle for our submission file, lets us know a number indicating what ratio of predictions were correct (0 is not one classification was correct, and 1 is perfect!). The confusion matrix does the same thing, but goes into a little more detail; this time it provides us with four values:\n* The number of times our classifier produced **true negatives** (TN) the model correctly predicts the negatives class\n* The number of times our classifier produced **true positives** (TP) the model correctly predicts the positive class\n* The number of times our classifier produced **false positives** (FP), a type I error the model incorrectly predicts the positive class\n* The number of times our classifier produced **false negatives** (FN), a type II error the model incorrectly predicts the negatives class\n\nwhich scikit-learn returns in the following format, hence the name matrix (note that there is no standard convention for arrangement of this matrix):\n\n<img src='https://miro.medium.com/max/1400/1*xMl_wkMt42Hy8i84zs2WGg.png'>\n\n\nThe *accuracy* is given by $\\frac{(TN + TP)}{(TN + TP + FP +FN)}$, in other words, the true values divided by all the values. And finally, another measure one may come across is the **$F_1$ score**, which is given by:\n\n$$ F_1 = 2\\frac{precision . recall}{precision + recall}$$\n\n\nwhere the *precision* is given by $\\frac{TP}{TP + FP}$, and *recall* by $\\frac{TP}{TP + FN}$.\n\nThese Wikipedia pages have excellent descriptions of the meaning of these terms: \n* [Confusion matrix](https://en.wikipedia.org/wiki/Confusion_matrix)\n* [False positives and false negatives](https://en.wikipedia.org/wiki/False_positives_and_false_negatives)\n* [Type I and type II errors](https://en.wikipedia.org/wiki/Type_I_and_type_II_errors)\n* [Receiver operating characteristic](https://en.wikipedia.org/wiki/Receiver_operating_characteristic)\n* [F1 score](https://en.wikipedia.org/wiki/F1_score)\n","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preparing a submission\n\nNow the same for test data.","metadata":{"_uuid":"68b784d190c7fccd394e67b7bc20f6dc58c57c5d"}},{"cell_type":"code","source":"df_test_features = pd.read_csv('../input/test_features.csv', index_col='match_id_hash')\n\nX_test = df_test_features.values\ny_test_pred = model.predict_proba(X_test)[:, 1]\n\ndf_submission = pd.DataFrame({'radiant_win_prob': y_test_pred}, \n                                 index=df_test_features.index)","metadata":{"_uuid":"bf9749b3551b6bcf1e8ce4bea73c32ca5bde1c32","execution":{"iopub.status.busy":"2022-07-14T12:29:40.177734Z","iopub.execute_input":"2022-07-14T12:29:40.177983Z","iopub.status.idle":"2022-07-14T12:29:40.706384Z","shell.execute_reply.started":"2022-07-14T12:29:40.177940Z","shell.execute_reply":"2022-07-14T12:29:40.705478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission.head()","metadata":{"_uuid":"8c13fa4c8690e51045e59040f4ef9a621434716c","execution":{"iopub.status.busy":"2022-07-14T12:29:40.708026Z","iopub.execute_input":"2022-07-14T12:29:40.708381Z","iopub.status.idle":"2022-07-14T12:29:40.724031Z","shell.execute_reply.started":"2022-07-14T12:29:40.708299Z","shell.execute_reply":"2022-07-14T12:29:40.723101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Save the submission file, it's handy to include current datetime in the filename. ","metadata":{"_uuid":"438f19efd3f319b862b73c908e91317f5acf1df2"}},{"cell_type":"code","source":"df_submission.to_csv('submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:40.725454Z","iopub.execute_input":"2022-07-14T12:29:40.725736Z","iopub.status.idle":"2022-07-14T12:29:40.775285Z","shell.execute_reply.started":"2022-07-14T12:29:40.725688Z","shell.execute_reply":"2022-07-14T12:29:40.774351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Cross-validation\n\n<img src='https://linzhenyuyuchen.github.io/img/grid_search_cross_validation.png'>","metadata":{"_uuid":"7aa8cfa79efd9253299c1bc52e1778a3660c7ec6"}},{"cell_type":"code","source":"from sklearn.model_selection import KFold\nn_fold = 3\ncv = KFold(n_splits=n_fold, random_state=17)","metadata":{"_uuid":"c19342d24b586d208d4aa167e6730f1c531dca32","execution":{"iopub.status.busy":"2022-07-14T12:29:40.776597Z","iopub.execute_input":"2022-07-14T12:29:40.777023Z","iopub.status.idle":"2022-07-14T12:29:40.781315Z","shell.execute_reply.started":"2022-07-14T12:29:40.776951Z","shell.execute_reply":"2022-07-14T12:29:40.780317Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import cross_val_score","metadata":{"_uuid":"dacc8bc320a3417b30fceff7e57d8dd954a9b138","execution":{"iopub.status.busy":"2022-07-14T12:29:40.782634Z","iopub.execute_input":"2022-07-14T12:29:40.782998Z","iopub.status.idle":"2022-07-14T12:29:40.799009Z","shell.execute_reply.started":"2022-07-14T12:29:40.782942Z","shell.execute_reply":"2022-07-14T12:29:40.797619Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Run cross-validation\n\nWe'll train 2 versions of the  `RandomForestClassifier` model - first with default capacity (trees are not limited in depth), second - with `min_samples_leaf`=3, i.e. each leave is obliged to have at least 3 instances. ","metadata":{"_uuid":"7e5e235b8d0b084fe5c3c3354e55eda59047f3fc"}},{"cell_type":"code","source":"%%time\n\nmodel_rf_cv = RandomForestClassifier(\n                               n_estimators=100, \n                               max_features=5,\n                               max_depth=5,\n                               min_samples_split=10,\n                               min_samples_leaf=10,\n                               n_jobs=-1, \n                               random_state=17\n                              )\n\n# calcuate ROC-AUC for each split\ncv_scores_rf = cross_val_score(model_rf_cv, X, y, cv=cv, scoring='accuracy')","metadata":{"_uuid":"0abb5505ab9dc21c389235fde06b16e1faca4ec5","execution":{"iopub.status.busy":"2022-07-14T12:29:40.800579Z","iopub.execute_input":"2022-07-14T12:29:40.801157Z","iopub.status.idle":"2022-07-14T12:29:45.743456Z","shell.execute_reply.started":"2022-07-14T12:29:40.801093Z","shell.execute_reply":"2022-07-14T12:29:45.742676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_scores_rf","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:45.744869Z","iopub.execute_input":"2022-07-14T12:29:45.745350Z","iopub.status.idle":"2022-07-14T12:29:45.750543Z","shell.execute_reply.started":"2022-07-14T12:29:45.745294Z","shell.execute_reply":"2022-07-14T12:29:45.749708Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Model 1 mean score:', cv_scores_rf.mean())","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:45.752473Z","iopub.execute_input":"2022-07-14T12:29:45.752725Z","iopub.status.idle":"2022-07-14T12:29:45.762652Z","shell.execute_reply.started":"2022-07-14T12:29:45.752684Z","shell.execute_reply":"2022-07-14T12:29:45.761900Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions = np.zeros(len(X_test))\naverage_accuracy = 0\n\nfor train_index, val_index in cv.split(X):\n    X_train_cv, X_val_cv = X[train_index], X[val_index]\n    y_train_cv, y_val_cv = y[train_index], y[val_index]\n        \n    model_rf_cv.fit(X_train_cv, y_train_cv)\n    \n    y_pred = model.predict_proba(X_val_cv)[:, 1]\n    \n    valid_accuracy = accuracy_score(y_val_cv, y_pred > 0.5)\n    \n    average_accuracy = average_accuracy + valid_accuracy\n    \n    predictions += model.predict_proba(X_test)[:, 1]\n    \npredictions = predictions / n_fold\naverage_accuracy = average_accuracy / n_fold","metadata":{"_uuid":"0fab6c78d32dc46c36fee562e7edffe503e35d12","execution":{"iopub.status.busy":"2022-07-14T12:32:32.703498Z","iopub.execute_input":"2022-07-14T12:32:32.703821Z","iopub.status.idle":"2022-07-14T12:32:37.106898Z","shell.execute_reply.started":"2022-07-14T12:32:32.703777Z","shell.execute_reply":"2022-07-14T12:32:37.106075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"average_accuracy","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:32:42.887884Z","iopub.execute_input":"2022-07-14T12:32:42.888193Z","iopub.status.idle":"2022-07-14T12:32:42.894736Z","shell.execute_reply.started":"2022-07-14T12:32:42.888142Z","shell.execute_reply":"2022-07-14T12:32:42.893622Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:50.082964Z","iopub.execute_input":"2022-07-14T12:29:50.083480Z","iopub.status.idle":"2022-07-14T12:29:50.089933Z","shell.execute_reply.started":"2022-07-14T12:29:50.083249Z","shell.execute_reply":"2022-07-14T12:29:50.089097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_submission = pd.DataFrame({'radiant_win_prob': predictions}, \n                                 index=df_test_features.index)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:50.091408Z","iopub.execute_input":"2022-07-14T12:29:50.091905Z","iopub.status.idle":"2022-07-14T12:29:50.102049Z","shell.execute_reply.started":"2022-07-14T12:29:50.091858Z","shell.execute_reply":"2022-07-14T12:29:50.101335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\nparam_grid = [\n{'n_estimators': [10, 25], 'max_features': [5, 10], \n 'max_depth': [10, 50, None], 'bootstrap': [True, False]}\n]\n\nforest = RandomForestClassifier(n_jobs=-1)\ngrid_search_forest = GridSearchCV(forest, param_grid, cv=3, scoring='accuracy')\ngrid_search_forest.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:29:50.103622Z","iopub.execute_input":"2022-07-14T12:29:50.104281Z","iopub.status.idle":"2022-07-14T12:31:26.547117Z","shell.execute_reply.started":"2022-07-14T12:29:50.104220Z","shell.execute_reply":"2022-07-14T12:31:26.546244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search_forest.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.548470Z","iopub.execute_input":"2022-07-14T12:31:26.548798Z","iopub.status.idle":"2022-07-14T12:31:26.555342Z","shell.execute_reply.started":"2022-07-14T12:31:26.548734Z","shell.execute_reply":"2022-07-14T12:31:26.554669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_search_forest.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.556389Z","iopub.execute_input":"2022-07-14T12:31:26.556672Z","iopub.status.idle":"2022-07-14T12:31:26.570115Z","shell.execute_reply.started":"2022-07-14T12:31:26.556616Z","shell.execute_reply":"2022-07-14T12:31:26.568880Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_best = grid_search_forest.best_estimator_.predict_proba(X_test)[:, 1]","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.571671Z","iopub.execute_input":"2022-07-14T12:31:26.572205Z","iopub.status.idle":"2022-07-14T12:31:26.689698Z","shell.execute_reply.started":"2022-07-14T12:31:26.572138Z","shell.execute_reply":"2022-07-14T12:31:26.688798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"grid_best","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.690923Z","iopub.execute_input":"2022-07-14T12:31:26.691245Z","iopub.status.idle":"2022-07-14T12:31:26.697429Z","shell.execute_reply.started":"2022-07-14T12:31:26.691191Z","shell.execute_reply":"2022-07-14T12:31:26.696711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import RandomizedSearchCV\nfrom pprint import pprint\n\n# Number of trees in random forest\nn_estimators = [int(x) for x in np.linspace(start = 20, stop = 200, num = 5)]\n# Number of features to consider at every split\nmax_features = ['auto', 'sqrt']\n# Maximum number of levels in tree\nmax_depth = [int(x) for x in np.linspace(1, 45, num = 3)]\n# Minimum number of samples required to split a node\nmin_samples_split = [5, 10]\n\n# Create the random grid\nrandom_grid = {'n_estimators': n_estimators,\n               'max_features': max_features,\n               'max_depth': max_depth,\n               'min_samples_split': min_samples_split}\n\npprint(random_grid)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.698784Z","iopub.execute_input":"2022-07-14T12:31:26.699115Z","iopub.status.idle":"2022-07-14T12:31:26.712091Z","shell.execute_reply.started":"2022-07-14T12:31:26.699022Z","shell.execute_reply":"2022-07-14T12:31:26.710892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_random = RandomizedSearchCV(estimator = forest, param_distributions = random_grid, n_iter = 5, cv = 3, verbose=2, random_state=42, n_jobs = -1, scoring='accuracy')\n# Fit the random search model\nrf_random.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:31:26.714413Z","iopub.execute_input":"2022-07-14T12:31:26.714913Z","iopub.status.idle":"2022-07-14T12:32:03.039562Z","shell.execute_reply.started":"2022-07-14T12:31:26.714864Z","shell.execute_reply":"2022-07-14T12:32:03.038406Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_random.best_estimator_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:32:03.040632Z","iopub.execute_input":"2022-07-14T12:32:03.040914Z","iopub.status.idle":"2022-07-14T12:32:03.047833Z","shell.execute_reply.started":"2022-07-14T12:32:03.040863Z","shell.execute_reply":"2022-07-14T12:32:03.046895Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rf_random.best_score_","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:32:03.049458Z","iopub.execute_input":"2022-07-14T12:32:03.050178Z","iopub.status.idle":"2022-07-14T12:32:03.061209Z","shell.execute_reply.started":"2022-07-14T12:32:03.050115Z","shell.execute_reply":"2022-07-14T12:32:03.060433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### exercise : use the best params, fit a 5 folds rf model","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:32:03.062193Z","iopub.execute_input":"2022-07-14T12:32:03.062615Z","iopub.status.idle":"2022-07-14T12:32:03.071065Z","shell.execute_reply.started":"2022-07-14T12:32:03.062558Z","shell.execute_reply":"2022-07-14T12:32:03.070063Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2022-07-14T12:32:03.072682Z","iopub.execute_input":"2022-07-14T12:32:03.073189Z","iopub.status.idle":"2022-07-14T12:32:03.083456Z","shell.execute_reply.started":"2022-07-14T12:32:03.073134Z","shell.execute_reply":"2022-07-14T12:32:03.082592Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<img src='https://miro.medium.com/max/1400/0*VYAbVhmGMpzUC8hH.jpeg'>","metadata":{}},{"cell_type":"code","source":"%%time\nfrom xgboost import XGBClassifier\n\nmodel = XGBClassifier(max_depth=5, \n                      learning_rate=0.01, \n                      n_estimators=100, \n                      subsample=0.8, \n                      colsample_bytree=0.8)\n\nmodel.fit(X_train, y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# exercise\n# grid search parameters with xgb","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom xgboost import XGBClassifier\n\nmodel = XGBClassifier(max_depth=5, \n                      learning_rate=0.01, \n                      n_estimators=100, \n                      subsample=0.8, \n                      colsample_bytree=0.8)\n\nmodel.fit(X_train, y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lightgbm import LGBMClassifier\nmodel = LGBMClassifier() # try to google the import parameters of lightgbm\n\nmodel.fit(X_train, y_train)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# combined the predictions of rf, xgb and lgbm","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}