{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"A lot of the code here was inspired from Paul Mooney's [notebook on TF-DF](https://www.kaggle.com/code/paultimothymooney/getting-started-with-tensorflow-decision-forests). Make sure to check that out as well.","metadata":{}},{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"A learning algorithm trains a machine learning model on a training dataset. The parameters of a learning algorithm–called \"hyper-parameters\"–control how the model is trained and impact its quality. Therefore, finding the best hyper-parameters is an important stage of modeling.\n\nAutomated tuning algorithms work by generating and evaluating a large number of hyper-parameter values. Each of those iterations is called a \"trial\". The evaluation of a trial is expensive as it requires to train a new model each time. At the end of the tuning, the hyper-parameter with the best evaluation is used.\n\nTo demonstrate automated hyper-parameter tuning in TF-DF we'll be working with the Tabular Playground Series Feb 2021 Kaggle Dataset. It is a tabular dataset with 300,000 rows and 26 columns in training (93.66 MiB .CSV training dataset + 58.85 MiB .CSV test set) that is suitable for training algorithms to solve regression problems.\n\nWe'll be predicting a continuous target based on a number of feature columns given in the data. All of the feature columns, cat0 - cat9 are categorical, and the feature columns cont0 - cont13 are continuous.\n\nBy studying this tutorial you will learn how to quickly and automatically tune hyper-parameters of a GradientBoostedTrees model to perform a regression task using tabular data. We'll also learn how to blend the predictions from multiple tuned TF-DF models to achieve higher scores in the leaderboard.","metadata":{}},{"cell_type":"markdown","source":"# Installing TensorFlow Decision Forests","metadata":{}},{"cell_type":"code","source":"# Display only the messages with ERROR, CRITICAL log levels\n!pip install tensorflow_decision_forests -U -qq","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing libraries","metadata":{}},{"cell_type":"code","source":"# Import Python packages\nimport os\nimport numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn import preprocessing\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.linear_model import LinearRegression\nimport tensorflow as tf\nimport tensorflow_decision_forests as tfdf\nprint(\"TensorFlow Decision Forests v\" + tfdf.__version__)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Helper functions","metadata":{}},{"cell_type":"code","source":"# Define helper functions for plotting training evaluation curves\n\ndef plot_tfdf_model_training_curves(model):\n    # This function was adapted from the following tutorial:\n    # https://www.tensorflow.org/decision_forests/tutorials/beginner_colab\n    logs = model.make_inspector().training_logs()\n    plt.figure(figsize=(12, 4))\n    plt.subplot(1, 2, 1)\n    # Plot RMSE vs number of trees\n    plt.plot([log.num_trees for log in logs], [log.evaluation.rmse for log in logs])\n    plt.xlabel(\"Number of trees\")\n    plt.ylabel(\"RMSE (out-of-bag)\")\n    plt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print list of all data and files attached to this notebook\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load the dataset and convert it in a tf.Dataset","metadata":{}},{"cell_type":"markdown","source":"The dataset contains a mix of numerical (e.g. cont0 - cont13) and categorical (e.g. cat0 - cat9) features. TF-DF supports all these feature types natively (differently than NN based models), therefore there is no need for preprocessing in the form of one-hot encoding or normalization. Also by default the task is set to Classification in TF-DF, we'll change that to Regression.","metadata":{}},{"cell_type":"code","source":"# load to pandas dataframe (for data exploration)\ntrain_df = pd.read_csv('/kaggle/input/tabular-playground-series-feb-2021/train.csv')\ntest_df = pd.read_csv('/kaggle/input/tabular-playground-series-feb-2021/test.csv')\n\n# load to tensorflow dataset (for model training)\ntrain_tfds = tfdf.keras.pd_dataframe_to_tf_dataset(train_df, label=\"target\", task=tfdf.keras.Task.REGRESSION)\ntest_tfds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df, task=tfdf.keras.Task.REGRESSION)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Exploratory data analysis","metadata":{}},{"cell_type":"code","source":"# print column names\nprint(train_df.columns)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# preview first few rows of data\ntrain_df.head(10)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# print basic summary statistics\ntrain_df.describe()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# check for missing values\nsns.heatmap(train_df.isnull(), cbar=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Automatic hyper-parameter tuning","metadata":{}},{"cell_type":"markdown","source":"Hyper-paramter tuning is enabled by specifying the tuner constructor argument of the model. The tuner object contains all the configuration of the tuner (search space, optimizer, trial and objective).","metadata":{}},{"cell_type":"code","source":"# This code block was adapted from the following tutorial:\n# https://www.tensorflow.org/decision_forests/tutorials/automatic_tuning_colab\n\n# Configure the tuner.\n\n# Create a Random Search tuner with 10 trials for demo purpose, tune with\n# higher number of trials in practice.\ntuner = tfdf.tuner.RandomSearch(num_trials=10)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**min_examples**: Minimum number of examples in a node.","metadata":{}},{"cell_type":"code","source":"# Define the search space.\n#\n# Adding more parameters generaly improve the quality of the model, but make\n# the tuning last longer.\n\ntuner.choice(\"min_examples\", [2, 5, 7, 10])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**categorical_algorithm**: How to learn splits on categorical attributes.\n* CART: CART algorithm. Find categorical splits of the form \"value \\in mask\". The solution is exact for binary classification, regression and ranking. It is approximated for multi-class classification. This is a good first algorithm to use. In case of overfitting (very small dataset, large dictionary), the \"random\" algorithm is a good alternative.\n* ONE_HOT: One-hot encoding. Find the optimal categorical split of the form \"attribute == param\". This method is similar (but more efficient) than converting converting each possible categorical value into a boolean feature. This method is available for comparison purpose and generally performs worse than other alternatives.\n* RANDOM: Best splits among a set of random candidate. Find the a categorical split of the form \"value \\in mask\" using a random search. This solution can be seen as an approximation of the CART algorithm. This method is a strong alternative to CART. This algorithm is inspired from section \"5.1 Categorical Variables\" of \"Random Forest\", 2001. Default: \"CART\".","metadata":{}},{"cell_type":"code","source":"tuner.choice(\"categorical_algorithm\", [\"CART\", \"RANDOM\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**growing_strategy**: How to grow the tree.\n* LOCAL: Each node is split independently of the other nodes. In other words, as long as a node satisfy the splits \"constraints (e.g. maximum depth, minimum number of observations), the node will be split. This is the \"classical\" way to grow decision trees.\n* BEST_FIRST_GLOBAL: The node with the best loss reduction among all the nodes of the tree is selected for splitting. This method is also called \"best first\" or \"leaf-wise growth\". See \"Best-first decision tree learning\", Shi and \"Additive logistic regression : A statistical view of boosting\", Friedman for more details. Default: \"LOCAL\".","metadata":{}},{"cell_type":"code","source":"# Some hyper-parameters are only valid for specific values of other\n# hyper-parameters. For example, the \"max_depth\" parameter is mostly useful when\n# \"growing_strategy=LOCAL\" while \"max_num_nodes\" is better suited when\n# \"growing_strategy=BEST_FIRST_GLOBAL\".\n\nlocal_search_space = tuner.choice(\"growing_strategy\", [\"LOCAL\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**max_depth**: Maximum depth of the tree. max_depth=1 means that all trees will be roots. Negative values are ignored. Default: 6.","metadata":{}},{"cell_type":"code","source":"local_search_space.choice(\"max_depth\", [3, 4, 5, 6, 8])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**max_num_nodes**: Maximum number of nodes in the tree. Set to -1 to disable this limit. Only available for growing_strategy=BEST_FIRST_GLOBAL. Default: None.","metadata":{}},{"cell_type":"code","source":"# merge=True indicates that the parameter (here \"growing_strategy\") is already\n# defined, and that new values are added to it.\nglobal_search_space = tuner.choice(\"growing_strategy\", [\"BEST_FIRST_GLOBAL\"], merge=True)\nglobal_search_space.choice(\"max_num_nodes\", [16, 32, 64, 128, 256])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**shrinkage**: Coefficient applied to each tree prediction. A small value (0.02) tends to give more accurate results (assuming enough trees are trained), but results in larger models. Analogous to neural network learning rate. Default: 0.1.","metadata":{}},{"cell_type":"code","source":"tuner.choice(\"shrinkage\", [0.02, 0.05, 0.10, 0.15])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**num_candidate_attributes_ratio**: Ratio of attributes tested at each node. If set, it is equivalent to num_candidate_attributes = number_of_input_features x num_candidate_attributes_ratio. The possible values are between ]0, and 1] as well as -1. If not set or equal to -1, the num_candidate_attributes is used. Default: -1.0.","metadata":{}},{"cell_type":"code","source":"tuner.choice(\"num_candidate_attributes_ratio\", [0.2, 0.5, 0.9, 1.0])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**split_axis**: What structure of split to consider for numerical features.\n* AXIS_ALIGNED: Axis aligned splits (i.e. one condition at a time). This is the \"classical\" way to train a tree. Default value.\n* SPARSE_OBLIQUE: Sparse oblique splits (i.e. splits one a small number of features) from \"Sparse Projection Oblique Random Forests\", Tomita et al., 2020. Default: \"AXIS_ALIGNED\".","metadata":{}},{"cell_type":"code","source":"tuner.choice(\"split_axis\", [\"AXIS_ALIGNED\"])\n\noblique_space = tuner.choice(\"split_axis\", [\"SPARSE_OBLIQUE\"], merge=True)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**sparse_oblique_normalization**: For sparse oblique splits i.e. split_axis=SPARSE_OBLIQUE. Normalization applied on the features, before applying the sparse oblique projections.\n* NONE: No normalization.\n* STANDARD_DEVIATION: Normalize the feature by the estimated standard deviation on the entire train dataset. Also known as Z-Score normalization.\n* MIN_MAX: Normalize the feature by the range (i.e. max-min) estimated on the entire train dataset. Default: None.","metadata":{}},{"cell_type":"code","source":"oblique_space.choice(\"sparse_oblique_normalization\",\n                     [\"NONE\", \"STANDARD_DEVIATION\", \"MIN_MAX\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**sparse_oblique_weights**: For sparse oblique splits i.e. split_axis=SPARSE_OBLIQUE. Possible values:\n* BINARY: The oblique weights are sampled in {-1,1} (default).\n* CONTINUOUS: The oblique weights are be sampled in [-1,1]. Default: None.","metadata":{}},{"cell_type":"code","source":"oblique_space.choice(\"sparse_oblique_weights\", [\"BINARY\", \"CONTINUOUS\"])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**sparse_oblique_num_projections_exponent**: For sparse oblique splits i.e. split_axis=SPARSE_OBLIQUE. Controls of the number of random projections to test at each node as num_features^num_projections_exponent. Default: None.","metadata":{}},{"cell_type":"code","source":"oblique_space.choice(\"sparse_oblique_num_projections_exponent\", [1.0, 1.5])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In addition to having a default set of hyper-parameters, TF-DF also provides you with a list of additional hyper-parameter choices to consider.","metadata":{}},{"cell_type":"code","source":"print(tfdf.keras.GradientBoostedTreesModel.predefined_hyperparameters())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"However, here we'll not be using the default hyper-parameter choices and tune them to get the best results instead. Also by default the task is set to Classification in TF-DF, we'll change that to Regression.","metadata":{}},{"cell_type":"code","source":"# Tune the model. Notice the `tuner=tuner`.\ngb_model = tfdf.keras.GradientBoostedTreesModel(tuner=tuner, task=tfdf.keras.Task.REGRESSION)\ngb_model.fit(x=train_tfds, verbose=2)","metadata":{"_kg_hide-output":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Display the tuning logs.\ntuning_logs = gb_model.make_inspector().tuning_logs()\ntuning_logs.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Best hyper-parameters.\ntuning_logs[tuning_logs.best].iloc[0]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot the model","metadata":{}},{"cell_type":"markdown","source":"We plot the evoluation of the best score during the training and then the tuning.","metadata":{}},{"cell_type":"code","source":"plot_tfdf_model_training_curves(gb_model)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 5))\nplt.plot(tuning_logs[\"score\"], label=\"current trial\")\nplt.plot(tuning_logs[\"score\"].cummax(), label=\"best trial\")\nplt.xlabel(\"Tuning step\")\nplt.ylabel(\"Tuning score\")\nplt.legend()\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate the model","metadata":{}},{"cell_type":"markdown","source":"For this dataset, submissions are scored on the root mean squared error. Hence we evaluate the model on that metrics.","metadata":{}},{"cell_type":"code","source":"gb_model.compile(metrics=[tf.keras.metrics.RootMeanSquaredError()])","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inspector = gb_model.make_inspector()\ninspector.evaluation()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"gb_model.evaluate(train_tfds)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Model type:\", inspector.model_type())\nprint(\"Objective:\", inspector.objective())\nprint(\"Evaluation:\", inspector.evaluation())","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Variable importances generally indicates how much a variable contributes to the model predictions or quality. Variable importance SUM_SCORE is sum of the split scores using a specific feature. The larger, the most important.","metadata":{}},{"cell_type":"code","source":"inspector.variable_importances()[\"SUM_SCORE\"]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# gb_model.summary()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Blending","metadata":{}},{"cell_type":"markdown","source":"A lot of the code here was inspired from Abhishek Thakur's [notebook on blending](https://www.kaggle.com/code/abhishek/blending-blending-blending). Make sure to check that out as well.","metadata":{}},{"cell_type":"markdown","source":"Fit a model with the tuned best hyper-parameters to the 5-folds cross-validation dataset and get predictions.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/tpsfeb21-folds/train_folds.csv\")\ndf_test = pd.read_csv(\"../input/tabular-playground-series-feb-2021/test.csv\")\nsample_submission = pd.read_csv(\"../input/tabular-playground-series-feb-2021/sample_submission.csv\")\n\nuseful_features = [c for c in df.columns if c not in (\"id\", \"kfold\")]\nobject_cols = [col for col in useful_features if 'cat' in col]\nuseful_features_test = [c for c in df.columns if c not in (\"id\", \"target\", \"kfold\")]\ndf_test = df_test[useful_features_test]\n\nfinal_test_predictions = []\nfinal_valid_predictions = {}\nscores = []\nfor fold in range(5):\n    xtrain = df[df.kfold != fold].reset_index(drop=True)\n    xvalid = df[df.kfold == fold].reset_index(drop=True)\n    xtest = df_test.copy()\n    \n    valid_ids = xvalid.id.values.tolist()\n\n    ytrain = xtrain.target\n    yvalid = xvalid.target\n    \n    xtrain = xtrain[useful_features]\n    xvalid = xvalid[useful_features]\n    \n    ordinal_encoder = preprocessing.OrdinalEncoder()\n    xtrain[object_cols] = ordinal_encoder.fit_transform(xtrain[object_cols])\n    xvalid[object_cols] = ordinal_encoder.transform(xvalid[object_cols])\n    xtest[object_cols] = ordinal_encoder.transform(xtest[object_cols])\n    \n    xtrain_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xtrain, label=\"target\", task=tfdf.keras.Task.REGRESSION)\n    xvalid_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xvalid, label=\"target\", task=tfdf.keras.Task.REGRESSION)\n    xtest_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xtest, task=tfdf.keras.Task.REGRESSION)\n    \n    model = tfdf.keras.GradientBoostedTreesModel(min_examples=2, categorical_algorithm=\"RANDOM\", growing_strategy=\"BEST_FIRST_GLOBAL\", max_num_nodes=32, shrinkage=0.05, num_candidate_attributes_ratio=0.9, split_axis=\"SPARSE_OBLIQUE\", sparse_oblique_normalization=\"MIN_MAX\", sparse_oblique_weights=\"CONTINUOUS\", sparse_oblique_num_projections_exponent=1.5, task=tfdf.keras.Task.REGRESSION)\n    model.compile(metrics=[tf.keras.metrics.RootMeanSquaredError()])\n    model.fit(x=xtrain_ds)\n    \n    preds_valid = model.predict(xvalid_ds)\n    test_preds = model.predict(xtest_ds)\n    final_test_predictions.append(test_preds)\n    final_valid_predictions.update(dict(zip(valid_ids, preds_valid)))\n    rmse = mean_squared_error(yvalid, preds_valid, squared=False)\n    print(fold, rmse)\n    scores.append(rmse)\n\nprint(np.mean(scores), np.std(scores))\nfinal_valid_predictions = pd.DataFrame.from_dict(final_valid_predictions, orient=\"index\").reset_index()\nfinal_valid_predictions.columns = [\"id\", \"pred_1\"]\nfinal_valid_predictions.to_csv(\"train_pred_1.csv\", index=False)\n\nsample_submission.target = np.mean(np.column_stack(final_test_predictions), axis=1)\nsample_submission.columns = [\"id\", \"pred_1\"]\nsample_submission.to_csv(\"test_pred_1.csv\", index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fit another model, with a separate set of best hyper-parameters obtained by the same tuning process described above, to the 5-folds cross-validation dataset and get predictions.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/tpsfeb21-folds/train_folds.csv\")\ndf_test = pd.read_csv(\"../input/tabular-playground-series-feb-2021/test.csv\")\nsample_submission = pd.read_csv(\"../input/tabular-playground-series-feb-2021/sample_submission.csv\")\n\nuseful_features = [c for c in df.columns if c not in (\"id\", \"kfold\")]\nobject_cols = [col for col in useful_features if 'cat' in col]\nuseful_features_test = [c for c in df.columns if c not in (\"id\", \"target\", \"kfold\")]\ndf_test = df_test[useful_features_test]\n\nfinal_test_predictions = []\nfinal_valid_predictions = {}\nscores = []\nfor fold in range(5):\n    xtrain = df[df.kfold != fold].reset_index(drop=True)\n    xvalid = df[df.kfold == fold].reset_index(drop=True)\n    xtest = df_test.copy()\n    \n    valid_ids = xvalid.id.values.tolist()\n\n    ytrain = xtrain.target\n    yvalid = xvalid.target\n    \n    xtrain = xtrain[useful_features]\n    xvalid = xvalid[useful_features]\n    \n    ordinal_encoder = preprocessing.OrdinalEncoder()\n    xtrain[object_cols] = ordinal_encoder.fit_transform(xtrain[object_cols])\n    xvalid[object_cols] = ordinal_encoder.transform(xvalid[object_cols])\n    xtest[object_cols] = ordinal_encoder.transform(xtest[object_cols])\n    \n    xtrain_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xtrain, label=\"target\", task=tfdf.keras.Task.REGRESSION)\n    xvalid_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xvalid, label=\"target\", task=tfdf.keras.Task.REGRESSION)\n    xtest_ds = tfdf.keras.pd_dataframe_to_tf_dataset(xtest, task=tfdf.keras.Task.REGRESSION)\n    \n    model = tfdf.keras.GradientBoostedTreesModel(min_examples=5, categorical_algorithm=\"RANDOM\", growing_strategy=\"BEST_FIRST_GLOBAL\", max_num_nodes=256, shrinkage=0.15, num_candidate_attributes_ratio=0.5, split_axis=\"AXIS_ALIGNED\", task=tfdf.keras.Task.REGRESSION)\n    model.compile(metrics=[tf.keras.metrics.RootMeanSquaredError()])\n    model.fit(x=xtrain_ds)\n    \n    preds_valid = model.predict(xvalid_ds)\n    test_preds = model.predict(xtest_ds)\n    final_test_predictions.append(test_preds)\n    final_valid_predictions.update(dict(zip(valid_ids, preds_valid)))\n    rmse = mean_squared_error(yvalid, preds_valid, squared=False)\n    print(fold, rmse)\n    scores.append(rmse)\n\nprint(np.mean(scores), np.std(scores))\nfinal_valid_predictions = pd.DataFrame.from_dict(final_valid_predictions, orient=\"index\").reset_index()\nfinal_valid_predictions.columns = [\"id\", \"pred_2\"]\nfinal_valid_predictions.to_csv(\"train_pred_2.csv\", index=False)\n\nsample_submission.target = np.mean(np.column_stack(final_test_predictions), axis=1)\nsample_submission.columns = [\"id\", \"pred_2\"]\nsample_submission.to_csv(\"test_pred_2.csv\", index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Merge the predictions from both the models into a single dataframe.","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv(\"../input/tpsfeb21-folds/train_folds.csv\")\ndf_test = pd.read_csv(\"../input/tabular-playground-series-feb-2021/test.csv\")\nsample_submission = pd.read_csv(\"../input/tabular-playground-series-feb-2021/sample_submission.csv\")\n\ndf1 = pd.read_csv(\"train_pred_1.csv\")\ndf2 = pd.read_csv(\"train_pred_2.csv\")\n\ndf_test1 = pd.read_csv(\"test_pred_1.csv\")\ndf_test2 = pd.read_csv(\"test_pred_2.csv\")\n\ndf = df.merge(df1, on=\"id\", how=\"left\")\ndf = df.merge(df2, on=\"id\", how=\"left\")\n\ndf_test = df_test.merge(df_test1, on=\"id\", how=\"left\")\ndf_test = df_test.merge(df_test2, on=\"id\", how=\"left\")\n\ndf.head()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Fit a Linear Regression model over the predictions that finds the coefficients. This is generally a smarter way to blend than simply averaging over the predictions.","metadata":{}},{"cell_type":"code","source":"useful_features = [\"pred_1\", \"pred_2\"]\ndf_test = df_test[useful_features]\n\nfinal_predictions = []\nscores = []\nfor fold in range(5):\n    xtrain =  df[df.kfold != fold].reset_index(drop=True)\n    xvalid = df[df.kfold == fold].reset_index(drop=True)\n    xtest = df_test.copy()\n\n    ytrain = xtrain.target\n    yvalid = xvalid.target\n    \n    xtrain = xtrain[useful_features]\n    xvalid = xvalid[useful_features]\n    \n    model = LinearRegression()\n    model.fit(xtrain, ytrain)\n    \n    preds_valid = model.predict(xvalid)\n    test_preds = model.predict(xtest)\n    final_predictions.append(test_preds)\n    rmse = mean_squared_error(yvalid, preds_valid, squared=False)\n    print(fold, rmse)\n    scores.append(rmse)\n\nprint(np.mean(scores), np.std(scores))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Make submission","metadata":{}},{"cell_type":"code","source":"sample_submission.target = np.mean(np.column_stack(final_predictions), axis=1)\nsample_submission.to_csv(\"submission.csv\", index=False)","metadata":{},"execution_count":null,"outputs":[]}]}