{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Rational\nDecision trees are most effective when a significant percentage of the features, observed individually, give an excellent information gain. \n\nHowever, the initial features we are given do not give significant information gain. For example, you cannot get valuable information just on the house's height or the basement's size.\nWe need some sophisticated features instead. That's precisely the thing that hidden layers in a neural network do - they provide us with better features.\n\nThe approach:\nWe train a neural network. Then we take the features from the last hidden layer and train with them a decision tree.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd \nimport tensorflow as tf\nfrom tensorflow.keras.models import Sequential\nfrom tensorflow.keras.layers import Dense\nfrom tensorflow.keras.activations import linear, relu, sigmoid\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import StandardScaler\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\n%matplotlib widget\nimport matplotlib.pyplot as plt\nimport xgboost as xgb\n\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data preparation","metadata":{}},{"cell_type":"markdown","source":"After a quick analysis over the data, I consider all the features significant enough","metadata":{}},{"cell_type":"code","source":"train_df = pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\ntest_df = pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_train = train_df.iloc[:1168, :].SalePrice\ny_cv = train_df.iloc[1168:, :].SalePrice","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Drop some columns:","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(['Id', 'SalePrice'], axis=1)\ntest_df = test_df.drop(['Id'], axis=1)\n\nlow_cardinality_cols = [cname for cname in train_df.columns if \n                                train_df[cname].nunique() < 9 and\n                                train_df[cname].dtype == \"object\"]\nnumeric_cols = [cname for cname in train_df.columns if \n                                train_df[cname].dtype in ['int64', 'float64']]\nmy_cols = low_cardinality_cols + numeric_cols\n\ntrain_df = train_df[my_cols]\ntest_df = test_df[my_cols]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Our data set consists categorical data, which will not be accepted by our algorithm. That is why we one-hot encode all the values. The training set has a different number of categorials vs the test set. This is why we need to align both sets.","metadata":{}},{"cell_type":"code","source":"train_dummy = pd.get_dummies(train_df)\ntest_dummy = pd.get_dummies(test_df)\ntrain_dummy, test_dummy = train_dummy.align(test_dummy, join='left', axis=1)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since we do not want to lose a whole column if data is missing, I am going to fill it with the mean value instead","metadata":{}},{"cell_type":"code","source":"my_imputer = SimpleImputer()\ntrain_imp = my_imputer.fit_transform(train_dummy)\ntest_imp = my_imputer.transform(test_dummy)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After one-hot encoding we end up with:","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\ntrain_imp = scaler.fit_transform(train_imp)\ntest_imp = scaler.transform(test_imp)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seperate the training set into\n* train set\n* cv set","metadata":{}},{"cell_type":"code","source":"df_train = train_imp[:1168, :] # 80% of training data set\ndf_cv = train_imp[1168:, :] # 20% of training data set","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training","metadata":{}},{"cell_type":"markdown","source":"In order to know whether the model underfits, we'd like to know the mean value of the price","metadata":{}},{"cell_type":"code","source":"look_mean = pd.read_csv('../input/house-prices-advanced-regression-techniques/train.csv')\nmean = look_mean[\"SalePrice\"].mean()\nmean","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The mean price is **180921**\n\n10% of 180 000 is **18 000**, which squared is exactly **3 240 000**. We are looking for a configuration that has **J error < 3 240 000**","metadata":{}},{"cell_type":"code","source":"model = Sequential(\n    [               \n        tf.keras.layers.Dense(200, activation='relu'),        \n        tf.keras.layers.Dense(100, activation='relu'),\n        tf.keras.layers.Dense(50, activation='relu'),\n        tf.keras.layers.Dense(25, activation='relu'),\n        tf.keras.layers.Dense(12, activation='relu', name='lastHidden'),\n        tf.keras.layers.Dense(1, activation='linear'),\n    ], name = \"my_model\" \n)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T16:06:24.199668Z","iopub.execute_input":"2022-08-11T16:06:24.200092Z","iopub.status.idle":"2022-08-11T16:06:24.214353Z","shell.execute_reply.started":"2022-08-11T16:06:24.200055Z","shell.execute_reply":"2022-08-11T16:06:24.213306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from keras import backend as K\n\noptimizer = keras.optimizers.Adam(learning_rate=0.1)\nmodel.compile(\n    loss=tf.keras.losses.MeanSquaredError(),\n    optimizer=optimizer)\n\nmodel.fit(df_train,y_train,epochs=100)\noptimizer.learning_rate.assign(0.01)\nmodel.fit(df_train,y_train,initial_epoch=100, epochs=250)\noptimizer.learning_rate.assign(0.001)\nmodel.fit(df_train,y_train,initial_epoch=250, epochs=350)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get the new representation of X_train","metadata":{}},{"cell_type":"code","source":"last_hidden_layer = keras.Model(inputs = model.input, outputs = model.get_layer('lastHidden').output)\ntrain_output = last_hidden_layer(df_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T16:08:24.846980Z","iopub.execute_input":"2022-08-11T16:08:24.847391Z","iopub.status.idle":"2022-08-11T16:08:24.865789Z","shell.execute_reply.started":"2022-08-11T16:08:24.847362Z","shell.execute_reply":"2022-08-11T16:08:24.864874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import GridSearchCV\n\nregressor = xgb.XGBRegressor()\n\nparam_grid = {\"max_depth\":    [5, 6, 7],\n              \"n_estimators\": [500, 600, 700]}\n\nsearch = GridSearchCV(regressor, param_grid).fit(train_output.numpy(), y_train)\n\nprint(\"The best hyperparameters are \",search.best_params_)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"regressor = xgb.XGBRegressor(max_depth = search.best_params_[\"max_depth\"],\n                            n_estimators = search.best_params_[\"n_estimators\"])\n\nregressor.fit(train_output, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T16:08:39.129757Z","iopub.execute_input":"2022-08-11T16:08:39.130147Z","iopub.status.idle":"2022-08-11T16:08:41.133692Z","shell.execute_reply.started":"2022-08-11T16:08:39.130114Z","shell.execute_reply":"2022-08-11T16:08:41.132780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cv_output = last_hidden_layer(df_cv)\nregressor.score(cv_output, y_cv)","metadata":{"execution":{"iopub.status.busy":"2022-08-11T16:08:44.031795Z","iopub.execute_input":"2022-08-11T16:08:44.032444Z","iopub.status.idle":"2022-08-11T16:08:44.054761Z","shell.execute_reply.started":"2022-08-11T16:08:44.032406Z","shell.execute_reply":"2022-08-11T16:08:44.053949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare for evaluation","metadata":{}},{"cell_type":"code","source":"test_output = last_hidden_layer(test_imp)\npredictions = regressor.predict(test_output)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"testtest = pd.read_csv('../input/house-prices-advanced-regression-techniques/test.csv')\ntestId = testtest['Id']","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"output = pd.DataFrame({'Id': testtest.Id, 'SalePrice': predictions})\noutput.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}